Upload date
All time
Last hour
Today
This week
This month
This year
Type
All
Video
Channel
Playlist
Movie
Duration
Short (< 4 minutes)
Medium (4-20 minutes)
Long (> 20 minutes)
Sort by
Relevance
Rating
View count
Features
HD
Subtitles/CC
Creative Commons
3D
Live
4K
360°
VR180
HDR
98,473 results
Your local LLM may be running a quantized file you never chose. Q4, Q8, GGUF formats, KV cache precision, and Ollama or LM ...
56,329 views
2w ago
Stop worrying about GPU memory constraints. Learn how AI model quantization reduces model size and speeds up inference ...
362 views
13d ago
Applied AI Course: https://arpitbhayani.me/applied-ai System Design for SDE-2 and above: https://arpitbhayani.me/masterclass ...
22,820 views
Two AI model files can both say “4-bit” and still have completely different quality, memory requirements and performance.
86,453 views
3w ago
A fast-paced explanation of quantization's major concepts - covering bit-width, symmetric/asymmetric, granularity, ...
177 views
I quantized Qwen3.8-27B — and used the process to look at what “4-bit” actually means in practice. Quantization is usually ...
15,414 views
Welcome to Module 8. This is the deployment module — the one where everything we have built over the last seven modules ...
14 views
deeplearning #computervision.
674 views
7 views
RebelUI: https://github.com/RealRebelAI/RebelUI/tree/main BUYMEACOFFEE: buymeacoffee.com/realrebelai #quantization ...
3,709 views
Qwen 3.8 27B quantized to 8.4GB via GSQ-RCO still loses 9 coding points, here's what the benchmarks actually prove vs. Claude.
16,783 views
14h ago
Quantization can make an LLM smaller, faster, and cheaper to run, but the tradeoff is more subtle than a single benchmark score.
199 views
2d ago
LLM quantization explained simply — how can a huge AI model become 4× smaller and use dramatically less GPU memory?
130 views
4w ago
In this video, I explain LLM quantization from first principles, decode those cryptic Hugging Face filenames, and show you how to ...
6,042 views
Two for one! 00:00 Intro 00:27 Logic Pro 07:34 Logic Pro Swing 09:27 Logic Pro Input Quantization 11:10 Ableton Live 15:30 ...
258 views
Ternary Bonsai 2-27B compresses Qwen 3.8 27B's language backbone to roughly 1.7 bits per weight, producing a model file of ...
120,358 views
6d ago
Website-https://harrisonconsoles.com/ Forum-https://forum.harrisonconsoles.com/ ...
473 views
1d ago
You found a model to run locally, then saw a dozen cryptic files: Q4_K_M, Q5_K_S, Q8_0, IQ3, plus separate AWQ, GPTQ, and ...
27 views
When a 70-billion-parameter reasoning model quantized down to 1.58 bits maintains near-perfect perplexity on Wikitext while ...
8 views
10d ago
Ternary weights, post-training quantization, and a 3-bit KV cache. Three papers in one evening at the Bangalore Paper Club by ...
11,791 views
Show more