ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

98,473 results

Devsplainers
LLM Quantization Explained: The Q4 Quality Trap

Your local LLM may be running a quantized file you never chose. Q4, Q8, GGUF formats, KV cache precision, and Ollama or LM ...

9:21
LLM Quantization Explained: The Q4 Quality Trap

56,329 views

2w ago

LearnITGuide Tutorials
AI Model Quantization Explained: Run Large Models on Small Hardware 2026

Stop worrying about GPU memory constraints. Learn how AI model quantization reduces model size and speeds up inference ...

7:19
AI Model Quantization Explained: Run Large Models on Small Hardware 2026

362 views

13d ago

Arpit Bhayani
Quantization Fundamentals - How LLMs are Served Efficiently with Low Memory - Inference Engineering

Applied AI Course: https://arpitbhayani.me/applied-ai System Design for SDE-2 and above: https://arpitbhayani.me/masterclass ...

19:17
Quantization Fundamentals - How LLMs are Served Efficiently with Low Memory - Inference Engineering

22,820 views

2w ago

RepoChad
I Tested Every Qwen3.8-27B Quant: Here’s the Best One For You

Two AI model files can both say “4-bit” and still have completely different quality, memory requirements and performance.

10:01
I Tested Every Qwen3.8-27B Quant: Here’s the Best One For You

86,453 views

3w ago

Levi McCarryster
I Tried to Explain Quantization

A fast-paced explanation of quantization's major concepts - covering bit-width, symmetric/asymmetric, granularity, ...

19:50
I Tried to Explain Quantization

177 views

2w ago

Chris Hay
I Quantized Qwen3.8-27B. “4-Bit” Doesn’t Mean What You Think

I quantized Qwen3.8-27B — and used the process to look at what “4-bit” actually means in practice. Quantization is usually ...

15:20
I Quantized Qwen3.8-27B. “4-Bit” Doesn’t Mean What You Think

15,414 views

2w ago

The AI-Native
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Hands-On

Welcome to Module 8. This is the deployment module — the one where everything we have built over the last seven modules ...

12:20
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Hands-On

14 views

3w ago

AGI Lambda
Quantization-Aware Training (QAT) | Deep Learning

deeplearning #computervision.

5:18
Quantization-Aware Training (QAT) | Deep Learning

674 views

2w ago

The AI-Native
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Lecture

Welcome to Module 8. This is the deployment module — the one where everything we have built over the last seven modules ...

14:39
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Lecture

7 views

3w ago

REBEL AI
Auto-Quantize Models with RebelUI | LOW VRAM Quantization

RebelUI: https://github.com/RealRebelAI/RebelUI/tree/main BUYMEACOFFEE: buymeacoffee.com/realrebelai #quantization ...

12:24
Auto-Quantize Models with RebelUI | LOW VRAM Quantization

3,709 views

13d ago

The Stack
Claude Grade Coding From An 8.4GB Local AI Model

Qwen 3.8 27B quantized to 8.4GB via GSQ-RCO still loses 9 coding points, here's what the benchmarks actually prove vs. Claude.

14:24
Claude Grade Coding From An 8.4GB Local AI Model

16,783 views

14h ago

Callstack
Quantization impact on Apex performance | Artur Morys-Magiera & Lech Kalinowski at Agent Conf 2026

Quantization can make an LLM smaller, faster, and cheaper to run, but the tradeoff is more subtle than a single benchmark score.

19:39
Quantization impact on Apex performance | Artur Morys-Magiera & Lech Kalinowski at Agent Conf 2026

199 views

2d ago

Rack Base
LLM Quantization Explained: How AI Models Get 4× Smaller

LLM quantization explained simply — how can a huge AI model become 4× smaller and use dramatically less GPU memory?

7:49
LLM Quantization Explained: How AI Models Get 4× Smaller

130 views

4w ago

Ebenezer Don
Quantization: How to Run LLMs That Shouldn’t Fit on Your Hardware

In this video, I explain LLM quantization from first principles, decode those cryptic Hugging Face filenames, and show you how to ...

7:42
Quantization: How to Run LLMs That Shouldn’t Fit on Your Hardware

6,042 views

4w ago

Jon Pike Music
Quantizing Isn't That Complicated! (Logic Pro and Ableton)

Two for one! 00:00 Intro 00:27 Logic Pro 07:34 Logic Pro Swing 09:27 Logic Pro Input Quantization 11:10 Ableton Live 15:30 ...

16:35
Quantizing Isn't That Complicated! (Logic Pro and Ableton)

258 views

2w ago

RepoChad
Qwen3.8-27B on 6GB VRAM: Bonsai 27B is HERE!

Ternary Bonsai 2-27B compresses Qwen 3.8 27B's language backbone to roughly 1.7 bits per weight, producing a model file of ...

10:46
Qwen3.8-27B on 6GB VRAM: Bonsai 27B is HERE!

120,358 views

6d ago

Harrison Audio
Quantize MIDI In Mixbus 12

Website-https://harrisonconsoles.com/ Forum-https://forum.harrisonconsoles.com/ ...

3:26
Quantize MIDI In Mixbus 12

473 views

1d ago

BigIron AI
Quantization Explained: How Much Quality Do You REALLY Lose?

You found a model to run locally, then saw a dozen cryptic files: Q4_K_M, Q5_K_S, Q8_0, IQ3, plus separate AWQ, GPTQ, and ...

10:44
Quantization Explained: How Much Quality Do You REALLY Lose?

27 views

4w ago

Compute Node
Why Sub-2-Bit Quantization Destroys Reasoning (The Perplexity Illusion)

When a 70-billion-parameter reasoning model quantized down to 1.58 bits maintains near-perfect perplexity on Wikitext while ...

8:27
Why Sub-2-Bit Quantization Destroys Reasoning (The Perplexity Illusion)

8 views

10d ago

Conscious Engines
Ternary Weights, 3-Bit KV Caches, and the Limits of Quantization | Bangalore Paper Club

Ternary weights, post-training quantization, and a 3-bit KV cache. Three papers in one evening at the Bangalore Paper Club by ...

44:53
Ternary Weights, 3-Bit KV Caches, and the Limits of Quantization | Bangalore Paper Club

11,791 views

10d ago

Show more