ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

38,624 results

RepoChad
I Tested Every Qwen3.8-27B Quant: Here’s the Best One For You

Two AI model files can both say “4-bit” and still have completely different quality, memory requirements and performance.

10:01
I Tested Every Qwen3.8-27B Quant: Here’s the Best One For You

87,825 views

4w ago

Arpit Bhayani
Quantization Fundamentals - How LLMs are Served Efficiently with Low Memory - Inference Engineering

Applied AI Course: https://arpitbhayani.me/applied-ai System Design for SDE-2 and above: https://arpitbhayani.me/masterclass ...

19:17
Quantization Fundamentals - How LLMs are Served Efficiently with Low Memory - Inference Engineering

23,551 views

2w ago

LearnITGuide Tutorials
AI Model Quantization Explained: Run Large Models on Small Hardware 2026

Stop worrying about GPU memory constraints. Learn how AI model quantization reduces model size and speeds up inference ...

7:19
AI Model Quantization Explained: Run Large Models on Small Hardware 2026

378 views

2w ago

Chris Hay
I Quantized Qwen3.8-27B. “4-Bit” Doesn’t Mean What You Think

I quantized Qwen3.8-27B — and used the process to look at what “4-bit” actually means in practice. Quantization is usually ...

15:20
I Quantized Qwen3.8-27B. “4-Bit” Doesn’t Mean What You Think

15,532 views

3w ago

Execute Automation
How I Ran Qwen3.8-27B GSQ-RCO on a 24GB Mac Mini !

In this video, I test Qwen3.8-27B GSQ-RCO IQ2_XS on an Apple M4 Mac Mini with 24GB RAM and explore what makes this ...

12:44
How I Ran Qwen3.8-27B GSQ-RCO on a 24GB Mac Mini !

17,358 views

3w ago

DeepWakeLabs
ISTA vs Unsloth: 7 Qwen3.8 27B Quants Tested—Which One Should You Download?

ISTA vs Unsloth: which Qwen3.8 27B quant should you actually download? I tested seven variants on an RTX 5090, and ISTA's ...

11:11
ISTA vs Unsloth: 7 Qwen3.8 27B Quants Tested—Which One Should You Download?

1,476 views

9h ago

Devsplainers
LLM Quantization Explained: The Q4 Quality Trap

Your local LLM may be running a quantized file you never chose. Q4, Q8, GGUF formats, KV cache precision, and Ollama or LM ...

9:21
LLM Quantization Explained: The Q4 Quality Trap

56,578 views

2w ago

Luke's Dev Lab
Qwen 3.8 27B GSQ RCO tested - 16GB Local LLM setup

In this video I will check out the GSQ RCO quant of Qwen 3.8 27B from ISTA DAS Lab Austria, can this quantization method ...

13:48
Qwen 3.8 27B GSQ RCO tested - 16GB Local LLM setup

67,761 views

3w ago

Bert Speaks About AI
Local AI Explained  Quants, INT8, Qwen3 8 and Model Sizes

How do you choose the right local AI model — and what do labels like 14B, INT8, Q6 and GGUF actually mean? In this video, I ...

10:18
Local AI Explained Quants, INT8, Qwen3 8 and Model Sizes

1,967 views

4w ago

REBEL AI
Auto-Quantize Models with RebelUI | LOW VRAM Quantization

RebelUI: https://github.com/RealRebelAI/RebelUI/tree/main BUYMEACOFFEE: buymeacoffee.com/realrebelai #quantization ...

12:24
Auto-Quantize Models with RebelUI | LOW VRAM Quantization

3,733 views

2w ago

The AI-Native
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Hands-On

Welcome to Module 8. This is the deployment module — the one where everything we have built over the last seven modules ...

12:20
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Hands-On

14 views

3w ago

SynthDad
The Quantizer Module That Plays Scales You've Never Heard — Xaoc Skopje

On the surface Skopje looks like a simple quantiser module. Two knobs, two channels, two outputs. But the real power isn't on the ...

17:26
The Quantizer Module That Plays Scales You've Never Heard — Xaoc Skopje

3,815 views

3w ago

The AI-Native
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Lecture

Welcome to Module 8. This is the deployment module — the one where everything we have built over the last seven modules ...

14:39
8-1 Optimization-Model-Quantization-Making-Big-Models-Fit-Real-Hardware Lecture

7 views

3w ago

Alex Hitt
NVIDIA Model Optimizer GitHub Explained: Quantization, Pruning & Faster LLM Inference

NVIDIA Model Optimizer GitHub by NVIDIA: https://github.com/NVIDIA/Model-Optimizer NVIDIA Model Optimizer helps engineers ...

8:44
NVIDIA Model Optimizer GitHub Explained: Quantization, Pruning & Faster LLM Inference

119 views

1d ago

Jon Pike Music
Quantizing Isn't That Complicated! (Logic Pro and Ableton)

Two for one! 00:00 Intro 00:27 Logic Pro 07:34 Logic Pro Swing 09:27 Logic Pro Input Quantization 11:10 Ableton Live 15:30 ...

16:35
Quantizing Isn't That Complicated! (Logic Pro and Ableton)

262 views

3w ago

Levi McCarryster
I Tried to Explain Quantization

A fast-paced explanation of quantization's major concepts - covering bit-width, symmetric/asymmetric, granularity, ...

19:50
I Tried to Explain Quantization

182 views

3w ago

AGI Lambda
Quantization-Aware Training (QAT) | Deep Learning

deeplearning #computervision.

5:18
Quantization-Aware Training (QAT) | Deep Learning

705 views

3w ago

Coding Horizon
Local AI Quantization Explained.

Four bit quantization can shrink a model's raw weight memory by about 75 percent. The catch is that “smaller” does not always ...

17:59
Local AI Quantization Explained.

433 views

1d ago

AUTOHOTKEY Gurus
Quantization Secrets: Rounding AI Numbers for Speed 🔢 | Hero Extract

Summary* The video explains how large language models represent words as numbers in vectors and use quantization by ...

9:05
Quantization Secrets: Rounding AI Numbers for Speed 🔢 | Hero Extract

179 views

2w ago

Physics & Engineering Insight
2: Quantization of the electromagnetic field and the harmonic oscillator
7:17
2: Quantization of the electromagnetic field and the harmonic oscillator

98 views

2w ago

Show more