ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

243,866 results

KodeKloud
LLM Quantization Explained

LLM quantization is how a 70B model that needs 140GB of memory gets small enough to run on a normal GPU. Every model you ...

4:18
LLM Quantization Explained

18,856 views

1 month ago

Alex Ziskind
Everything looks fine at 4-bit

I quantized one model 8 ways to find the exact level it starts making things up. Take your personal data back with Incogni!

18:26
Everything looks fine at 4-bit

132,628 views

3 months ago

Zachary Huang
Give me 30 min, I will make Quantization click forever

Text:* https://github.com/The-Pocket/PocketFlow-Tutorial-Video-Generator/blob/main/docs/llm/quantization.md 0:00:00 ...

32:42
Give me 30 min, I will make Quantization click forever

10,931 views

9 months ago

Tales Of Tensors
LLM Quantization Explained: GPTQ, AWQ, QLoRA, GGUF and More

00:00 Introduction to LLM Quantization 02:15 What is Quantization? 04:45 Post-Training Quantization (PTQ) vs. QAT 07:30 GPTQ ...

30:14
LLM Quantization Explained: GPTQ, AWQ, QLoRA, GGUF and More

3,861 views

6 months ago

Tim Carambat
How Do We Get MASSIVE Model To Run On Device? Quantization Explained.

Every time I do a video about a model I get a comment saying "Well you never said what it takes to run it!" Well since I am not ...

26:41
How Do We Get MASSIVE Model To Run On Device? Quantization Explained.

14,906 views

5 months ago

CodeLucky
Quantization Explained: How to Run Large AI Models on Small Devices

Ever wondered how massive Large Language Models (LLMs) can run on your laptop or phone? The secret is Quantization!

4:05
Quantization Explained: How to Run Large AI Models on Small Devices

282 views

8 months ago

Cloud and Coffee with Navnit
Day 28: Product Quantization (PQ) Explained: HNSW vs IVF vs PQ vs LSH – Which Should You Use?

Are you struggling with high-dimensional data in your vector database? In this video, we dive deep into Product Quantization (PQ) ...

6:56
Day 28: Product Quantization (PQ) Explained: HNSW vs IVF vs PQ vs LSH – Which Should You Use?

1,035 views

7 months ago

Arpit Bhayani
Quantization Fundamentals - How LLMs are Served Efficiently with Low Memory - Inference Engineering

Applied AI Course: https://arpitbhayani.me/applied-ai System Design for SDE-2 and above: https://arpitbhayani.me/masterclass ...

19:17
Quantization Fundamentals - How LLMs are Served Efficiently with Low Memory - Inference Engineering

21,892 views

13 days ago

Bijan Bowen
GLM-4.7 218B Cerebras REAP – Local Quantization Testing (1-bit, 3-bit & 4-bit)

Timestamps: 00:00 - Intro 01:35 - First Look 03:11 - REAP Technical Look 05:58 - Q1 Quant Testing 07:16 - Q3 Quant Testing ...

20:24
GLM-4.7 218B Cerebras REAP – Local Quantization Testing (1-bit, 3-bit & 4-bit)

7,841 views

8 months ago

Learn Signal Processing
Sampling & Quantization Explained in 5 Minutes (DSP Basics)

In this video, we explain the fundamental differences between sampling and quantization in digital signal processing (DSP).

6:31
Sampling & Quantization Explained in 5 Minutes (DSP Basics)

1,524 views

7 months ago

RepoChad
I Tested Every Qwen3.8-27B Quant: Here’s the Best One For You

Two AI model files can both say “4-bit” and still have completely different quality, memory requirements and performance.

10:01
I Tested Every Qwen3.8-27B Quant: Here’s the Best One For You

85,636 views

3 weeks ago

DeepManim
What is quantization aware training ?

This video explains how to shrink massive neural networks to fit on mobile devices without sacrificing their performance. You will ...

3:28
What is quantization aware training ?

1,244 views

4 months ago

NeuralNine
From 15GB to 4.7GB: Quantizing AI Models Locally

Need some help with a project or some consulting? Contact me here: https://www.neuralnine.com/services The Python Bible ...

13:42
From 15GB to 4.7GB: Quantizing AI Models Locally

8,939 views

5 months ago

MusicTechHelpGuy
Understanding the Grid & Quantization Values [Part 9]

In this video, I show you how to understand the grid, common rhythmic values and quantization values used in music production.

18:07
Understanding the Grid & Quantization Values [Part 9]

1,834 views

8 months ago

Liv4IT
📦 LLM Quantization Explained: FP32, FP16, INT8, INT4, GPTQ, AWQ & GGUF

LLM Quantization Explained: FP32, FP16, INT8, INT4, GPTQ, AWQ & GGUF What exactly is LLM quantization, and why is it so ...

20:52
📦 LLM Quantization Explained: FP32, FP16, INT8, INT4, GPTQ, AWQ & GGUF

333 views

1 month ago

PY
The Engineering Behind LLM Inference: Quantization

Every token an LLM generates costs a full read of the model's weights out of memory: for a large model, hundreds of gigabytes ...

20:51
The Engineering Behind LLM Inference: Quantization

5,851 views

2 months ago

ML Guy
Why 4-Bit AI Models Still Work (Quantization Explained)

If you are reading the description, you found the hidden quantizer Most people skip this part, so here is your technical treat: ...

4:51
Why 4-Bit AI Models Still Work (Quantization Explained)

467 views

3 months ago

scrollypedia
Quantization vs Distillation: How Big AI Models Get Small

Frontier AI models are almost too big to use — a 70B model needs ~140 GB of memory just to hold its weights. So how do these ...

7:24
Quantization vs Distillation: How Big AI Models Get Small

300 views

3 months ago

PyTorch
Brevitas Quantization Library - Pablo Monteagudo Lago, AMD

Brevitas Quantization Library - Pablo Monteagudo Lago, AMD Brevitas is an open‑source PyTorch library from AMD designed to ...

30:20
Brevitas Quantization Library - Pablo Monteagudo Lago, AMD

185 views

5 months ago

Micro Learning
Unsloth Dynamic NVFP4 Explained | 4-Bit LLM Quantization for NVIDIA Blackwell, vLLM & SGLang

Learn how Unsloth Dynamic NVFP4 revolutionizes 4-bit Large Language Model (LLM) inference on NVIDIA Blackwell GPUs.

10:48
Unsloth Dynamic NVFP4 Explained | 4-Bit LLM Quantization for NVIDIA Blackwell, vLLM & SGLang

240 views

2 months ago

Show more