Upload date
All time
Last hour
Today
This week
This month
This year
Type
All
Video
Channel
Playlist
Movie
Duration
Short (< 4 minutes)
Medium (4-20 minutes)
Long (> 20 minutes)
Sort by
Relevance
Rating
View count
Features
HD
Subtitles/CC
Creative Commons
3D
Live
4K
360°
VR180
HDR
243,866 results
LLM quantization is how a 70B model that needs 140GB of memory gets small enough to run on a normal GPU. Every model you ...
18,856 views
1 month ago
I quantized one model 8 ways to find the exact level it starts making things up. Take your personal data back with Incogni!
132,628 views
3 months ago
Text:* https://github.com/The-Pocket/PocketFlow-Tutorial-Video-Generator/blob/main/docs/llm/quantization.md 0:00:00 ...
10,931 views
9 months ago
00:00 Introduction to LLM Quantization 02:15 What is Quantization? 04:45 Post-Training Quantization (PTQ) vs. QAT 07:30 GPTQ ...
3,861 views
6 months ago
Every time I do a video about a model I get a comment saying "Well you never said what it takes to run it!" Well since I am not ...
14,906 views
5 months ago
Ever wondered how massive Large Language Models (LLMs) can run on your laptop or phone? The secret is Quantization!
282 views
8 months ago
Are you struggling with high-dimensional data in your vector database? In this video, we dive deep into Product Quantization (PQ) ...
1,035 views
7 months ago
Applied AI Course: https://arpitbhayani.me/applied-ai System Design for SDE-2 and above: https://arpitbhayani.me/masterclass ...
21,892 views
13 days ago
Timestamps: 00:00 - Intro 01:35 - First Look 03:11 - REAP Technical Look 05:58 - Q1 Quant Testing 07:16 - Q3 Quant Testing ...
7,841 views
In this video, we explain the fundamental differences between sampling and quantization in digital signal processing (DSP).
1,524 views
Two AI model files can both say “4-bit” and still have completely different quality, memory requirements and performance.
85,636 views
3 weeks ago
This video explains how to shrink massive neural networks to fit on mobile devices without sacrificing their performance. You will ...
1,244 views
4 months ago
Need some help with a project or some consulting? Contact me here: https://www.neuralnine.com/services The Python Bible ...
8,939 views
In this video, I show you how to understand the grid, common rhythmic values and quantization values used in music production.
1,834 views
LLM Quantization Explained: FP32, FP16, INT8, INT4, GPTQ, AWQ & GGUF What exactly is LLM quantization, and why is it so ...
333 views
Every token an LLM generates costs a full read of the model's weights out of memory: for a large model, hundreds of gigabytes ...
5,851 views
2 months ago
If you are reading the description, you found the hidden quantizer Most people skip this part, so here is your technical treat: ...
467 views
Frontier AI models are almost too big to use — a 70B model needs ~140 GB of memory just to hold its weights. So how do these ...
300 views
Brevitas Quantization Library - Pablo Monteagudo Lago, AMD Brevitas is an open‑source PyTorch library from AMD designed to ...
185 views
Learn how Unsloth Dynamic NVFP4 revolutionizes 4-bit Large Language Model (LLM) inference on NVIDIA Blackwell GPUs.
240 views
Show more