ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

737,768 results

Stanford Online
Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 8 - LLM Evaluation

For more information about Stanford's graduate programs, visit: https://online.stanford.edu/graduate-education November 21, ...

1:49:25
Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 8 - LLM Evaluation

192,382 views

9mo ago

IBM Technology
LLM as a Judge: Scaling AI Evaluation Strategies

Ready to become a certified watsonx AI Assistant Engineer? Register now and use code IBMTechYT20 for 20% off of your exam ...

6:09
LLM as a Judge: Scaling AI Evaluation Strategies

51,699 views

1y ago

Dave Ebbelaar
How to Systematically Setup LLM Evals (Metrics, Unit Tests, LLM-as-a-Judge)

... 1:54 Understanding LLM Evaluations 4:54 Core Challenges in LLM Development 7:54 Importance of Iteration and Improvement ...

55:02
How to Systematically Setup LLM Evals (Metrics, Unit Tests, LLM-as-a-Judge)

114,322 views

1y ago

Evidently AI
LLM evaluation methods and metrics

What are the different methods to run automated LLM evaluations? 00:38 Ground truth-based vs. open-ended evals 00:53 ...

5:10
LLM evaluation methods and metrics

9,455 views

1y ago

IBM Technology
What are Large Language Model (LLM) Benchmarks?

Want to play with the technology yourself? Explore our interactive demo → https://ibm.biz/BdKetJ Learn more about the ...

6:21
What are Large Language Model (LLM) Benchmarks?

25,974 views

2y ago

Peter Yang
Complete Beginner's Course on AI Evaluations in 50 Minutes (2025) | Aman Khan

Copy my best AI workflows to save time and automate busywork: https://www.behindthecraft.com Subscribe to my practical AI ...

51:48
Complete Beginner's Course on AI Evaluations in 50 Minutes (2025) | Aman Khan

58,811 views

1y ago

Google Cloud Tech
The agent evaluation revolution

This video introduces a new series on testing AI agents, focusing on why traditional evaluation methods fall short for autonomous ...

8:32
The agent evaluation revolution

28,634 views

9mo ago

Matthew Berman
How to Setup LLM Evaluations Easily (Tutorial)

Learn more about Amazon Bedrock evaluations at http://bit.ly/45vM2hU Join My Newsletter for Regular AI Updates ...

17:55
How to Setup LLM Evaluations Easily (Tutorial)

21,249 views

1y ago

Logical Lenses
RAG Evaluation: Precision, Recall, Faithfulness, RAGAS Explained Clearly

Evaluating RAG systems requires more than checking how similar an answer is to a reference. Traditional metrics like BLEU, ...

12:19
RAG Evaluation: Precision, Recall, Faithfulness, RAGAS Explained Clearly

29,034 views

9mo ago

Databricks
Evaluating LLM-based Applications

Evaluating LLM-based applications can feel like more of an art than a science. In this workshop, we'll give a hands-on introduction ...

33:50
Evaluating LLM-based Applications

46,025 views

3y ago

Generative AI at MIT
LLM Evaluation Basics: Datasets & Metrics

This is an introduction to evaluating Large Language Models (LLMs), which covers what a dataset is, how we measure ...

5:18
LLM Evaluation Basics: Datasets & Metrics

17,781 views

3y ago

NeuralNine
Evaluate LLMs in Python with DeepEval

Today we learn how to easily and professionally evaluate LLMs in Python using DeepEval.

34:23
Evaluate LLMs in Python with DeepEval

19,242 views

1y ago

AI Engineer
Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop) — Taylor Jordan Smith

Accuracy scores and leaderboard metrics look impressive—but production-grade AI requires evals that reflect real-world ...

32:28
Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop) — Taylor Jordan Smith

17,370 views

1y ago

Show more