Measured shelfAI & Machine Learning
AI and machine learning videos, ranked by how much of them is useful
- Videos measured
- 207
- Median padding
- 37%
- Useful part starts
- 1:07
- Typical runtime
- 22 min
Ranked from 207 measured videos last recalculated 2 Sep 2026
This is the category where the gap between reputation and substance is widest, because the subject rewards two very different videos that look identical from the outside.
The first kind spends its runtime on analogy. Neurons are like brain cells, attention is like paying attention, the model is like a very good autocomplete. It is pleasant, it is often accurate, and after forty minutes you cannot do anything you could not do before. The second kind uses one analogy to get you onto the ramp and then shows the actual arithmetic. Both get the same thumbnail treatment, and both are described as “explained”.
The measurement that separates them is density: what share of the runtime introduces something the previous minute did not already contain. Repetition of an idea in three different metaphors scores low here, and that is deliberate — it is the single most common way a fifteen-minute idea becomes a fifty-minute video in this subject.
Entries also carry a difficulty level, because “explained simply” and “explained” are not the same shelf, and the fastest way to waste an hour is to pick the wrong one for where you already are.
-
01
Let's build GPT: from scratch, in code, spelled out.
0:001:56:20- Density
- 91
- Padding
- 11%
- Useful from
- 1:28
The gold-standard deep-dive into the mechanics of building a transformer language model from scratch.
-
02
Deep Dive into LLMs like ChatGPT
0:003:31:24- Density
- 86
- Padding
- 14%
- Useful from
- 1:28
The gold-standard technical overview of how LLMs are actually built.
-
03
But what is a convolution?
0:0023:01- Density
- 84
- Padding
- 24%
- Useful from
- 1:41
A masterclass build-up of convolution — from dice probabilities to a genuinely surprising O(n log n) FFT algorithm.
-
04
Attention in transformers, step-by-step | Deep Learning Chapter 6
0:0026:10- Density
- 83
- Padding
- 21%
- Useful from
- 1:40
A masterfully clear, dense walkthrough of the attention mechanism that rewards careful watching with real technical understanding.
-
05
But what is a neural network? | Deep learning chapter 1
0:0018:40- Density
- 81
- Padding
- 26%
- Useful from
- 2:39
A rare, genuinely from-scratch, rigorous and re-derivable explanation of what a neural network's math actually is.
-
06
Transformers, the tech behind LLMs | Deep Learning Chapter 5
0:0027:14- Density
- 80
- Padding
- 26%
- Useful from
- 1:26
A masterclass primer on transformer internals — dense, rigorous, and exactly what its title promises.
-
07
But how do AI images and videos actually work? | Guest video by Welch Labs
0:0037:20- Density
- 82
- Padding
- 22%
- Useful from
- 3:28
A dense, physics-grounded explainer that actually teaches how diffusion models and CLIP combine to generate AI images and video — rare, high-value technical content.
-
08
But what is quantum computing? (Grover's Algorithm)
0:0036:54- Density
- 82
- Padding
- 24%
- Useful from
- 0:52
A rigorous, honest deep dive that builds Grover's algorithm from first principles — 3Blue1Brown at its best.
-
09
What is backpropagation really doing? | Deep learning chapter 3
0:0012:47- Density
- 80
- Padding
- 33%
- Useful from
- 0:53
A masterclass in building genuine intuition for backpropagation without a single formula — top-tier education.
-
10
[1hr Talk] Intro to Large Language Models
0:0059:48- Density
- 80
- Padding
- 25%
- Useful from
- 3:31
A dense, ad-free, masterclass-level explainer of how LLMs are built, used, and attacked — among the best general AI education videos available.
-
11
Backpropagation calculus | Deep Learning Chapter 4
0:0010:18- Density
- 80
- Padding
- 25%
- Useful from
- 0:32
A tight, rigorous derivation of backprop's chain-rule math — dense, durable, exactly what the title promises.
-
12
Gradient descent, how neural networks learn | Deep Learning Chapter 2
0:0020:33- Density
- 79
- Padding
- 28%
- Useful from
- 1:52
A masterclass in building genuine intuition for gradient descent, honest about the method's real limitations.
-
13
Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!
0:0036:15- Density
- 79
- Padding
- 23%
- Useful from
- 1:20
A rigorous, worked-numbers walkthrough of transformer internals that actually teaches how ChatGPT-style models work, not just what they do.
-
14
How I use LLMs
0:002:11:12- Density
- 86
- Padding
- 19%
- Useful from
- 0:35
An essential, high-utility masterclass on how to actually think about and interact with LLMs.
-
15
How might LLMs store facts | Deep Learning Chapter 7
0:0022:43- Density
- 82
- Padding
- 22%
- Useful from
- 0:58
A rigorous, self-contained walkthrough of how transformer MLP blocks might encode facts, capped with a genuinely illuminating dive into superposition and near-orthogonality in high…
-
16
The Most Important Algorithm in Machine Learning
0:0040:08- Density
- 78
- Padding
- 27%
- Useful from
- 1:28
A rigorous, ground-up derivation of backpropagation that earns its title without needing hype.
-
17
Intelligence is collective, not artificial — Prof. Michael I. Jordan (UC Berkeley / Inria)
0:001:17:10- Density
- 80
- Padding
- 24%
- Useful from
- 3:20
A rigorous, refreshing academic critique of AI hype and a proposal for a more robust, collectivist engineering framework.
-
18
Why Does Diffusion Work Better than Auto-Regression?
0:0020:18- Density
- 79
- Padding
- 32%
- Useful from
- 0:32
A rigorous, original explanation of why diffusion models generate images faster than autoregressive ones — genuinely worth the watch.
-
19
The Essential Main Ideas of Neural Networks
0:0018:54- Density
- 76
- Padding
- 29%
- Useful from
- 1:54
A rare, fully worked walkthrough of the actual math inside a neural network — StatQuest at its clearest.
-
20
Deep RL Bootcamp Lecture 4B Policy Gradients Revisited
0:0034:55- Density
- 76
- Padding
- 28%
- Useful from
- 1:08
A superb, intuitive deep dive into policy gradients with a real code walkthrough — genuinely teaches, nothing being sold.
-
21
Flow-Matching vs Diffusion Models explained side by side
0:0016:08- Density
- 75
- Padding
- 32%
- Useful from
- 0:33
A dense, well-structured technical breakdown that delivers exactly what its title promises: diffusion vs flow matching, math and all.
-
22
Diffusion Models | DDPM Explained
0:0029:29- Density
- 81
- Padding
- 20%
- Useful from
- 1:18
A rigorous, self-derived walkthrough of DDPM math that earns its 'explained' title with real depth, not just paper restatement.
-
23
Diffusion Models: DDPM | Generative AI Animated
0:0032:06- Density
- 79
- Padding
- 22%
- Useful from
- 1:37
A dense but genuinely rigorous derivation of DDPM from theory to working code — one of the clearer deep explainers on the topic.
-
24
Diffusion Models From Scratch | Score-Based Generative Models Explained | Math Explained
0:0038:11- Density
- 79
- Padding
- 27%
- Useful from
- 1:20
A rigorous, derivation-heavy walkthrough that genuinely explains where the diffusion model equations come from, not just what they are.
-
25
Gradient Descent, Step-by-Step
0:0023:54- Density
- 76
- Padding
- 28%
- Useful from
- 0:39
A genuinely rigorous, well-paced derivation of gradient descent that earns its step-by-step title.
Gistil's own measurement. Not a YouTube rating, and not the channel's position. Ranking is ours; the videos belong to their channels.
How this ranking is decided
Every video here was measured on the same axes: how much of its runtime carries information, how much is padding, and the second at which it starts paying off. The order is by value density, not by views, recency or how much we liked the channel. What the numbers mean →