Measured shelfAI & Machine Learning

AI and machine learning videos, ranked by how much of them is useful

Median padding
37%
Useful part starts
1:08
Typical runtime
24 min

last recalculated 16 Sep 2026

This is the category where the gap between reputation and substance is widest, because the subject rewards two very different videos that look identical from the outside.

The first kind spends its runtime on analogy. Neurons are like brain cells, attention is like paying attention, the model is like a very good autocomplete. It is pleasant, it is often accurate, and after forty minutes you cannot do anything you could not do before. The second kind uses one analogy to get you onto the ramp and then shows the actual arithmetic. Both get the same thumbnail treatment, and both are described as “explained”.

The measurement that separates them is density: what share of the runtime introduces something the previous minute did not already contain. Repetition of an idea in three different metaphors scores low here, and that is deliberate — it is the single most common way a fifteen-minute idea becomes a fifty-minute video in this subject.

Entries also carry a difficulty level, because “explained simply” and “explained” are not the same shelf, and the fastest way to waste an hour is to pick the wrong one for where you already are.

  1. 01

    Let's reproduce GPT-2 (124M)

    Worth it AndrejKarpathy 4:01:26 advanced

    Density
    84
    Padding
    27%
    Useful from
    3:34

    A code-complete, first-principles GPT-2 reproduction — one of the most valuable hands-on deep learning tutorials available.

    Start at 3:34 →

  2. 02

    But what is a convolution?

    Worth it 3Blue1Brown 23:01 intermediate

    Density
    84
    Padding
    24%
    Useful from
    1:41

    A masterclass build-up of convolution — from dice probabilities to a genuinely surprising O(n log n) FFT algorithm.

    Start at 1:41 →

  3. 03

    Attention in transformers, step-by-step | Deep Learning Chapter 6

    Worth it 3Blue1Brown 26:10 intermediate

    Density
    83
    Padding
    21%
    Useful from
    1:40

    A masterfully clear, dense walkthrough of the attention mechanism that rewards careful watching with real technical understanding.

    Start at 1:40 →

  4. 04

    But what is a neural network? | Deep learning chapter 1

    Worth skimming 3Blue1Brown 18:40 beginner

    Density
    81
    Padding
    26%
    Useful from
    2:39

    A rare, genuinely from-scratch, rigorous and re-derivable explanation of what a neural network's math actually is.

    Start at 2:39 →

  5. 05

    Transformers, the tech behind LLMs | Deep Learning Chapter 5

    Worth it 3Blue1Brown 27:14 intermediate

    Density
    80
    Padding
    26%
    Useful from
    1:26

    A masterclass primer on transformer internals — dense, rigorous, and exactly what its title promises.

    Start at 1:26 →

  6. 06

    Deep Dive into LLMs like ChatGPT

    Worth it AndrejKarpathy 3:31:24 intermediate

    Density
    82
    Padding
    27%
    Useful from
    1:07

    One of the clearest, most complete plain-English explanations available of how ChatGPT-style LLMs are actually built and why they behave the way they do.

    Start at 1:07 →

  7. 07

    Let's build GPT: from scratch, in code, spelled out.

    Worth skimming AndrejKarpathy 1:56:20 advanced

    Density
    82
    Padding
    35%
    Useful from
    14:11

    A masterclass build-along: real working GPT code, the actual mechanics behind ChatGPT, with almost no filler.

    Start at 14:11 →

  8. 08

    But how do AI images and videos actually work? | Guest video by Welch Labs

    Worth skimming 3Blue1Brown 37:20 advanced

    Density
    82
    Padding
    22%
    Useful from
    3:28

    A dense, physics-grounded explainer that actually teaches how diffusion models and CLIP combine to generate AI images and video — rare, high-value technical content.

    Start at 3:28 →

  9. 09

    But what is quantum computing? (Grover's Algorithm)

    Worth it 3Blue1Brown 36:54 intermediate

    Density
    82
    Padding
    24%
    Useful from
    0:52

    A rigorous, honest deep dive that builds Grover's algorithm from first principles — 3Blue1Brown at its best.

    Start at 0:52 →

  10. 10

    What is backpropagation really doing? | Deep learning chapter 3

    Worth it 3Blue1Brown 12:47 intermediate

    Density
    80
    Padding
    33%
    Useful from
    0:53

    A masterclass in building genuine intuition for backpropagation without a single formula — top-tier education.

    Start at 0:53 →

  11. 11

    [1hr Talk] Intro to Large Language Models

    Worth it AndrejKarpathy 59:48 beginner

    Density
    80
    Padding
    25%
    Useful from
    3:31

    A dense, ad-free, masterclass-level explainer of how LLMs are built, used, and attacked — among the best general AI education videos available.

    Start at 3:31 →

  12. 12

    Backpropagation calculus | Deep Learning Chapter 4

    Worth it 3Blue1Brown 10:18 intermediate

    Density
    80
    Padding
    25%
    Useful from
    0:32

    A tight, rigorous derivation of backprop's chain-rule math — dense, durable, exactly what the title promises.

    Start at 0:32 →

  13. 13

    Gradient descent, how neural networks learn | Deep Learning Chapter 2

    Worth it 3Blue1Brown 20:33 intermediate

    Density
    79
    Padding
    28%
    Useful from
    1:52

    A masterclass in building genuine intuition for gradient descent, honest about the method's real limitations.

    Start at 1:52 →

  14. 14

    Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!

    Worth it StatQuest with Josh Starmer 36:15 intermediate

    Density
    79
    Padding
    23%
    Useful from
    1:20

    A rigorous, worked-numbers walkthrough of transformer internals that actually teaches how ChatGPT-style models work, not just what they do.

    Start at 1:20 →

  15. 15

    Let's build the GPT Tokenizer

    Worth it AndrejKarpathy 2:13:35 advanced

    Density
    82
    Padding
    31%
    Useful from
    4:22

    An exceptionally dense, hands-on build of a real GPT tokenizer that also explains — with concrete transcript-verifiable examples — why tokenization is the hidden cause of most weird LLM…

    Start at 4:22 →

  16. 16

    How might LLMs store facts | Deep Learning Chapter 7

    Worth it 3Blue1Brown 22:43 intermediate

    Density
    82
    Padding
    22%
    Useful from
    0:58

    A rigorous, self-contained walkthrough of how transformer MLP blocks might encode facts, capped with a genuinely illuminating dive into superposition and near-orthogonality in high…

    Start at 0:58 →

  17. 17

    The Most Important Algorithm in Machine Learning

    Worth it Artem Kirsanov 40:08 intermediate

    Density
    78
    Padding
    27%
    Useful from
    1:28

    A rigorous, ground-up derivation of backpropagation that earns its title without needing hype.

    Start at 1:28 →

  18. 18

    Why Does Diffusion Work Better than Auto-Regression?

    Worth it Algorithmic Simplicity 20:18 intermediate

    Density
    79
    Padding
    32%
    Useful from
    0:32

    A rigorous, original explanation of why diffusion models generate images faster than autoregressive ones — genuinely worth the watch.

    Start at 0:32 →

  19. 19

    The Essential Main Ideas of Neural Networks

    Worth it StatQuest with Josh Starmer 18:54 beginner

    Density
    76
    Padding
    29%
    Useful from
    1:54

    A rare, fully worked walkthrough of the actual math inside a neural network — StatQuest at its clearest.

    Start at 1:54 →

  20. 20

    Deep RL Bootcamp Lecture 4B Policy Gradients Revisited

    Worth it AI Prism 34:55 intermediate

    Density
    76
    Padding
    28%
    Useful from
    1:08

    A superb, intuitive deep dive into policy gradients with a real code walkthrough — genuinely teaches, nothing being sold.

    Start at 1:08 →

  21. 21

    Flow-Matching vs Diffusion Models explained side by side

    Worth it AI Coffee Break with Letitia 16:08 advanced

    Density
    75
    Padding
    32%
    Useful from
    0:33

    A dense, well-structured technical breakdown that delivers exactly what its title promises: diffusion vs flow matching, math and all.

    Start at 0:33 →

  22. 22

    Diffusion Models | DDPM Explained

    Worth it ExplainingAI 29:29 advanced

    Density
    81
    Padding
    20%
    Useful from
    1:18

    A rigorous, self-derived walkthrough of DDPM math that earns its 'explained' title with real depth, not just paper restatement.

    Start at 1:18 →

  23. 23

    Diffusion Models: DDPM | Generative AI Animated

    Worth it Deepia 32:06 advanced

    Density
    79
    Padding
    22%
    Useful from
    1:37

    A dense but genuinely rigorous derivation of DDPM from theory to working code — one of the clearer deep explainers on the topic.

    Start at 1:37 →

  24. 24

    Diffusion Models From Scratch | Score-Based Generative Models Explained | Math Explained

    Worth it Outlier 38:11 advanced

    Density
    79
    Padding
    27%
    Useful from
    1:20

    A rigorous, derivation-heavy walkthrough that genuinely explains where the diffusion model equations come from, not just what they are.

    Start at 1:20 →

  25. 25

    The AI Progress Chart Everyone Is Misreading — Beth Barnes & David Rein

    Worth it MachineLearningStreetTalk 1:53:27 advanced

    Density
    78
    Padding
    28%
    Useful from
    3:58

    A rare, unusually candid deep-dive by the researchers themselves into how the AI 'progress chart' is built, where it breaks, and why the public keeps over-reading it.

    Start at 3:58 →

Gistil's own measurement. Not a YouTube rating, and not the channel's position. Ranking is ours; the videos belong to their channels.

How this ranking is decided

Every video here was measured on the same axes — how much of its runtime carries information, how much is padding, and the second at which it starts paying off. The order is by value density, not by views, recency or how much we liked the channel. What the numbers mean →

← Every measured shelf