Measured shelfAI & Machine Learning

Transformer explainers, ranked by where the analogy ends

Videos measured
27
Median padding
35%
Useful part starts
1:28
Typical runtime
28 min

Ranked from 27 measured videos last recalculated 4 Aug 2026

Every video on this shelf has to solve the same problem: attention is a mechanism made of matrix multiplication, and matrix multiplication is not something you can watch. So each of them picks a metaphor — a lookup table, a search engine, a room where words vote on each other — and the entire quality of the video is decided by what happens after the metaphor.

The ones that rank highly here spend the metaphor quickly and then show the actual operation: what the three projections are, what shape the things being multiplied have, why the scaling term is there, what the mask does. The ones lower down keep the metaphor going for forty minutes, restated in progressively more elaborate ways, and never show a number.

Both types are titled “transformers explained”. The density measurement is what separates them, because a metaphor restated is by definition not new information, and it is the single most common way this topic becomes an hour of video.

Entries carry a difficulty level, and on this topic it matters more than anywhere else on the site: a video that is perfect for someone who already knows what a dot product does is a wasted hour for someone who does not, and the reverse is worse.

  1. 01

    Let's build GPT: from scratch, in code, spelled out.

    Worth it AndrejKarpathy 1:56:20 intermediate

    Density
    91
    Padding
    11%
    Useful from
    1:28

    The gold-standard deep-dive into the mechanics of building a transformer language model from scratch.

    Start at 1:28 →

  2. 02

    Deep Dive into LLMs like ChatGPT

    Worth it AndrejKarpathy 3:31:24 intermediate

    Density
    86
    Padding
    14%
    Useful from
    1:28

    The gold-standard technical overview of how LLMs are actually built.

    Start at 1:28 →

  3. 03

    Attention in transformers, step-by-step | Deep Learning Chapter 6

    Worth it 3Blue1Brown 26:10 intermediate

    Density
    83
    Padding
    21%
    Useful from
    1:40

    A masterfully clear, dense walkthrough of the attention mechanism that rewards careful watching with real technical understanding.

    Start at 1:40 →

  4. 04

    Transformers, the tech behind LLMs | Deep Learning Chapter 5

    Worth it 3Blue1Brown 27:14 intermediate

    Density
    80
    Padding
    26%
    Useful from
    1:26

    A masterclass primer on transformer internals — dense, rigorous, and exactly what its title promises.

    Start at 1:26 →

  5. 05

    Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!

    Worth it StatQuest with Josh Starmer 36:15 intermediate

    Density
    79
    Padding
    23%
    Useful from
    1:20

    A rigorous, worked-numbers walkthrough of transformer internals that actually teaches how ChatGPT-style models work, not just what they do.

    Start at 1:20 →

  6. 06

    Large Language Models explained briefly

    Worth it 3Blue1Brown 7:58 intermediate

    Density
    75
    Padding
    26%
    Useful from
    0:33

    A masterfully compressed, accurate primer on how LLMs actually work — dense, honest, no sales pitch.

    Start at 0:33 →

  7. 07

    Transformers Step-by-Step Explained (Attention Is All You Need)

    Worth skimming ByteByteGo 10:04 intermediate

    Density
    70
    Padding
    39%
    Useful from
    2:40

    A tight, genuinely educational explainer of Transformer attention with a real worked example, lightly interrupted by a disclosed sponsor read.

    Start at 2:40 →

  8. 08

    Transformer Neural Networks - EXPLAINED! (Attention is all you need)

    Worth it CodeEmporium 13:05 intermediate

    Density
    68
    Padding
    31%
    Useful from
    1:54

    A dense, well-structured conceptual walkthrough of transformer architecture that earns its 'EXPLAINED' title with zero filler.

    Start at 1:54 →

  9. 09

    Illustrated Guide to Transformers Neural Network: A step by step explanation

    Worth it The AI Hacker 15:01 intermediate

    Density
    72
    Padding
    31%
    Useful from
    1:03

    A tight, accurate conceptual tour of the Transformer's internals — solid teaching, though it retreads familiar illustrated-guide territory rather than breaking new ground.

    Start at 1:03 →

  10. 10

    Transformers explained | The architecture behind LLMs

    Worth it AI Coffee Break with Letitia 19:48 intermediate

    Density
    71
    Padding
    33%
    Useful from
    0:38

    A genuinely dense, accurate transformer explainer that earns its title with real mechanics, not hype.

    Start at 0:38 →

  11. 11

    Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman

    Worth skimming Lex Clips 8:38 intermediate

    Density
    69
    Padding
    34%
    Useful from
    2:44

    A sharp, dense breakdown of why the Transformer works — one of the clearer plain-language explanations of its design philosophy, if brief.

    Start at 2:44 →

  12. 12

    Transformers, explained: Understand the model behind GPT, BERT, and T5

    Worth it Google Cloud Tech 9:11 intermediate

    Density
    67
    Padding
    34%
    Useful from
    1:25

    A genuinely solid, jargon-light explainer of transformer architecture that earns its title without ever really selling anything.

    Start at 1:25 →

  13. 13

    Transformers for beginners | What are they and how do they work

    Worth it AssemblyAI 19:59 beginner

    Density
    68
    Padding
    29%
    Useful from
    0:31

    Solid, math-grounded beginner explainer of transformer internals, lightly bookended by the channel's own API plug.

    Start at 0:31 →

  14. 14

    Transformers Explained | Simple Explanation of Transformers

    Worth it codebasics 57:31 intermediate

    Density
    67
    Padding
    27%
    Useful from
    1:36

    A patient, analogy-heavy but genuinely thorough walkthrough of Transformer internals — worth the long runtime if you already know your deep learning basics.

    Start at 1:36 →

  15. 15

    How does AI actually work? Transformers explained

    Worth it AI Search 32:21 intermediate

    Density
    66
    Padding
    30%
    Useful from
    0:31

    A solid, honestly-titled conceptual explainer of Transformer architecture, weakened only by redundant recaps and a mid-video sponsor detour.

    Start at 0:31 →

  16. 16

    Transformer Architecture Explained 'Attention Is All You Need'

    Worth it ByteMonk 12:49 intermediate

    Density
    62
    Padding
    34%
    Useful from
    0:47

    A clear, well-paced conceptual primer on Transformer attention — not groundbreaking, but a genuinely solid explainer worth the 13 minutes for newcomers to the architecture.

    Start at 0:47 →

  17. 17

    Transformers, explained: Understand the model behind ChatGPT

    Worth skimming Leon Petrou 24:07 beginner

    Density
    61
    Padding
    28%
    Useful from
    3:29

    A clear, accessible mental model of how Transformers work end-to-end — skips the real attention math (no Q/K/V) but is genuinely useful as a conceptual primer.

    Start at 3:29 →

  18. 18

    What are Large Language Models (LLMs)?

    Worth skimming Google for Developers 5:30 beginner

    Density
    56
    Padding
    41%
    Useful from
    0:32

    A tight, honest beginner explainer of LLMs and prompt design — light on depth but dense and accurate for its length.

    Start at 0:32 →

  19. 19

    Everything You Need To Know About Large Language Models (LLMs)

    Worth skimming Matthew Berman 25:20 beginner

    Density
    58
    Padding
    39%
    Useful from
    0:32

    A solid, broad beginner's overview of LLM mechanics and history, only lightly diluted by a sponsor segment for AI Camp.

    Start at 0:32 →

  20. 20

    The Transformer architecture

    Worth skimming Hugging Face 2:45 beginner

    Density
    50
    Padding
    47%
    Useful from
    0:59

    A clean, honest, high-level primer that sets up the series without pretending to teach the deep mechanics yet.

    Start at 0:59 →

  21. 21

    What are Transformers (Machine Learning Model)?

    Worth skimming IBM Technology 5:51 beginner

    Density
    51
    Padding
    46%
    Useful from
    1:17

    A clear, accurate but fairly standard conceptual primer on transformers — solid intro, low novelty.

    Start at 1:17 →

  22. 22

    Transformer Explained

    Worth skimming Caleb Writes Code 6:55 intermediate

    Density
    55
    Padding
    40%
    Useful from
    2:07

    A solid conceptual primer on transformer limitations and fixes, but it openly admits it skips the actual mechanics the title implies.

    Start at 2:07 →

  23. 23

    Large Language Models Explained Simply (In 13 Minutes)

    Worth skimming The Gradient Descent 12:57 beginner

    Density
    50
    Padding
    45%
    Useful from
    3:21

    A clear, if conceptually shallow, LLM 101 explainer padded with light jokes and capped by a short affiliate plug.

    Start at 3:21 →

  24. 24

    Large Language Models | How Large Language Models Work? | Introduction to LLM | Simplilearn

    Worth skimming Simplilearn 15:47 beginner

    Density
    46
    Padding
    48%
    Useful from
    2:46

    A solid, if generic, beginner overview of how LLMs and transformers work, padded with a short in-house course pitch.

    Start at 2:46 →

  25. 25

    How Large Language Models Work

    Skip IBM Technology 5:34 beginner

    Density
    48
    Padding
    51%
    Useful from
    2:04

    A clear, competent beginner overview of LLM mechanics from IBM, though fairly generic and light on real depth.

    Start at 2:04 →

Gistil's own measurement. Not a YouTube rating, and not the channel's position. Ranking is ours; the videos belong to their channels.

How this ranking is decided

Every video here was measured on the same axes: how much of its runtime carries information, how much is padding, and the second at which it starts paying off. The order is by value density, not by views, recency or how much we liked the channel. What the numbers mean →

← All measured ai & machine learning videos