Measured shelfAI & Machine Learning
Transformer explainers, ranked by where the analogy ends
- Videos measured
- 27
- Median padding
- 35%
- Useful part starts
- 1:28
- Typical runtime
- 28 min
Ranked from 27 measured videos last recalculated 4 Aug 2026
Every video on this shelf has to solve the same problem: attention is a mechanism made of matrix multiplication, and matrix multiplication is not something you can watch. So each of them picks a metaphor — a lookup table, a search engine, a room where words vote on each other — and the entire quality of the video is decided by what happens after the metaphor.
The ones that rank highly here spend the metaphor quickly and then show the actual operation: what the three projections are, what shape the things being multiplied have, why the scaling term is there, what the mask does. The ones lower down keep the metaphor going for forty minutes, restated in progressively more elaborate ways, and never show a number.
Both types are titled “transformers explained”. The density measurement is what separates them, because a metaphor restated is by definition not new information, and it is the single most common way this topic becomes an hour of video.
Entries carry a difficulty level, and on this topic it matters more than anywhere else on the site: a video that is perfect for someone who already knows what a dot product does is a wasted hour for someone who does not, and the reverse is worse.
-
01
Let's build GPT: from scratch, in code, spelled out.
0:001:56:20- Density
- 91
- Padding
- 11%
- Useful from
- 1:28
The gold-standard deep-dive into the mechanics of building a transformer language model from scratch.
-
02
Deep Dive into LLMs like ChatGPT
0:003:31:24- Density
- 86
- Padding
- 14%
- Useful from
- 1:28
The gold-standard technical overview of how LLMs are actually built.
-
03
Attention in transformers, step-by-step | Deep Learning Chapter 6
0:0026:10- Density
- 83
- Padding
- 21%
- Useful from
- 1:40
A masterfully clear, dense walkthrough of the attention mechanism that rewards careful watching with real technical understanding.
-
04
Transformers, the tech behind LLMs | Deep Learning Chapter 5
0:0027:14- Density
- 80
- Padding
- 26%
- Useful from
- 1:26
A masterclass primer on transformer internals — dense, rigorous, and exactly what its title promises.
-
05
Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!
0:0036:15- Density
- 79
- Padding
- 23%
- Useful from
- 1:20
A rigorous, worked-numbers walkthrough of transformer internals that actually teaches how ChatGPT-style models work, not just what they do.
-
06
Large Language Models explained briefly
0:007:58- Density
- 75
- Padding
- 26%
- Useful from
- 0:33
A masterfully compressed, accurate primer on how LLMs actually work — dense, honest, no sales pitch.
-
07
Transformers Step-by-Step Explained (Attention Is All You Need)
0:0010:04- Density
- 70
- Padding
- 39%
- Useful from
- 2:40
A tight, genuinely educational explainer of Transformer attention with a real worked example, lightly interrupted by a disclosed sponsor read.
-
08
Transformer Neural Networks - EXPLAINED! (Attention is all you need)
0:0013:05- Density
- 68
- Padding
- 31%
- Useful from
- 1:54
A dense, well-structured conceptual walkthrough of transformer architecture that earns its 'EXPLAINED' title with zero filler.
-
09
Illustrated Guide to Transformers Neural Network: A step by step explanation
0:0015:01- Density
- 72
- Padding
- 31%
- Useful from
- 1:03
A tight, accurate conceptual tour of the Transformer's internals — solid teaching, though it retreads familiar illustrated-guide territory rather than breaking new ground.
-
10
Transformers explained | The architecture behind LLMs
0:0019:48- Density
- 71
- Padding
- 33%
- Useful from
- 0:38
A genuinely dense, accurate transformer explainer that earns its title with real mechanics, not hype.
-
11
Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman
0:008:38- Density
- 69
- Padding
- 34%
- Useful from
- 2:44
A sharp, dense breakdown of why the Transformer works — one of the clearer plain-language explanations of its design philosophy, if brief.
-
12
Transformers, explained: Understand the model behind GPT, BERT, and T5
0:009:11- Density
- 67
- Padding
- 34%
- Useful from
- 1:25
A genuinely solid, jargon-light explainer of transformer architecture that earns its title without ever really selling anything.
-
13
Transformers for beginners | What are they and how do they work
0:0019:59- Density
- 68
- Padding
- 29%
- Useful from
- 0:31
Solid, math-grounded beginner explainer of transformer internals, lightly bookended by the channel's own API plug.
-
14
Transformers Explained | Simple Explanation of Transformers
0:0057:31- Density
- 67
- Padding
- 27%
- Useful from
- 1:36
A patient, analogy-heavy but genuinely thorough walkthrough of Transformer internals — worth the long runtime if you already know your deep learning basics.
-
15
How does AI actually work? Transformers explained
0:0032:21- Density
- 66
- Padding
- 30%
- Useful from
- 0:31
A solid, honestly-titled conceptual explainer of Transformer architecture, weakened only by redundant recaps and a mid-video sponsor detour.
-
16
Transformer Architecture Explained 'Attention Is All You Need'
0:0012:49- Density
- 62
- Padding
- 34%
- Useful from
- 0:47
A clear, well-paced conceptual primer on Transformer attention — not groundbreaking, but a genuinely solid explainer worth the 13 minutes for newcomers to the architecture.
-
17
Transformers, explained: Understand the model behind ChatGPT
0:0024:07- Density
- 61
- Padding
- 28%
- Useful from
- 3:29
A clear, accessible mental model of how Transformers work end-to-end — skips the real attention math (no Q/K/V) but is genuinely useful as a conceptual primer.
-
18
What are Large Language Models (LLMs)?
0:005:30- Density
- 56
- Padding
- 41%
- Useful from
- 0:32
A tight, honest beginner explainer of LLMs and prompt design — light on depth but dense and accurate for its length.
-
19
Everything You Need To Know About Large Language Models (LLMs)
0:0025:20- Density
- 58
- Padding
- 39%
- Useful from
- 0:32
A solid, broad beginner's overview of LLM mechanics and history, only lightly diluted by a sponsor segment for AI Camp.
-
20
The Transformer architecture
0:002:45- Density
- 50
- Padding
- 47%
- Useful from
- 0:59
A clean, honest, high-level primer that sets up the series without pretending to teach the deep mechanics yet.
-
21
What are Transformers (Machine Learning Model)?
0:005:51- Density
- 51
- Padding
- 46%
- Useful from
- 1:17
A clear, accurate but fairly standard conceptual primer on transformers — solid intro, low novelty.
-
22
Transformer Explained
0:006:55- Density
- 55
- Padding
- 40%
- Useful from
- 2:07
A solid conceptual primer on transformer limitations and fixes, but it openly admits it skips the actual mechanics the title implies.
-
23
Large Language Models Explained Simply (In 13 Minutes)
0:0012:57- Density
- 50
- Padding
- 45%
- Useful from
- 3:21
A clear, if conceptually shallow, LLM 101 explainer padded with light jokes and capped by a short affiliate plug.
-
24
Large Language Models | How Large Language Models Work? | Introduction to LLM | Simplilearn
0:0015:47- Density
- 46
- Padding
- 48%
- Useful from
- 2:46
A solid, if generic, beginner overview of how LLMs and transformers work, padded with a short in-house course pitch.
-
25
How Large Language Models Work
0:005:34- Density
- 48
- Padding
- 51%
- Useful from
- 2:04
A clear, competent beginner overview of LLM mechanics from IBM, though fairly generic and light on real depth.
Gistil's own measurement. Not a YouTube rating, and not the channel's position. Ranking is ours; the videos belong to their channels.
How this ranking is decided
Every video here was measured on the same axes: how much of its runtime carries information, how much is padding, and the second at which it starts paying off. The order is by value density, not by views, recency or how much we liked the channel. What the numbers mean →