Виміряна полицяШІ та машинне навчання

Transformer explainers, ranked by where the analogy ends

Медіана зайвого
36%
Корисна частина починається
1:36
Типовий хронометраж
29 min

востаннє перераховано 16 вер. 2026

Every video on this shelf has to solve the same problem: attention is a mechanism made of matrix multiplication, and matrix multiplication is not something you can watch. So each of them picks a metaphor — a lookup table, a search engine, a room where words vote on each other — and the entire quality of the video is decided by what happens after the metaphor.

The ones that rank highly here spend the metaphor quickly and then show the actual operation: what the three projections are, what shape the things being multiplied have, why the scaling term is there, what the mask does. The ones lower down keep the metaphor going for forty minutes, restated in progressively more elaborate ways, and never show a number.

Both types are titled “transformers explained”. The density measurement is what separates them, because a metaphor restated is by definition not new information, and it is the single most common way this topic becomes an hour of video.

Entries carry a difficulty level, and on this topic it matters more than anywhere else on the site: a video that is perfect for someone who already knows what a dot product does is a wasted hour for someone who does not, and the reverse is worse.

  1. 01

    Let's reproduce GPT-2 (124M)

    Варте часу AndrejKarpathy 4:01:26 advanced

    Щільність
    84
    Зайве
    27%
    Корисне з
    3:34

    Повністю робоче відтворення GPT-2 за первинними принципами — один із найцінніших практичних туторіалів з глибокого навчання.

    Почати з 3:34 →

  2. 02

    Attention in transformers, step-by-step | Deep Learning Chapter 6

    Варте часу 3Blue1Brown 26:10 intermediate

    Щільність
    83
    Зайве
    21%
    Корисне з
    1:40

    A masterfully clear, dense walkthrough of the attention mechanism that rewards careful watching with real technical understanding.

    Почати з 1:40 →

  3. 03

    Transformers, the tech behind LLMs | Deep Learning Chapter 5

    Варте часу 3Blue1Brown 27:14 intermediate

    Щільність
    80
    Зайве
    26%
    Корисне з
    1:26

    Майстер-клас із внутрішньої будови трансформаторів — насичений, суворий і саме такий, як обіцяє назва.

    Почати з 1:26 →

  4. 04

    Let's build GPT: from scratch, in code, spelled out.

    Проглянути AndrejKarpathy 1:56:20 advanced

    Щільність
    82
    Зайве
    35%
    Корисне з
    14:11

    A masterclass build-along: real working GPT code, the actual mechanics behind ChatGPT, with almost no filler.

    Почати з 14:11 →

  5. 05

    Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!

    Варте часу StatQuest with Josh Starmer 36:15 intermediate

    Щільність
    79
    Зайве
    23%
    Корисне з
    1:20

    A rigorous, worked-numbers walkthrough of transformer internals that actually teaches how ChatGPT-style models work, not just what they do.

    Почати з 1:20 →

  6. 06

    Large Language Models explained briefly

    Варте часу 3Blue1Brown 7:58 intermediate

    Щільність
    75
    Зайве
    26%
    Корисне з
    0:33

    Майстерно стислий і точний вступний посібник про те, як насправді працюють LLM — насичений, чесний, без рекламних трюків.

    Почати з 0:33 →

  7. 07

    Transformers Step-by-Step Explained (Attention Is All You Need)

    Проглянути ByteByteGo 10:04 intermediate

    Щільність
    70
    Зайве
    39%
    Корисне з
    2:40

    A tight, genuinely educational explainer of Transformer attention with a real worked example, lightly interrupted by a disclosed sponsor read.

    Почати з 2:40 →

  8. 08

    Transformer Neural Networks - EXPLAINED! (Attention is all you need)

    Варте часу CodeEmporium 13:05 intermediate

    Щільність
    68
    Зайве
    31%
    Корисне з
    1:54

    A dense, well-structured conceptual walkthrough of transformer architecture that earns its 'EXPLAINED' title with zero filler.

    Почати з 1:54 →

  9. 09

    Illustrated Guide to Transformers Neural Network: A step by step explanation

    Варте часу The AI Hacker 15:01 intermediate

    Щільність
    72
    Зайве
    31%
    Корисне з
    1:03

    A tight, accurate conceptual tour of the Transformer's internals — solid teaching, though it retreads familiar illustrated-guide territory rather than breaking new ground.

    Почати з 1:03 →

  10. 10

    Transformers explained | The architecture behind LLMs

    Варте часу AI Coffee Break with Letitia 19:48 intermediate

    Щільність
    71
    Зайве
    33%
    Корисне з
    0:38

    A genuinely dense, accurate transformer explainer that earns its title with real mechanics, not hype.

    Почати з 0:38 →

  11. 11

    Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman

    Проглянути Lex Clips 8:38 intermediate

    Щільність
    69
    Зайве
    34%
    Корисне з
    2:44

    A sharp, dense breakdown of why the Transformer works — one of the clearer plain-language explanations of its design philosophy, if brief.

    Почати з 2:44 →

  12. 12

    Transformers, explained: Understand the model behind GPT, BERT, and T5

    Варте часу Google Cloud Tech 9:11 intermediate

    Щільність
    67
    Зайве
    34%
    Корисне з
    1:25

    A genuinely solid, jargon-light explainer of transformer architecture that earns its title without ever really selling anything.

    Почати з 1:25 →

  13. 13

    Transformers for beginners | What are they and how do they work

    Варте часу AssemblyAI 19:59 beginner

    Щільність
    68
    Зайве
    29%
    Корисне з
    0:31

    Solid, math-grounded beginner explainer of transformer internals, lightly bookended by the channel's own API plug.

    Почати з 0:31 →

  14. 14

    Transformers Explained | Simple Explanation of Transformers

    Варте часу codebasics 57:31 intermediate

    Щільність
    67
    Зайве
    27%
    Корисне з
    1:36

    A patient, analogy-heavy but genuinely thorough walkthrough of Transformer internals — worth the long runtime if you already know your deep learning basics.

    Почати з 1:36 →

  15. 15

    How does AI actually work? Transformers explained

    Варте часу AI Search 32:21 intermediate

    Щільність
    66
    Зайве
    30%
    Корисне з
    0:31

    A solid, honestly-titled conceptual explainer of Transformer architecture, weakened only by redundant recaps and a mid-video sponsor detour.

    Почати з 0:31 →

  16. 16

    Transformer Architecture Explained 'Attention Is All You Need'

    Варте часу ByteMonk 12:49 intermediate

    Щільність
    62
    Зайве
    34%
    Корисне з
    0:47

    A clear, well-paced conceptual primer on Transformer attention — not groundbreaking, but a genuinely solid explainer worth the 13 minutes for newcomers to the architecture.

    Почати з 0:47 →

  17. 17

    Transformers, explained: Understand the model behind ChatGPT

    Проглянути Leon Petrou 24:07 beginner

    Щільність
    61
    Зайве
    28%
    Корисне з
    3:29

    A clear, accessible mental model of how Transformers work end-to-end — skips the real attention math (no Q/K/V) but is genuinely useful as a conceptual primer.

    Почати з 3:29 →

  18. 18

    What are Large Language Models (LLMs)?

    Проглянути Google for Developers 5:30 beginner

    Щільність
    56
    Зайве
    41%
    Корисне з
    0:32

    A tight, honest beginner explainer of LLMs and prompt design — light on depth but dense and accurate for its length.

    Почати з 0:32 →

  19. 19

    Everything You Need To Know About Large Language Models (LLMs)

    Проглянути Matthew Berman 25:20 beginner

    Щільність
    58
    Зайве
    39%
    Корисне з
    0:32

    A solid, broad beginner's overview of LLM mechanics and history, only lightly diluted by a sponsor segment for AI Camp.

    Почати з 0:32 →

  20. 20

    The Transformer architecture

    Проглянути Hugging Face 2:45 beginner

    Щільність
    50
    Зайве
    47%
    Корисне з
    0:59

    A clean, honest, high-level primer that sets up the series without pretending to teach the deep mechanics yet.

    Почати з 0:59 →

  21. 21

    What are Transformers (Machine Learning Model)?

    Проглянути IBM Technology 5:51 beginner

    Щільність
    51
    Зайве
    46%
    Корисне з
    1:17

    A clear, accurate but fairly standard conceptual primer on transformers — solid intro, low novelty.

    Почати з 1:17 →

  22. 22

    Transformer Explained

    Проглянути Caleb Writes Code 6:55 intermediate

    Щільність
    55
    Зайве
    40%
    Корисне з
    2:07

    A solid conceptual primer on transformer limitations and fixes, but it openly admits it skips the actual mechanics the title implies.

    Почати з 2:07 →

  23. 23

    Large Language Models Explained Simply (In 13 Minutes)

    Проглянути The Gradient Descent 12:57 beginner

    Щільність
    50
    Зайве
    45%
    Корисне з
    3:21

    A clear, if conceptually shallow, LLM 101 explainer padded with light jokes and capped by a short affiliate plug.

    Почати з 3:21 →

  24. 24

    Large Language Models | How Large Language Models Work? | Introduction to LLM | Simplilearn

    Проглянути Simplilearn 15:47 beginner

    Щільність
    46
    Зайве
    48%
    Корисне з
    2:46

    A solid, if generic, beginner overview of how LLMs and transformers work, padded with a short in-house course pitch.

    Почати з 2:46 →

  25. 25

    How Large Language Models Work

    Пропустити IBM Technology 5:34 beginner

    Щільність
    48
    Зайве
    51%
    Корисне з
    2:04

    A clear, competent beginner overview of LLM mechanics from IBM, though fairly generic and light on real depth.

    Почати з 2:04 →

Gistil's own measurement. Not a YouTube rating, and not the channel's position. Порядок наш; відео належать своїм каналам.

Як вирішується цей порядок

Кожне відео тут виміряно за тими самими осями — скільки хронометражу несе інформацію, скільки з нього зайвого і на якій секунді воно починає окупатися. Порядок — за щільністю користі, а не за переглядами, свіжістю чи тим, наскільки нам сподобався канал. Що означають числа →

← Усі виміряні відео в категорії «ші та машинне навчання»