Zmierzona półkaAI i uczenie maszynowe

Transformer explainers, ranked by where the analogy ends

Mediana zbędnych treści
36%
Przydatna część zaczyna się
1:36
Typowy czas trwania
29 min

ostatnio przeliczono 16 wrz 2026

Every video on this shelf has to solve the same problem: attention is a mechanism made of matrix multiplication, and matrix multiplication is not something you can watch. So each of them picks a metaphor — a lookup table, a search engine, a room where words vote on each other — and the entire quality of the video is decided by what happens after the metaphor.

The ones that rank highly here spend the metaphor quickly and then show the actual operation: what the three projections are, what shape the things being multiplied have, why the scaling term is there, what the mask does. The ones lower down keep the metaphor going for forty minutes, restated in progressively more elaborate ways, and never show a number.

Both types are titled “transformers explained”. The density measurement is what separates them, because a metaphor restated is by definition not new information, and it is the single most common way this topic becomes an hour of video.

Entries carry a difficulty level, and on this topic it matters more than anywhere else on the site: a video that is perfect for someone who already knows what a dot product does is a wasted hour for someone who does not, and the reverse is worse.

  1. 01

    Let's reproduce GPT-2 (124M)

    Warto AndrejKarpathy 4:01:26 advanced

    Gęstość
    84
    Zbędne
    27%
    Przydatne od
    3:34

    A code-complete, first-principles GPT-2 reproduction — one of the most valuable hands-on deep learning tutorials available.

    Zacznij od 3:34 →

  2. 02

    Attention in transformers, step-by-step | Deep Learning Chapter 6

    Warto 3Blue1Brown 26:10 intermediate

    Gęstość
    83
    Zbędne
    21%
    Przydatne od
    1:40

    A masterfully clear, dense walkthrough of the attention mechanism that rewards careful watching with real technical understanding.

    Zacznij od 1:40 →

  3. 03

    Transformers, the tech behind LLMs | Deep Learning Chapter 5

    Warto 3Blue1Brown 27:14 intermediate

    Gęstość
    80
    Zbędne
    26%
    Przydatne od
    1:26

    A masterclass primer on transformer internals — dense, rigorous, and exactly what its title promises.

    Zacznij od 1:26 →

  4. 04

    Let's build GPT: from scratch, in code, spelled out.

    Do przejrzenia AndrejKarpathy 1:56:20 advanced

    Gęstość
    82
    Zbędne
    35%
    Przydatne od
    14:11

    A masterclass build-along: real working GPT code, the actual mechanics behind ChatGPT, with almost no filler.

    Zacznij od 14:11 →

  5. 05

    Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!

    Warto StatQuest with Josh Starmer 36:15 intermediate

    Gęstość
    79
    Zbędne
    23%
    Przydatne od
    1:20

    A rigorous, worked-numbers walkthrough of transformer internals that actually teaches how ChatGPT-style models work, not just what they do.

    Zacznij od 1:20 →

  6. 06

    Large Language Models explained briefly

    Warto 3Blue1Brown 7:58 intermediate

    Gęstość
    75
    Zbędne
    26%
    Przydatne od
    0:33

    A masterfully compressed, accurate primer on how LLMs actually work — dense, honest, no sales pitch.

    Zacznij od 0:33 →

  7. 07

    Transformers Step-by-Step Explained (Attention Is All You Need)

    Do przejrzenia ByteByteGo 10:04 intermediate

    Gęstość
    70
    Zbędne
    39%
    Przydatne od
    2:40

    A tight, genuinely educational explainer of Transformer attention with a real worked example, lightly interrupted by a disclosed sponsor read.

    Zacznij od 2:40 →

  8. 08

    Transformer Neural Networks - EXPLAINED! (Attention is all you need)

    Warto CodeEmporium 13:05 intermediate

    Gęstość
    68
    Zbędne
    31%
    Przydatne od
    1:54

    A dense, well-structured conceptual walkthrough of transformer architecture that earns its 'EXPLAINED' title with zero filler.

    Zacznij od 1:54 →

  9. 09

    Illustrated Guide to Transformers Neural Network: A step by step explanation

    Warto The AI Hacker 15:01 intermediate

    Gęstość
    72
    Zbędne
    31%
    Przydatne od
    1:03

    A tight, accurate conceptual tour of the Transformer's internals — solid teaching, though it retreads familiar illustrated-guide territory rather than breaking new ground.

    Zacznij od 1:03 →

  10. 10

    Transformers explained | The architecture behind LLMs

    Warto AI Coffee Break with Letitia 19:48 intermediate

    Gęstość
    71
    Zbędne
    33%
    Przydatne od
    0:38

    A genuinely dense, accurate transformer explainer that earns its title with real mechanics, not hype.

    Zacznij od 0:38 →

  11. 11

    Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman

    Do przejrzenia Lex Clips 8:38 intermediate

    Gęstość
    69
    Zbędne
    34%
    Przydatne od
    2:44

    A sharp, dense breakdown of why the Transformer works — one of the clearer plain-language explanations of its design philosophy, if brief.

    Zacznij od 2:44 →

  12. 12

    Transformers, explained: Understand the model behind GPT, BERT, and T5

    Warto Google Cloud Tech 9:11 intermediate

    Gęstość
    67
    Zbędne
    34%
    Przydatne od
    1:25

    A genuinely solid, jargon-light explainer of transformer architecture that earns its title without ever really selling anything.

    Zacznij od 1:25 →

  13. 13

    Transformers for beginners | What are they and how do they work

    Warto AssemblyAI 19:59 beginner

    Gęstość
    68
    Zbędne
    29%
    Przydatne od
    0:31

    Solid, math-grounded beginner explainer of transformer internals, lightly bookended by the channel's own API plug.

    Zacznij od 0:31 →

  14. 14

    Transformers Explained | Simple Explanation of Transformers

    Warto codebasics 57:31 intermediate

    Gęstość
    67
    Zbędne
    27%
    Przydatne od
    1:36

    A patient, analogy-heavy but genuinely thorough walkthrough of Transformer internals — worth the long runtime if you already know your deep learning basics.

    Zacznij od 1:36 →

  15. 15

    How does AI actually work? Transformers explained

    Warto AI Search 32:21 intermediate

    Gęstość
    66
    Zbędne
    30%
    Przydatne od
    0:31

    A solid, honestly-titled conceptual explainer of Transformer architecture, weakened only by redundant recaps and a mid-video sponsor detour.

    Zacznij od 0:31 →

  16. 16

    Transformer Architecture Explained 'Attention Is All You Need'

    Warto ByteMonk 12:49 intermediate

    Gęstość
    62
    Zbędne
    34%
    Przydatne od
    0:47

    A clear, well-paced conceptual primer on Transformer attention — not groundbreaking, but a genuinely solid explainer worth the 13 minutes for newcomers to the architecture.

    Zacznij od 0:47 →

  17. 17

    Transformers, explained: Understand the model behind ChatGPT

    Do przejrzenia Leon Petrou 24:07 beginner

    Gęstość
    61
    Zbędne
    28%
    Przydatne od
    3:29

    A clear, accessible mental model of how Transformers work end-to-end — skips the real attention math (no Q/K/V) but is genuinely useful as a conceptual primer.

    Zacznij od 3:29 →

  18. 18

    What are Large Language Models (LLMs)?

    Do przejrzenia Google for Developers 5:30 beginner

    Gęstość
    56
    Zbędne
    41%
    Przydatne od
    0:32

    A tight, honest beginner explainer of LLMs and prompt design — light on depth but dense and accurate for its length.

    Zacznij od 0:32 →

  19. 19

    Everything You Need To Know About Large Language Models (LLMs)

    Do przejrzenia Matthew Berman 25:20 beginner

    Gęstość
    58
    Zbędne
    39%
    Przydatne od
    0:32

    A solid, broad beginner's overview of LLM mechanics and history, only lightly diluted by a sponsor segment for AI Camp.

    Zacznij od 0:32 →

  20. 20

    The Transformer architecture

    Do przejrzenia Hugging Face 2:45 beginner

    Gęstość
    50
    Zbędne
    47%
    Przydatne od
    0:59

    A clean, honest, high-level primer that sets up the series without pretending to teach the deep mechanics yet.

    Zacznij od 0:59 →

  21. 21

    What are Transformers (Machine Learning Model)?

    Do przejrzenia IBM Technology 5:51 beginner

    Gęstość
    51
    Zbędne
    46%
    Przydatne od
    1:17

    A clear, accurate but fairly standard conceptual primer on transformers — solid intro, low novelty.

    Zacznij od 1:17 →

  22. 22

    Transformer Explained

    Do przejrzenia Caleb Writes Code 6:55 intermediate

    Gęstość
    55
    Zbędne
    40%
    Przydatne od
    2:07

    A solid conceptual primer on transformer limitations and fixes, but it openly admits it skips the actual mechanics the title implies.

    Zacznij od 2:07 →

  23. 23

    Large Language Models Explained Simply (In 13 Minutes)

    Do przejrzenia The Gradient Descent 12:57 beginner

    Gęstość
    50
    Zbędne
    45%
    Przydatne od
    3:21

    A clear, if conceptually shallow, LLM 101 explainer padded with light jokes and capped by a short affiliate plug.

    Zacznij od 3:21 →

  24. 24

    Large Language Models | How Large Language Models Work? | Introduction to LLM | Simplilearn

    Do przejrzenia Simplilearn 15:47 beginner

    Gęstość
    46
    Zbędne
    48%
    Przydatne od
    2:46

    A solid, if generic, beginner overview of how LLMs and transformers work, padded with a short in-house course pitch.

    Zacznij od 2:46 →

  25. 25

    How Large Language Models Work

    Pomiń IBM Technology 5:34 beginner

    Gęstość
    48
    Zbędne
    51%
    Przydatne od
    2:04

    A clear, competent beginner overview of LLM mechanics from IBM, though fairly generic and light on real depth.

    Zacznij od 2:04 →

Gistil's own measurement. Not a YouTube rating, and not the channel's position. Ranking jest nasz; filmy należą do swoich kanałów.

Jak ustala się ten ranking

Każdy film tutaj zmierzono na tych samych osiach — ile z czasu trwania niesie informację, ile to zbędne treści i w której sekundzie film zaczyna się opłacać. Kolejność wynika z gęstości wartości, a nie z liczby wyświetleń, świeżości czy tego, jak bardzo podobał nam się kanał. Co oznaczają te liczby →

← Wszystkie zmierzone filmy w kategorii ai i uczenie maszynowe