Prateleira medidaIA e aprendizado de máquina

Transformer explainers, ranked by where the analogy ends

Mediana de enrolação
36%
A parte útil começa
1:36
Duração típica
29 min

recalculado pela última vez 16 de set. de 2026

Every video on this shelf has to solve the same problem: attention is a mechanism made of matrix multiplication, and matrix multiplication is not something you can watch. So each of them picks a metaphor — a lookup table, a search engine, a room where words vote on each other — and the entire quality of the video is decided by what happens after the metaphor.

The ones that rank highly here spend the metaphor quickly and then show the actual operation: what the three projections are, what shape the things being multiplied have, why the scaling term is there, what the mask does. The ones lower down keep the metaphor going for forty minutes, restated in progressively more elaborate ways, and never show a number.

Both types are titled “transformers explained”. The density measurement is what separates them, because a metaphor restated is by definition not new information, and it is the single most common way this topic becomes an hour of video.

Entries carry a difficulty level, and on this topic it matters more than anywhere else on the site: a video that is perfect for someone who already knows what a dot product does is a wasted hour for someone who does not, and the reverse is worse.

  1. 01

    Let's reproduce GPT-2 (124M)

    Vale a pena AndrejKarpathy 4:01:26 advanced

    Densidade
    84
    Enrolação
    27%
    Útil a partir de
    3:34

    A code-complete, first-principles GPT-2 reproduction — one of the most valuable hands-on deep learning tutorials available.

    Começar em 3:34 →

  2. 02

    Attention in transformers, step-by-step | Deep Learning Chapter 6

    Vale a pena 3Blue1Brown 26:10 intermediate

    Densidade
    83
    Enrolação
    21%
    Útil a partir de
    1:40

    A masterfully clear, dense walkthrough of the attention mechanism that rewards careful watching with real technical understanding.

    Começar em 1:40 →

  3. 03

    Transformers, the tech behind LLMs | Deep Learning Chapter 5

    Vale a pena 3Blue1Brown 27:14 intermediate

    Densidade
    80
    Enrolação
    26%
    Útil a partir de
    1:26

    A masterclass primer on transformer internals — dense, rigorous, and exactly what its title promises.

    Começar em 1:26 →

  4. 04

    Let's build GPT: from scratch, in code, spelled out.

    Vale passar os olhos AndrejKarpathy 1:56:20 advanced

    Densidade
    82
    Enrolação
    35%
    Útil a partir de
    14:11

    A masterclass build-along: real working GPT code, the actual mechanics behind ChatGPT, with almost no filler.

    Começar em 14:11 →

  5. 05

    Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!

    Vale a pena StatQuest with Josh Starmer 36:15 intermediate

    Densidade
    79
    Enrolação
    23%
    Útil a partir de
    1:20

    A rigorous, worked-numbers walkthrough of transformer internals that actually teaches how ChatGPT-style models work, not just what they do.

    Começar em 1:20 →

  6. 06

    Large Language Models explained briefly

    Vale a pena 3Blue1Brown 7:58 intermediate

    Densidade
    75
    Enrolação
    26%
    Útil a partir de
    0:33

    A masterfully compressed, accurate primer on how LLMs actually work — dense, honest, no sales pitch.

    Começar em 0:33 →

  7. 07

    Transformers Step-by-Step Explained (Attention Is All You Need)

    Vale passar os olhos ByteByteGo 10:04 intermediate

    Densidade
    70
    Enrolação
    39%
    Útil a partir de
    2:40

    A tight, genuinely educational explainer of Transformer attention with a real worked example, lightly interrupted by a disclosed sponsor read.

    Começar em 2:40 →

  8. 08

    Transformer Neural Networks - EXPLAINED! (Attention is all you need)

    Vale a pena CodeEmporium 13:05 intermediate

    Densidade
    68
    Enrolação
    31%
    Útil a partir de
    1:54

    A dense, well-structured conceptual walkthrough of transformer architecture that earns its 'EXPLAINED' title with zero filler.

    Começar em 1:54 →

  9. 09

    Illustrated Guide to Transformers Neural Network: A step by step explanation

    Vale a pena The AI Hacker 15:01 intermediate

    Densidade
    72
    Enrolação
    31%
    Útil a partir de
    1:03

    A tight, accurate conceptual tour of the Transformer's internals — solid teaching, though it retreads familiar illustrated-guide territory rather than breaking new ground.

    Começar em 1:03 →

  10. 10

    Transformers explained | The architecture behind LLMs

    Vale a pena AI Coffee Break with Letitia 19:48 intermediate

    Densidade
    71
    Enrolação
    33%
    Útil a partir de
    0:38

    A genuinely dense, accurate transformer explainer that earns its title with real mechanics, not hype.

    Começar em 0:38 →

  11. 11

    Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman

    Vale passar os olhos Lex Clips 8:38 intermediate

    Densidade
    69
    Enrolação
    34%
    Útil a partir de
    2:44

    A sharp, dense breakdown of why the Transformer works — one of the clearer plain-language explanations of its design philosophy, if brief.

    Começar em 2:44 →

  12. 12

    Transformers, explained: Understand the model behind GPT, BERT, and T5

    Vale a pena Google Cloud Tech 9:11 intermediate

    Densidade
    67
    Enrolação
    34%
    Útil a partir de
    1:25

    A genuinely solid, jargon-light explainer of transformer architecture that earns its title without ever really selling anything.

    Começar em 1:25 →

  13. 13

    Transformers for beginners | What are they and how do they work

    Vale a pena AssemblyAI 19:59 beginner

    Densidade
    68
    Enrolação
    29%
    Útil a partir de
    0:31

    Solid, math-grounded beginner explainer of transformer internals, lightly bookended by the channel's own API plug.

    Começar em 0:31 →

  14. 14

    Transformers Explained | Simple Explanation of Transformers

    Vale a pena codebasics 57:31 intermediate

    Densidade
    67
    Enrolação
    27%
    Útil a partir de
    1:36

    A patient, analogy-heavy but genuinely thorough walkthrough of Transformer internals — worth the long runtime if you already know your deep learning basics.

    Começar em 1:36 →

  15. 15

    How does AI actually work? Transformers explained

    Vale a pena AI Search 32:21 intermediate

    Densidade
    66
    Enrolação
    30%
    Útil a partir de
    0:31

    A solid, honestly-titled conceptual explainer of Transformer architecture, weakened only by redundant recaps and a mid-video sponsor detour.

    Começar em 0:31 →

  16. 16

    Transformer Architecture Explained 'Attention Is All You Need'

    Vale a pena ByteMonk 12:49 intermediate

    Densidade
    62
    Enrolação
    34%
    Útil a partir de
    0:47

    A clear, well-paced conceptual primer on Transformer attention — not groundbreaking, but a genuinely solid explainer worth the 13 minutes for newcomers to the architecture.

    Começar em 0:47 →

  17. 17

    Transformers, explained: Understand the model behind ChatGPT

    Vale passar os olhos Leon Petrou 24:07 beginner

    Densidade
    61
    Enrolação
    28%
    Útil a partir de
    3:29

    A clear, accessible mental model of how Transformers work end-to-end — skips the real attention math (no Q/K/V) but is genuinely useful as a conceptual primer.

    Começar em 3:29 →

  18. 18

    What are Large Language Models (LLMs)?

    Vale passar os olhos Google for Developers 5:30 beginner

    Densidade
    56
    Enrolação
    41%
    Útil a partir de
    0:32

    A tight, honest beginner explainer of LLMs and prompt design — light on depth but dense and accurate for its length.

    Começar em 0:32 →

  19. 19

    Everything You Need To Know About Large Language Models (LLMs)

    Vale passar os olhos Matthew Berman 25:20 beginner

    Densidade
    58
    Enrolação
    39%
    Útil a partir de
    0:32

    A solid, broad beginner's overview of LLM mechanics and history, only lightly diluted by a sponsor segment for AI Camp.

    Começar em 0:32 →

  20. 20

    The Transformer architecture

    Vale passar os olhos Hugging Face 2:45 beginner

    Densidade
    50
    Enrolação
    47%
    Útil a partir de
    0:59

    A clean, honest, high-level primer that sets up the series without pretending to teach the deep mechanics yet.

    Começar em 0:59 →

  21. 21

    What are Transformers (Machine Learning Model)?

    Vale passar os olhos IBM Technology 5:51 beginner

    Densidade
    51
    Enrolação
    46%
    Útil a partir de
    1:17

    A clear, accurate but fairly standard conceptual primer on transformers — solid intro, low novelty.

    Começar em 1:17 →

  22. 22

    Transformer Explained

    Vale passar os olhos Caleb Writes Code 6:55 intermediate

    Densidade
    55
    Enrolação
    40%
    Útil a partir de
    2:07

    A solid conceptual primer on transformer limitations and fixes, but it openly admits it skips the actual mechanics the title implies.

    Começar em 2:07 →

  23. 23

    Large Language Models Explained Simply (In 13 Minutes)

    Vale passar os olhos The Gradient Descent 12:57 beginner

    Densidade
    50
    Enrolação
    45%
    Útil a partir de
    3:21

    A clear, if conceptually shallow, LLM 101 explainer padded with light jokes and capped by a short affiliate plug.

    Começar em 3:21 →

  24. 24

    Large Language Models | How Large Language Models Work? | Introduction to LLM | Simplilearn

    Vale passar os olhos Simplilearn 15:47 beginner

    Densidade
    46
    Enrolação
    48%
    Útil a partir de
    2:46

    A solid, if generic, beginner overview of how LLMs and transformers work, padded with a short in-house course pitch.

    Começar em 2:46 →

  25. 25

    How Large Language Models Work

    Pular IBM Technology 5:34 beginner

    Densidade
    48
    Enrolação
    51%
    Útil a partir de
    2:04

    A clear, competent beginner overview of LLM mechanics from IBM, though fairly generic and light on real depth.

    Começar em 2:04 →

Gistil's own measurement. Not a YouTube rating, and not the channel's position. O ranking é nosso; os vídeos pertencem a seus canais.

Como esse ranking é decidido

Todo vídeo aqui foi medido pelos mesmos critérios — quanto da duração carrega informação, quanto é enrolação, e em que segundo ele começa a valer a pena. A ordem é pela densidade de valor, não por visualizações, data ou o quanto gostamos do canal. O que os números significam →

← Todos os vídeos medidos em ia e aprendizado de máquina