Scaffale misuratoIA e machine learning
Transformer explainers, ranked by where the analogy ends
- Riempitivo mediano
- 36%
- La parte utile inizia
- 1:36
- Durata tipica
- 29 min
ultimo ricalcolo 16 set 2026
Every video on this shelf has to solve the same problem: attention is a mechanism made of matrix multiplication, and matrix multiplication is not something you can watch. So each of them picks a metaphor — a lookup table, a search engine, a room where words vote on each other — and the entire quality of the video is decided by what happens after the metaphor.
The ones that rank highly here spend the metaphor quickly and then show the actual operation: what the three projections are, what shape the things being multiplied have, why the scaling term is there, what the mask does. The ones lower down keep the metaphor going for forty minutes, restated in progressively more elaborate ways, and never show a number.
Both types are titled “transformers explained”. The density measurement is what separates them, because a metaphor restated is by definition not new information, and it is the single most common way this topic becomes an hour of video.
Entries carry a difficulty level, and on this topic it matters more than anywhere else on the site: a video that is perfect for someone who already knows what a dot product does is a wasted hour for someone who does not, and the reverse is worse.
-
01
Let's reproduce GPT-2 (124M)
0:004:01:26- Densità
- 84
- Riempitivo
- 27%
- Utile da
- 3:34
A code-complete, first-principles GPT-2 reproduction — one of the most valuable hands-on deep learning tutorials available.
-
02
Attention in transformers, step-by-step | Deep Learning Chapter 6
0:0026:10- Densità
- 83
- Riempitivo
- 21%
- Utile da
- 1:40
Una spiegazione magistralmente chiara e densa del meccanismo di attenzione, che ripaga chi guarda con attenzione con una comprensione tecnica reale.
-
03
Transformers, the tech behind LLMs | Deep Learning Chapter 5
0:0027:14- Densità
- 80
- Riempitivo
- 26%
- Utile da
- 1:26
Un'introduzione magistrale ai meccanismi interni dei transformer — densa, rigorosa ed esattamente ciò che il titolo promette.
-
04
Let's build GPT: from scratch, in code, spelled out.
0:001:56:20- Densità
- 82
- Riempitivo
- 35%
- Utile da
- 14:11
A masterclass build-along: real working GPT code, the actual mechanics behind ChatGPT, with almost no filler.
-
05
Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!
0:0036:15- Densità
- 79
- Riempitivo
- 23%
- Utile da
- 1:20
Una spiegazione rigorosa e basata su numeri concreti degli aspetti interni dei transformer che insegna davvero come funzionano i modelli in stile ChatGPT, non solo cosa fanno.
-
06
Large Language Models explained briefly
0:007:58- Densità
- 75
- Riempitivo
- 26%
- Utile da
- 0:33
Un primer magistralmente condensato e accurato su come funzionano davvero gli LLM: denso, onesto, senza pitch commerciale.
-
07
Transformers Step-by-Step Explained (Attention Is All You Need)
0:0010:04- Densità
- 70
- Riempitivo
- 39%
- Utile da
- 2:40
Una spiegazione efficace e realmente didattica dell'attenzione nei Transformer, con un esempio concreto svolto passo passo, interrotta brevemente da un messaggio sponsorizzato dichiarato.
-
08
Transformer Neural Networks - EXPLAINED! (Attention is all you need)
0:0013:05- Densità
- 68
- Riempitivo
- 31%
- Utile da
- 1:54
Una spiegazione concettuale densa e ben strutturata dell'architettura transformer che si guadagna il titolo 'EXPLAINED' senza alcun riempitivo.
-
09
Illustrated Guide to Transformers Neural Network: A step by step explanation
0:0015:01- Densità
- 72
- Riempitivo
- 31%
- Utile da
- 1:03
Un tour concettuale preciso e ben strutturato del funzionamento interno del Transformer: una spiegazione solida, anche se ripercorre un territorio già noto nel genere delle guide illustrate…
-
10
Transformers explained | The architecture behind LLMs
0:0019:48- Densità
- 71
- Riempitivo
- 33%
- Utile da
- 0:38
Una spiegazione del transformer genuinamente densa e accurata che si guadagna il titolo con meccaniche reali, non con hype.
-
11
Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman
0:008:38- Densità
- 69
- Riempitivo
- 34%
- Utile da
- 2:44
Un'analisi acuta e densa del perché il Transformer funziona — una delle spiegazioni più chiare in linguaggio semplice della sua filosofia progettuale, seppur breve.
-
12
Transformers, explained: Understand the model behind GPT, BERT, and T5
0:009:11- Densità
- 67
- Riempitivo
- 34%
- Utile da
- 1:25
Una spiegazione davvero solida e priva di gergo eccessivo dell'architettura transformer che è all'altezza del suo titolo senza mai realmente vendere nulla.
-
13
Transformers for beginners | What are they and how do they work
0:0019:59- Densità
- 68
- Riempitivo
- 29%
- Utile da
- 0:31
Solida spiegazione per principianti degli aspetti interni del transformer, fondata sulla matematica, con una breve pubblicità dell'API del canale a inizio e fine.
-
14
Transformers Explained | Simple Explanation of Transformers
0:0057:31- Densità
- 67
- Riempitivo
- 27%
- Utile da
- 1:36
Una spiegazione paziente, ricca di analogie ma genuinamente approfondita del funzionamento interno dei Transformer — vale la lunga durata se si conoscono già le basi del deep learning.
-
15
How does AI actually work? Transformers explained
0:0032:21- Densità
- 66
- Riempitivo
- 30%
- Utile da
- 0:31
Una spiegazione concettuale solida e dal titolo onesto dell'architettura Transformer, indebolita solo da riepiloghi ridondanti e da una deviazione sponsorizzata a metà video.
-
16
Transformer Architecture Explained 'Attention Is All You Need'
0:0012:49- Densità
- 62
- Riempitivo
- 34%
- Utile da
- 0:47
Un'introduzione concettuale chiara e ben ritmata all'attenzione del Transformer — non rivoluzionaria, ma una spiegazione davvero solida che vale i 13 minuti per chi è nuovo all'architettura.
-
17
Transformers, explained: Understand the model behind ChatGPT
0:0024:07- Densità
- 61
- Riempitivo
- 28%
- Utile da
- 3:29
Un modello mentale chiaro e accessibile di come funzionano i Transformer dall'inizio alla fine: salta la vera matematica dell'attenzione (niente Q/K/V), ma è genuinamente utile come…
-
18
What are Large Language Models (LLMs)?
0:005:30- Densità
- 56
- Riempitivo
- 41%
- Utile da
- 0:32
Una spiegazione onesta ed essenziale per principianti su LLM e progettazione dei prompt: poco approfondita ma densa e accurata per la sua brevità.
-
19
Everything You Need To Know About Large Language Models (LLMs)
0:0025:20- Densità
- 58
- Riempitivo
- 39%
- Utile da
- 0:32
Una panoramica solida e ampia sui meccanismi e sulla storia degli LLM per principianti, solo leggermente diluita da un segmento sponsorizzato per AI Camp.
-
20
The Transformer architecture
0:002:45- Densità
- 50
- Riempitivo
- 47%
- Utile da
- 0:59
Un'introduzione pulita e onesta di alto livello che imposta la serie senza pretendere di insegnare ancora i meccanismi approfonditi.
-
21
What are Transformers (Machine Learning Model)?
0:005:51- Densità
- 51
- Riempitivo
- 46%
- Utile da
- 1:17
Un'introduzione concettuale chiara e accurata ma piuttosto standard sui transformer — solida ma con poca originalità.
-
22
Transformer Explained
0:006:55- Densità
- 55
- Riempitivo
- 40%
- Utile da
- 2:07
Una solida introduzione concettuale ai limiti dei transformer e alle relative soluzioni, ma ammette apertamente di saltare i meccanismi effettivi suggeriti dal titolo.
-
23
Large Language Models Explained Simply (In 13 Minutes)
0:0012:57- Densità
- 50
- Riempitivo
- 45%
- Utile da
- 3:21
Una spiegazione chiara ma concettualmente superficiale sugli LLM per principianti, arricchita da battute leggere e conclusa con una breve promozione in affiliazione.
-
24
Large Language Models | How Large Language Models Work? | Introduction to LLM | Simplilearn
0:0015:47- Densità
- 46
- Riempitivo
- 48%
- Utile da
- 2:46
Una panoramica per principianti solida, anche se generica, su come funzionano gli LLM e i transformer, arricchita (e appesantita) da una promozione di un corso interno.
-
25
How Large Language Models Work
0:005:34- Densità
- 48
- Riempitivo
- 51%
- Utile da
- 2:04
Una panoramica chiara e competente sui meccanismi degli LLM pensata per principianti, offerta da IBM, anche se piuttosto generica e poco approfondita.
Gistil's own measurement. Not a YouTube rating, and not the channel's position. La classifica è nostra; i video appartengono ai rispettivi canali.
Come viene decisa questa classifica
Ogni video qui è stato misurato sugli stessi parametri — quanto della sua durata porta informazioni, quanto è riempitivo, e il secondo in cui inizia a ripagare. L'ordine è dato dalla densità di valore, non dalle visualizzazioni, dalla recenza o da quanto ci piaceva il canale. Cosa significano i numeri →