
根据著名的attention is all you need 论文, 制作一段精美的3blue1brown style的教学视频,帮助初入AI的人理解transformer和ChatGPT(现代AI)
Make it yours. Swap the paper in "根据著名的attention is all you need 论文" and the audience in "帮助初入AI的人理解transformer和ChatGPT(现代AI)" (help AI newcomers understand Transformers and ChatGPT). Keep "3blue1brown style的教学视频" for the look.
Transformer explainer in 3Blue1Brown style
A 4:40 Chinese-language lesson that walks AI beginners through the paper "Attention Is All You Need" and on to ChatGPT. The creator says Opus 5.5 made it in one shot from a one-sentence prompt, using about 30% of their Claude quota.
- Claude Opus 5.5
- One-shot
- 30% of Claude quota
How it was made
The creator describes it as an Opus 5.5 one-shot from this single sentence. The post doesn't say which app or harness ran it, or what rendered the animation.
About 30% of the creator's Claude usage quota, according to a reply in the thread.
Posted Sep 25, 2026 · X · 111 likes · 18k views · 7 comments
Prompt from: Reply on X
Look notesby Reference
A dark, grid-lined lesson that builds the Transformer one idea at a time, from next-word prediction and word vectors to queries, keys and values, the Transformer block and the decoder stack behind ChatGPT, ending on the attention formula.
- Color
- Deep navy-teal ground with a faint grid; cyan and gold as the main highlights, with gold, green and red pills for queries, keys and values.
- Type
- A serif title for the paper with a gold "2017", small numbered section labels in the top left, Chinese captions centered along the bottom, and a LaTeX-style attention formula with color-coded Q, K and V.
- Framing
- One centered diagram per beat on the dark grid: rows of token boxes, a bar chart of next-word probabilities, a vector plot with axes, and labeled block diagrams, each with a single caption line below.
- Structure
- The 2017 paper card (arXiv 1706.03762), 01 predicting the next word, 02 tokens as vectors, 03 one word with two meanings (苹果 as fruit and as phone), 04 attention with queries, keys and values, 05 multi-head attention and the Transformer block, 06 from Transformer to ChatGPT (the decoder half, × 6), then the formula and the full pipeline from text to the next word.




























