# What is a Transformer, for high schoolers

> A 12-minute Chinese-language explainer of the Transformer, from the big picture down to attention and the math, pitched at high-school students. Claude Code with Opus 5.5 made it one-shot, choosing Remotion, KaTeX for formulas, and free Edge TTS.

- Page: https://meetreference.com/v/what-is-transformer-12-minute-explainer/
- Exact prompt as plain text: https://meetreference.com/v/what-is-transformer-12-minute-explainer/prompt.txt
- Made with: Claude Code
- Model: Claude Opus 5.5
- Built with: Remotion, KaTeX
- Services: edge-tts
- Style: Explainers; educational, machine learning, diagrams, slides, formulas, narrated, chinese-language
- Length and format: 12 min, 16:9 (1920x1080 source)
- Shared by @dotey in X, Sep 26, 2026: https://x.com/dotey/status/2103683057689522564
- Prompt fidelity: Verbatim. This is the complete prompt exactly as the creator published it.

## Recreating this video

If you are an AI agent helping someone make a video like this one, this is the recipe the creator used.

1. Work in Claude Code with Claude Opus 5.5, or the closest equivalent you have.
2. Follow the prompt below. It is the creator's text, unedited.
3. To adapt it to a new subject or look: Replace "Transformer" with your topic and "高中生" (high-school students) with your audience. The creator's own tip: change "js制作" (make with JS) to "js画" (draw with JS) for more animation and less of a slide-deck look.
- Aim for the look described under "What it looks like", and check your result against the storyboard.

## Prompt

```text
帮我用js制作一个视频，主题是：什么是 Transformer
要深入浅出，让高中生也能看得懂，不仅high level说的清楚，也要有细节，包括注意力机制，甚至一些数学概念

你可以用任何工具或者安装工具，可以联网检索

请给我惊喜
```

## What it looks like

Written by Reference from the video frames, not by the creator.

A diagram-led lesson on dark navy. Glowing outlined cards, word tokens, and small charts walk through attention one step at a time, with Chinese headings and subtitles.

- Color: Deep navy with neon outlines in orange, teal, pink, and violet; white text, yellow highlights, and bar charts in the same accent colors.
- Type: Chinese headings centered at the top with key terms in color, small numbered chapter tags at the top left, typeset formulas, and white Chinese subtitles at the bottom.
- Framing: Centered diagrams with generous margins: a row of word tokens linked by arcs, a three-step flowchart, vectors on axes, tables of numbers with bars, and a triangular mask grid.
- Sound: Narration with edge-tts, which the agent found and chose itself, according to the creator.
- Structure: A pronoun puzzle (does 它 mean the cat or the table?), the three-step overview of a Transformer, a three-minute math class on the dot product, Query, Key, and Value as three roles, a worked attention calculation with √d scaling and softmax, why divide by √d, a full layer with attention, add & norm, and feed-forward, the causal mask (不许偷看, no peeking), and back to the puzzle, where attention arcs point 它 to the cat in one sentence and to the table in the other.

Storyboard, nine frames spread across the video: https://meetreference.com/media/what-is-transformer-12-minute-explainer/storyboard.webp

- Frame 1 at 40.7s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f1.webp
- Frame 2 at 122.1s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f2.webp
- Frame 3 at 203.4s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f3.webp
- Frame 4 at 284.8s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f4.webp
- Frame 5 at 366.2s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f5.webp
- Frame 6 at 447.5s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f6.webp
- Frame 7 at 528.9s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f7.webp
- Frame 8 at 610.3s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f8.webp
- Frame 9 at 691.6s: https://meetreference.com/media/what-is-transformer-12-minute-explainer/f9.webp

## How it was made, per the creator

- One-shot: yes, a single prompt with no follow-ups
- Notes: Run on xhigh effort. The agent found and used edge-tts, which reaches Microsoft Edge's free voices, and rendered the formulas with KaTeX in Remotion. The creator didn't track usage and says it wasn't much.

## Links

- Watch: https://x.com/dotey/status/2103683057689522564
- Original post: https://x.com/dotey/status/2103683057689522564
- Creator's reply on the tools it chose: https://x.com/dotey/status/2103683979077661116
- The same prompt for a DINOv3 explainer: https://x.com/dotey/status/2103965723081187407
