
"One Claude Code session on Claude Opus 5.5 at max effort. Sourced script in near-ASD-STE100 English, a free local voice (Kokoro), word timing with Whisper, Remotion for the animation, procedural sound, automated QA." "The whole procedure lives in one runbook. A /video command in Claude Code points at it, and generated skills give Codex and Grok the same entry point, without a second copy of the rules."
Make it yours. Use the first description as a brief: name your topic and keep "Sourced script in near-ASD-STE100 English, a free local voice (Kokoro), word timing with Whisper, Remotion for the animation, procedural sound, automated QA." The blog post spells out each stage of the pipeline.
How AI coding agents actually work
A 4-minute narrated explainer that climbs five levels, from pressing Enter to KV caches, drawn as a neon HUD with word-timed captions. One Claude Code session on Opus 5.5 researched and scripted it, voiced it with local Kokoro, and animated it in Remotion.
“Nobody recorded a word or drew a frame.”
You'll need
- Remotion, which is free for individuals and companies of up to three people under its licence
- Kokoro TTS and MLX Whisper running locally on Apple Silicon, or ElevenLabs instead
- FFmpeg, and Python 3 with numpy for the sound
How it was made
The creator didn't publish the prompt or the runbook behind his /video command; these are his own descriptions of the setup. Opus 5.5 ran at max effort. Five research subagents built a 264-quote fact sheet that a script checked against saved sources. The script is 38 beats in 7 scenes in one JSON file, with markup that separates caption text from spoken text. Kokoro voiced each beat, MLX Whisper word times drive the Remotion cues, and the effects and music bed are synthesized in numpy. QA covered cue checks, loudness, black or frozen frames, and contact sheets the agent read; a second fact pass changed four lines. The full render took 12 minutes 26 seconds for 7,967 frames.
$0 for the voice and the renderer licence, plus local compute and the agent subscription, according to the post.
Posted Oct 4, 2026 · Blog
Prompt from: blog post, in the TL;DR and "One command, three agents"
Karpathy's post on bespoke explainer videos, discussed in the article
Look notesby Reference
A dark synthwave HUD over a glowing perspective grid. Neon diagrams explain an agent: a glowing orb labeled BRAIN and a small robot labeled BODY, a terminal reading a prompt as tokens, a loop wheel, a context-window meter, prompt-caching bars, stacked transformer layers and a closing "whole machine" sum.
- Color
- Tokyo Night navy and near-black with neon violet, cyan, pink, orange and green accents over a magenta grid floor.
- Type
- A blocky, pixel-edged display face for titles ("HOW AI CODING AGENTS ACTUALLY WORK"), small monospace HUD labels and token chips, and burned-in captions at the bottom with the key word picked out in color.
- Framing
- A fixed HUD frame with corner brackets, a header strip with the level tag and a timecode, and a vanja.io mark top right. Each diagram sits center stage above the grid, with the caption bar centered at the bottom.
- Sound
- Narration by the local Kokoro voice af_heart, with synthesized effects and a music bed that ducks under the voice, according to the post.
- Structure
- Title card, level 1 brain and body, level 2 a terminal turning a request into tokens, level 3 the agent loop, level 4 memory and scale (the context window and prompt caching), level 5 inside the model (transformer layers and reasoning effort), then the whole machine: model plus harness plus loop.






























