Claude Opus 5.5 cannot output a single pixel, yet people keep posting motion-graphics videos "made by Opus 5.5". The trick is that the model writes a program that draws every frame, and you render that program into an MP4. This post is the full, reproducible pipeline we used for a 12-second 1280x720 promo: the exact prompt, the API call, a 25-line frame renderer, the ffmpeg command, and what it actually cost.
Disclosure up front: we ran this through apimodels.app, the API gateway I work on. apimodels.app is a multi-model API gateway: one API key and an OpenAI-compatible endpoint for about 150 image, video, audio and language models, including claude-opus-5-5. Every step below works the same against any OpenAI-compatible endpoint that serves Opus 5.5; only the base URL changes.
The idea in one sentence
Ask Opus 5.5 for an HTML page with a deterministic renderFrame(t) function, not for "a video", then let a headless browser call that function once per frame.
Opus 5.5 outputs text. A page that animates itself with requestAnimationFrame looks fine in a browser, but you cannot export it frame-accurately, because what you see depends on timing. A page whose every frame is a pure function of t can be rendered frame 0, frame 1 … frame 359 in any order, and each call returns the same picture. That is exactly what a video encoder needs.
What you need
- An API key for an endpoint that serves
claude-opus-5-5(OpenAI chat-completions format). - Node.js 20+ and
puppeteer-core(npm i puppeteer-coreworks; we use the Chrome already installed on the machine). - ffmpeg 6+.
- About seven minutes of patience for the model call. More on that below.
Step 1: write the prompt as a frame contract
The rules in the first half of this prompt are what make the output renderable. The story comes second, the style last. This is the prompt we sent, unedited:
You are writing a program that renders a video. Output one complete, self-contained HTML file and nothing else.
Rules for the program (these make it renderable frame by frame):
1. A single <canvas> of exactly 1280x720, no CSS scaling, black page background.
2. Expose a global function window.renderFrame(t) that draws the frame at time t (seconds, 0 <= t < 12) from scratch. Everything on screen must be a pure function of t: no Date.now, no performance.now, no Math.random (use a seeded PRNG if you need noise), no accumulated state between calls, no requestAnimationFrame loop.
3. Also expose window.DURATION = 12 and window.FPS = 30. When the page is opened normally, play it once in a loop with requestAnimationFrame by calling renderFrame, so a human can preview it.
4. No external assets, fonts, images or libraries. Use system-ui for text.
5. The rhythm is 120 BPM: a beat every 0.5 s. Put every major cut, text entrance and impact on a beat time (a multiple of 0.5). List the beat plan as a comment at the top of the script.
The story (tell this, then decide the visuals):
A 12-second promo for "APIMODELS", an API gateway. Problem, then turn, then payoff:
- 0-3 s: a developer's screen is cluttered with many different API keys and SDK logos flying in from every side, overlapping, getting chaotic (draw them as simple rounded labels such as "image", "video", "LLM", "audio", "key_1", "key_2", "sdk", not real company logos).
- 3-4.5 s: everything snaps together on a beat and collapses into one glowing key.
- 4.5-9 s: from that one key, lines fan out to a grid of model cards that light up one per beat (label them "145 models", "image", "video", "LLM", "audio", "one endpoint").
- 9-12 s: clean end card: the word APIMODELS, the line "One key. Every model.", and "apimodels.app" small underneath; hold still for the last second.
Style: dark background, one accent colour (electric violet #7c5cff) plus white, motion with easing (easeOutCubic / easeInOutQuad written by you), subtle depth (scale and blur fall-off), no clutter in the final 3 seconds.
Before the code, do not explain. Return only the HTML.
Three details mattered more than the wording:
-
Ban every source of non-determinism by name.
Date.now,performance.now,Math.randomand accumulated state are the usual culprits. The model followed all four. - Fix the rhythm in numbers. "120 BPM, cuts on multiples of 0.5 s, write the beat plan as a comment" produced a 24-beat plan, and every cut landed on one.
- Say what must not appear. We asked for plain rounded labels instead of real company logos, which keeps the result usable.
Step 2: call Opus 5.5 with streaming and no small max_tokens
curl https://api.apimodels.app/v1/chat/completions \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d "$(jq -n --rawfile p prompt.md \
'{model:"claude-opus-5-5", stream:true, messages:[{role:"user", content:$p}]}')" \
--no-buffer > stream.txt
Two things will bite you here if you skip them:
- Opus 5.5 thinks before it writes, and thinking counts as output. On this prompt the first character of the answer arrived after 110 seconds, and the full reply took 6 minutes 37 seconds. Stream the request so the connection stays alive. A client with a 60-second timeout will give up long before the first token.
-
Do not send a small
max_tokens. The reply used 31,786 output tokens for about 16,000 characters of HTML, most of it thinking. Amax_tokens: 4096cap cuts the file off mid-script. If you omit the field, apimodels.app uses the model's own ceiling instead of inventing a small default.
Join the delta.content pieces from stream.txt and save the result as promo.html. Open it in a browser first: it should play its own preview loop.
Step 3: render every frame with headless Chrome
// render.mjs: pnpm add puppeteer-core (or npm i puppeteer-core)
import puppeteer from 'puppeteer-core'
import { mkdirSync, writeFileSync } from 'node:fs'
mkdirSync('frames', { recursive: true })
const browser = await puppeteer.launch({
executablePath: '/Applications/Google Chrome.app/Contents/MacOS/Google Chrome',
})
const page = await browser.newPage()
await page.setViewport({ width: 1280, height: 720 })
await page.goto('file://' + process.cwd() + '/promo.html')
await page.evaluate(() => { window.requestAnimationFrame = () => 0 }) // stop the preview loop
const { DURATION, FPS } = await page.evaluate(() => ({ DURATION, FPS }))
for (let i = 0; i < DURATION * FPS; i++) {
const png = await page.evaluate((t) => {
window.renderFrame(t)
return document.querySelector('canvas').toDataURL('image/png')
}, i / FPS)
writeFileSync(`frames/f${String(i).padStart(4, '0')}.png`, Buffer.from(png.split(',')[1], 'base64'))
}
await browser.close()
On Linux, point executablePath at your Chromium binary. 360 frames took 14 seconds on a laptop, and the page threw no errors.
Step 4: encode with ffmpeg
ffmpeg -framerate 30 -i frames/f%04d.png -c:v libx264 -pix_fmt yuv420p -crf 20 -movflags +faststart promo.mp4
# optional soundtrack, keep the BPM you gave the model:
# ffmpeg -i promo.mp4 -i beat.mp3 -c:v copy -shortest promo-with-audio.mp4
yuv420p and faststart are what make the file play everywhere, including in browsers and on phones. The result was a 1.8 MB, 12-second, 30 fps MP4.
Step 5: check frames at the story beats
Before judging the video, pull frames at the moments the brief describes and put them side by side: 2.3 s (the clutter), 3.5 s (one key), 7.2 s (cards lighting up) and 11.3 s (end card).
All four matched on the first attempt. To change something, ask by timecode and beat ("at 7.5 s the sixth card should light up with a pulse") and ask for the whole file back, so the frame contract stays intact.
What it cost
| Item | Value |
|---|---|
| Model | claude-opus-5-5 |
| First token / total time | 110 s / 6 min 37 s |
| Output tokens | 31,786 (mostly thinking) |
| Cost at Anthropic list price ($4 / $20 per 1M tokens) | about $0.64 |
| Cost at apimodels.app list price ($2.40 / $12 per 1M tokens) | about $0.38 |
| Rendering and encoding | local, free |
Output tokens dominate. A clip that needs 30,000 to 80,000 tokens of thinking and code costs roughly $0.36 to $0.96 per attempt at the lower rate, so budget for two or three attempts per finished clip.
When this approach is the wrong tool
Code-rendered video is good at kinetic typography, UI walkthroughs, animated charts, logo reveals, explainers and social templates you can re-render with new text in seconds. It is the wrong tool for anything that has to look filmed: photoreal people, natural camera motion, skin, fabric, physics. For those shots use a text-to-video model and let Opus 5.5 write the shot list and prompts instead.
It also has no audio (add it with ffmpeg), and the model never runs its own page. Treat the HTML as untested code: render it and look at the frames before you trust it.
Where to go from here
- The model page for Claude Opus 5.5 on apimodels.app has the API parameters and this run's files.
- We collected 60+ Opus 5.5 video prompts with the original posts and methods, sorted into motion design, product promos, explainers, music videos and 3D.
What would you render first with a renderFrame(t) contract: a chart, a logo reveal, or a UI walkthrough? I'm curious which kind of clip breaks the approach.
This post was written from our own run's code and numbers; AI helped with structure and editing.


Top comments (0)