Here's a prompt that makes coding agents build a procedural octopus that unscrews a jar lid from the inside, in Three.js, in one HTML file. Run it with whatever agent you use and send me the session. You get three things back:
- Your run next to everyone else's. Same prompt, different agents, models, and setups, side by side: time, bugs, where a person had to step in, and token cost. A fair answer to "is my setup actually better" on one hard task.
- A spot in the gallery. The results post will show every octopus. Readers pick the best one, and that run gets the top of the post.
- Your name in the write-up, with a link to you, whether your run was the cleanest or the funniest failure.
Last week I ran it through four models myself and wrote up what each got wrong. The comments were better than the post: Arhan and Kartik both pointed out that one run per model proves nothing. So this time everyone runs it.
All four runs are side by side here, so you can see what yours will land next to.
The question underneath, if you care
In those four runs, the agents that built themselves a way to check the scene early (step through the animation phase by phase, read the lid angle) ended with nothing broken. The one that checked late shipped with console warnings.
Four runs can't separate "this model is careful" from "this run happened to check early". Arhan suggested the fix: forget the model for a moment, log when each run first ran a real check, and see whether earlier checks mean fewer bugs left at the end, across everyone's runs.
To make that a real test, half of the runs get one extra paragraph in the prompt asking for the check up front. If that paragraph matters more than which model you picked, it should show.
How to take part
- Pick your prompt by the day you were born. Odd day: prompt A. Even day: prompt B. This keeps people from choosing the one they think will win.
- Use any agent and any model you normally use: Claude Code, Codex, Cursor, Pi, something else.
- Start in a new, empty folder (a folder with an old run in it means the agent edits that file), paste the prompt exactly as written, and let it work. Don't help it. If you have to step in, that's fine; don't fix things by hand.
-
Send the session. With the KeepPlain plugin:
/keepplain:buildin Claude Code,$keepplain:buildin Codex, orkeepplain buildin a terminal for any of them. No plugin? Upload the session file at keepplain.com/new. You see a preview before anything leaves your machine, and keys and tokens are redacted locally. Then publish it.
Runs of the same prompt link to each other on the site automatically, so every published run shows up next to the others. That only works if the prompt is pasted unchanged. If you share a live link to your octopus (GitHub Pages, CodePen, anything), add it to your comment for the gallery.
Then drop a link in the comments, along with your agent, model, and effort setting.
About your own setup
Something I had to think about before posting this: your agent isn't a clean room. Claude Code reads your global CLAUDE.md, any memory it keeps, your plugins and hooks. Codex reads your global AGENTS.md. Two people with the same model aren't running the same thing.
I don't think that's a problem to hide. It's how people actually work, and the question is what happens in real setups. But please mention in your comment if you have global instructions like "always test first." One thing does matter: memory that can find earlier runs of this exact prompt. If your agent has KeepPlain or any other session memory connected, disconnect it for the run, or it may look up how other people's octopus went and skip their mistakes. The same goes for KeepPlain's personal rules: switch them off at keepplain.com/rules; they add lessons from your past sessions.
Prompt A
Build a single self-contained HTML file: a real-time 3D scene of an octopus escaping from a glass jar by unscrewing the lid from the inside.
## Scene
- A glass jar (cylinder with a threaded neck) sits on a wooden table in a dim room. One warm light source from the side, soft shadows.
- The jar is filled with water. Inside is an octopus, fully procedural: a soft mantle, two eyes, eight tentacles with suckers. No external models or textures.
- The lid is a metal screw cap sitting on the jar's thread.
## Behavior — this is the core of the task
The octopus must actually unscrew the lid. Not a canned animation: the lid's rotation must be driven by the tentacles' contact with it.
1. Idle: the octopus rests at the bottom, breathing, tentacles drifting.
2. Exploration: it rises, several tentacles reach up and touch the lid's underside and rim.
3. Unscrewing: tentacles grip the lid and apply torque. Grip → rotate a fraction of a turn → release → re-grip. Each cycle advances the lid along the thread. The lid should rise as it turns (thread pitch), needing roughly 3–4 full turns to come free.
4. Escape: the lid tilts and falls off onto the table (with a small bounce and sound-free thud), the octopus pulls itself out through the neck and slides onto the table.
5. Loop back to idle in a new jar, or offer a "Reset" button.
Tentacles should be driven by inverse kinematics (or an equivalent chain solver) with visible per-segment bending. Suckers should press flat against surfaces they touch. The octopus body should squash through the neck, not clip through it.
## Interactivity
- Drag to orbit the camera, scroll to zoom.
- Space: pause / resume.
- A small on-screen readout: current phase name, lid rotation in degrees, and lid height.
## Visual quality
Aim for something people will stop scrolling for. Things that matter: refraction and caustics of the glass and water, the octopus skin (subtle color pulsing like chromatophores), suckers that read as suckers, believable wood and metal. Choose a palette and lighting mood deliberately.
## Constraints
- One .html file. Three.js is allowed via a CDN import map (use r165 or newer, ES modules). No other libraries, no external assets, no fetch.
- Must run at 60 fps on a mid-range laptop in Chrome. Keep geometry modest; spend effort where it's visible.
- No placeholder comments like "// add details here". Everything you describe must be implemented.
- Don't ask questions. Make reasonable decisions and ship the complete file in one response.
Prompt B
Prompt A, word for word, plus this at the very end:
## Before you build the scene
Build a way to check it first: expose the current phase, lid angle, lid height and which tentacles are gripping on window, and add a way to jump to any phase on demand. Use it to step through the whole sequence in a real browser before you call it done.
Only that paragraph differs, so if B runs end up cleaner, that's the reason.
What I'll do with the runs
I'll write a small checker that opens every final file and tests the same things for all of them: the lid never turns backwards, it comes off, tentacles don't wrap around the octopus's own body, the console is clean. The checker won't be shown to any agent. Then one chart: minute of the first real check against bugs left at the end, A and B in different colours, all models pooled.
Once there are enough runs to say something, I'll publish it here, including the runs that make my guess look wrong.
My own runs go in too, same rules.
If you don't want to publish a session, you can still run it and tell me in the comments what happened. That counts for something, just not for the chart.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more