DEV Community

Cover image for Same prompt, four models: what Opus, Sonnet, Astra and Sol each got wrong
Evgeny Shevtsov
Evgeny Shevtsov

Posted on

Same prompt, four models: what Opus, Sonnet, Astra and Sol each got wrong

I ran one prompt through four models last week. Opus 5.5 and Sonnet 5.5 in Claude Code, GPT-6 Astra and GPT-6.1 Sol in Codex. The task: a single HTML file where a procedural octopus escapes a glass jar by unscrewing the lid from the inside.

Astra was done in 14 minutes. Sonnet took 72. Opus hit the most bugs of the four and still finished with nothing open, while Sonnet's session had zero recorded bugs and its final build had console warnings.

It's one run per model, so please don't read this as a benchmark. I watched four sessions end to end and changed my prompt because of them. That's all this is.

The task

The prompt was long. This is the line that made it hard:

The octopus must actually unscrew the lid. Not a canned animation: the lid's rotation must be driven by the tentacles' contact with it.

Any model can make an octopus wave near a lid. Here the lid only turns if tentacles grip it, so they have to grab, twist, let go, grab again, and the lid has to climb the thread as it turns. If that logic is wrong, a screenshot usually still looks fine.

Everything else was set dressing. Dim room, wooden table, water in the jar, eight tentacles with suckers, no external models, orbit controls, a pause button and a readout of the phase and the lid angle. And "a single self-contained HTML file", which comes back later.

The results

Model Agent Time Bugs it hit What it delivered
GPT-6 Astra Codex, high effort 14 min 1 Full sequence verified: ~58 s, 3.5 turns of the lid
Opus 5.5 Claude Code 25 min 2 87 KB file, full loop tested through a local server
GPT-6.1 Sol Codex, high effort 29 min 0 Full cycle at 56–60 fps, pause and reset tested
Sonnet 5.5 Claude Code 72 min 0 recorded 112 KB file, console warnings left at the end

The times include each agent testing its own work. The Codex runs were at high effort, and the Claude Code runs didn't record one. Different tools, one run each.

Timeline of the four sessions: when each started checking the running scene, and where it found bugs

GPT-6 Astra, 14 minutes

Astra put everything in one file and treated the lid as a body that can only rotate. Grip forces from the tentacles turn it, and the thread pitch turns rotation into height.

Before writing code, it looked up how similar builds had gone. One of those was my Opus run from below, which didn't find its worst bug until it could step through the animation phase by phase. Eleven minutes in, Astra built itself the same kind of check: a local server plus timed checkpoints which replay the sequence and read the simulation state out of the DOM.

The checkpoints paid off two minutes later. Over the 3.5 turns, the tentacles were encircling the octopus's own body instead of letting go and grabbing a fresh spot within reach. Watching it, you'd see busy arms.

It rewrote how each arm picks its anchor, ran the checkpoints again and was done at minute 14.

Opus 5.5, 25 minutes

Opus built the scene in parts and glued them into a 1,735-line file. The first thing that broke was the skin shader, which wouldn't compile because a variable was called patch. That's a reserved word in GLSL. Rename, move on.

What happened next is why I'm writing this. Instead of waiting for the animation to play out in real time, Opus added temporary hooks so it could jump to any phase and check the numbers. The first jump showed the lid sometimes turning backwards while being unscrewed. I would not have caught that by watching it once.

Once that was fixed, it stepped through the rest: the lid coming off, falling on the table, the octopus squeezing through the neck. It measured 0.68 ms of simulation and 1.1 ms of rendering per frame and deleted the hooks before handing over the file.

GPT-6.1 Sol, 29 minutes

Sol started by picking a palette (brass, walnut, teal glass, a coral octopus) and then wrote a 309-line file where friction between tentacles and lid drives the thread. From minute 16, it read the live state off the canvas, so phase, lid angle, torque and grips were all visible to it.

No bugs recorded. Most of the remaining time went into the look. Light colour, wood grain, light again, then two passes on how the body squeezes through the neck. Its own note at the end says it should have been told the fps target and the lighting mood up front.

The result is solid. The lid comes off at 1260°, bounces on the table, and the octopus ends up next to the jar without clipping through the glass.

Sonnet 5.5, 72 minutes

Sonnet split the work into 14 files and wrote a little build script to concatenate them. It opened a bare stub in the browser at minute 34, before the octopus existed.

It then went deep on things nobody else touched. It profiled uneven frame pacing, found the lid never quite came to rest after falling, stepped through the bounces, and even checked the scene on a 375×812 phone viewport.

No failure was recorded in that session. It did run into a duplicate JavaScript declaration twice, though, and its last check still showed that declaration and some WebGL feedback-loop warnings in the console. Everything rendered and the readout worked, so it shipped like that. It also never ran the whole sequence live without stepping, or on an integrated GPU.

What I'm taking from this

I stopped trusting the bug count. Opus logged two and ended clean because it found both. Sonnet logged none and ended with warnings, because nothing forced it to look at them. The number tells you what got caught.

Every real bug in these four runs turned up while an agent was poking at the running scene. Nobody found one by reading their own code back. Astra and Opus built that kind of tooling early and found their bugs early. Sonnet opened a browser early too, but its checks were mostly about performance, and the warnings got through.

The "self-contained" thing annoyed me more than it should have. Both Claude Code runs pulled Three.js from a CDN anyway (r170 for Opus, r167 for Sonnet), and Opus never opened its file straight from disk. I didn't check the Codex files for this. Either way, if a word matters to you, spell out what it means.

And whatever the prompt left vague is where each run spent its extra time. For Sol that was lighting. For Sonnet it was frame pacing and the falling lid. Neither was wrong to care; I just would have picked something else.

So this goes into the first prompt next time. Every one of the four sessions points to at least one of these lines:

Before you build the scene, build a way to check it:
expose phase, lid angle and grip state on window,
and let me fast-forward to each phase on demand.
"Self-contained" means no network: inline three.js.
Test the file opened from disk, not only from a server.
Target: 60 fps on an integrated GPU.
Lighting: one warm side light, dim room. Do not retune it.
Enter fullscreen mode Exit fullscreen mode

The lines I'd add to the first prompt, merged from all four sessions

My guess is that with these lines, all four would finish faster and closer together. I'll rerun it and see.

The sessions

All four are public, with the moments that counted pulled out and the raw transcript a click away. I keep them on coders.talk, which I'm building:

Has a "clean" run ever turned out to be your worst one? And for something physics-heavy like this, which model would you hand it to first?

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

Thanks for saying up front that this is one run per model. The line that stuck with me is that if the lid logic is wrong, a screenshot usually still looks fine. That's the whole problem with judging agent output by eye.

On the timings: agent run time varies a lot from run to run (retries, which file it reads first, when it decides to test). With one sample each, 14 vs 29 minutes could easily swap on a rerun. Even three runs per model, reported as a range, would show whether "Astra is fast" holds up or was a lucky path.

Your timeline suggests something testable, though: did the runs that started stepping through the animation phase by phase earlier also finish with fewer open bugs? If so, that's a prompt change worth more than the model choice.