DEV Community

Vishal Habib
Vishal Habib

Posted on AI-assisted

My grader got every real answer right. Four red teams still broke it.

I built a grader that checks whether a quantum circuit an AI agent wrote is the circuit it was asked for. It uses no model in the loop. It builds the circuit's full matrix (or output state) and compares it with the reference to within 1e-9.

Then I ran two models through it and had an agent attack it with the code open, four times. The two results tell different stories, and the gap between them is the point of this post.

The eval result

45 tasks: Bell states, controlled rotations, a 3-qubit QFT, phase oracles, "use only CNOTs" constraints. 30 pilot tasks plus 15 blind ones written before any grader run on them.

Sonnet Haiku
Passed 45/45 28/45
Said "success" and was wrong 0 13 of 41

Sonnet hit the ceiling, so the set is too easy for it. Haiku's gap was mostly not physics. On 11 of its 13 wrong "success" claims, it had left the OPENQASM 2.0; line off the file. With the line added (a post hoc check, labeled as one), it passes 39/45 and only 2 claims are wrong. A real pipeline would still reject those files. The agent failure that mattered most was the boring one: output format.

The red team result

Before each round I wrote down the bar: zero wrong circuits graded pass, each one proven wrong by a separate check that doesn't use the grader's code.

Round Wrong circuits graded pass What got through
1 16 Unenforced task rules, a reset hidden inside initialize, tolerances too loose, a gate checked by name only, the Python runner tampered with
2 3 The Python runner again, plus a QASM 2 parsing gap
3 3 The Python runner again
4 5 A real gate renamed barrier, which the grader skipped as a no-op

That's 27, and I found one more while fixing round 4. 14 fixes so far. The bar has never been met, so the README doesn't say the grader passed a red team.

The fact I keep coming back to: after every round I re-graded both model runs, and not one verdict changed. None of the 90 real answers used any of these tricks. The grader was right on every honest input and wrong on 28 adversarial ones.

Three things I'd tell a team building an eval

1. "Right on real answers" and "safe to gate on" are different claims. If a person reads the score, the honest-input result is what you need. If an agent optimizes against the score (a reward signal, a CI gate, an agent grading its own work), the adversarial result is the one that matters. Agents do find these routes. Decide which claim you're making before you write the headline.

2. After two failed patches, change the design. Rounds 1–3 all beat the Python runner from inside the same process: rewriting the output file, swapping the save function, reading the runner's secret off stdin first. Every patch closed one route and the next round found another. What ended it was a design change: the submission runs in an OS sandbox with no file writes, no network and stdin closed, and the circuit only comes back as the runner's last line of output. Round 4 didn't touch it.

3. Fix the bug class, not the instance. Round 4 was round 1's bug in a new spot. In round 1, a gate called cx that was secretly a SWAP passed a "use only CNOTs" task, because the check trusted the name. I fixed that check to test what the gate does. I didn't go looking for every other place that trusted a name. The list of operations to ignore (barrier, delay) still matched by name, so a real X gate named barrier got deleted before grading. The same mistake applied to measure: a do-nothing gate with that name satisfied a "measure every qubit" task. Now, when a finding comes in, I search for every place that makes the same assumption before I call it fixed.

What I'd put on the slide

45 tasks, exact grading. Haiku: 13 of 41 "success" claims were wrong, mostly from a missing header line. The grader was right on all 90 real answers. Four adversarial rounds still got 27 wrong circuits graded pass; all are fixed and now regression tests. It has not yet passed a red team.

It's a less tidy slide than "45/45." It's the one that tells you where the grader can be trusted.


Code, tasks, preregistration and every red-team finding: circuit-claim-check. Everything else I build: github.com/vishalhabib99.

Top comments (0)