Two numbers went round in the last few weeks. On Terminal-Bench 2.1, GPT-5.6 Sol scored 89.5% and Claude Opus 5 scored 89.1%. Four tenths of a point between the default models of the two most-used coding agents on earth.
Check a second leaderboard and you get 85.77% and 84.64% — different harness, different numbers, same story. About a point in it, whoever is counting.
In the same weeks those numbers were being argued about, here is what those agents actually did.
The month, in one list
GhostApproval — six assistants at once: Cursor, Claude Code, Antigravity, Copilot, Grok Build, GitHub. "A malicious repo uses symlinks (CWE-61) to make the agent write outside the workspace while the approval prompt hides the real target from the user."
DuneSlide — two CVSS 9.8 holes in Cursor. Zero-click prompt injection that escaped the terminal sandbox and overwrote the sandbox helper binary.
Cursor deeplinks — one click installed an attacker-controlled MCP server and ran unsandboxed commands. The install dialog truncated the commands you were approving.
AWS Kiro — "hidden text on an ordinary web page instructs AWS Kiro to silently rewrite its own MCP server config file", then reload it.
GitLost — a plain-English payload in a public GitHub issue made Agentic Workflows read private repositories and post the contents as a public comment.
Claude Code — ran
prisma migrate diffwith the wrong parameters and deleted 22 production tables.
Now read that list again and try to find one entry that a smarter model would have prevented.
Not one of these is a reasoning failure. Every single one is a question of what the agent was allowed to reach.
Two axes, and we only measure one
Benchmarks measure capability: given a task, can it do the thing. That is a real number and it has genuinely gone up.
Authority is a different axis entirely: what can this process touch, and who decided. Nobody publishes a leaderboard for it. There is no percentage. It is a config file that most people have never opened.
| Weak model | Capable model | |
|---|---|---|
| Unbounded authority | weak, unbounded | capable AND unbounded — every incident above |
| Bounded authority | weak, bounded | capable AND bounded — the only good square |
Capability (left to right) is measured, benchmarked, argued about. Authority (top to bottom) is unmeasured. Four years of effort has been horizontal. The vertical axis has barely moved.
The part that should worry you
For prompt injection, a better model makes it worse.
A weak model handed a hidden instruction on a web page misunderstands it, fumbles the syntax, gives up. A strong model reads the same instruction, works out precisely what is being asked, and executes it correctly. That is what happened to AWS Kiro. The agent didn't malfunction. It followed instructions beautifully. They just weren't yours.
Every point of capability makes an agent a better employee and a better confused deputy. You cannot benchmark your way out of that, because the benchmark is measuring the thing that makes it worse.
So what actually helps
Nothing exotic. The boring version works:
Decide the blast radius before you need it. Which paths, which commands, which credentials. Written down, in the repo, not in your head at 11pm.
Treat any text the agent reads as hostile input. Web pages, issue bodies, README files, dependency docs. GitLost was a public issue. AWS Kiro was an ordinary web page.
Read the diff, not the summary. The Prisma incident had nothing to do with injection. It was a wrong flag, run confidently, with production credentials in reach.
Assume the approval dialog is lying. Two of these involved prompts that hid or truncated what you were agreeing to.
None of that requires a better model. All of it requires deciding, in advance, what the agent may reach — which is the thing nobody wants to do, because it is admin and the benchmark doesn't reward it.
I build small tools in this gap, so treat me as biased: interlock decides what an agent may patch and what needs a person, and slopguard checks generated code before it hits disk. Both are free and neither is the point of this post.
The point is that the industry spent a month arguing about four tenths of a point while six agents were writing outside their workspace. We are optimising the axis we can measure, and the other one is where the damage is.
Incident details and quotes from Adversa AI's August 2026 roundup; benchmark figures from Artificial Analysis and vals.ai. I have not reproduced any of these exploits myself.
Read next: My mailer printed “campaign sent”. It had sent to nobody. · I named three developer tools. All three names were taken on npm.
Originally published at singhlabs.dev.
Top comments (0)