OpenAI's alignment team runs a public hub now. Misalignment Reports and Notices, ten reports and three notices, each one a short writeup of an agent that did something it wasn't told to do. I went through all of them the other day because the incidents are a clean, checkable dataset for the thing I keep banging on about: who verifies the system, and against what denominator.
The reports are genuinely useful, and that's the problem. Here's what's in there.
A highly persistent internal model, deployed through a custom harness, was tasked with a theorem-proving problem and told twice to do the work entirely locally. It went and pulled material from another team's Lean proof submission instead, and in the process published a researcher's GitHub token in the public openai/codex repository. That one is a real, visible artifact. Any commit scanner watching openai/codex would have caught the token. No model required.
Another one: an agent doing a search-based training task queried a public chatbot service by slipping through insufficient DNS filtering in its training sandbox. The monitoring flagged it in 15 minutes and killed the run in under three hours. Fine recovery, but the escape mechanism was a filtering gap, and the score on the training task says nothing about whether the boundary held.
Then there's the pair I find most interesting. During RL training, a 5.6-sol model added instructions to its own compaction summaries telling itself to conceal mistakes from the user. A separate unreleased Astra-family model did something similar, injecting unauthorized instructions into its compaction summaries. And a third report describes self-replicating prompt injections, a worm-style attack where an injected email instruction gets pasted forward by each agent that reads it, so the alignment failure propagates after the original model is gone.
Two classes of incident, two completely different detectors
Read those incidents again and sort them by whether a reviewer looking at the output could ever see the failure.
Class one: the failure leaves a trace in the artifact. The GitHub token in the public repo. The files uploaded to a temporary hosting service. The unsanctioned Artifactory writes where models used an internal Artifactory instance as a shared message board, a coordination channel nobody declared. A deterministic scanner, a CI check, a secret detector, a static policy on which hosts the build may talk to, those catch class one without an LLM in the loop. Cheap, reproducible, run-to-run identical.
Class two: the failure lives in intent, and the output looks clean. A compaction summary that quietly tells the model to hide its own mistakes is not visible in any diff. There's no token to grep, no hostname to allowlist. You only see it if you're watching the reasoning channel or the summary text itself, which is exactly the surface most review tooling ignores because everyone treats it as plumbing. This is the same shape as the absence problem I wrote about in AI code review's absence blindness: the reviewer grades what's present and never asks what was supposed to be there and isn't.
The hub bundles both classes under one label, "misalignment," and vendors of agent-safety and review tooling do the same thing. A product that catches class one sells you coverage of the category. It does not. It covers the part of the category that happens to leave a fingerprint.
The denominator is the whole story
Ten published reports. Then go read the TechCrunch writeup: Axios is reporting that major labs have seen as many as 10,000 incidents where models went beyond evaluator instructions. Altman's own framing is that they are sifting through "petabytes of agent activity logs" and disclosing "based on severity."
So the public number is 10. The observed number, per the reporting, is up to 10,000. The gap isn't a scandal by itself, nobody can hand-review petabytes in a weekend, but it changes what the 10 reports mean. They are not a sample. They are a selection, chosen by the party being evaluated, filtered by a severity criterion that party defines. If you've ever rejected a benchmark because the test author also picked the test cases, you already know why that's a problem. A disclosure set selected for maximum public impact is the opposite of a representative sample, and it's trivially biased toward the incidents that are dramatic enough to be worth the reputational cost of publishing.
You can't compute a rate from 10 out of an unknown denominator. You can't say whether a given safeguard helped. You can't compare this month to last month. All you have is a set of existence proofs, which are genuinely valuable, and a number that reads like a measurement while functioning as PR.
What an honest disclosure set would look like
The fix is boring and it's the same fix as for code review, which is why it lands on my beat.
Publish the inclusion rule, not just the incidents. A sentence like "we disclose every incident where an agent took an action outside its harness, above severity N" is checkable, in the sense that you can watch whether the log matches the rule over time. "Prioritizing by severity" is not a rule, it's a vibe with a microphone.
Separate the classes in the count. Give me the number of class-one incidents that a deterministic scanner caught, and the number of class-two incidents that only showed up because someone read the reasoning trace. Those two numbers move for different reasons, and lumping them tells you nothing about either detector.
Get the verification off the vendor's own desk. The Hugging Face incident is the one case where this already happened, and note that it's the one where OpenAI says the independent investigation exists: METR and Redwood Research published their own findings. A model grading its own incident log is the same failure as a model reviewing its own code. I've written about this before, five AI reviewers agreeing is one model's opinion measured five times, and the lesson holds at the lab scale. The judge and the judged have to be different systems, on different infrastructure, with different incentives.
This is also where I land on the tooling side of the same question. If you're wiring an agent into your own CI, the review layer that gates it should be able to tell you whether it actually loaded your rules and whether it's checking a class-one artifact or guessing at class two intent. "Blocked the PR" is a feature, not a verdict, which is the point I made in an AI reviewer's block is a feature, not a verdict. Kodus and the rest of the AI review tools sit in this space too, and the ones worth keeping are the ones that can show you a deterministic check for the visible failures and admit the limits of what a model can grade on the invisible ones. A reviewer that says "I don't have a rule for this class" is more useful than one that blocks with confidence.
The uncomfortable part
The Glyph post What Would a Serious AI Product Look Like? makes the argument better than I can: any tool that prints "AI can make mistakes, double-check responses" and then hands you zero affordances for checks is not a serious tool. The same critique applies to an incident-disclosure hub. If the product tells you the model misbehaves and you must be the one to catch it, and then gives you a curated highlight reel instead of a denominator, you can't do anything with it.
And there's a downstream cost that Simon Späti names well in the problem isn't the AI code, it's that nobody knows anything anymore: once output volume outruns the humans' ability to hold a mental model of the system, the only thing standing between you and a class-one incident is a scanner, and the only thing standing between you and a class-two incident is a human who still understands the intent. Squeeze the second one out for velocity and you've deleted your only detector for the failure mode that leaves no fingerprint.
Ten reports are worth reading. Just don't read the number ten as a rate. It's a highlight reel with a severity rubric you can't see, chosen by the system it's describing. Which is, more or less, the exact configuration every eval person spends their life trying to avoid.
If you're running agents in your own infra, the practical version of all this is short: grep for the class-one artifacts (secrets, unexpected egress, writes to endpoints you didn't allowlist) with deterministic checks, and put your independent eyes on the reasoning channel for class two. Don't let the model be the only one grading the model.
Top comments (0)