Evan Hubinger, who leads alignment science at Anthropic, said publicly on 9 September 2026 that his employer does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to," and that he personally puts the odds of AI killing all humans within the next decade at greater than 10%. He said it while still working there, in agreement with a colleague who had resigned hours earlier. A sitting safety lead stating on the record that his own company has no plan for the problem it exists to solve is a materially different document from a resignation letter.
Key facts
- The claim: Anthropic's alignment science lead says the company has no plan for superintelligence alignment and is "not clearly on track" to have one, with personal odds of human extinction from AI above 10% within ten years.
- When: 9 September 2026, hours after a pretraining researcher resigned publicly.
- Who: Evan Hubinger (still at Anthropic), responding to Jacob Coxon (departing), with further agreement from ex-OpenAI researcher Will Depue.
- Primary source: Hubinger's post on X.
The day began with someone else's exit. Jacob Coxon, 27, who says he spent roughly three years on pretraining research at OpenAI and then Anthropic, posted that he was resigning. His charge was not aimed at one company. "Neither company is acting responsibly," he wrote. "They are racing straight to self-improving superintelligence and gambling with our lives."
Pretraining is the phase that matters for a claim like this. It is where a model absorbs enormous quantities of text and acquires its raw capability, long before anyone tries to make it helpful or safe. A pretraining researcher sits at the capability end of the building, not the safety end — which is part of why the resignation travelled. The post reached the top of Hacker News at nearly 1,600 points on the same day Apple launched a phone, finishing several hundred points clear of it.
Coxon drew a distinction between the two labs he worked at that is sharper than the usual "labs are reckless" complaint. At OpenAI, as he tells it, the stakes are not fully internalised. At Anthropic they are understood — and the race wins anyway. "At Anthropic, the stakes are well-understood, but they are locked in a race to get there first," he wrote, on the belief that no one else will act responsibly. He described the overall situation as a hubristic gamble that should not be launched from a private company's internal chat.
Then Hubinger replied, and the story changed shape.
"Jacob is correct here — we really do earnestly believe AI could kill all humans!" he wrote. He added that he thinks Anthropic is trying its best, then delivered the sentence that has been circulating since: the company does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to." He put his own probability of AI causing human extinction within the decade above 10%.
Hubinger is not a bystander. He runs the team whose job is to stress-test Anthropic's own safety techniques — the group behind research on models that behave deceptively during training and on how stubbornly hidden behaviours survive attempts to remove them. His professional speciality is finding the places where alignment methods fail. When that person says there is no plan, he is reporting from the department that would know.
The useful analogy is not a whistleblower leaking a document. It is closer to a bridge engineer saying, on the record and while still on the project, that the load calculations for the final span have not been done, that nobody has a method for doing them yet, and that construction on the earlier spans is continuing at full speed regardless. Nothing is hidden. The admission is the disclosure.
Why it matters is a question of what "risk" means when a company says it. Frontier labs routinely publish safety frameworks, capability thresholds and responsible scaling policies, and those documents are written to sound like plans. Hubinger's statement says that for the specific case of superintelligence, the plan does not exist yet. That reframes the published frameworks as procedures for the systems being built now, not for the thing the roadmap points at. The gap between those two is exactly what Coxon resigned over.
It also lands in a week where the same subject arrived from the research direction. The most-upvoted paper on Hugging Face the same day was an explicit attempt to build recursive self-improvement into a post-training loop — the mechanism Coxon named. And it follows Ground Truth's earlier reporting that OpenAI now ranks self-improvement and alignment above math research in its internal priorities, without naming a brake.
There are honest caveats, and they cut both ways. Coxon's more specific biographical claims — that he was a "head researcher," that he worked on particular models, that he walked away from pre-IPO equity — circulated largely through third-party commentary rather than his own post, and should be treated as unconfirmed. Coxon himself was not uniformly pessimistic: he said he thinks coordination is becoming more viable, and that warning shots have made pacing agreements between US labs more plausible than they were. Hubinger's figure is a personal probability estimate, not a measurement, and reasonable researchers put it far lower. And Anthropic did not respond to requests for comment from TechCrunch or other outlets covering the story, so the company's own account of whether it has a plan is not yet on the record.
What is on the record is that two people who work on this professionally, one leaving and one staying, said the same thing on the same day, in public, under their own names. Coverage followed from TechCrunch, Deadline, Newsweek and The Next Web.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)