Ten days ago, Sam Altman promised that external safety evaluators would get "desks, badges and laptops, with the right to publish what they found." On Tuesday, OpenAI announced its framework for third-party safety assessments. The specifics shrink the commitment significantly.
Altman's September 12 post said evaluators would be embedded inside the company with real access and publishing rights. Tuesday's blog post uses softer language: evaluators "may be brought into the offices for the most sensitive work." It names no confirmed partner. It sets no access terms. It lays out seven principles that sound good in principle, "strong independence mechanisms," "scientific rigor", but doesn't specify what those mean or how they'll be enforced.
OpenAI is in talks with METR and Redwood Research, two groups that have done work with the company before. That's not nothing. Both have credibility in the safety research space. But the move from "here's who we're committing to, here's what access looks like" to "we're in talks" is a step backward from what was promised.
The timing matters. This announcement comes after TechCrunch and CNBC pieces questioning whether embedded evaluators can actually stay independent when the lab controls office space, access, what work is "sensitive enough" for in-person review, and which findings see daylight. OpenAI's response to that skepticism was to dial back the commitment, not deepen it.
Anthropic is doing something similar but with a bigger price tag. Anthropic's parallel move embeds Accenture evaluators at a reported cost of at least $1 billion over five years. That's expensive, but it's also concrete: Accenture has real leverage when you're paying them a billion dollars. OpenAI's setup is hazier.
The substantive part of Tuesday's post is useful. OpenAI's Preparedness Framework covers risk categories including chemical and biological risks, cybersecurity, and AI self-improvement. The company identified four priority areas for external review: assessment of safety cases spanning training and deployment, evaluation of critical safeguards, review of capability evaluations tied to its Preparedness Framework, and independent investigation of misalignment incidents. That lays out a real menu of what external review could look like.
But here's the tension: laying out what you could do and committing to what you will do are different things. Altman said evaluators would get "desks, badges and laptops, with the right to publish what they found." The post now says evaluators "may be brought into the offices for the most sensitive work," names no partner, and sets no access terms.
This is what walking back a commitment looks like when you can't quite say you're walking it back. You announce a framework, use careful language, and wait for people to move on to the next story. The gap between "we're embedding independent evaluators" and "we might bring you in for sensitive work" is the space where real oversight goes to die.
Top comments (0)