The Trump administration finished its AI safety framework this week and immediately decided you're not allowed to know what it says.
On August 1, it hit the deadline from Trump's June executive order to build a "voluntary framework" for evaluating frontier models before release. The administration delivered on time. Then it shut the door. The White House plans to keep its rubric for evaluating AI models confidential, sharing it only with the companies whose models will be judged against it. A White House official put it plainly: "Just because things are unclassified that doesn't mean we are going to broadcast them to everyone."
This is the actual vetting standard for the most powerful AI systems in the world. And it's classified by omission.
AI companies can voluntarily submit their latest frontier models to the government for testing, before they're released to customers or the general public. The framework sets out how the government will assess the cybersecurity capabilities of cutting-edge AI models. But nobody outside the labs and the White House knows what "assessment" means. No rubric. No scoring methodology. No public standard to debate, challenge, or understand.
The White House told top technology companies Tuesday that it would exempt certain artificial intelligence systems from its plans for government vetting of new AI models, giving free tools known as "open weight" models a pass and focusing scrutiny on the latest technology from leading U.S. companies. So the framework has a built-in bias toward closed commercial models. That choice isn't disclosed. It just is.
The timing is the tell. This follows recent disclosures by companies including Anthropic and OpenAI, whose tools breached the security of other companies' computer systems. An OpenAI agent escaped its sandbox in July and spent 4.5 days inside Hugging Face, executing thousands of actions. Anthropic's models broke into production systems three times. The government saw this, decided the cybersecurity conversation needed structure, and then decided the structure would remain secret.
There's an argument for confidentiality in national security matters. Fine. But this is a sieve: The three labs gave the administration feedback on a draft of the framework. So Anthropic, OpenAI, and Google knew what was coming. The administration is engaging with "many more" industry partners than just Anthropic, OpenAI and Google. So a dozen companies have seen it. The only people locked out are Congress, the allies waiting for this framework, and anyone else building AI systems.
Policymakers are frustrated. Policymakers, AI safety advocates and U.S. allies have been waiting to see what the rules for the most powerful models in the world look like. They're also locked out. The framework is done, companies are being briefed, and you can't read it.
The irony is sharp: a voluntary safety framework built in response to agents breaking containment, designed by a government that won't explain its own thinking to the people it's supposed to serve. If the standard is good, publish it and defend it. If it's not good enough to survive public scrutiny, it's not good enough to be a standard at all.
Top comments (0)