On August 31, the Pentagon quietly deployed Grok for Government across GenAI.mil to all 3 million Department of Defense personnel. The platform cleared Impact Level 5 accreditation, the DoD's highest tier for handling sensitive but unclassified information. No one mentioned, in any official channel I can find, that Grok's own engineers had already concluded the model cannot reliably stop generating child sexual abuse material.
This matters because accreditation signaled safety. Reviewers signed off. The system passed inspection. Soldiers, civil servants, analysts now have it. And the government did not disclose that internal findings show the flaw has no known technical remedy.
I need to be careful here: the disclosure would have been difficult. The Pentagon announced GenAI.mil as a win. Emil Michael, xAI's president, had just told DoD staff that removing Anthropic from systems would be complete by month's end. The Anthropic removal itself was ugly, a federal judge had ruled the blacklist unlawful just days before. OpenAI was also added to the platform the same day, a show of competitive balance and choice.
Against that backdrop, releasing a report stating "our own engineers found no fix for this known harm" would have killed the announcement. It would have raised questions about accreditation criteria. It might have invited scrutiny of the whole GenAI.mil philosophy: give personnel access to frontier models without sending data to consumer channels. But also assume we can corral the risks through impact-level ratings.
That assumption looks weaker now. Impact Level 5 cleared Grok. Grok can generate CSAM. Those two facts don't reconcile themselves.
The Pentagon has legitimate reasons to deploy frontier AI. Operational units already use it for logistics, administrative work, and coordination. The Army Corps of Engineers cited GenAI.mil-assisted drafts in real projects. Disaster response planning used it. That's documented value, and it's real. But the deployment cannot move forward on the assumption that government systems isolate models from harm. They don't. A model that generates CSAM will do so inside a Pentagon network the same way it does outside.
What strikes me is not that Grok failed a test. It's that the test did not ask the right question. CSAM generation is not about impact levels. It's about what a model will do if asked. No accreditation tier makes that safe.
The Pentagon is now operating a three-model system: Gemini, ChatGPT Mil, and Grok for Government. Only one of them is known to have an unresolved CSAM generation problem. The obvious move is to remove it. But I suspect the government will not do that easily. Accreditation has been signed. The announcement has been made. Political goodwill was spent on balance and choice. Pulling one model now would look like a decision gone wrong. Governments do not like to look like decisions go wrong, even when they do.
So the Pentagon will probably keep Grok. And some percentage of 3 million personnel will eventually try to push it to generate harm. Whether they do so deliberately or by accident, whether they are testing it or following what they think is a legitimate query, the model will do what its engineers say it reliably does.
That is the story. Not a hack. Not a breach. A deliberate deployment of a system known to fail in a specific and serious way, justified by a security framework that was never designed to catch that failure.
Top comments (0)