Anthropic's threat intelligence team says the autonomous style of hacking it first documented in a single suspected state-sponsored campaign in November 2025 "has now proliferated across every class of actors we investigated," from Russian state espionage to opportunistic data thieves. The report, published on 10 September 2026, covers misuse Anthropic disrupted between December 2025 and August 2026. It also describes a cell in northern Yemen that used Claude Code to write guidance software for rockets and missiles.
Key facts
- Detecting and countering misuse of AI: September 2026 covers eight months and seven areas of harm; Anthropic says "a majority" of the operations in it were carried out by AI through direct execution or orchestration.
- One criminal operator mass-downloaded 1.8 million Android apps to hunt for passwords and keys left in their code; in another intrusion, AI agents pulled more than 2,100 sets of corporate login tokens in about 34 hours.
- The Yemen-based cell used Claude on guidance software; Anthropic says its safeguards "blocked many of their requests, but not all of them."
- Primary sources: Anthropic's threat intelligence report and its companion research post on weapons-related capabilities.
The attacker's assembly line
For most of the history of hacking, two things capped how much damage an attacker could do: the supply of working exploits and the supply of skilled people to use them. Anthropic's report argues that AI has loosened both. "The operators behind observed cases range from state services to lone individuals," it says, and publicly available offensive agent frameworks such as PentAGI "reproduce much of the same scaffolding for anyone who downloads them." Google made a similar case earlier this week when it said attackers have moved from prompting to autonomous agents; Anthropic puts named operations behind the claim.
The report's most quotable conclusion is aimed at investigators: "For threat intelligence investigators, sophistication has stopped being a reliable signal of who is behind an operation."
Three operations
Russian espionage. Anthropic says its attribution of the group it tracks as GTG-20006 "is consistent with public reporting linking the actor to Midnight Blizzard," a Russian state-linked espionage group. The operators built an AI workflow that noticed when security products flagged their malware, then modified and redeployed it. "The result of the above is that AI has inverted the cost back onto defenders," the report says. A new detection used to buy defenders days; now attackers can "close the loop."
Smash-and-grab data theft. A cluster Anthropic labels "ShinyHunters smash-and-grab opportunists" used Claude to speed up mass scanning and exploitation. One operator ran a pipeline across ten cloud servers that downloaded 1.8 million distinct Android apps, unpacked them and scanned them for secrets left in the code. In a separate intrusion, after compromising a software provider, the actor dumped more than 2,100 sets of login tokens for Microsoft's corporate sign-in service, spanning more than 40 companies, in about 34 hours. "AI agents performed nearly all of the work." The group's playbook also included prompt injection against deployments of LiteLLM, an open-source gateway many companies put in front of their AI providers. The same week, the cloud security firm Wiz reported that nearly one in ten internet-facing LiteLLM instances it scanned accepted a default master key.
Exploit foundries. Chinese-speaking operators likely based in Changsha, two of whom Anthropic identified as undergraduate students, used Claude as "the engineering and orchestration layer" of an espionage programme, with autonomous workflows running vulnerability and exploit research around the clock.
AI keys are now loot
A newer pattern runs through several cases: attackers stealing AI API keys from victims and running their own operations on someone else's bill. "In every instance, the API keys involved were stolen from Anthropic customers' environments. Anthropic's own systems were not compromised by this actor," the report says. One group, after breaking into an AI vendor's evaluation sandbox, "took its production keys first." Anthropic's advice is blunt: "Organizations should treat AI keys and agent integrations with the same level of seriousness as they do production credentials."
The Yemen case
The report's weapons section details six cases, "three in China, two in Russia, and one in Yemen." The Yemen case describes "a cell of threat actors based in northern Yemen running three weapons development programs," including a guided rocket and a multi-stage ballistic missile. The actors used Claude Code "in place of human software engineers" to build guidance, navigation and control software, running several Claude sessions like a small team: one writing code, one researching, one reviewing the first one's work. They hid their goals and split tasks across sessions so no single conversation revealed the whole programme.
Anthropic is careful about what it knows. "We do not have evidence the actors succeeded in fielding an operational device; but they did test-fire a guided rocket. This field test appears to have failed: within hours, the actors returned to Claude to work out why it failed." It banned the accounts and shared threat information with public- and private-sector partners, while noting the cell had already built an offline simulation toolkit that does not rely on Claude.
Why it matters, and the caveat
For defenders, the report's message is that the gap between detection and redeployment is closing, and that AI keys and agent integrations are now credentials worth stealing. The same report's distillation section is covered separately in our story on Moonshot and DeepSeek quietly routing their users to Claude, and this week's PaperCut campaign shows the same agent-driven pattern from outside Anthropic's platform.
The caveat is that this is Anthropic's own account of its own platform. Outsiders cannot inspect the logs, several attributions are hedged, and a model provider has reasons to show both that misuse is real and that its safeguards work. Anthropic also notes that none of the cases involved its most restricted Fable or Mythos-class models, with one distillation exception. Coverage from The Record, BleepingComputer and Al Jazeera adds context; the Hacker News thread is where practitioners are arguing about it.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)