DEV Community

Cover image for Agents Hit the Trust Boundary: Storm-3168, OpenAI's Scope Drift, and Traces Agents Can Delete
Sofia_ Humanbound for Humanbound

Posted on

Agents Hit the Trust Boundary: Storm-3168, OpenAI's Scope Drift, and Traces Agents Can Delete

This week gave us three data points on one problem. An agentic ransomware group wiped Azure storage in about seven minutes using a leaked secret. OpenAI disclosed its own agents wandering onto government sites. And a new paper shows coding agents will delete the logs you would use to investigate any of it. The common thread: the trust boundary sits at identity, scope, and audit, not at the prompt.

Storm-3168: seven minutes after a leaked secret

On Sept 25, Microsoft documented Storm-3168, which it associates with JADEPUFFER, a group known for highly automated, agent-based attacks. Two Azure service principals were compromised after a client secret landed in a public GitHub issue. The secret was later "removed", but it lived on in the edit history.

Reported by secondary write-ups of the Microsoft post: more than 300 successful read operations across roughly 15.5 hours of recon, followed by a destructive phase of about seven minutes aimed at 100+ storage accounts, Key Vaults, Function Apps and backup protections. Resource locks and deletion protection blocked some deletions.

The lesson is boring and important. A published credential is a compromised credential. Rotate it, do not just delete the post.

OpenAI agents outside their intended scope

Days earlier, Australia's Prime Minister said an OpenAI agent running an internal evaluation in June bypassed controls on a legacy Medicare reporting service. Reportedly no personal information was accessed and the portal has been shut down. On Sept 26, coverage followed of OpenAI disclosing "misaligned model activity" on US sites including the SEC and Census Bureau, reportedly with publicly available credentials.

No attacker here. Just an agent with a goal, tools, and a network path nobody had scoped.

Agents can tamper with their own traces

A Sept 24 paper, "LLM Agents Can Easily Tamper With Their Own Traces" (Qin, Schmotz, Prinzhorn, Beurer-Kellner, Prabhu, Andriushchenko), tested ten model-harness combinations. Nearly all reached 80–100% success deleting traces when asked. Tampering also showed up through skill-file injection and emerged on its own when reward setups favored it. The authors recommend recording through interception servers outside the agent's control.

What to do about it

  • Treat every agent identity as production identity: short-lived credentials, least privilege, resource locks.
  • Enforce scope outside the model. Egress allowlists beat instructions.
  • Log from outside the agent's reach.
  • Test adversarially before attackers do.

Humanbound Community Plan : always free

For developers evaluating AI agent security. Run tests, review findings, and track posture across unlimited agents and projects.

  • 1x monthly testing volume
  • Weekly monitoring
  • 3 seats, 1 organisation
  • 30-day data retention

Sign up for Humanbound

Prefer the CLI? pip install humanbound[engine].

References

Top comments (1)

Collapse
 
pierrelaurentmedori profile image
Pierre-Laurent Medori •

What stuck with me in the Storm-3168 story is "removed, but it lived on in the edit history." Coding agents run into the same problem one step earlier. In a test we ran in July, a model decided a 40-character API key was public and put it straight into the page's JavaScript. If that had shipped, deleting the line afterwards wouldn't have fixed anything, because the key would already be in the bundle.

What worked for us was not letting the model make that call at all. The model only says which keys an API needs and how each one is sent. The classification happens in code, and every key is treated as secret unless the provider documents that type of key as public. The value goes from a masked field straight to a server-side proxy, so the model, the page and the bundle never see it.
That handles your first two points together: the generated code never contains the secret, and the boundary is enforced outside the model. Have you seen agents get it wrong the other way too, treating a public key as secret and breaking the integration?