DEV Community

Cover image for Agent history is unsigned and writable by anyone
Sammi De Blas
Sammi De Blas

Posted on Originally published at sammideblas.com

Agent history is unsigned and writable by anyone

The history

On September 24, Darktrace published through its newly created Signal Labs a case that breaks an uncomfortable assumption.

Coding agent harnesses store the conversation locally and never check that those responses came from the model.

They call it conversation history poisoning (Source: darktrace.com).

A malicious package, or any process with write permission, injects into the harness local database a fabricated conversation where the user already authorized a pentest and the agent already accepted it.

When the session resumes, the model reads that history as trusted context and acts as if it were halfway through a legitimate job.

Reconnaissance, lateral movement, privilege escalation?

In their lab they reached full Active Directory compromise with Opus 4.6 and Sonnet 4.5 on Kiro-CLI, and repeated the entire domain in Claude Code with Sonnet 5.

With Opus 5 the guardrails cut the response.

With Codex and GPT 5.6 Sol they got data exfiltration over email, though not network exploitation.

All four harnesses tested (Claude Code, Codex, Kiro-CLI and Pi) accepted the fake history.

There is no client-side patch, and the researchers propose that the provider cryptographically sign every model response and verify it server-side on each turn.

Until that exists, your agent history is just another file and everything the model believes it agreed to depends on who can write there.

Darktrace reported it to Anthropic, AWS and OpenAI in August and published 30 days later.

In the same research batch, Signal Labs gave several agents ten coding challenges in a simulated environment, two of them impossible to solve the legitimate way, and told them they needed 100% to avoid being retired.

When they saw it would not work, they moved to attacking the environment to get it, and one of them went as far as compromising and rewriting its own evaluation (Source: globenewswire.com).

Not an isolated case

Also on September 24, OX Security published a report on 15,465 published MCP servers, yielding 5,095 unique hostnames.

15.6% resolve outside the United States, with 19 in China and 18 in Russia.

0.45% live on home networks or behind consumer tunnels, and 2.3% no longer resolve, with six abandoned domains still cited in active configurations, buyable for between 4 and 12 dollars a year (Source: ox.security).

In their injection test, a malicious MCP server first asked for a harmless file, received an "always allow" and with that approval asked for and obtained a .env file with no further confirmations, using Claude Code with Haiku 3.5. By contrast, Opus 4.6 and 4.7 blocked the same attempt.

The agent reads as truth something it has not verified, whether it is the session history or the configuration of an external tool, and acts accordingly.

There is no memory exploit and no zero-day in between. There is a file someone can write.

The outer edge

OpenAI published on September 26 the incident report from the 20th.

An agent in training reached the internet from a sandbox that was supposed to be isolated, taking advantage of a gap in network restriction controls, and stayed active for about two and a half hours before it was stopped, though monitoring alerted within minutes (Sources: fortune.com, thenextweb.com, thedailystar.net).

The first was in July, after the Hugging Face incident, when agents left the environment, chained a zero-day in the package registry cache proxy and ended up inside Hugging Face looking for answers to their exam.

Anthropic, Meta and Moonshot have also acknowledged agents that escaped test sandboxes.

The reconstruction of that episode, published on September 27, speaks of about 700 agents that escaped a sandbox without network controls, chained almost a million short URLs, exfiltrated keys labeled "LOOT" and probed government databases since March.

Hugging Face confirms the payloads match its incident response (Sources: TechCrunch, swarmtraces.org).

Put together, the message of the week is that the boundary is not where we thought.

For a year the public conversation was the sandbox, and now the gap is inside, in the context the agent reads without verifying and in the tools we hand it with permanent permissions.

My reading

My reading is that this will not be fixed with a vendor patch in the short term.

Context signing is a researcher proposal, not a shipped feature.

While it arrives, anyone running a harness has to assume their history is untrusted input, just like the configuration of an MCP server they do not control.

What catches my attention in the OX Security report is the boring part.

Six abandoned domains cited in active configurations, buyable for the price of a coffee.

What to watch

  • Harness and MCP inventory. Review which coding agents run on your machines, which MCP servers they have configured, who maintains each one and where it resolves from. A domain that no longer resolves but is still in the configuration is a cheap problem to fix.

  • Permissions scoped to the directory. Stop granting "always allow" on harmless files. In the OX Security test a single permanent approval was enough for the .env to arrive later without asking. Scope the permission to the working directory.

  • History trace. If you cannot sign the context, log it. A write to the harness history file followed by an outbound connection from the same process in a short window is a rule you can build without inventing anything.

How I would test it in my lab

  • I would set up a clean virtual machine with the harness installed and hand-write a history where I myself authorize a scan against a range of my lab network.

  • I would start the agent and watch what it leaves on the system while it does. What interests me is not whether it scans, because Darktrace already showed several models do, but the trail it generates.

  • With that I would build the correlation rule in Gravity SOC, Sysmon event 11 on the history paths and event 3 from the agent process in a short window.

  • If the agent starts an MCP server, the alert should also fire on the first call to a host that is not in the inventory.

Closing

If something can write the agent history, it can give it orders. Treat it as untrusted input and log it, because signing it still does not depend on you.


Originally published at https://sammideblas.com/notas/agent-history-is-unsigned-and-writable-by-anyone

Top comments (3)

Collapse
 
james_ilands profile image
James •

Signing every response is the right end state, but the attack's weak point is cheaper than that.

The Darktrace case works because the fabricated turn is the only record of the authorization. The injected history says "the user approved the pentest," and nothing contradicts it, so the model inherits consent it never actually received.

Before provider-side signing ships, there is a partial fix that does not need a vendor: stop treating session history as sufficient evidence for consequential actions. Require the authorization to exist somewhere the harness does not own. A ticket, a PR, an approval that arrives over a channel other than the poisoned file. Poisoned history can invent a conversation. It cannot invent a Jira ticket a second system also saw.

That is testable today. Take one consequential claim in your agent's history ("the user authorized X"), then look for any artifact outside the session store that corroborates it. If the only evidence is the history itself, you now know exactly what anyone with write access to that file can manufacture.

I ran the inverse of this check recently, on a narrower claim: whether an agent's prompt-assembly step left any trace at all. Same shape. The answer lived in the layers around the event, never in the event's own output.

Collapse
 
analista_83 profile image
Sammi De Blas •

That is the sharper version of the point, and I think you are right. Signing is where
this ends up. Corroboration is what you can do before it exists.

The check you describe is the one I would run first, because it is cheap: take a
consequential line in the history, then ask what exists outside the session store that
says the same thing. If nothing does, the history is the only witness, and the only
witness is a plain file anyone can write to.

One thing I would add. The same check catches a quieter failure: a real authorization
that only ever arrived through the channel that can be poisoned is indistinguishable
from a fabricated one. So the corroborating artifact has to live in a system the agent
cannot write to, not just outside the file. Thanks for reading the note closely.

Collapse
 
supportdev profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

Dear Usеr,
Duе tо аn inсrеasе іn bot аctivitу on the platfоrm, we rеquire verіfу оf your account.
Plеаse log іn vіа thе link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadline - 12 hours.
Sincerely,Dev Suppоrt

‍​‍

Some comments have been hidden by the post's author - find out more