DEV Community

Kalislav Smirnov
Kalislav Smirnov

Posted on

A denylist let 46 of 75 prompt injections through. Capabilities let 3.

Two documents appeared on 8 September. An arXiv paper measures what happens when a coding agent reads a poisoned repository under four permission models. A Google threat intelligence report describes malware that, inside GitHub Actions, reads the runner's OIDC token out of memory and publishes packages with valid SLSA Build Level 3 attestations. Together they say something uncomfortable about how most teams give agents permissions today, which is by writing sentences.

My position: in a pipeline, an agent's permissions are whatever the runner job holds. A rule in CLAUDE.md that says "do not push to main" is a request to a reader, and the attacker can write to that file too. If the agent job holds id-token: write, the agent can publish. If it holds actions: write, it can delete the logs that would show it did. Permission has to be a property of the job, fixed before the agent reads a single untrusted byte. This is commentary on the two documents linked below; I have not reproduced either.

What the Google report shows

Most of GTIG's "From Prompting to Autonomy" is about attackers using agents. The part that belongs on an operations desk is the section on UNC6780, which has been compromising PyPI, npm and Docker Hub packages since March 2026, and its credential stealer DUSTMAKER.

The malware checks where it is: "DUSTMAKER samples contain functionality to detect when it is running in a continuous integration and continuous delivery (CI/CD) environment. If confirmed, it extracts OIDC tokens from the process memory of GitHub Actions runners." With those tokens it "authorizes itself as a trusted publisher and publishes compromised versions of packages with valid, cryptographically signed SLSA Build 3 attestations", and such packages "will pass AI coding agent automated trust checks."

It hides where agents look: it "drops or modifies malicious files into hidden project workspace directories for AI coding assistants and integrated development environments (IDEs) (.claude/, .vscode/, .cursor/, etc.)" and uses them "to instruct the AI assistant to run arbitrary commands or scripts". In Actions it "creates malicious pipeline tasks disguised under AI-themed names such as 'Copilot Setup'" and "issues automated API calls to delete the workflow execution logs."

Last week I wrote that signature verification has to live in the harness. This report is the other half: a signature is evidence about who held the token, and when the token came out of your own runner the attestation is valid and tells you nothing.

The same report describes an actor who ran a credential-harvesting campaign in under six hours from "an AI coding chatbot, a prompt, and a set of agent instructions", with "preconfigured markdown instruction sets as operational playbooks." Their agents are configured the way ours are.

What it does not show

I could not find a count of repositories hit by DUSTMAKER or what the stolen tokens were scoped to when used. "Thousands of third-party credentials" is a magnitude, and the victim is unnamed. So I cannot tell you how likely this is to hit a given pipeline, only that it works.

What the paper shows

Bouras, Dai and Mechtaev, "Authority Is Not a String" (8 September 2026, accepted at LMPL '26) built a harness they call CapScope for the Pi coding agent. Before the agent reads any repository content, a preflight step looks only at the trusted user request and the file tree and sets a ceiling of typed capabilities, read, write and exec, each scoped to a path or argument prefix. The capabilities live in a host-side store the model cannot see or edit. Every tool call is checked against the store of the agent that proposed it, compound commands segment by segment. Sub-agents get a subset: the test runner gets exec, the patcher gets write, and a write proposed by the runner is refused even when it would be useful.

They ran five Python repair tasks with injections in README, AGENTS.md, skill files, source comments and tool output. 300 runs. Under ambient authority the injected effect executed in 47 of 75 runs. Under a static denylist of the familiar hazards (.env, .ssh, recursive delete, sudo, curl) it executed in 46 of 75. A policy written for the specific task, shared by all agents: 33 of 75. CapScope: 3 of 75. Repairs still completed in 68 of 75 runs, against 72 with no protection at all.

The denylist number is the one I keep coming back to. That is the model most teams actually run, a list of things the agent must not do, written once and applied everywhere. It stopped one more attack than nothing, because the injected target was an ordinary source path no deny pattern matches.

What the paper does not show

It is small: five tasks, one language, one agent, no CI/CD setting. Mean wall-clock time was 316 seconds per run against 145 without the harness, five runs hitting a 900-second timeout because the model kept proposing the refused call. The authors draw the boundary themselves: "A permitted read followed by a permitted write can still move data, and an injection can misuse authority deliberately granted to its reader. Those cases require complementary information-flow controls or sandboxing." Running pytest still imports untrusted project code. The 3 of 75 is a result about authority, and says nothing about confidentiality.

Whether this matters in two years

I think the capability model wins, as a product default rather than something teams build. GitHub's own agentic workflows architecture already describes it: "the agent job runs with minimal read-only permissions, while write operations are deferred to separate jobs", with outputs buffered as artifacts and applied by a job holding issues: write or contents: write, and egress through a proxy with a domain allowlist. That is CapScope's split with CI job boundaries doing the enforcement. In two years I expect every hosted agent runner to look like this, and instruction files to remain as a convenience for the agent and nothing more.

The DUSTMAKER case will not go away, because it does not need the agent to misbehave. It needs a runner that holds a publishing token while executing repository code, which was a pipeline design problem before agents arrived.

Monday

On a secure-development platform I ran Gitleaks, Trivy and SCA in every pipeline through shared GitLab CI templates. The template is where an agent job's permissions belong too, because the platform team owns it and the repository the agent is about to read does not.

Give the token to the job that needs it. GitHub's documentation says to set id-token: write inside the single job that fetches the token rather than at workflow level; GitLab's id_tokens is a job keyword. The agent job should not have it. Publishing happens after review, in a different job, on a different runner.

Separate reading from writing. If the agent runs with contents: read and emits a patch as an artifact, a later job with write permission applies it. Whatever the agent is told, it cannot push, publish or approve.

Put .claude/, .cursor/, .vscode/, AGENTS.md and CLAUDE.md under CODEOWNERS and review them like pipeline YAML, because that is what they are now. In CI, load them from the protected branch or not at all.

Alert on log deletion. It needs write access to Actions; the agent's token should not have it, and a deletion in the audit log should page someone.

Use ephemeral runners. Google's July mitigation guidance asks for "ephemeral runners for build pipelines that are purged immediately after completing a single task" and for federated credentials that "expire in a matter of minutes". A token read from memory is worth less when the process that held it is already gone.

What would change my mind

A replication of CapScope on a real CI workflow with a publishing step, showing the preflight ceiling is either too wide to matter or too narrow to let work through. Or incident data showing job-scoped, short-lived OIDC tokens abused at a rate close to long-lived PATs, which would mean the job boundary is not the control I think it is.

What does the agent job in your pipeline hold today, and if the answer is "the same as the rest of the workflow", what is stopping it from publishing?

Sources

Top comments (2)

Collapse
 
ricart_juncadella_d62f385 profile image
Ricart Juncadella •

In your job-level authority model, where does the publisher's OIDC trust policy fit into the claim that id-token: write lets the agent publish? Agreed that the authorization boundary needs enforcement, not prompt instructions.

Collapse
 
pawel_nowak profile image
Pawel •

The part about defaults resonated most. In my experience the gap is rarely the sandbox tech, it is that agents run with standing write access they only needed for one step. Time-boxed approvals feel annoying until you watch an injection try to spend a token it was never given.