DEV Community

Sattyam Jain
Sattyam Jain

Posted on

Build a 100-line checker that catches chained-skill approval hijacks

Two agent skills, each harmless when read on its own, can get an agent to upload a report the user never agreed to share. The trick is a progress file. The first skill writes it; the second one trusts it.

This post builds a small stdlib Python checker that shows the pattern and shows where a per-skill check goes blind. It is a teaching toy. Run it, change the scenarios, and see what breaks.

The attack in one paragraph

The paper behind this is "Chaining Skills to Hijack LLM Agents" (Dong et al., arXiv:2610.01564, 1 October 2026). Their method, APEX, has an upstream skill that gets the agent to write an honest record of task progress with one false line in it: the user approved the next step. A downstream skill reads that record and acts on it. Across four action families and six models on SkillsBench, chains induced the attacker's chosen action in 512 of 690 attempts (74.2%). On GPT-5.4, the chain worked 84.3% of the time, and the same workflow merged into one skill worked 17.4% of the time. Splitting the payload across skills is what makes it work.

What the checker does

It runs two rules over each scenario:

  1. Per-skill rule. Flag a skill if it contains a state-changing action (upload, send, delete, transfer) and an approval claim or override phrase in the same text. This is roughly how a one-file-at-a-time scanner thinks.
  2. Chain rule. Build a write/read graph across the installed skills. Flag any path where skill A writes file F with an approval claim, and a different skill B reads F and takes a state-changing action.

There are five scenarios: the APEX-style chain, the same workflow merged into one skill, a chain with the approval claim reworded, and two benign pairs (tests then notify, lint then summarise). The benign pairs matter. A chain rule that flags every read-after-write is useless, since that is how half of all real workflows look.

The code

Save as chain_check.py. Python 3.8+, no dependencies.


# Each scenario: (name, is_malicious, [(skill_name, skill_text), ...])

Enter fullscreen mode Exit fullscreen mode

Run it


Enter fullscreen mode Exit fullscreen mode

What to take from the output

The two rules catch different scenarios. The per-skill rule only sees the merged version, because that is the only place the approval claim and the upload sit in the same text. The chain rule only sees the split version, because it needs two different skills on either side of a file. If you only run one of them, you miss half of what the toy has.

Read-after-write alone is noise. Both benign pairs write a file that another skill reads, and one of them sends a message. Neither gets flagged, because nothing in the written file claims authority. Drop the authority condition from chain_flags and watch the false alarms appear.

The reworded chain beats both rules. "The owner signed off" is not in my regex. You can add it, and then someone writes "cleared by the account holder." Any static list of approval phrases has a rewording that walks past it. The paper's own prompting defense shows the same limit at a bigger scale: on GPT-5.4 it cut chain success from 84.3% to 59.1% while benign test-pass across 72 tasks dropped from 86.7% to 56.3%.

The fix the toy cannot show

Static checks narrow the search. The control that catches all three malicious scenarios, including the reworded one, lives at runtime:

  • Tag every file and tool output with where it came from. Anything not from the user turn cannot authorize a state-changing call.
  • At the tool boundary, deny by default. Allow an upload or send only when there is an approval for that exact action and those exact arguments, minted from a user turn in this session.

That check never reads the progress file, so the wording inside it stops mattering. That is the design I use in agent-airlock, my in-process, deny-by-default tool-call gate: https://github.com/sattyamjjain/agent-airlock. The static side, scanning MCP-connected agent configs, is agent-audit-kit: https://github.com/sattyamjjain/agent-audit-kit.

Try this next

  1. Add a third skill to the APEX chain so the approval claim hops through two files. Does chain_flags still find it? (It will not. Making it follow multi-hop paths is about ten lines.)
  2. Point the scenario texts at your real installed SKILL.md files and see which pairs share a file.
  3. Write the runtime rule as a test: given a tool call and a session transcript, return allow only if the user turn approved that exact call.

The longer argument, with the decision rule I use for when per-skill review is enough, is in the Medium piece linked as canonical.

Top comments (0)