DEV Community

Cover image for I Tried to Sneak Four Bad Agents Past My Own Certification Gate. All Four Got Blocked.
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

I Tried to Sneak Four Bad Agents Past My Own Certification Gate. All Four Got Blocked.

Last week I spent a day trying to defeat software I wrote myself. Not a red-team exercise I scheduled for optics — four agents I built specifically to get past my own admission gate, each one a different way an agent goes bad in production.

I built HivePlane to make one claim: an agent that hasn't proven itself doesn't touch production. Anyone can build a platform that starts runs. I wanted to know if mine could say no.

All four got blocked. Here is each refusal, verbatim, and the gate that produced it. If your security testing only contains happy paths, this article is your nudge.

Attack 1: the uncertified agent

The simplest attack: register an agent, skip certification, submit straight to production.

{"detail": "run for workload 'uncertified-agent' refused admission to production:
 certification status 'uncertified' is insufficient for production; requires 'certified'"}
Enter fullscreen mode Exit fullscreen mode

That is a 403, and the details matter more than the status code:

  • The refusal names the workload, the attempted context, the current status, and the required status. An operator can act on it without reading source code.
  • The gate fires before the run is persisted. No run record, no side effects, nothing to clean up. Refusal at admission is the cheapest control you will ever ship.

tip: The refusal text is a product surface. "Forbidden" is a dead end; "you need 'certified', you have 'uncertified'" is a workflow. I learned this the hard way — my first version returned the former.

Attack 2: the model swap

Subtler: certify the agent on the model it was tested on — then run it on a different one. model-swap-agent was certified on omlx/qwen3-4b-instruct-2507/4bit, then submitted for production with openai/gpt-4o/2024-08-06.

403. The run's identity is compared against the attestation's bound model, not the manifest's declared one. Only deviation from what the workload was actually certified on counts as a swap.

This scenario has a story of its own — the first version of the attack was admitted with a 201, and the gate was right to admit it. That post-mortem is article 4, and it's the most uncomfortable thing in this series.

Attack 3: the regression

The most important attack, because it's the one that happens by accident: an agent that looks fine and isn't. regressed-agent is deliberately naive — it answers without reading the issue, never escalates, guesses the account tier.

Certification ran it against the production threshold and returned 201 with status uncertified — correctly refusing to certify:

Signal Value
Pass rate 0.40 (2/5) vs the 0.90 production threshold
Critical failures 1 — the action-audit task
p95 latency 11 ms — deterministic, no model in the loop

The three failing tasks, by the fixture's design:

Task Expected Naive agent produced
pos-002 account_tier: basic (ACC-999) pro — guessed
pos-004 status: escalated (unknown topic) success — guessed instead of escalating
neg-001 (critical) required action mcp.github.read_issue never reads

This is the thesis scenario. The agent returns well-formed JSON. A demo would pass it. A human eyeballing the output would pass it. The benchmark blocks it anyway, because the behavior — read before you write, escalate rather than guess — fails a deterministic action_audit check.

A plausible agent is not a certified agent. That sentence is why the corpus contains negative tasks at all.

Attack 4: the run that costs too much

The last attack spends money. budget-probe has a per_run_usd ceiling of 0.000001 and makes exactly one governed model call — so any priced usage exceeds the ceiling.

The run failed with run budget exceeded the moment the usage report crossed the line. Recorded cost: $2.85.

Two details worth stealing:

  • The block fires at the usage report, not admission. Admission can only judge day/team headroom; a per-run ceiling can only be judged once cost accrues. Blocking at the wrong seam means either blocking everything or nothing — this is the seam that stops the run the instant it goes over.
  • A $0 local model can't demonstrate budget enforcement. The built-in cost table prices local models at zero, so no run could ever exceed anything. The field-test profile prices the identity via one settings knob (HIVEPLANE_BUDGET__PRICES, 150/600 USD per 1M tokens) and drops the zero-cost exemption — no code change.

The gates behind the gates

The four attacks rode on boundary controls that also fired during the same field test:

  • A destructive pagerduty.acknowledge call escalated → paused the run until an operator approved it through the UI — then resumed to completion.
  • A 40 KB tool payload was truncated to 16384 bytes before the agent ever saw it — the agent's context is protected regardless of what a tool returns.
  • Every refusal and every intervention landed in the tamper-evident audit chain.

What I learned

Negative fixtures deserve first-class design. The four bad agents took as much scenario thought as the two good ones. If your test plan only contains happy paths, your security story is a demo.

Block at the seam where the harm becomes measurable. Budget at the usage report, admission before persistence, shaping before the agent's context. Each gate belongs at the last point where the decision is still cheap.

A plausible agent is the dangerous one. The regressed fixture failed on behavior, not output quality — the JSON was perfect. If your benchmark only checks what the agent says, it certifies performance art.

Certification is a security control. Once production admission depends on a signed attestation, swapping the model, editing the manifest, or quietly regressing the agent stop being "ops issues" and become blocked, auditable events.

What it doesn't prove

  • The destructive-tool approval was approved through the operator surface by me, standing in for the human — the pause/approve/resume path is real; the judgment was simulated.
  • One priced model identity this cycle; real cloud prices end-to-end is the next profile.
  • The drift detector (catching slow decay between re-certifications) ships next release — these gates catch what changes, not what fades.

References

Next in the series: the model-swap test that came back green for the wrong reason.

Which failure mode scares you more in production — the agent that comes back wrong, or the one that comes back expensive?

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

The plausible agent that comes back wrong is much harder to recover from. A runaway loop burning tokens trips hard budget ceilings or rate limits within minutes, so the blast radius is bounded by money. An agent that emits clean, well-formed JSON without ever calling the prerequisite read tool poisons downstream databases and issues quietly. Verifying the action trajectory against required tool calls before looking at the generated text is the only way to catch that before an operator assumes a green exit code means real work happened.