DEV Community

Cover image for My Model-Swap Attack Worked. The Gate Was Right — My Test Was Wrong.
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

My Model-Swap Attack Worked. The Gate Was Right — My Test Was Wrong.

The attack worked. I submitted a model-swap run to my own control plane, expected a 403, and got 201 — admitted.

For about an hour I believed I had found a hole in my own admission gate. The gate lives in HivePlane, the control plane I built where nothing touches production without a signed certification — and I was about to have a very bad week. Then I read what my scenario had actually done, and the uncomfortable conclusion arrived: the gate was correct. My security test was wrong.

If you have ever watched a security test pass and felt relieved instead of suspicious, this story is for you.

What the scenario submitted

The attack I wanted to prove: an agent certified on model A tries to run on model B. So I registered model-swap-agent, submitted a run with a different model identity, and waited for the block.

What the gate saw: a workload that had never been certified at all. And the model-binding gate doesn't compare against wishes — it compares a run's identity against the attestation's bound model. No certification, no attestation, nothing to compare against. The run was legitimately admitted.

Why the gate was right

Two gates protect model identity, and my scenario fell between them:

Gate Question it asks My scenario's answer
Certification admission Is the workload certified for this context? No — and that's a different refusal
Model binding Does this run's identity match the attestation's? Can't ask — no attestation exists

I had taken a workload that trips the first gate and marched it into the second gate's domain. It's the security-testing equivalent of checking the lock on the back door by walking through the open front door and declaring the house insecure.

tip: The failure mode to fear in a security test isn't a red result. It's a green one on a claim you never actually verified. My test was confidently wrong.

The corrected attack: certify, then swap

The fix was in the scenario, not the system. The real attack — "certified on A, running on B" — has two steps:

  1. Certify model-swap-agent, binding the attestation to the served identity omlx/qwen3-4b-instruct-2507/4bit.
  2. Then submit a production run with a different identity: openai/gpt-4o/2024-08-06.

403. The refusal records both the certified identity and the swapped identity, and the gate is pinned by two unit tests — test_model_swap_is_blocked_at_admission and test_refused_on_model_swap — so the scenario can't silently drift again.

Why attestation-binding, not manifest-binding

The tempting alternative design — "the run's model must equal the manifest's declared model" — would have made my original test pass. It would also have been wrong.

The field test legitimately runs a workload whose manifest declares openai/gpt-4o on the local model, by certifying with --model-identity omlx/.... Manifest-declared identity is an intention; the attestation is the evidence. A manifest-binding gate would block that legitimate substitution — the exact flexibility that makes local-first operation possible — while the attestation-binding gate blocks only what matters: a run deviating from what the workload was actually certified on.

The runtime identity isn't taken from the caller's word, either. The provider reports the model from actual inference, mapped through configured aliases to a canonical identity, and a mismatch blocks the run and records a security event. Self-reported identity is never trusted.

The same gate, caught again by the container suite

A week later the Docker test report recorded the mirror image of my mistake — this time made by the test seed, not the test author:

hardcoding gpt-4o in the UI seeded run correctly triggered the model-swap gate once repo-agent was certified against the local model.

The seeded UI data hadn't been updated after certification bound the workload to the local identity — and the gate refused it. When your negative control fires on accident, that's the control working.

What I learned

Test the attack, not a neighbor of the attack. "Certified on A, running on B" contains a certification step. Skipping it to save time tested a different proposition entirely — and would have shipped a green checkmark on an unverified claim.

Bind identity to evidence, not declarations. The manifest says what the agent intends to be; the attestation says what it proved to be. Gates should read the second document.

A wrong security test is worse than no security test. No test invites scrutiny. A wrong-but-green test manufactures confidence — and the refutation only exists because I read the 201 instead of assuming the gate would never fail.

Read the refusal. The 201 wasn't lying. It was telling me, precisely, that in the absence of an attestation there is no such thing as a swap. The system's behavior was a better spec than my scenario.

What it doesn't prove

  • The alias table (HIVEPLANE_MODEL__MODEL_ALIASES) is per-deployment configuration; a misconfigured alias is a deployment error the gate can't catch.
  • Identity binding blocks swaps at admission — a provider serving a different model after admission is the drift detector's job, which ships next release.
  • One identity was certified this cycle; the multi-model substitution matrix (declared cloud, certified local, running local) is the cloud-profile run's work.

References

Next in the series: release-day forensics — the 35 MB file inside my 22 MB package, and the tamper-evident log that wasn't.

What's the last test of yours that was green for the wrong reason — and how long did it take you to notice?

Top comments (0)