DEV Community

Edison Flores
Edison Flores

Posted on

First third-party interop evidence: live Agent Trust Cards over RFC 9421 (conformance v1.4.0)

What happened

On 2026-10-02, the UTA conformance suite captured its first third-party interoperability evidence: three live exchanges between the MarketNow verifier and AgentBouncer (beta) — an external agent-trust gateway — over RFC 9421 signed HTTP, with the x-uta-trust header as a covered component of the request signature.

It shipped the next day as @marketnow/uta-conformance@1.4.0 (npm, 2026-10-03T02:36Z) and commit 25696d48 on main.

Three paths were exercised. The three outcomes are the whole point:

Path Card Envelope (AgentBouncer) Trust verdict (UTA) Composed result
unsigned control — verified: false, reason: "no_signature" not reached rejected at the envelope layer
allowed ATC-2026-003701 signature valid, replay checked PASS — 8/8 controls, caps match allowed
denied ATC-2026-003702 signature valid, replay checked, allowed: true POLICY_FAIL — action mcp:tools/call requires network.egress=allowlist, card grants none denied

Why the denial is the interesting row

Read the third row carefully: AgentBouncer accepted the request at its own layer, and the composed system still denied it. That is the experiment working, not failing.

The two cards were sent with the same request body — the content-digest header is byte-identical in both live requests (sha-256=:vRx4QTR9SRY9lK3tmE4cM1/3IGpPXZ6ij1TcUI5mYmY=:), same endpoint, same signature coverage. The only variable is the trust token in the covered x-uta-trust header (2,227 vs 2,220 base64url bytes). One card grants network.egress=allowlist; the other grants nothing. Everything else held constant, the verdict flipped from PASS to POLICY_FAIL.

That isolation is what makes the evidence meaningful. It shows the denial is attributable to trust policy, not transport — and that a gateway which passes the envelope cannot smuggle through a card that the trust layer rejects. Composed verdicts, clean attribution, fail-closed at the layer that owns the decision.

Covered component, briefly

RFC 9421 signs HTTP messages by listing header fields in Signature-Input. The x-uta-trust header is in that list, which means the trust token is bound into the request signature: you cannot sign a request and swap the card afterwards, because the signature check and the card bytes are verified against the same covered set. The card travels as base64url(JSON of the full ATC/1.0 document) in that header; AgentBouncer verifies the envelope (signature, timestamp, replay, project key/policy) and — per the vendor's own clarification on 2026-10-01 — does not parse the ATC or enforce its capabilities. The UTA verdict is produced independently by the MarketNow verifier.

Each layer rejects for its own reasons. The unsigned control died at the envelope (no_signature) without the trust layer ever being reached. The denied card survived the envelope and died at policy. The allowed card passed both. Three outcomes, three different failure/success layers, no ambiguity about who decided what.

What is claimed — and what is not

Claimed: two live exchanges captured on 2026-10-02 against AgentBouncer's public test endpoint, envelope-verified by the counterparty, trust-verified by UTA, published as evidence vectors in vectors/third-party/, with the full request headers, digests, verdicts, and timing in the traces file. Every byte of that is checkable at the URLs below.

Not claimed:

  • These are evidence artifacts, not scored vectors. The 14-vector scored set is unchanged, byte-for-byte. A live counterparty is neither deterministic nor self-contained, so it cannot be part of the scored set — the scored set must stay runnable offline by a stranger with a git clone and nothing else.
  • This is not an endorsement in either direction. The counterparty is a beta product at a public test endpoint. We are documenting what happened on the wire, not certifying the vendor.
  • AgentBouncer did not verify the Agent Trust Card itself. Envelope binding and trust verification are different jobs; this exchange exercised both, held by different parties.

Verify it yourself

Everything below is live:

# the interop index (schema, both evidence vectors, verdicts)
curl -s https://www.marketnow.site/uta/conformance/vectors/third-party/_interop-index.json

# the two evidence files (card ids, counterparty verdicts, UTA verdicts)
curl -s https://www.marketnow.site/uta/conformance/vectors/third-party/agentbouncer-transport-allowed.json
curl -s https://www.marketnow.site/uta/conformance/vectors/third-party/agentbouncer-transport-denied.json

# the full traces: request headers, content-digest, token sizes, unsigned control
curl -s https://www.marketnow.site/uta/conformance/vectors/third-party/traces_20261002_final.json

# the scored set is still 14/14 and unchanged
npx @marketnow/uta-conformance

# the npm release and the commit
npm view @marketnow/uta-conformance dist-tags.latest   # 1.4.0
Enter fullscreen mode Exit fullscreen mode

Or read them in the repo: uta-monorepo/packages/conformance/vectors/third-party/.

Where this fits

The suite's job has been the same since v1.2.0: separate real verifiers from fake ones. v1.3.0 made the accept side unbounded (generator + published test CA). v1.3.3 added two-sided validity windows and per-stage scoring — a runner that fires the right verdict for the wrong reason now scores wrong. v1.4.0 adds a different axis entirely: not "does your runner verify cards correctly" (still 14 vectors, still the job) but "does the protocol actually travel between a signer's stack and someone else's gateway, over a standard signed-HTTP transport, without the trust meaning getting lost in translation."

That question only has evidence-based answers. This is the first entry.

State of everything else (2026-10-03)

  • MCP Registry: still listed and active — io.github.alicelabs-llc/marketnow v1.15.0, both the streamable-http remote and the marketnow-mcp npm stdio package. marketnow-mcp: 2,131 downloads last month.
  • M8ven rating badge went into the README on 2026-10-01 — C, 74/100. Publishing the score we got, not the score we wanted.
  • alethech 0.9.1 is on PyPI (since 2026-09-30) — the verifiable agent-memory protocol, with the mutation-guard test paths from the last post.
  • Five dependabot CI-action bumps were merged to main the same day v1.4.0 shipped. CI green.

Next on the conformance roadmap: more counterparties (the evidence format is deliberately counterparty-agnostic), and the runner-under-test pinning rule — the reference scorer is now a tested component, not a trusted one: its bytes are digest-pinned and its behavior must reproduce the pinned answer key across five surfaces. More on that when the next evidence lands.

Top comments (0)