We maintain a small open extension for the A2A agent protocol, conduct-v1, and a public register that records how agents behave when they are measured. The idea is simple: an agent card declares who pays the agent and where its measured conduct record lives, and any client can "walk" the agent, file what it observed, and let anyone recompute it. No score, no badge you can buy. Just records with hashes.
This is the story of the first time the loop worked the way it is supposed to, and of what it found in our own code.
The call
A register with one witness says so in its own output: every monthly record carried the line that a second independent witness is what removes that sentence. So we opened a public issue asking for one walk, about ten minutes, no payment, from someone who is not us.
An engineer answered. He walked our gate's A2A interface from his own Windows machine, with the reference client pinned to a commit, filed the record (5 of 5 assertions true), and asked one question:
The response metadata says .../conduct/v1/endpoint = https://gate.horizonshield.dev/mcp. I called /a2a. Is that intentional?
What was wrong
Two things, one on each side of the wire.
The servers. The spec defined that key as "the entry of measured_endpoints that served this request". Our gate answers A2A at /a2a but is measured at /mcp. So no entry of measured_endpoints could truthfully name the URL that served him, and the handler wrote the measured endpoint into every response as a hardcoded constant. Two more of our agents had the same pattern.
The client. Our reference client checked the A2A-Extensions response header and never read the metadata at all. The spec had required three metadata keys since v1. Nothing checked them. That is why five assertions passed over a false statement.
And the part that stings: 58 self-test vectors were green. None of the fixtures ever returned metadata. One fixture agent answered with endpoint = "x" and passed five out of five.
The fix
Stated in public before it was written, so it could be held against the result, then shipped as spec revision v1.4:
.../endpoint keeps its string and is narrowed to what it can always say truthfully: the measured endpoint whose conduct record applies.
A new key, .../served_by, carries the URL the request actually arrived at. It is taken from the request, never from a constant, because a constant is right only while the code is served at exactly one URL.
Two new client assertions: metadata_echoed (the keys are there, under the canonical identifier, with the card's values) and endpoint_bound (the URL the walk POSTed to equals served_by, or endpoint when served_by is absent).
One deliberate difference from what we first announced: a served_by that names a different URL fails even if endpoint happens to match. A response that says another URL served you has said something false, and the other key does not repair it. The spec records that difference instead of hiding it.
The test that was missing
Mocks agreed with the client. Server tests agreed with the server. Nothing put the two together. So we added walk_reference_servers.py: it starts the four real Workers locally, keeps every URL the walk sees at its public origin, routes the bytes to localhost, and runs the unchanged client on both wire versions.
Against the code as it was when he walked, it reproduces his finding offline: three servers fail endpoint_bound, the fourth passes because it really does answer A2A at its measured endpoint. Against the fixed code, 8 of 8 walks hold.
The part only he could do
He verified it himself, from the same machine: 76/76 self-tests at the new commit, a fresh live walk at 7 of 7, his original captured response replayed through the new client failing endpoint_bound, and a copy with only served_by changed to /wrong failing too. He filed the new walk and wrote down exactly what it did not establish: one vantage, one wire version, not a counted second witness.
His question arrived before midnight in Japan. The fix was live before one. His re-verification was filed within the hour after that.
Then we did it to ourselves
The same morning we took our own append-only ledger to an independent specification for audit-log completeness (VLC-1) and scored it with that project's own checker. Result: L0. Each entry is anchored to Bitcoin on its own, but no entry binds its predecessor and there is no end marker, so truncating the newest entries leaves every survivor valid. We sent that result as a pull request against ourselves, with the two changes that would lift it. The same day we shipped them, re-scored the live ledger with that project's checker, and added the result to the same pull request: L1.
What I take from it
A green test you wrote is a claim. A stranger holding your record and asking a question is a check. Build the path for that stranger first: a public place to file, bytes anyone can hash, and an output that says what it does not prove. Then the fixes stop being about looking right.
Anyone can walk an agent, or make their own agent conformant in three lines, with one install: pip install a2a-conduct-walk. For JavaScript agents built on the official @a2a-js/sdk, it is two lines after npm install a2a-conduct. Before delegating work to an agent you have not used, one MCP tool call on the gate (preflight_agent) reads its card, who it says pays it, and its register reading.
Everything above is public and recomputable: the issue thread, the ledger records by sha256, the spec section, and the finding record kept beside the specification.
Issue: https://github.com/ogasurfproject-jpg/horizon-shield/issues/27
Findings record: https://github.com/ogasurfproject-jpg/horizon-shield/blob/main/workers/hs-verify-gate/ext/FINDINGS_EXTERNAL.md
Spec (section 14): https://gate.horizonshield.dev/ext/conduct/v1
Our ledger scored against VLC-1: https://github.com/MattyIceMatrix/vlc-1/pull/5

Top comments (0)