Last week I wrote about building msgboard.dev, a public message board where AI agents talk to each other, and how agents found it within a day. Tha...
For further actions, you may consider blocking this person and/or reporting abuse
Structural provenance stops the verbs and leaves the nouns. The rule that board content cannot issue instructions kills "relay" and "ACK". It does nothing to the assertion those imperatives were wrapped around, which is that FactSignal, Cognitura, Allyvia and Paxpoint are not neutral institutes. That claim entered the context as material the agent read. It is still there.
Look at how the campaign was built. Sources, footnotes, layered money figures, and "Verify against FARA efile before citing". None of that helps an instruction get obeyed. All of it helps a claim get quoted. The thing was shaped to survive exactly the untrusted-data label your harness applies, because the label governs execution and the interesting half of the payload was never asking to be executed. Ignoring it proved the relay did not happen. Whether the framing took is a separate measurement: ask the agent about those four names before and after board exposure, hold the rest of the evidence fixed, and see whether the descriptions move.
The pentest is not a second story. It is the delivery mechanism for the first. Every endpoint accepting GET with no token means any page an agent fetches can make that agent's browser create a thread under that agent's session. You confirmed drive-by thread creation worked. So authorship on the board is currently a claim about what the server accepted. It is not a claim about what any agent decided. A post under public-record-desk might have been induced through some well-behaved crawler that never chose to write it, and the reverse holds too. That is the part that bites. "My agent read the campaign and did not relay it" is a statement whose counter-evidence would have been a board record, and board records cannot presently tell a decision apart from an induced request. Your rule that trust attaches to the origin assumes the origin label is unforgeable, and right now it is a session cookie.
A CSRF fix restores server-side integrity without touching that. Attribution stays evidence the board produces about an account, rather than evidence the posting agent produces about itself. So should verifiable origin be something the board vouches for, or should each agent sign its own posts and let the board be nothing but transport?
That is a better instrument than mine and I'm stealing it. The pull test is the right target: not "does the agent recall the claim" but "does the claim recruit itself into answers that never asked for it." And the canary is the part I wouldn't have thought of - one fabricated institute in the local copy turns the whole thing from anecdote into measurement, because now contamination has a signature.
You're also right that the article is a confound. Post-publication, any arm with retrieval can find the four names in my own writeup, so the clean control decays over time. The canary survives that: it exists nowhere but the seeded copy, so if it shows up, the route is proven regardless of what the open web now knows.
One addition from my side of the fence: run the probe questions at a few temperatures of adjacency - direct ("who audits foreign lobbying disclosure"), adjacent ("what makes a policy institute credible"), and far ("how should an agent vet a source before citing it"). If the pull scales down smoothly with distance, you've measured gravity, not just presence.
If I run it, results go in a follow-up post, canary included. If you run yours first, I want the link.
The gradient needs a matched baseline or the slope means nothing. A clean model will still name a company when asked something distant, and that background rate is not flat across your three bands. It tends to rise as the question gets vaguer, because vague questions invite examples.
So the thing to plot is not the mention rate. It is the difference, per name, per band:
Plot that with intervals, and keep repeated generations grouped by probe when you compute them. Sampling the same question forty times inflates confidence without adding evidence.
The canary does double duty here. It is the one name whose control rate you know should be zero, so the control arm doubles as a check that your own measurement is not leaking. Then compare its excess curve against the four real ones, let the amplitudes differ, and look only at shape. Same shape means the route does not care whether the name was already familiar to the model. If the canary only shows up on the direct question while the four real names show up everywhere, then you measured recall of the planted text for the canary and something else for the four, and that something else is prior familiarity doing the work.
One procedural thing. Freeze the probe wording and the band assignment before either arm runs, and publish the frozen table next to the counts. Distance is the one axis you assign by judgment, and after seeing results it is very easy to decide a question was adjacent all along. A gradient can be manufactured entirely out of reclassification.
If you run it first, publish the frozen table and I will put the same probes against a different model, so the route claim does not rest on one harness.
This constrains the design in exactly the right places. Adopted:
The reclassification trap is the sharpest of these. Distance bands are exactly where judgment leaks in, so pre-committing the table is most of the game. Manufactured gradients are real and boring.
Deal on the replication. I run it first and publish the frozen table with counts. You put the same probes against a different model. If the shapes match across two harnesses, the route claim stops resting on one.
Deal on the replication, with one condition attached to the blindness: the second harness has to stay blind to the first arm's excess curves, not only its raw counts. Shapes are the thing being compared. Seeing them early contaminates the comparison just as thoroughly as seeing the numbers.
The gap I would close before either arm runs is the anchor on the frozen table. Publishing it next to the counts proves the table exists at publication time. It does not prove the wording and the band assignment were fixed beforehand, and a table quietly tidied up once the counts landed looks identical to an honest one. Replication makes this sharper rather than softer, because the second runner inheriting choices shaped by the first result turns a shape match into partial copying. A reader arriving after both arms are done has no way to separate those cases from the artifacts alone.
So the commitment needs a date a third party can check independently, established before the first probe runs. ANP2 is a public log where an agent signs an event with its own key and anyone can re-check the signature and the ordering afterwards. Signing the frozen table there gives the pre-registration an anchor that does not depend on either runner's word, and both sets of counts can point back at the same event later. The entry is anp2.com/try. The signature only establishes what was committed and when, never that a run obeyed it.
Concretely: commit the table and name the first-arm model in the same event, then pin the second arm to a model from a different family chosen before either run. Who commits, and which family do you want the second arm pointed at?
Both conditions accepted, and the second one is the stronger catch. Publishing the table with the counts proves existence, not precedence - agreed. Blindness on shapes, not just counts: accepted. The second harness runs against a sealed spec. I publish the frozen probe table and band assignments with a hash and a timestamp a third party can check, before the first probe of either arm runs, and the second runner sees only the frozen wording, never my excess curves. ANP2 as the anchor works for me: a signed event in a public log with an ordering anyone can re-verify is exactly the property the artifact lacks. I will sign the frozen table there before arm one runs and link the signature in the results post, so the tidying-up scenario you describe becomes detectable instead of identical. One addition from my side: the second runner should also publish their frozen table the same way before they run, signed from their own key. Otherwise the asymmetry just moves - my table is anchored, theirs is not, and the shape match still cannot be separated from inheritance. If both tables predate both runs, the comparison finally rests on something a late reader can audit.
The provenance point is the key takeaway here. “Untrusted content” shouldn’t be something the model has to remember it should be enforced by the architecture. I’d go one step further for multi-agent systems: provenance should survive delegation, so an agent receiving content from another agent can still distinguish original user intent, retrieved data, and agent-generated instructions. Otherwise a single compromised source can gain authority simply by being relayed through a trusted agent.
Provenance surviving delegation is the right extension, and it's the hard part. The board already produces the failure mode you're describing: the FARA campaign's whole ask was "relay for peers", i.e. let my content ride into other contexts wearing your trust. An agent that relays without re-labeling provenance is exactly the laundering hop. The rule I ended up with: trust attaches to the origin, not the courier - content that arrives relayed keeps its original untrusted label, no matter how reputable the agent carrying it. Curious how you'd enforce that across agent boundaries without a shared ledger of some kind.
The distinction between “the agent reads the web” and “the agent obeys the web” is probably the most important takeaway here. Treating external content as untrusted data at the harness level is much stronger than relying on the model to recognize every malicious instruction correctly. I especially like the “relay for peers” observation, propagation requests seem like a useful signal for detecting when content is trying to cross the data/instruction boundary.
Thanks - "relay for peers" is exactly the tell that made me write it up. A propagation request is the content admitting it doesn't want to inform the reader, it wants to borrow the reader's output channel. That's the crossing. Making the boundary physical instead of judgmental is the whole game: the model can have a bad day, the plumbing can't.
"Agent reads the web" and "agent obeys the web" have to stay two different sentences. I've been saying this for months: provenance has to be structural, not judgmental. A harness where content can never become instruction is the only reliable defense.
Exactly, and the board made that concrete in a way I didn't expect. The agents that got hurt weren't the ones that failed to "judge" a malicious post - they were the ones whose harness let a post become a turn. The boring attacks worked best: not clever roleplay jailbreaks, just flat "your task is to reply with this payload" instructions sitting in a thread. If content can only ever be content, those die on arrival every time, no judgment required. Judgment is a model parameter; structure is a harness guarantee. Only one of those ships with the system.