DEV Community

Cover image for The Data Was Public. The Agent Path Wasn't. So His Mock Became My Documentation.
Self-Correcting Systems
Self-Correcting Systems

Posted on AI-assisted

The Data Was Public. The Agent Path Wasn't. So His Mock Became My Documentation.

Two commands against the same data. Run them yourself:

$ curl -sS --get 'https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production' \
       --data-urlencode 'query=count(*)'

{"query":"count(*)","result":120,"syncTags":["s1:dmd/mg"],"ms":11}
Enter fullscreen mode Exit fullscreen mode

120 documents, no key, no account. Now the path my agent actually uses:

$ curl -sS -o /dev/null -w '%{http_code}\n' \
    'https://api.sanity.io/v1/context/organizations/<ORG>/mcp/self-correcting-systems'

401
Enter fullscreen mode Exit fullscreen mode

Same underlying dataset. One route is publicly queryable. The route my agent takes goes through
an authenticated Context MCP endpoint, which needs an organization API token — not a project
token, which Sanity explicitly does not accept there. I had not provisioned Pouya any access to my
organization, so cloning the repo gave him no way to authenticate the path the agent actually uses.

I am being careful about "same" here: an MCP endpoint can be configured with its own sources and a
groqFilter, so it is not guaranteed to expose the same 120 documents the public route does. Same
source, different access path, different interface contract. That distinction turns out to be the
bug.

That gap is the whole story, and I did not know it was there until the contributor who fell into
it told me.

What happened

A developer named Pouya read a post of mine, decided my diagnosis was wrong, cloned the repo and
opened a pull request. First outside contribution the project has ever had.

PR #1   https://github.com/keniel13-ui/ask-the-record/pull/1
        PouyaZX4:fix/groq-schema-routing
opened  2026-09-22T12:33:52Z
merged  2026-09-24T00:53:58Z   ->  36.3 hours
        3 commits, 1 file, +25 / -12, merge 4a2940f
Enter fullscreen mode Exit fullscreen mode

It is merged, so it will not appear in the default pull request list, which shows open ones only.
A reviewer told me the repo had no PRs at all for exactly that reason, which is a small instance of
the same mistake this whole post is about.

His diagnosis was reasonable. My agent queries a Sanity dataset with GROQ, and he thought the
failures came from the model guessing at document types and inventing query syntax. His fix: put
the real schema and worked examples into the tool description, so the model is told the shape
instead of guessing it.

Sound idea. I would have tried the same thing.

Then I ran his examples.

The two queries

Both from the pull request description. Here is the one for a claim, against the live dataset:

*[_type == "claim" && _id == "claim-ledger-population"][0].{status, expiryStatus}

-> HTTP 400  "attribute expected"
Enter fullscreen mode Exit fullscreen mode

A stray dot between [0] and {. Remove it and it works:

*[_type == "claim" && _id == "claim-ledger-population"][0]{status, expiryStatus}

-> {"status": "standing", "expiryStatus": "no_expiry_set"}
Enter fullscreen mode Exit fullscreen mode

And the one for a patch record:

*[_type == "patch" && findingRef._ref == "finding-b1"][0]

-> null
Enter fullscreen mode Exit fullscreen mode

Null for two separate reasons. There is no findingRef field on my patch documents — it is
findings, and it is an array. And the ids are case sensitive: finding-B1, not finding-b1.

Written correctly:

*[_type == "patch" && "finding-B1" in findings[]._ref][0]{sha, inMain}

-> {"sha": "dd1a654", "inMain": false}
Enter fullscreen mode Exit fullscreen mode

Six field names that do not exist

The patch did not just carry two broken examples. It wrote a schema into the tool description,
and the schema did not match my data:

the patch told the model what the database actually has
claim.subject nothing
claim.statement text
finding.finders foundBy
finding.verifiedReceipts commentId or commentIds, depending on the record
patch.findingRef findings[]
patch.commitHash sha

Check it yourself, no key required:

curl -sS --get 'https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production' \
  --data-urlencode 'query=*[_id=="finding-B1"][0]'
Enter fullscreen mode Exit fullscreen mode

That is the record his example referenced. An earlier draft of this post ran the same command
against finding-B8 instead, because B8 carries commentIds and B1 carries the singular
commentId, which made my table read cleaner. That is picking the sample that proves the point,
in a post about not doing that, and a reviewer caught it before it shipped.

So a patch written to stop a model from inventing schema would have installed a mismatched schema
into the exact place the model reads it from. Worse than the bug it fixed, because a guess made at
runtime is a guess, while a guess sitting in the tool description arrives with the authority of
documentation.

Why the names were wrong, which is not what I assumed

I wrote a draft of this post that said I did not know how the patch was produced and was not going
to guess. That was the right call, because when I asked, the answer was nothing I would have
guessed. In his words, published with his permission:

The dataset itself is public of course, but I wasn't able to replicate the full MCP setup on my
end at the time I tested it so I was hitting that 401 on the Context MCP endpoint. So I just
tested the query generation logic against an offline mock schema on LM Studio instead, which is
where those field names came from.

Nothing was hallucinated. He hit the 401 at the top of this post.

He could read every one of my 120 documents in a browser. He could not run my agent, because the
agent does not take the public route — it goes through an org-scoped MCP endpoint needing an
organization credential I had not provided and was not going to publish with the repository.

So he did the careful thing. He built an offline mock of the schema and tested the query
generation logic against that. And when he built the mock he chose better names than mine:

When creating that schema, I used clean, self-describing semantic names (commitHash instead of
sha, verifiedReceipts instead of commentIds, finders instead of foundBy) because explicit naming
makes it way easier for the SLMs to understand what fields represent and prevents confusion.

He is right about that, by the way. commitHash is a better name than sha. foundBy is worse
than finders. Every wrong name in that table is the name a competent person would pick if they
were designing the schema rather than reading it.

One of them is worse than a naming difference. verifiedReceipts does not just rename
commentIds — it asserts something the data cannot support. A comment id says a comment exists at
that address. It says nothing about whether anyone verified it. That distinction is the entire
subject of the record it was describing.

The defect is mine

The mock is not the bug. The mock was a reasonable response to a 401.

The bug is that my repository advertised a reproduction path that is not the one the system
uses.
The README says "Public dataset" and gives the URL at the top of this post. That is true,
and it is what I put in an earlier version of this very post — a curl, no key, check it
yourself. It reproduces the data. It does not reproduce the agent, which reaches the same
records through an endpoint that returns 401 to everyone who is not me.

A contributor following my README lands in a place where every record is visible and nothing is
runnable. The only way forward is to build a stand-in. And a stand-in built at that boundary does
not stay at the boundary — his mock's field names travelled from a local LM Studio test into the
tool description, which is where my harness explicitly tells the model what the schema is.

That is the whole difference in one line. A runtime guess is visibly a guess. Put the same guess in
the tool description and I have promoted it into authoritative guidance.

I have shipped the same class of error. I described my own validator to someone as checking
whether a cited record matched what was retrieved. It does not. I was describing the system I
meant to build, from memory, instead of opening the file.

What actually resolved it

Not an argument. Two commands.

I pulled his branch, ran both examples against the live dataset, and sent him exactly what came
back: the 400, the null, and the corrected queries with their real output. No opinion about his
approach, no debate about the diagnosis.

He pushed a revision in about two hours. Every field name corrected, both queries rewritten:

claim  claim-ledger-population -> standing / no_expiry_set
patch  for finding-B1         -> dd1a654 / false
finding finding-B8            -> commentIds ['3ee98'], status unbuilt
Enter fullscreen mode Exit fullscreen mode

All three still return exactly that today. Note the third is B8, not the B1 I told you to curl
above — B1 carries the singular commentId "3eanf" and status implemented. Those are two
different records and I have mixed them up once already while writing this.

Still one stale reference in a routing rule, so I sent that too, with the null it produced. He
fixed it that morning. I re-ran the three patterns, checked the file still parsed and still had no
third-party imports, and merged it.

Three commits, thirty six hours, between two people who have never met. And one detail I like:
the pull request body still shows [0].{status, expiryStatus}, the broken stray-dot form, after
the merged code had moved to the corrected one. Documentation can keep a false schema alive after
the executable stops using it, which is the same failure as this whole post, one layer up.

The thing worth taking

If you maintain anything an outsider might contribute to, run this check:

Can someone who clones your repo actually execute the path your system takes, or only the path
your README documents?

For me those were different, and the gap was invisible from the inside because I hold the token.
Everything worked on my machine for a reason I never had to think about.

When they differ, a contributor's only option is a mock. They will build a good one — Pouya's
names were better than mine. And then the mock's assumptions become your documentation, because
the tool description is documentation, and the model does not know it was written against a
fixture.

There is a third option, and it is better than either of the two I first wrote down. I said the
fix was to make the real path reachable, or to say in the README that it is not. Both are weak,
because both still leave the contributor inventing the contract.

If outsiders cannot execute a privileged dependency, ship them a reproducible contract for it.
A fixture generated from the real schema, checked into the repo, lets someone test against the same
interface without ever receiving my organization credential. The pipeline becomes

production schema -> generated contract fixture -> contributor harness
Enter fullscreen mode Exit fullscreen mode

instead of

README -> unreachable MCP -> contributor invents a substitute schema
Enter fullscreen mode Exit fullscreen mode

The problem was never that Pouya mocked the boundary. It is that my repository gave him no
canonical boundary to mock, so he had to design one — and a well-designed guess is still a guess.

That generalises past Sanity and past agents. Anything an outsider cannot run — a private API, a
payment sandbox, an internal queue, an OAuth service — has this shape. If you do not own the
stand-in, your contributors will build one, and theirs will encode their assumptions instead of
yours.


Thanks to Pouya for the patch, for taking two rounds of corrections without once arguing the
diagnosis, and for answering the question about where those names came from when he could easily
have let me publish a guess instead. The merge is 4a2940f. Every query above can be run by
anyone against the public dataset — which, as it turns out, is exactly the point.

Top comments (0)