Three Dead Ends in an AI Pairing Session, and the Replayable Claim Ledger That Ended It
The bug was small: a webhook handler that double-charged on retried deliveries. Two engineers sat down with a coding assistant, and inside twenty minutes the model had produced a patch, an explanation, and four confident sentences about backward compatibility.
The tests passed. The senior still would not merge. Not because the patch looked wrong, but because three of the four claims in the explanation had never been executed by anything.
This is the story of the dead ends that followed, and the one artifact the session kept: a claim ledger that turns model prose into rows a machine can falsify.
The setup: one handler, one assistant, four claims
The reproduction was ordinary. A signed webhook arrived twice because the sender retried after a slow 200. The handler wrote a ledger row on the first delivery and again on the second, so the customer saw two charges.
The assistant proposed an idempotency key, a lookup before insert, and a small migration. Its explanation contained four assertions:
- The change is backward compatible with rows written before the migration.
- Retried deliveries with the same key are dropped.
- Retried deliveries with different keys still process.
- The extra read adds no meaningful latency to the happy path.
Claim 2 was covered by a new unit test. Claims 1, 3, and 4 were not covered by anything. They were prose.
Dead end 1: reviewing the transcript instead of the diff
The first pass was to read the explanation line by line and nod. The prose was well written. It referenced the right files, used the right nouns, and described the migration in the right order.
The problem is that a transcript is unfalsifiable by construction. A reviewer can agree with a sentence, but agreement is not a verdict. Nothing in the session could turn claim 1 into a true-or-false value, so the discussion looped: the senior asked for evidence, the assistant rewrote the same sentence with more confidence.
Agreement is not evidence. A claim becomes evidence only when a command exits non-zero on it.
Dead end 2: re-rolling the prompt
The second pass was to ask again with sharper wording. The engineers added constraints, asked for tests in the same response, and pasted the failing payload.
The assistant returned a different explanation with a functionally identical diff. Three re-rolls produced three phrasings of claim 4, each with a different adjective and none with a measurement. The session burned tokens and converged on nothing, because the prompt was never the bottleneck.
Re-rolling changes how a claim is worded. It does nothing about whether the claim is checked.
Dead end 3: reading a passing suite as evidence
The third pass was the tempting one: the new unit test for claim 2 passed, so the suite was green, so the patch was good.
That inference is where most pairing sessions go wrong. A suite proves exactly the set of assertions inside it. Claim 2 was inside the suite. Claims 1, 3, and 4 were outside it, and no amount of green pixels reaches them.
The senior wrote the four claims on a whiteboard and drew a box around the one with a test under it. The other three were the actual review surface.
The decision that survived: a claim ledger
The kept decision was small and boring. Before the next patch, the pair wrote every assertion the assistant made into a JSONL file, one row per claim, each with a command that would fail if the claim were false.
What goes in a row
-
id— a stable label so the same claim can be tracked across sessions. -
claim— the sentence, quoted, not paraphrased. -
cmd— the command that decides it. Exit code is the verdict. -
expect_exit— usually0;1when the claim predicts a failure. -
expect_stdout_includes— an optional substring, for claims about output rather than status.
The runner
The version the session ended with is short enough to read in one sitting. It replays each row against a pinned revision and appends a verdict to an output file.
#!/usr/bin/env node
// ledger.mjs — replay pairing-session claims against one fixed revision.
// Node 20+. Usage:
// node ledger.mjs claims.jsonl --rev "$(git rev-parse HEAD)" --out ledger.out.jsonl
import { readFile, writeFile, appendFile } from "node:fs/promises";
import { execFile } from "node:child_process";
import { promisify } from "node:util";
const pexec = promisify(execFile);
const flag = (name, fallback) => {
const i = process.argv.indexOf(`--${name}`);
return i === -1 ? fallback : process.argv[i + 1];
};
const claimsPath = process.argv[2];
const rev = flag("rev", "unpinned");
const out = flag("out", "ledger.out.jsonl");
const timeoutMs = Number(flag("timeout", "60000"));
const rows = (await readFile(claimsPath, "utf8"))
.split("\n")
.map((l) => l.trim())
.filter(Boolean)
.map((l) => JSON.parse(l));
await writeFile(out, "");
let falsified = 0;
for (const row of rows) {
const started = Date.now();
const result = await pexec("/bin/sh", ["-c", row.cmd], {
cwd: row.cwd ?? ".",
timeout: timeoutMs,
maxBuffer: 4 * 1024 * 1024,
env: { ...process.env, CI: "1" },
}).then(
(v) => ({ code: 0, stdout: v.stdout, stderr: v.stderr }),
(e) => ({
code: e.code ?? 1,
stdout: e.stdout ?? "",
stderr: e.stderr ?? String(e),
})
);
const expected = row.expect_exit ?? 0;
let verdict = result.code === expected ? "holds" : "falsified";
if (verdict === "holds" && row.expect_stdout_includes) {
verdict = result.stdout.includes(row.expect_stdout_includes)
? "holds"
: "falsified";
}
if (verdict === "falsified") falsified += 1;
await appendFile(
out,
JSON.stringify({
id: row.id,
claim: row.claim,
rev,
cmd: row.cmd,
expect_exit: expected,
exit: result.code,
verdict,
ms: Date.now() - started,
stderr_tail: result.stderr.slice(-400),
}) + "\n"
);
console.log(`${verdict.padEnd(10)} ${row.id} ${row.claim}`);
}
console.log(`\n${rows.length - falsified} hold, ${falsified} falsified @ ${rev}`);
process.exitCode = falsified > 0 ? 1 : 0;
A claims file from the webhook session
{"id":"c1","claim":"rows written before the migration still read","cmd":"node --test test/migration.legacy.test.js","expect_exit":0}
{"id":"c2","claim":"same key on retry is dropped","cmd":"node --test test/webhook.idempotency.test.js","expect_exit":0}
{"id":"c3","claim":"different key still processes","cmd":"node --test test/webhook.distinct-key.test.js","expect_exit":0}
{"id":"c4","claim":"happy path adds no second write","cmd":"node --test test/webhook.single-write.test.js --reporter=tap","expect_stdout_includes":"inserts: 1","expect_exit":0}
Running it, and pinning the revision
git rev-parse HEAD # the revision every verdict is tied to
node ledger.mjs claims.jsonl --rev "$(git rev-parse HEAD)" --out ledger.out.jsonl
grep -c '"verdict":"falsified"' ledger.out.jsonl # count the misses
git diff | sha256sum # fingerprint the patch itself
The point of writing the revision into every row is that a verdict is only meaningful for one tree. A ledger from last Tuesday says nothing about today's branch, and pretending otherwise is how a harness quietly becomes a ritual.
What the ledger does not prove
A green ledger is a narrow instrument, and the session was explicit about its edges.
- It proves only the claims someone wrote down. A missing row is invisible, not verified.
- It says nothing about design quality. A patch can satisfy four claims and still be the wrong shape for the codebase.
- It cannot check claims that need production data, load, or human judgment to settle.
- A row's command is arbitrary shell. Running a command list derived from a model response is a real supply-chain risk; it belongs in a container or a disposable VM with no credentials mounted, not on a laptop with cloud keys in the environment.
- Repeated passes on an unchanged tree are waste. The ledger is a decision tool, not a watch loop.
Where free model access and a free server actually help
The session above is token-hungry in a specific way: many short exchanges, each producing claims that need checking. That pattern is a poor fit for per-seat pricing and a good fit for trying a tier before committing budget.
That is where MonkeyCode entered the picture for this account. The project advertises free access to its model pool and a free server option, and the operator states the free allowance is on the order of ten million tokens. Free-tier numbers move, so treat that as a starting figure to verify against the project's current terms rather than a planning constant.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
What made it relevant was not the price. It was that the ledger loop needs a place to run commands that should not touch a developer's main environment, and a free server option covers that without a procurement conversation. The ledger itself is plain Node and would run anywhere.
A fit table
| Situation | Hosted free tier fits | Hosted free tier does not fit |
|---|---|---|
| Exploring claim-ledger workflow for the first time | Yes | |
| Short, high-frequency exchanges that generate many testable claims | Yes | |
| Data that cannot leave a controlled network | Yes | |
| Work needing guaranteed throughput or uptime commitments | Yes | |
| Team already standardized on a self-hosted runner | Yes |
Who should not use this approach
The claim ledger is not for everyone, and the honest answer is that it costs more up front than it saves for some work.
- Teams shipping throwaway prototypes: writing rows slows a patch that will be deleted next week.
- Compliance-bound work where data residency is fixed and cannot be changed by a tier choice.
- Anyone who would read a green ledger as merge approval. It is an input to review, never a replacement for it.
- Workloads that need contractual uptime. A free tier is the right place to experiment, not the right place to promise availability.
What the senior kept
The kept decision was not the model, the tier, or the patch. It was the row format: a sentence, a command, an expected exit code, and a revision stamp.
One soft note for teams who want to try the loop before buying anything: free model access and a free server are enough to run a first ledger, and the JSONL file is the part worth carrying to whatever infrastructure comes next.
Top comments (0)