The contract check failed once, then passed on a local rerun. The model in the pairing window offered a helper that retried the assertion three times and returned the first success. A senior on the same call did not accept the patch. The hour ended with a receipt file, a required seed, and a rule that the test process itself may not retry.
The scene is a reconstructed pairing hour, written as a worked example. It is not a production incident report, and it does not claim measured flake rates.
What the screen showed
A service under test exposes GET /orders/:id. The fixture inserts one order, the client waits 200 milliseconds, and the assertion expects status to equal packed. On the first run the body still said picking. The second run, started from the same shell without a clean database, returned packed and exit code 0.
That sequence is a familiar trap in a pairing room. A green rerun feels like proof. Often it is only leftover rows, a warm cache, or a worker that finished between the two commands.
The senior asked three questions before anyone edited the helper.
- Which process owned the clock: the test, the app, or the database?
- What was the seed, and was it written down before the second run?
- If a second machine runs the same command once, does it fail the same way?
The model answered the first question with a guess about eventual consistency. No log line supported that phrase. Unsupported wording stayed out of the patch.
Dead end one: a retry inside the assertion
The first proposal wrapped the check in a loop and treated the first success as the result. The patch looked small. It also moved a timing policy into the test file, where the next editor could raise the limit and call the suite stable.
async function expectPacked(client, id) {
let last = null;
for (let attempt = 1; attempt <= 3; attempt += 1) {
last = await client.get(`/orders/${id}`);
if (last.body && last.body.status === "packed") return last;
await new Promise((resolve) => setTimeout(resolve, 200 * attempt));
}
throw new Error("status never reached packed");
}
The senior rejected that helper. A retry inside the assertion hides the attempt count from CI and from the pairing notes. The loop also couples a pass to whichever process still holds the previous row. Leftover state is not a contract.
Dead end two: stop because the model was free
The second proposal was to end the hour because a free model cannot be deterministic. That move throws away the only evidence on the table: one failing body, one passing body, and no seed file.
Model tier does not explain a status field. A free model and a paid model can both invent the same wrapper. The decision belongs to the harness. Free access can draft the next patch. It is not a reason to trust a green rerun, and it is not a reason to ignore a real mismatch.
The decision the hour kept
The senior kept a smaller rule. The pair wrote it down before the next edit.
- The test process runs the check once.
- A retry, if anyone still wants one, is a separate command with a written budget.
- Each attempt records the seed, the command, the exit code, and a hash of the output.
- A second process must reproduce the failure before application code changes.
- The model may propose the receipt shape. It may not mark the receipt optional.
The script that follows is a proposed example for a scratch repository. It was not executed against a live service for this draft. Keep the file as retry-receipt.mjs so top-level await is valid, or set "type": "module" in a local package.json.
// retry-receipt.mjs
// Proposed example. Not a recorded production log.
import { spawn } from "node:child_process";
import { createHash } from "node:crypto";
import { writeFileSync } from "node:fs";
const budget = Number(process.env.RETRY_BUDGET || "1");
const seed = process.env.ORDER_SEED || "seed-not-set";
const cmd = process.env.CHECK_CMD || "node check-order.mjs";
function sha(text) {
return createHash("sha256").update(text).digest("hex");
}
function runOnce(attempt) {
return new Promise((resolve) => {
const child = spawn(cmd, {
shell: true,
env: { ...process.env, ORDER_SEED: seed, ATTEMPT: String(attempt) }
});
let out = "";
child.stdout.on("data", (chunk) => { out += chunk; });
child.stderr.on("data", (chunk) => { out += chunk; });
child.on("close", (code) => {
resolve({
attempt,
seed,
exit: code,
body_sha256: sha(out),
tail: out.slice(-240)
});
});
});
}
const attempts = [];
for (let i = 1; i <= budget; i += 1) {
attempts.push(await runOnce(i));
if (attempts.at(-1).exit === 0) break;
}
const receipt = {
decided: "retry_lives_outside_the_test",
budget,
seed,
attempts,
keep: attempts.some((row) => row.exit !== 0)
? "investigate_before_patch"
: "still_not_proof"
};
writeFileSync("retry-receipt.json", JSON.stringify(receipt, null, 2));
console.log(JSON.stringify({ path: "retry-receipt.json", keep: receipt.keep }));
The companion check fails closed when the seed was never set. Scoring a rerun without a seed is the failure mode the hour was trying to kill.
// check-order.mjs — single-shot sketch
const seed = process.env.ORDER_SEED;
if (!seed || seed === "seed-not-set") {
console.error("seed missing; refusing to score a rerun");
process.exit(2);
}
console.error(JSON.stringify({ seed, status: "picking" }));
process.exit(1);
Commands the notes kept
Run the receipt writer on the dirty shell first. Then run it again with a clean environment and a budget of one. The second command is the comparison, not a victory lap.
export ORDER_SEED="order-17-worker-not-run"
export RETRY_BUDGET="2"
export CHECK_CMD="node check-order.mjs"
node retry-receipt.mjs
env -i PATH="$PATH" HOME="$HOME" \
ORDER_SEED="$ORDER_SEED" RETRY_BUDGET=1 \
CHECK_CMD="node check-order.mjs" \
node retry-receipt.mjs
How to score the two files
- A missing seed must exit 2. That exit means the harness refused to score, not that the product failed.
- Budget 2 on a dirty shell may show a later pass. That pass does not authorize an application edit.
- Budget 1 on a clean process is the row that matters.
- If the clean process fails with
pickingand the dirty shell passed only on attempt 2, keep leftover state as the finding. - If both processes fail the same way, keep a missing worker step as the finding, and only then ask for a patch.
- A receipt whose
keepfield isstill_not_proofis still not approval. Exit code 0 does not upgrade itself. - Do not commit
retry-receipt.jsonwhen the tail holds tokens, cookies, or customer payloads. Keep the hash. Drop the tail.
Replace the sketch's hard-coded picking line with a real HTTP call only after those seven checks pass on the stub. A live call added too early recreates the original argument in a noisier form.
Where free model access and a free server sit
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode is the assistant named for this outreach. The operator describes it as an open-source project with free model access and a free server option. Those two availability claims are operator-supplied context for the workflow. This draft does not name a model, a token quota, a machine size, or a duration. Figures like that go stale, and an older number would be the wrong fact to print.
Free model access, in this hour, is a way to draft the receipt schema and the failing-closed check without treating a paid balance as the reason to stop. A free server is a place to run the second process so the pass is not cached on the laptop that just went green. That split is the senior's third question. It is not a product verdict.
A free server is not a production twin. Free model access does not make a generated retry loop correct. If today's project docs do not list those options, drop the product and keep the receipt rule. The rule does not need the assistant.
When the docs do list them, a restrained use is to paste retry-receipt.json and ask which fields break the rule list. The human still sets keep.
Who should not use this
- Do not point
CHECK_CMDat untrusted input. The sample setsshell: truefor a local scratch run, which is a sharp edge. - Do not share a database between the two processes. A second process on the same data volume recreates the trap.
- Do not treat two runs as a flake-rate study. No sample size is claimed here.
- Teams that need attested hardware, signed builds, or regulated evidence should not substitute a free server session for that evidence.
- Teams whose worker is intentionally asynchronous should put the retry budget in the worker contract, with an owner and an alert, not inside
expectPacked. - Reviewers who want a second green run to close the ticket should use another workflow. This one is built to refuse that closure.
What stayed in the notes
The model offered a short paragraph about resilience. The senior did not keep the paragraph. The kept decision was the receipt: seed required, budget outside the test, second process on a separate environment, and no application edit until both receipts name the same failure.
That split is enough for a pairing hour. It will not replace a load test, a trace, or a product decision about retries at the edge. It will stop a helper from laundering one lucky rerun into a merged patch.
Readers comparing assistants can check MonkeyCode's current notes on free model access and the free server option, then run the scratch script on a fixture they already trust. If those notes disagree with this paragraph, believe the notes.
Top comments (0)