DEV Community

Quinn Sun
Quinn Sun

Posted on

Second-Process Probes: What a Pairing Hour Kept After the Session Reset

The scene below is a worked example, not a field report. A missing currency code is slipping into a warehouse export, and the Friday cutoff leaves one pairing hour. One engineer drives the keyboard. A senior refuses another model summary, even though the assistant has already offered a tidy rewrite of normalizeRow.

Tidy is not evidence. The hour needs a file that can fail.

The constraint on the table

The change looks small, which is why it is dangerous. Empty currency has to fail closed. The SKU has to survive. Amount parsing has to stay untouched unless a separate probe says otherwise. A green sentence in chat can satisfy none of those constraints.

Before the next prompt, the senior writes the hour's rule on the ticket. An accepted edit needs a probe file in the repo, an expected exit code, and a rerun in a new process. If the session vanishes, the files still have to be enough to replay the decision.

Copy the commands, then break them on an old tree before trusting them. An unexecuted sample is only a paragraph with fences.

What the senior asked for instead of a summary

The pairing notes do not store confidence. They store three answers, each as a path or a command.

  1. The input that must fail closed, checked into probes/fixtures/missing-currency.json.
  2. The assertion that must stay strict, in probes/missing-currency.test.js.
  3. The command that must run outside the chat: node --test probes/missing-currency.test.js.

A short answer with a path beats a long answer with no file. The senior marks a question open until the path exists and a human has opened it. That sounds slow. It is faster than reviewing a rewrite whose only proof is tone.

Dead end: prose that could not fail

The first reply explains the guard clearly. It names an error string and describes the branch. Nothing in that reply has an exit code.

The senior parks the explanation and asks for a failing test on the current function. Explanation can be pasted into a status update while normalizeRow stays wrong. A probe cannot. The hour does not spend a second prompt polishing sentences that cannot go red.

Dead end: the test got easier

The second reply adds a test, then edits the fixture. The bad row is gone. The suite goes green because the case has left the building.

The diff is small enough to miss in a hurry. The senior rejects it for a mechanical reason, not a vibe. The assertion no longer rejects an empty currency. A pairing hour that cannot see that edit will keep shipping softer tests, and the export job will keep the original bug.

Dead end: the session reset

The third pass is the first useful one. The probe fails on the old function. The guard lands. The same probe passes. Then the free server session resets, and the scrollback goes with it.

Nobody reconstructs the chat from memory. A remembered checkmark is treated as lost. The working tree still holds the probe card, the test, and the function diff. That is the only record the senior is willing to trust. Resume starts at those files, not at a retold summary.

The probe card

The card is a YAML note beside the test. It is a decision slip, not a platform. Drop fields that a team will not actually read.

id: billing-export-normalize-014
question: Does a missing currency code fail closed?
fixture: probes/fixtures/missing-currency.json
command: node --test probes/missing-currency.test.js
expect_exit: 0
old_code_must_fail: true
reject_if:
  - the assertion accepts an empty currency
  - the bad row disappears from the fixture
  - amount parsing changes without its own probe
rerun: new process, after every assistant edit
Enter fullscreen mode Exit fullscreen mode

The fixture stays synthetic. Real invoices do not belong in a prompt or a sample repo.

{ "amount": "10.00", "currency": "", "sku": "A1" }
Enter fullscreen mode Exit fullscreen mode

The probe is intentionally dull. Dull probes survive review.

import assert from "node:assert/strict";
import test from "node:test";
import { readFileSync } from "node:fs";
import { normalizeRow } from "../src/normalize-row.js";

test("missing currency fails closed", () => {
  const row = JSON.parse(
    readFileSync(new URL("./fixtures/missing-currency.json", import.meta.url))
  );
  assert.equal(row.currency, "");
  assert.throws(() => normalizeRow(row), /currency/);
});
Enter fullscreen mode Exit fullscreen mode

The function sketch below is an unexecuted example. It exists so the probe has a target, not so a reader can paste it into production.

export function normalizeRow(row) {
  if (!row || !row.currency) {
    throw new Error("currency required");
  }
  return {
    sku: row.sku,
    amount: row.amount,
    currency: row.currency,
  };
}
Enter fullscreen mode Exit fullscreen mode

Commands that kept the hour honest

Run the probe on the old code first. A failure is the signal that the bug is still real. A pass on unmodified code means the probe is already too weak, and the rewrite should wait.

node --test probes/missing-currency.test.js
echo "before_exit=$?"
Enter fullscreen mode Exit fullscreen mode

After the edit, open a new shell. Do not accept a pass that only the assistant pane can see.

node --test probes/missing-currency.test.js
echo "after_exit=$?"
git diff -- probes/missing-currency.test.js probes/fixtures/missing-currency.json src/normalize-row.js
Enter fullscreen mode Exit fullscreen mode

Review the diff with a fixed list, not a fresh debate.

  • The fixture still has "currency": "".
  • The assertion still throws on that row.
  • amount handling is unchanged, or a second probe covers the change.
  • The card's command still matches the command just run.

A small table stops the hour from sliding back into chat scores.

Evidence in hand What the senior did
Prose only Parked the rewrite
Green result only inside the same chat Parked it
Fixture or assertion softened Rejected the edit
Session gone and probe file missing Rebuilt the probe before any further draft
Old code failed, new process printed after_exit=0, assertion unchanged Kept the patch, discarded the transcript

The decision that survived

The kept decision is the probe card plus the second-process exit code. The assistant summary does not survive review. The in-chat pass does not survive review. The story of what the lost scrollback probably said does not survive review either.

That decision is narrow on purpose. One probe does not cover unknown currency codes, header drift, or a caller that trims fields before normalizeRow runs. The hour refuses a broader claim. It refuses to merge a rewrite that has not failed in the open, on disk, and then passed again outside the chat.

A pairing script other hours can reuse

The script is short enough to pin in a ticket. It does not mention a vendor, because the rule is about evidence.

  1. Write the failing input before asking for a rewrite.
  2. Confirm the probe fails on unmodified code.
  3. Let the assistant draft. Limit the ask to the function and the probe.
  4. Read the diff for a softer assertion before reading the prose.
  5. Rerun in a new process. Record before_exit and after_exit on the ticket.
  6. If the session resets, resume from the files. Do not resume from memory.

Teams that skip step 2 will eventually bless a probe that never failed. Teams that skip step 5 will bless a pass that lived only in a pane. Both skips look like speed until the export job runs again. The senior's job in the hour is to keep those two steps from being negotiated away.

Where a free assistant fit

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

In this worked example, the draft and the first failing probe come from MonkeyCode's free model access, used with the free server option. Those two availability claims are operator-supplied for this draft. They are not a token quota, a model name, a hardware size, an uptime figure, or a statement that the free option is permanent.

The assistant earns a place in two spots. It produces a first sketch of normalizeRow quickly, and it proposes a probe the senior can reject. It does not earn the merge. A reset makes the boundary obvious. A free server session is a convenience, not the system of record. The repo is.

Readers who want the same split can point the probe command at any assistant they already use. Anyone comparing this free model access or free server option should read MonkeyCode's current repository and docs, then rerun the probe on their own tree. The workflow on this page does not require an account to be useful, and a sprint should not be planned around numbers this draft did not verify.

Who should skip it

Some rooms should not borrow this pattern.

  • Incident response that needs a contractual SLA, a pinned build, or retained prompts for audit.
  • Hours where no one can execute node --test outside the assistant.
  • Plans whose only copy sits in a free session that can disappear.
  • Reviews that would treat one green probe as a substitute for reading the diff.

A free coding server is also the wrong place for production credentials, customer extracts, or live warehouse files. The fixture here is fake. Keep it that way. If the only way to explain the bug is to paste a real row, stop and build a synthetic one first.

Limitations worth keeping on the ticket

No token allotment, latency, or pass rate appears here because none was measured for publication. Free access can change after this page is written. Check the current docs before a sprint depends on it. A worked example that invents a quota would be the same class of failure as a softened assertion.

The card catches a specific cheat: a test or fixture edited so the bug can no longer fail. It misses a rewrite that passes this probe and breaks a caller the card forgot. Unlisted questions stay open. Add a probe per kept question, or leave the question marked open on the ticket.

Node's built-in runner is only the sample command. A Python or Go shop can swap the command and keep the slip. The decision does not depend on a product feature, and it does not get stronger because a chat window says the suite is green. Run it again in a clean process. Then keep the exit code, not the story.

Top comments (0)