DEV Community

Quinn Wang
Quinn Wang

Posted on

Accept Nothing You Cannot Replay

The first fifteen minutes of an AI coding session fail when you accept a diff you cannot replay the next morning. I keep watching the same quiet collapse, and it is rarely the login wall, the model menu, or the speed of the first token. You paste a goal, three files change, and nobody can name the command that was red before the edit landed. Would you merge a teammate's patch if the only remaining witness was a chat transcript and a hopeful shrug?

That shrug is the friction that actually eats the quarter hour, long before any architecture debate earns the right to speak. A mechanic who swaps parts before reading the fault code sometimes gets lucky, then cannot explain the repair. An assistant that edits before a named check is that mechanic, with your repository as the workshop and the diff as the invoice. If the first fifteen minutes do not end with a frozen failure, what exactly did you purchase with that time?

I do not need those minutes for a settings tour, a green spinner, or praise about how a tool understood the codebase. I want a working directory, one check command, a non-zero exit, and a short tail of the output that made the failure obvious. I do not want a rewritten module or a new dependency until that red exit has a name I can type again. If the check cannot fail on purpose before the patch, what are we asking any model to fix?

The one fix that changed the session for me was a replay card, a small file written before any edit is allowed. It freezes the command, the directory, the git revision, a fingerprint of the dirty tree, and the exit code you just watched. Later, that card is the only context worth pasting, and the same command is the only verdict worth trusting. If the card cannot be rerun in a clean shell, the session was a conversation, not a repair you can defend.

The script below is a local workflow you can copy, not a score I am pretending to have measured on your laptop. It writes .replay/card.env when you pass capture, and it reruns the frozen command when you pass verify. Nothing in the file calls a model, which is deliberate, because the witness should not depend on the assistant that proposed the change. Read it before you trust it, and keep the card out of any commit that might carry secrets from the failing run.

#!/usr/bin/env bash
set -euo pipefail

# Local proposal: bash, optional git, and shasum. Not a measured product benchmark.
card_dir="${REPLAY_DIR:-.replay}"
card="$card_dir/card.env"
mkdir -p "$card_dir"

tree_fingerprint() {
  if git rev-parse --is-inside-work-tree >/dev/null 2>&1; then
    printf '%s %s\n' \
      "$(git rev-parse HEAD)" \
      "$(git status --porcelain=v1 | shasum -a 256 | awk '{print $1}')"
  else
    find . -type f -not -path './.replay/*' -print0 \
      | sort -z | xargs -0 shasum -a 256 | shasum -a 256 | awk '{print $1}'
  fi
}

capture() {
  local cmd="$1" log exit_code
  log="$(mktemp)"
  set +e
  bash -lc "$cmd" >"$log" 2>&1
  exit_code=$?
  set -e
  {
    printf 'captured_at=%q\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
    printf 'workdir=%q\n' "$PWD"
    printf 'tree_fp=%q\n' "$(tree_fingerprint)"
    printf 'cmd=%q\n' "$cmd"
    printf 'exit_before=%q\n' "$exit_code"
  } >"$card"
  tail -n 40 "$log" >"$card_dir/stderr-before.txt"
  rm -f "$log"
  printf 'captured exit_before=%s into %s\n' "$exit_code" "$card"
  if [[ "$exit_code" -eq 0 ]]; then
    echo "warning: freeze a failing command before you ask for a patch" >&2
    exit 2
  fi
}

verify() {
  # shellcheck disable=SC1090
  source "$card"
  local log exit_code
  [[ -n "${cmd:-}" && -n "${exit_before:-}" ]] || {
    echo "card is incomplete" >&2
    exit 3
  }
  log="$(mktemp)"
  set +e
  bash -lc "$cmd" >"$log" 2>&1
  exit_code=$?
  set -e
  tail -n 40 "$log" >"$card_dir/stderr-after.txt"
  rm -f "$log"
  printf 'exit_before=%s exit_after=%s\n' "$exit_before" "$exit_code"
  if [[ "$exit_before" -eq 0 ]]; then
    echo "card was not a failing check" >&2
    exit 3
  fi
  if [[ "$exit_code" -ne 0 ]]; then
    echo "still red; do not treat the diff as the fix" >&2
    exit 1
  fi
  echo "same command flipped to zero"
}

case "${1:-}" in
  capture) capture "${2:?command required}" ;;
  verify) verify ;;
  *) echo "usage: replay-card.sh capture '<command>' | verify" >&2; exit 64 ;;
esac
Enter fullscreen mode Exit fullscreen mode

A session that respects the card starts with a command you already believe should pass, then captures the moment it does not. I usually point it at one file-scoped check, because a full suite hides the failure inside a crowd of unrelated green lines. After the card exists, and only after exit_before is non-zero, I allow an edit of any kind. The verify step is intentionally dumb: same words, same shell, and a demand that the recorded red exit become zero.

chmod +x replay-card.sh
./replay-card.sh capture 'npm test -- src/parser.test.js'
# card must exist, and exit_before must be non-zero, before any edit
./replay-card.sh verify
git diff --stat
Enter fullscreen mode Exit fullscreen mode

If verify is still red, the patch is only a hypothesis wearing a confident tone, and the chat does not get a vote. I treat that red line as the session still being inside its opening quarter hour, no matter how many files already changed. A green verify is permission to read the diff carefully, not permission to stop asking whether the check was weak.

Notice what the fingerprint is not allowed to do during the verdict, even though the card bothers to store it. A real fix must change the tree, so locking the porcelain hash would freeze the bug and call that stillness purity. I keep tree_fp as a breadcrumb for the morning after, the way a mechanic keeps the old fault code on a clipboard. The lock is the command string plus the exit flip, not the beauty of the diff and not the confidence of the chat.

This is where a free model session earns a narrow job, and where I should be plain about the relationship before going further. Disclosure: This article was prepared as part of MonkeyCode's product outreach. I would paste the command and the stderr tail there, and I would ask only for a patch aimed at that card. I would run verify on the free server option only as a second shell, not as proof the laptop was wrong.

I am not stating a quota, a model name, a hardware shape, a duration, or a promise that either free option stays fixed. The second shell is a witness rather than a trophy, and it counts only if it can build what the check command needs. If that server lacks the toolchain, green is noise and red is a portability finding, not an insult aimed at the model. Would a passing run on an empty image prove anything except that you skipped the hard dependency on purpose?

There is a temptation, once verify prints its one success line, to treat the quarter hour as finished and the review as optional. That temptation is how a flipped exit gets promoted into a feeling of safety that the card never actually earned. The card proves that one command changed its mind under one shell, which is more than a transcript and less than a careful reader. I still read the diff, and I still ask whether the check was too weak to notice a wrong fix.

Some failures should never enter this workflow, and pretending otherwise is how a tidy script becomes a leak. Do not capture a command that prints tokens, customer rows, or production URLs into .replay, and do not commit that directory by accident. Skip the method when the check is non-deterministic, because a flip to zero might be luck wearing a green coat. Skip it when you need signed provenance, an air-gapped policy, or a human gesture that no shell can honestly replay.

Teams with a gated CI job can keep the card local, then let CI be the stranger who reruns the same words. If your check needs a browser, a paid sandbox, or a secret you must not paste, leave the remote server out of the story. A local red-to-green flip is already a better first fifteen minutes than a chat that rewrote three files and wished you luck. Why hand the only copy of the failure to a window you might close before lunch?

If you already have a free model session and a spare server, refuse the first diff that arrives without a card. The useful question is not whether the assistant sounds sure of itself before you have looked at the exit. The useful question is whether tomorrow morning you can rerun the same words and watch that recorded red become zero. I would rather lose a clever patch than keep a repair I cannot replay on a clean shell the next day.

Top comments (0)