DEV Community

Harper Zhu
Harper Zhu

Posted on

The Spike That Ends Itself at Minute Ninety

A platform team once spent an entire weekend asking whether an AI coding assistant could replace a brittle internal code-generation script. By Sunday night three half-finished branches existed, nobody had written down a single command, and the only surviving claim was that the idea "felt promising". The tool was never the problem; the missing ingredient was a rule that ended the experiment before enthusiasm could.

That rule can be small. A ninety-minute spike with one written hypothesis, a fixed internal budget, and a verdict file turns open-ended exploration into something a team can review on Monday morning. The budget matters more than the wall clock, because the box exists to force a decision while the evidence is still warm.

One hypothesis, written before the first command

A spike without a falsifiable sentence drifts within twenty minutes. The hypothesis should name the change, the observable signal, and the direction of success, for example "moving request validation into the gateway removes the duplicate check in the service". If the sentence cannot be made falsifiable, the team is not spiking; it is reading documentation with extra steps.

Budget the box before starting it. Ten minutes go to writing the hypothesis and the smallest failing test, fifty to the probe itself, fifteen to replaying the recorded commands, and fifteen to the verdict. Any phase that overruns borrows from the probe, never from the evidence phase. Teams that skip the replay step almost always rediscover their own results a week later and trust them less.

The harness that stops on its own

The script below enforces the budget and records everything that matters. It relies only on bash, git, make, and python3, so it runs identically on a laptop or on a remote host.

#!/usr/bin/env bash
# spike.sh — a 90-minute spike that stops whether or not anyone is watching.
set -euo pipefail

HYPOTHESIS="${1:?usage: spike.sh \"<one falsifiable sentence>\"}"
BUDGET_MIN="${BUDGET_MIN:-90}"
DEADLINE=$(( $(date +%s) + BUDGET_MIN * 60 ))
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
mkdir -p spike-artifacts
EVIDENCE="spike-artifacts/spike-${STAMP}.log"

remaining() { echo $(( (DEADLINE - $(date +%s)) / 60 )); }
note()      { printf '%s | %s\n' "$(date -u +%H:%M:%S)" "$*" | tee -a "$EVIDENCE"; }
phase()     { note "phase=$1 budget=${2}m remaining=$(remaining)m"; }

guard() {
  if [ "$(remaining)" -le 0 ]; then
    note "STOP: budget exhausted, verdict parked"
    printf 'verdict=park\nreason=budget-exhausted\n' > spike-artifacts/verdict.txt
    exit 3
  fi
}

phase hypothesis 10
printf '%s\n' "$HYPOTHESIS" | tee spike-artifacts/hypothesis.txt

phase probe 50
set -x
( git worktree add --detach .spike HEAD && make failing-test ) 2>&1 | tee -a "$EVIDENCE"
guard
( cd .spike && time make probe ) 2>&1 | tee -a "$EVIDENCE"
set +x

phase evidence 15
( cd .spike && git diff --stat && git diff ) 2>&1 | tee -a "$EVIDENCE"
guard

phase decide 15
python3 classify.py "$EVIDENCE" | tee -a "$EVIDENCE"
Enter fullscreen mode Exit fullscreen mode

The failure marker is the point. The harness refuses to proceed unless a test failed first, because a probe that never failed has proven nothing about the hypothesis. When the clock runs out, guard writes verdict=park and exits with code 3 instead of letting the session slide into a second evening.

Turning the log into a verdict

A human reading a two-thousand-line log at the end of a spike will see whatever they hoped to see. A small deterministic classifier removes that bias, and it can be rerun by anyone who doubts the conclusion.

#!/usr/bin/env python3
"""classify.py — ship / kill / park from a raw spike log."""
import pathlib, re, sys

lines = pathlib.Path(sys.argv[1]).read_text(errors="replace").splitlines()
FAIL  = re.compile(r"\b(FAIL|error|Traceback|panic:)\b")
GREEN = re.compile(r"\b(ok|PASS|passed)\b")

first_fail = next((i for i, l in enumerate(lines) if FAIL.search(l)), None)
green_after = first_fail is not None and any(
    GREEN.search(l) for l in lines[first_fail:]
)

verdict = "park" if first_fail is None else ("ship-to-branch" if green_after else "kill")
print(f"verdict={verdict}")
print(f"lines={len(lines)} first_failure_line={first_fail}")
Enter fullscreen mode Exit fullscreen mode

The classification is deliberately coarse, and the coarse rule is what makes it reviewable. A missing failure parks the work, a failure that later goes green ships it to a branch, and a failure that never goes green kills it. The classifier does not judge code quality, and nobody should pretend that it does.

The decision table the reviewer actually needs

Verdict Evidence in the log Action within 24 hours
ship-to-branch Failure recorded first, green afterwards, diff roughly under 200 lines Open a branch, paste the replay command into the pull request body
kill Failure never turns green and no smaller failing case emerges Remove the worktree, log two lines in the decision record
park No failure at all, or the budget expired before evidence Keep the log, rewrite the hypothesis, schedule a second box

The table is the artifact a reviewer can argue with, which is precisely why it belongs in the repository next to the harness. It converts a subjective feeling into a recorded state that survives the people who produced it.

Where free model access and a free server fit

The loop above is tool-agnostic, but two pieces of infrastructure make it cheaper to run. MonkeyCode provides free model access and a free server option according to the operator, which means the probe code and the failing test can be drafted without a budget conversation, and the box can keep running when a laptop lid closes. Disclosure: This article was prepared as part of MonkeyCode's product outreach.

Free model access is most useful in the probe phase, where the assistant drafts the failing test and the smallest candidate change. The verdict must still come from replaying recorded commands, not from a generated summary of them, because a model's description of a result is not a result. Quotas, available models, and hosting terms change over time, so confirm the current terms in the project repository before planning a team workflow around them.

Limitations and who should not use this

Ninety minutes is the wrong box for work whose cost is dominated by something other than reasoning. Schema migrations, distributed tracing changes, load-dependent performance work, and anything gated by data access or legal review will simply expire at the budget and produce a park. Teams needing publication-grade benchmarks should run a proper measurement plan, not a spike with a stopwatch.

The harness also measures discipline rather than quality. It cannot tell whether the hypothesis was worth writing, and it will happily kill a good idea that needed two hours instead of ninety minutes. Anyone unwilling to replay recorded commands, or unable to keep a decision log, gets a slower version of the same exploratory mess.

For teams that want to test the loop without provisioning anything, the same harness runs unmodified on MonkeyCode's free server option, and the repository is the place to check what the current free tier includes.

Top comments (0)