A remote model call is not spike evidence. A failed local fixture is enough to kill the idea. This plan tests one hypothesis inside ninety minutes.
Conclusion first
Ship only when the local gate passes first. Kill the work if that gate is still red at minute forty. A free remote runner cannot repair a broken check.
The single hypothesis
Write the hypothesis down before any client starts. A bad row must end the process with exit code 2. A later text reply must not override that exit.
The hypothesis stays fixed for the whole block. Do not add a second claim at minute fifty. A second claim belongs in a different spike.
The ninety-minute clock
Use one uninterrupted block for the whole spike. Do not extend the block to chase a cleaner log. Split it into four slices and record the start time.
- Minutes 0 to 15 lock the hypothesis and the kill rules.
- Minutes 15 to 40 run the local fixture and save both exits.
- Minutes 40 to 70 allow one remote pass only if the gate is green.
- Minutes 70 to 90 write the decision sheet and stop all edits.
Minute forty is a hard cut for eligibility. A missing receipt at that minute is a kill. The remote slice does not reopen a missed local slice.
Kill rules before tools
Write these rules in the note before you import a client.
- Kill the spike if the bad file exits zero.
- Kill the spike if the good file exits non-zero.
- Kill the spike if a secret-shaped string enters the prompt file.
- Kill the spike if the remote step starts on a red gate.
- Kill the spike if minute forty arrives with no receipt.
These five kill rules stay binary on purpose. A partial pass still counts as a kill. Debate belongs in the next spike, not in this block.
Local fixture
The program below is only a proposed check. This listing is not a recorded production run. This listing is not a latency or cost benchmark.
Run it locally and keep the raw files. Save the script as spike_gate.py before you time the block. Do not tune it after the clock starts.
#!/usr/bin/env python3
"""Proposed spike fixture. Not a recorded production result."""
import json
import sys
REQUIRED = ("id", "status", "source")
ALLOWED = {"ok", "skip"}
def check(row):
if not isinstance(row, dict):
return "not-object"
missing = [key for key in REQUIRED if key not in row]
if missing:
return "missing:" + ",".join(missing)
if row["status"] not in ALLOWED:
return "bad-status"
return ""
def main():
payload = json.load(sys.stdin)
rows = payload if isinstance(payload, list) else [payload]
errors = []
for index, row in enumerate(rows):
reason = check(row)
if reason:
errors.append({"index": index, "reason": reason})
if errors:
json.dump({"ok": False, "errors": errors}, sys.stdout)
return 2
json.dump({"ok": True, "count": len(rows)}, sys.stdout)
return 0
if __name__ == "__main__":
raise SystemExit(main())
Two files, two exits
Create one valid payload and one invalid payload. The invalid payload must fail closed every time. Do not edit the checker to match a wish.
printf '%s\n' '{"id":"a1","status":"ok","source":"local"}' > good.json
printf '%s\n' '{"id":"b2","status":"maybe"}' > bad.json
python3 spike_gate.py < good.json > good.out; echo "good_exit:$?"
python3 spike_gate.py < bad.json > bad.out; echo "bad_exit:$?"
Record the two exit lines in the note. The good run must print an exit of 0. The bad run must print an exit of 2. Any other exit pair ends the spike immediately.
Also confirm the bad output names the missing field. A bare failure string is weaker review evidence. The error object should include index and reason.
Receipt before any network call
Hash the inputs and outputs before a remote step. The receipt is the artifact a reviewer can rerun. A chat log is not a valid substitute.
python3 - << 'PY'
import hashlib, json, pathlib
names = ["good.json", "bad.json", "good.out", "bad.out"]
rows = []
for name in names:
blob = pathlib.Path(name).read_bytes()
rows.append({
"file": name,
"sha256": hashlib.sha256(blob).hexdigest(),
"bytes": len(blob),
})
pathlib.Path("receipt.json").write_text(json.dumps(rows, indent=2))
print("receipt_files", len(rows))
PY
If a file is missing, the script should fail. Do not catch that error and then continue. A missing file means the local slice is incomplete.
Rerun check
A second local run should match the saved bad output. Drift means the checker changed during the block. Drift after minute forty is a kill, even if both exits stay correct.
python3 spike_gate.py < bad.json > bad2.out
cmp bad.out bad2.out && echo "rerun:same" || echo "rerun:drift"
Decision table at minute forty
Apply this decision table once at minute forty. Do not renegotiate a red row after a hopeful reply.
| Local exits | Prompt file | Clock | Decision |
|---|---|---|---|
| good 0, bad 2 | no secret-shaped text | under 40 min | one remote pass allowed |
| bad exits 0 | any | any | kill, the gate is open |
| good exits non-zero | any | any | kill, the fixture is wrong |
| any | secret-shaped text | any | kill, do not send the file |
| any | any | no receipt by minute 40 | kill, the clock was lost |
Count the red rows before you count the green cell. One red row dominates the entire decision sheet. The green cell allows a pass, not a launch.
Optional remote pass
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
When a remote tool is allowed in
MonkeyCode enters only after the local receipt exists. Operator notes for this draft say free model access is available. The same notes say a free server option is available.
This spike does not assume a token quota. It also does not assume a model name. It does not assume a fixed hardware shape.
Those access details can change without a notice. Read the current project page before you depend on them. Do not treat a blog claim as a contract.
What the pass may decide
Use the remote pass for a single classification. The question is whether the error JSON matches the hypothesis. The question is not how to bypass the checker.
A free server can host that one clean shell. It should not hold keys, tokens, or private rows. Copy only the bad output and the hypothesis line.
Leave good.json local if it resembles real data. Prefer synthetic ids such as a1 and b2. Keep customer rows and live tokens off the remote shell.
Bounded prompt
Save this template as prompt.txt after the secret scan. It is a template, not a captured session. Replace the placeholder with the raw bad output.
Hypothesis: a bad row must exit 2.
Classify the fixture JSON as MATCH, MISMATCH, or UNCLEAR.
Do not propose a bypass. Do not rewrite the checker.
JSON:
<paste bad.out>
Scan that prompt file before sending it anywhere.
if grep -E 'AKIA|sk-|BEGIN PRIVATE|api_key' prompt.txt; then
echo "secret_scan:hit"
else
echo "secret_scan:clean"
fi
A hit is a kill, even when the pattern is a false positive. Clear the file and rebuild it from the fixture. Do not send the original as a one-off exception.
Four-line close
Stop all code edits once minute seventy starts. Write four lines and leave the tree alone.
- Record the hypothesis id and the clock start.
- Record both local exits and the receipt hash prefix.
- Record whether the remote pass ran, and the gate reason.
- Record ship or kill, with one reason only.
Ship means the checker is worth a repo commit. Kill means the idea stops in this branch. Both outcomes still need the same four lines.
A kill without a receipt is just an abandoned note. A ship without both exit lines is also incomplete. The clock does not grant extra minutes for a prettier summary.
What this does not prove
This spike does not score remote model quality. It does not measure latency, uptime, or price. It does not prove a free tier will remain available.
It does not clear production traffic or regulated data. A green pair of exits proves these two files only.
A remote classification label can still be wrong. Treat MATCH as a hint, not a merge approval. Re-run the local fixture after every checker edit.
If the new exit disagrees with the hint, stop. Keep the exit code and drop the hint. Then mark the spike killed for conflicting evidence.
Who should not use it
Skip this plan when you need a load test. Skip it when you need a promised token quota. Skip it when payloads hold personal data you cannot share.
Skip it when reviewers cannot see the raw exit codes. Skip it when the task has two hypotheses. Skip it when ninety minutes is not actually free.
What to retain
Keep the checker, both JSON files, both outputs, and the receipt. Keep the four-line note in the same directory. Delete transcripts that do not change the decision.
The next spike should open the receipt first. After a green gate, confirm current MonkeyCode access terms yourself. Then run one bounded pass on a clean shell.
Stop when the clock hits ninety, even if the reply looks unfinished. A finished receipt beats an unfinished chat. The gate, not the reply, decides ship or kill.
Top comments (0)