You get the page at 2:17 a.m. because the webhook handler stopped answering, and the only notes sit on a laptop that already slept. The tunnel that pointed at your kitchen table is gone, and a forgotten job is still retrying a paid model against logs that no longer exist. This is not a story about a clever prompt, because the morning review has nothing it can replay. It is a postmortem about a missing timeline, and about the small host you should have left awake.
The first failure is ordinary, which is why it survives code review and then dies at night. A three-service demo receives a burst of error callbacks, and your handler writes each raw body under a temporary directory before it summarizes anything. The laptop is the server, which feels reasonable at dusk, the way a sticky note feels reasonable until someone opens the window. At 1:40 a.m. the lid closes after an update prompt, the process dies without a flush, and the last file is cut in half.
By 2:05 a.m. a second runner still holds yesterday's API key and keeps asking a paid model to explain bodies it cannot see. You are paged at 2:17, the tunnel target is gone, and the spend notice is already queued for breakfast. The file you needed is a zero-length leftover, so the review becomes a debate about memory rather than a replay of evidence. That is the incident, and the rest of this piece is the durable fix rather than a recap of the panic.
What stacked
Contributing factors stack the way wet firewood stacks, quiet in the afternoon and useless when you need a flame. You treated a developer laptop as an always-on host, even though sleep, updates, and cafe wireless are features of a laptop rather than accidents. You let every log line become a model call, so a retry loop could spend money without leaving a record a colleague could open. You stored the timeline in a temporary directory, a hallway rather than a cabinet, and never split secrets from outbound text.
None of those choices looked dramatic at six in the evening, because each one worked during the demo you watched with your own eyes. Together they made the night impossible to reconstruct, which is worse than a slow page, because a slow page at least leaves a body. A postmortem that only retells the panic teaches nothing, so the fix has to change what done means for this handler. Done is not a chat transcript that vanishes when a browser tab closes or a vendor session expires.
The order that survives sleep
A generated handler can look finished in the afternoon and still forget to flush, which is how a tidy demo becomes an empty file. Done is a replayable file that another person can run tomorrow without your laptop, without your paid key, and without the original payloads. You keep events on a host that stays awake, append one JSON object per event, and flush before any network call. Only after that write succeeds do you send a model a redacted summary, and if the call fails you still hold the file.
Free model access and a free server option can carry that rehearsal, but they are a spare room, not a bank vault. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Here MonkeyCode is the open-source place where you park a tiny recorder on the free server option and send redacted classification calls. This piece states no measured quota, named model, hardware shape, or promise that zero cost lasts, because those details change.
If the free path is down, the recorder must still write the file, or you have rebuilt the same incident under a different name. The artifact below is a proposed recorder, labeled as an example you can read and adapt, not as a benchmark executed for this article. It appends one JSON object per event, redacts obvious secret-shaped fields, and refuses to pretend a network call happened before the write. You can run it on any small Linux host, including a free server when the current offer matches this modest shape.
#!/usr/bin/env python3
"""Proposed incident recorder. Example only, not a measured run."""
import json
import re
import sys
from datetime import datetime, timezone
from pathlib import Path
SECRET = re.compile(r"(?i)(api[_-]?key|token|password|secret|authorization)")
PATH = Path(sys.argv[1] if len(sys.argv) > 1 else "timeline.jsonl")
def redact(value):
if isinstance(value, dict):
return {
k: ("[redacted]" if SECRET.search(k) else redact(v))
for k, v in value.items()
}
if isinstance(value, list):
return [redact(v) for v in value]
if isinstance(value, str) and SECRET.search(value):
return "[redacted]"
return value
def append_event(kind, detail):
event = {
"ts": datetime.now(timezone.utc).isoformat(),
"kind": kind,
"detail": redact(detail),
}
PATH.parent.mkdir(parents=True, exist_ok=True)
with PATH.open("a", encoding="utf-8") as handle:
handle.write(json.dumps(event, ensure_ascii=False) + "\n")
handle.flush()
return event
def check_timeline(path):
raw = Path(path).read_text(encoding="utf-8")
if raw and not raw.endswith("\n"):
raise SystemExit("trailing partial line")
lines = [line for line in raw.splitlines() if line]
for index, line in enumerate(lines, start=1):
event = json.loads(line)
blob = json.dumps(event.get("detail", {}))
if "sk-" in blob or "Bearer " in blob:
raise SystemExit(f"secret residue on line {index}")
print(f"checked {len(lines)} events")
if __name__ == "__main__":
sample = {
"route": "/hooks/errors",
"status": 502,
"api_key": "sk-live-do-not-keep",
}
print(json.dumps(append_event("webhook.failed", sample), indent=2))
check_timeline(PATH)
You can also replay the file with the network disabled, which is the test that would have saved the morning review. You start the recorder only after the directory lives on real disk, not on a memory filesystem that disappears when the machine reboots. A short shell ritual makes that habit visible, and it gives the later review a clock you can quote without guessing. Create the directory, run the recorder against a sample failure, and assert that the secret field was replaced before you celebrate.
mkdir -p "$HOME/incident-repro" && cd "$HOME/incident-repro"
python3 recorder.py timeline.jsonl
python3 - <<'PY'
import json
last = open("timeline.jsonl", encoding="utf-8").read().strip().splitlines()[-1]
event = json.loads(last)
assert event["detail"]["api_key"] == "[redacted]"
assert event["kind"] == "webhook.failed"
print("replay-ok", event["ts"])
PY
If the assertion fails, you have a defect in the recorder, which is a kinder failure than a leaked key inside a model transcript. The sample redactor catches secret-shaped keys, not every token buried in a sentence, so the daily check looks for residue it might have missed. The handler change is smaller than people expect, and that smallness is why teams skip it after the demo looks green. Before the incident, the handler called the model inside the request, so a slow vendor could hold the webhook until the sender retried.
Handler, clock, and partial lines
After the fix, the handler only appends and returns, and a separate process classifies when it finds unlabeled lines. That split is the difference between a doorbell and a conversation that refuses to end, even when the vendor is slow. Partial lines are the quiet bug inside an otherwise careful recorder, and they deserve a named check of their own. A process killed mid-write can leave a trailing fragment, and a naive reader will throw and abandon the whole night.
def accept_webhook(body: dict) -> dict:
saved = append_event("webhook.failed", body)
# Classify later, offline, against the file. Do not call out here.
return {"accepted": True, "ts": saved["ts"]}
Your replay should skip nothing silently, count a trailing incomplete line, and fail the daily check if that count is greater than zero. That rule turns a crash into a visible gap instead of a parser exception nobody on the team is willing to own. Clock choice belongs in the postmortem because two runners will disagree if you trust local wall time. Write UTC timestamps, and do not sort events by the order a chat window happened to print them during the night.
When you reconstruct, you sort the file, then you look for gaps longer than your webhook timeout. A long gap is a missing host, not a mysterious model failure, and that name stops a larger-model purchase meant to fix sleep. Read the teaching timeline as a pattern, not as a personal metric, because no private invoice belongs in a public note. At 1:40 the host sleeps and the handler dies mid-write, so the last event never becomes a complete line.
At 1:51 retries begin from a second runner that still has a paid key and no local copy of the bodies. At 2:17 you are paged, but the tunnel target is gone, so you cannot fetch the request that woke everyone up. At 8:10 the spend notice arrives, and the only artifact is a zero-length file in a temporary directory. The durable sequence reverses that order: write, flush, redact, then classify, and never classify a body you have not saved.
Classification is the only step that needs a model, and the question should be narrow enough to survive a weaker or interrupted call. You take the last twenty redacted lines, ask for a short label such as timeout, auth, or unknown, and store that label beside the file. Free model access is enough for that narrow question when the current offer includes it, because the model is not your system of record. If that call never returns, the file is still the incident, and the label can wait until a person is awake.
Who should walk away
Who should not use this approach matters as much as who should, and the boundary is about data and promises rather than taste. Do not place customer secrets, health data, or payment events on a free server you do not control under a written agreement. Do not treat free model access as an unlimited batch pipe, and do not assume the offer, the queue, or the region will match next month. Do not use this recorder as your only copy if a regulator expects retention, encryption, and an access log you can hand to an auditor.
Teams with a formal on-call promise should keep the timeline on infrastructure they already monitor, then copy the redact-then-classify order. Hobby projects and internal rehearsals fit this spare room, provided a human still opens the file before any paid key returns to the retry path. If you need a region pin, a signed retention period, or a support contract, this free path is the wrong room. The regression you want is boring, which is a compliment in incident work, because excitement usually means a missing fact.
Once a day, replay yesterday's file with the network disabled and assert that every secret-shaped field is redacted and that every line parses. If that check fails, you have a code defect, not an incident mystery, and you can fix it before the next page. If the free server disappears, you move the same directory to any other small host and you do not rewrite the format around a vendor. The fix is the file format and the order of operations, not the brand painted on the spare room.
If you want scratch space for that rehearsal, the free server option and free model access are a fair place for redacted timelines and short labels. Confirm the live limits in the project docs before you depend on them, and keep the paid key parked until a human has opened the file.
Top comments (0)