It is 3 a.m. in the lab, and a batch of 412 prompt jobs dies at item 271. The terminal holds a partial traceback; the disk holds nothing. The next run starts at prompt one and pays for two hundred and seventy calls a second time. This workshop is built around that single failure and the smallest artifact that prevents it: an append-only ledger plus a key derived from the job itself.
The session runs ninety minutes. Students leave able to kill a process with kill -9 in the middle of a batch and then prove, with a script, that the resumed run produced neither a gap nor a duplicate.
What the lab assumes
Every student needs a machine with Python 3.9 or newer, two input files, and an HTTP JSON model endpoint configured through two environment variables. Nothing else is installed. No queue, no database, no orchestration layer, because the point of the exercise is that the durable part is a text file.
For the reference run the endpoint points at MonkeyCode. The project's operator describes it as providing free model access and a free server option, which matters here for a mundane reason: a ninety-minute lab where the long-running process has somewhere to live beats a lab where half the room is still provisioning. Disclosure: This article was prepared as part of MonkeyCode's product outreach.
Treat every number about that free tier as moveable. The operator states the free tier carries up to 10 million tokens, and that a free server option is available; quotas, model availability, and server terms change, so read the project's current documentation rather than a blog post from any month, including this one. The code below is a teaching artifact written for the lab. I have not benchmarked it, and the timings are teaching budgets rather than measured results.
The artifact: one line per finished job
The design is deliberately boring. Each completed job appends one JSON line to ledger.jsonl, flushed and fsynced before the runner moves on. The job's identity is a hash of its own payload, so the same job presented twice produces the same key. On restart the runner reads the ledger into a set and skips anything already in it.
jobs.jsonl looks like this:
{"id": "j-001", "prompt": "Summarise this changelog entry in one line: ..."}
{"id": "j-002", "prompt": "Rewrite this error message for a non-technical reader: ..."}
The runner is about forty lines. The request and response shapes depend on the endpoint, so students adapt call_model first and leave the ledger logic alone.
# runner.py -- teaching artifact, not production code
import hashlib, json, os, time, urllib.request
from pathlib import Path
LEDGER = Path("ledger.jsonl")
JOBS = Path("jobs.jsonl")
MODEL_URL = os.environ["MODEL_URL"]
API_KEY = os.environ.get("MODEL_KEY", "")
def key_of(job: dict) -> str:
payload = json.dumps(job, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(payload.encode()).hexdigest()[:16]
def load_done() -> set:
done = set()
if not LEDGER.exists():
return done
for line in LEDGER.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
try:
record = json.loads(line)
except json.JSONDecodeError:
continue # torn final write; let the job run again
if record.get("status") == "ok":
done.add(record["key"])
return done
def append(record: dict) -> None:
with LEDGER.open("a", encoding="utf-8") as fh:
fh.write(json.dumps(record, sort_keys=True) + "\n")
fh.flush()
os.fsync(fh.fileno())
def call_model(prompt: str, timeout: int = 60) -> str:
body = json.dumps({"prompt": prompt}).encode()
req = urllib.request.Request(
MODEL_URL,
data=body,
headers={"Content-Type": "application/json",
"Authorization": f"Bearer {API_KEY}"},
)
with urllib.request.urlopen(req, timeout=timeout) as resp:
return json.loads(resp.read())["text"]
def run_one(job: dict) -> None:
k = key_of(job)
last_error = None
for attempt in range(4):
try:
out = call_model(job["prompt"])
append({"key": k, "job": job.get("id"), "status": "ok", "output": out})
return
except Exception as exc: # noqa: BLE001 -- lab code on purpose
last_error = exc
time.sleep(2 ** attempt)
append({"key": k, "job": job.get("id"),
"status": "error", "error": str(last_error)})
def main() -> None:
done = load_done()
for line in JOBS.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
job = json.loads(line)
if key_of(job) in done:
continue
run_one(job)
if __name__ == "__main__":
main()
Notice that load_done only counts records whose status is ok. Error records still land in the ledger for forensics, but they stay eligible for a later run. That single conditional is the difference between a retry policy and a silent data hole.
The ninety minutes
| Minutes | Segment |
|---|---|
| 0-10 | The 3 a.m. crash: what a lost batch actually costs |
| 10-25 | Read the ledger format, run forty jobs end to end |
| 25-45 | Exercise 1: kill -9 and resume |
| 45-60 | Exercise 2: the torn write |
| 60-75 | Exercise 3: two workers, one ledger |
| 75-90 | Debrief: what belongs in a record and what does not |
Exercise 1: crash and resume
Students start the runner in the background, wait for a few seconds of work, and kill it hard.
python runner.py &
PID=$!
sleep 25
kill -9 $PID
wc -l ledger.jsonl # some completed jobs, no summary
python runner.py # resumes
The proof step is the part that matters. Counting lines is not evidence, because duplicates also add lines.
python - <<'PY'
import collections, json
keys = [json.loads(l)["key"] for l in open("ledger.jsonl") if l.strip()]
c = collections.Counter(keys)
print("lines", len(keys), "unique", len(c))
print("duplicates", [k for k, v in c.items() if v > 1])
PY
A clean run reports matching counts and an empty duplicate list. When a pair does appear, it is usually because the killed process had already sent the request and the restart sent it again before any ledger line existed; the fix is a claim file, which is Exercise 3.
Exercise 2: the torn write
kill -9 can land in the middle of fh.write, leaving a truncated last line. Students reproduce this deterministically without racing anything.
head -c -25 ledger.jsonl > ledger.tmp && mv ledger.tmp ledger.jsonl
python runner.py
The JSON parse fails on the final line, load_done skips it, and the corresponding job reruns. Nobody has to reason about it in the abstract; they watch the line count grow by one and the duplicate check stay clean. The lesson worth writing on the board: a reader that tolerates garbage is more valuable than a writer that promises never to produce it.
Exercise 3: two workers, one ledger
Two terminals, one ledger, no coordination.
python runner.py & python runner.py &
wait
Both processes read the same "done" set, both see the same missing key, and both call the model. The ledger now holds two records for one job. The repair is an exclusive claim created before the network call.
def claim(k: str) -> bool:
os.makedirs("claims", exist_ok=True)
try:
fd = os.open(f"claims/{k}", os.O_CREAT | os.O_EXCL | os.O_WRONLY)
except FileExistsError:
return False
os.close(fd)
return True
Then have the class break it again by killing a worker between the claim and the ledger append. The claim file survives, the job never reruns, and the batch quietly loses a result. That trade-off is the real content of the exercise: exclusivity buys you no duplicates and costs you a stale-claim policy, usually a timestamp older than a few minutes that may be reclaimed.
When a ledger is not worth the trouble
| Situation | Ledger? | Reason |
|---|---|---|
| Forty jobs, two seconds each | No | Rerunning the batch is cheaper than maintaining state |
| Four hundred jobs, a minute each | Yes | Losing halfway through costs hours |
| Side effects outside the ledger | Only with idempotent side effects | A key stops duplicate calls, not duplicate emails |
| Regulated or personal data | Not in plaintext | Hash the prompt for the key; store raw text only where policy allows |
| The only copy lives on a free host | Risky | Convenience is not durability |
The last row deserves the most discussion. A free server option is a fine place to run a ninety-minute lab, and a poor place to keep the only record of a week-long run. Ledgers belong on storage you control, with the compute pointed at it.
Who should skip this
If the work is interactive, with a human reading each answer before the next call, the ledger costs more attention than it saves. If the pipeline writes to a customer system, hashing the input is not enough, because the side effect itself has to be idempotent. If the batch carries personal data, a plaintext JSONL file on a shared machine is a compliance problem, not an engineering one. And if the team has no place to keep the ledger, the honest answer is to fix that before writing the runner.
Early finishers should spend the remaining time attacking the artifact rather than admiring it: kill the process during fsync, put the ledger on a network path two machines share, and watch what the append does. Breaking the file teaches more than the happy path ever will.
The real output of the workshop is not runner.py. It is the habit of asking, before a long run starts, what the world looks like when it dies at item 271. Free model access and a free server option lower the cost of setting that lab up, so check the current limits before you schedule the session, and keep the fixture set small enough that a rerun is never a crisis.
Top comments (0)