DEV Community

Charlie Zhu
Charlie Zhu

Posted on

The Checkpoint Workshop: Resumable Batch Runs in 90 Minutes

It is 3 a.m. in the lab, and a batch of 412 prompt jobs dies at item 271. The terminal holds a partial traceback; the disk holds nothing. The next run starts at prompt one and pays for two hundred and seventy calls a second time. This workshop is built around that single failure and the smallest artifact that prevents it: an append-only ledger plus a key derived from the job itself.

The session runs ninety minutes. Students leave able to kill a process with kill -9 in the middle of a batch and then prove, with a script, that the resumed run produced neither a gap nor a duplicate.

What the lab assumes

Every student needs a machine with Python 3.9 or newer, two input files, and an HTTP JSON model endpoint configured through two environment variables. Nothing else is installed. No queue, no database, no orchestration layer, because the point of the exercise is that the durable part is a text file.

For the reference run the endpoint points at MonkeyCode. The project's operator describes it as providing free model access and a free server option, which matters here for a mundane reason: a ninety-minute lab where the long-running process has somewhere to live beats a lab where half the room is still provisioning. Disclosure: This article was prepared as part of MonkeyCode's product outreach.

Treat every number about that free tier as moveable. The operator states the free tier carries up to 10 million tokens, and that a free server option is available; quotas, model availability, and server terms change, so read the project's current documentation rather than a blog post from any month, including this one. The code below is a teaching artifact written for the lab. I have not benchmarked it, and the timings are teaching budgets rather than measured results.

The artifact: one line per finished job

The design is deliberately boring. Each completed job appends one JSON line to ledger.jsonl, flushed and fsynced before the runner moves on. The job's identity is a hash of its own payload, so the same job presented twice produces the same key. On restart the runner reads the ledger into a set and skips anything already in it.

jobs.jsonl looks like this:

{"id": "j-001", "prompt": "Summarise this changelog entry in one line: ..."}
{"id": "j-002", "prompt": "Rewrite this error message for a non-technical reader: ..."}
Enter fullscreen mode Exit fullscreen mode

The runner is about forty lines. The request and response shapes depend on the endpoint, so students adapt call_model first and leave the ledger logic alone.

# runner.py -- teaching artifact, not production code
import hashlib, json, os, time, urllib.request
from pathlib import Path

LEDGER = Path("ledger.jsonl")
JOBS = Path("jobs.jsonl")
MODEL_URL = os.environ["MODEL_URL"]
API_KEY = os.environ.get("MODEL_KEY", "")


def key_of(job: dict) -> str:
    payload = json.dumps(job, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(payload.encode()).hexdigest()[:16]


def load_done() -> set:
    done = set()
    if not LEDGER.exists():
        return done
    for line in LEDGER.read_text(encoding="utf-8").splitlines():
        if not line.strip():
            continue
        try:
            record = json.loads(line)
        except json.JSONDecodeError:
            continue  # torn final write; let the job run again
        if record.get("status") == "ok":
            done.add(record["key"])
    return done


def append(record: dict) -> None:
    with LEDGER.open("a", encoding="utf-8") as fh:
        fh.write(json.dumps(record, sort_keys=True) + "\n")
        fh.flush()
        os.fsync(fh.fileno())


def call_model(prompt: str, timeout: int = 60) -> str:
    body = json.dumps({"prompt": prompt}).encode()
    req = urllib.request.Request(
        MODEL_URL,
        data=body,
        headers={"Content-Type": "application/json",
                 "Authorization": f"Bearer {API_KEY}"},
    )
    with urllib.request.urlopen(req, timeout=timeout) as resp:
        return json.loads(resp.read())["text"]


def run_one(job: dict) -> None:
    k = key_of(job)
    last_error = None
    for attempt in range(4):
        try:
            out = call_model(job["prompt"])
            append({"key": k, "job": job.get("id"), "status": "ok", "output": out})
            return
        except Exception as exc:  # noqa: BLE001 -- lab code on purpose
            last_error = exc
            time.sleep(2 ** attempt)
    append({"key": k, "job": job.get("id"),
            "status": "error", "error": str(last_error)})


def main() -> None:
    done = load_done()
    for line in JOBS.read_text(encoding="utf-8").splitlines():
        if not line.strip():
            continue
        job = json.loads(line)
        if key_of(job) in done:
            continue
        run_one(job)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Notice that load_done only counts records whose status is ok. Error records still land in the ledger for forensics, but they stay eligible for a later run. That single conditional is the difference between a retry policy and a silent data hole.

The ninety minutes

Minutes Segment
0-10 The 3 a.m. crash: what a lost batch actually costs
10-25 Read the ledger format, run forty jobs end to end
25-45 Exercise 1: kill -9 and resume
45-60 Exercise 2: the torn write
60-75 Exercise 3: two workers, one ledger
75-90 Debrief: what belongs in a record and what does not

Exercise 1: crash and resume

Students start the runner in the background, wait for a few seconds of work, and kill it hard.

python runner.py &
PID=$!
sleep 25
kill -9 $PID
wc -l ledger.jsonl    # some completed jobs, no summary
python runner.py      # resumes
Enter fullscreen mode Exit fullscreen mode

The proof step is the part that matters. Counting lines is not evidence, because duplicates also add lines.

python - <<'PY'
import collections, json
keys = [json.loads(l)["key"] for l in open("ledger.jsonl") if l.strip()]
c = collections.Counter(keys)
print("lines", len(keys), "unique", len(c))
print("duplicates", [k for k, v in c.items() if v > 1])
PY
Enter fullscreen mode Exit fullscreen mode

A clean run reports matching counts and an empty duplicate list. When a pair does appear, it is usually because the killed process had already sent the request and the restart sent it again before any ledger line existed; the fix is a claim file, which is Exercise 3.

Exercise 2: the torn write

kill -9 can land in the middle of fh.write, leaving a truncated last line. Students reproduce this deterministically without racing anything.

head -c -25 ledger.jsonl > ledger.tmp && mv ledger.tmp ledger.jsonl
python runner.py
Enter fullscreen mode Exit fullscreen mode

The JSON parse fails on the final line, load_done skips it, and the corresponding job reruns. Nobody has to reason about it in the abstract; they watch the line count grow by one and the duplicate check stay clean. The lesson worth writing on the board: a reader that tolerates garbage is more valuable than a writer that promises never to produce it.

Exercise 3: two workers, one ledger

Two terminals, one ledger, no coordination.

python runner.py & python runner.py &
wait
Enter fullscreen mode Exit fullscreen mode

Both processes read the same "done" set, both see the same missing key, and both call the model. The ledger now holds two records for one job. The repair is an exclusive claim created before the network call.

def claim(k: str) -> bool:
    os.makedirs("claims", exist_ok=True)
    try:
        fd = os.open(f"claims/{k}", os.O_CREAT | os.O_EXCL | os.O_WRONLY)
    except FileExistsError:
        return False
    os.close(fd)
    return True
Enter fullscreen mode Exit fullscreen mode

Then have the class break it again by killing a worker between the claim and the ledger append. The claim file survives, the job never reruns, and the batch quietly loses a result. That trade-off is the real content of the exercise: exclusivity buys you no duplicates and costs you a stale-claim policy, usually a timestamp older than a few minutes that may be reclaimed.

When a ledger is not worth the trouble

Situation Ledger? Reason
Forty jobs, two seconds each No Rerunning the batch is cheaper than maintaining state
Four hundred jobs, a minute each Yes Losing halfway through costs hours
Side effects outside the ledger Only with idempotent side effects A key stops duplicate calls, not duplicate emails
Regulated or personal data Not in plaintext Hash the prompt for the key; store raw text only where policy allows
The only copy lives on a free host Risky Convenience is not durability

The last row deserves the most discussion. A free server option is a fine place to run a ninety-minute lab, and a poor place to keep the only record of a week-long run. Ledgers belong on storage you control, with the compute pointed at it.

Who should skip this

If the work is interactive, with a human reading each answer before the next call, the ledger costs more attention than it saves. If the pipeline writes to a customer system, hashing the input is not enough, because the side effect itself has to be idempotent. If the batch carries personal data, a plaintext JSONL file on a shared machine is a compliance problem, not an engineering one. And if the team has no place to keep the ledger, the honest answer is to fix that before writing the runner.

Early finishers should spend the remaining time attacking the artifact rather than admiring it: kill the process during fsync, put the ledger on a network path two machines share, and watch what the append does. Breaking the file teaches more than the happy path ever will.

The real output of the workshop is not runner.py. It is the habit of asking, before a long run starts, what the world looks like when it dies at item 271. Free model access and a free server option lower the cost of setting that lab up, so check the current limits before you schedule the session, and keep the fixture set small enough that a rerun is never a crisis.

Top comments (0)