DEV Community

Charlie Zhu
Charlie Zhu

Posted on

The Dry-Run Card Workshop

A borrowed classroom on a weeknight has a particular silence. Twelve chairs face a projector, the building Wi-Fi is already tired, and the only machine that can run the class exercise belongs to a free server option nobody on the teaching team administers. The model route is free as well, which is why it was chosen. The transcript that seemed to prove it worked came from a different prompt, on a different day, in a different account.

That transcript is not a fixture. A dry-run card is the scrap a shop clips to a job before the first cut, and it records whether a known prompt, a known expectation, and a known runner still agree. This workshop spends ninety minutes teaching that habit, with a Python probe pairs can rerun from an empty directory.

MonkeyCode enters the hour only as one live route an operator may already have, with free model access and a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The exercise stays useful if that name is removed. Any documented HTTP route, and any short-lived remote shell the instructor is allowed to touch, can fill the same two blanks.

A bench, not a promise

Free capacity resembles a loaner lathe. It lets a small group start without buying iron, and it can be claimed by someone else, slowed by a queue, or withdrawn when the offer changes. This draft names neither models nor token allowances, and it names neither hardware nor an end date. Those figures move, and no primary source for them was attached here, so the card treats them as observations students write in the morning rather than as facts copied from an older post.

A portfolio page thrown together in an afternoon can look finished while the runner underneath was never probed. The pressure to ship that kind of surface is easy to recognize. The remedy in this room is narrower than a debate about taste: a fit check small enough to finish before the building closes.

How the ninety minutes are cut

The first ten minutes are rules, spoken once and written on the board. Keys stay in the shell, never in the repository, and production hostnames stay out. Homework and customer text stay out of prompts. The only statuses the room will accept are offline_pass, live_pass, live_skip, and fail.

A skip is a result. It means the live bench was not configured, or the network refused the favor. A skip is not a pass wearing a softer word.

The following fifteen minutes build the card as data rather than as a slide. Each pair agrees on a task id, a prompt, a short expected substring, a timeout, and a field for the runner. The prompt in the worked example asks for the digits of a sum a human can check without a model. That modesty is intentional, because a workshop that opens with a framework migration will spend the hour debugging the exercise instead of the probe.

From minute 25 to minute 50, pairs type the probe and run the offline path. The script is a worked example for Python 3.11 with only the standard library. It was not executed against a live MonkeyCode deployment for this article, and the offline branch is the part students can rerun on a train. The live branch is opt-in, and a missing variable fails closed.

#!/usr/bin/env python3
"""Dry-run card probe. Offline fixture always runs. Live mode is opt-in.

Teaching example only. Not executed against a live vendor deployment
when these notes were written. The /probe path is a seam, not an API claim.
"""

from __future__ import annotations

import hashlib
import json
import os
import shlex
import subprocess
import sys
import urllib.error
import urllib.request
from datetime import datetime, timezone

CARD = {
    'task_id': 'sum-two-ints',
    'prompt': 'Reply with only the digits of 17+25.',
    'expect': '42',
    'timeout_s': 20,
}


def prompt_hash(text):
    return hashlib.sha256(text.encode('utf-8')).hexdigest()[:12]


def offline_model(prompt):
    if prompt == CARD['prompt']:
        return '42'
    return 'unsure'


def live_model(prompt):
    base = os.environ['MODEL_BASE_URL'].rstrip('/')
    token = os.environ['MODEL_API_TOKEN']
    body = json.dumps({'prompt': prompt, 'max_tokens': 16}).encode('utf-8')
    req = urllib.request.Request(
        base + '/probe',
        data=body,
        headers={
            'Authorization': 'Bearer ' + token,
            'Content-Type': 'application/json',
            'Accept': 'application/json',
        },
        method='POST',
    )
    with urllib.request.urlopen(req, timeout=CARD['timeout_s']) as resp:
        payload = json.loads(resp.read().decode('utf-8'))
    return str(payload.get('text', '')).strip()


def runner_check():
    raw = os.environ.get('RUNNER_CMD', '').strip()
    if not raw:
        return {'status': 'live_skip', 'detail': 'RUNNER_CMD unset'}
    argv = shlex.split(raw)
    try:
        done = subprocess.run(
            argv,
            check=False,
            capture_output=True,
            text=True,
            timeout=CARD['timeout_s'],
        )
    except (OSError, subprocess.TimeoutExpired) as exc:
        return {'status': 'fail', 'detail': type(exc).__name__}
    text = (done.stdout or '').strip()
    ok = done.returncode == 0 and CARD['expect'] in text
    return {
        'status': 'live_pass' if ok else 'fail',
        'exit_code': done.returncode,
        'detail': text[:80],
    }


def judge(text):
    return 'pass' if CARD['expect'] in text else 'fail'


def main():
    mode = os.environ.get('PROBE_MODE', 'offline')
    error = ''
    try:
        if mode == 'live':
            answer = live_model(CARD['prompt'])
            model_status = 'live_' + judge(answer)
            runner = runner_check()
        elif mode == 'offline':
            answer = offline_model(CARD['prompt'])
            model_status = 'offline_' + judge(answer)
            runner = {'status': 'live_skip', 'detail': 'offline mode'}
        else:
            answer = ''
            model_status = 'fail'
            error = 'bad_mode'
            runner = {'status': 'live_skip', 'detail': 'bad mode'}
    except (KeyError, urllib.error.URLError, TimeoutError, json.JSONDecodeError) as exc:
        answer = ''
        model_status = 'fail'
        error = type(exc).__name__
        runner = {'status': 'live_skip', 'detail': 'model probe failed closed'}

    card = {
        'task_id': CARD['task_id'],
        'prompt_hash': prompt_hash(CARD['prompt']),
        'mode': mode,
        'model_status': model_status,
        'answer_excerpt': answer[:40],
        'error': error,
        'runner': runner,
        'recorded_at': datetime.now(timezone.utc).strftime('%Y-%m-%dT%H:%M:%SZ'),
    }
    json.dump(card, sys.stdout, indent=2)
    sys.stdout.write('\n')
    model_ok = model_status.endswith('pass')
    runner_ok = runner['status'] in ('live_pass', 'live_skip')
    return 0 if model_ok and runner_ok else 1


if __name__ == '__main__':
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Students run the offline cut before they touch a network. A healthy card shows model_status of offline_pass and a runner status of live_skip. That second field stops a pair from claiming they tested a free server when they only tested a dictionary in a file.

python3 dry_run_card.py > offline-card.json
echo exit:$?
python3 -m json.tool offline-card.json > /dev/null
Enter fullscreen mode Exit fullscreen mode

The live cut stays in the operator's own shell. The example host is intentionally invalid, and pairs substitute the base URL from documentation they opened that morning. If MonkeyCode is the route in use, that documentation is the source for the path, the auth header, and whether a free server is currently offered. The script's /probe path is a teaching seam, not a claim about a real API.

export PROBE_MODE=live
export MODEL_BASE_URL=https://example.invalid/v1
export MODEL_API_TOKEN=replace-me
export RUNNER_CMD='python3 -c print(17+25)'
python3 dry_run_card.py > live-card.json
echo exit:$?
Enter fullscreen mode Exit fullscreen mode

During class, RUNNER_CMD may be that local Python process, so a pair without SSH can still finish. A real free server, when the morning's notes say one exists, should be an instructor-supplied wrapper such as a BatchMode SSH command written on the board from that day's access sheet. This article does not invent the host, the user, or the flags. A wrapper that echoes the token, or that sends student homework to the runner, fails the shop rules even if the sum comes back right.

From minute 50 to minute 70, pairs sabotage their own fixture. They change the expected substring, unset the token, and point RUNNER_CMD at a command that exits nonzero. The model field and the runner field are allowed to fail separately. A courteous sentence that happens to contain the digits does not mean the free server executed anything, and a server that prints those digits does not mean the model route is awake.

unset MODEL_API_TOKEN
PROBE_MODE=live python3 dry_run_card.py > failed-card.json
echo exit:$?
Enter fullscreen mode Exit fullscreen mode

During that window the instructor walks the aisle with a paper copy of the card fields, not with a laptop full of answers. One pair will point the runner at a pipeline and discover that shlex refused the meaning they wanted from a pipe character. That refusal is the correct lesson. A free server is already a hallway the class does not own, and smuggling a shell pipeline into it on the same afternoon undoes the card.

Another pair will want to paste a class assignment into the prompt to make the demo feel real. The instructor sends them back to 17+25. Realism can wait until the gauge is boringly repeatable. A workshop that starts with an interesting bug usually ends with an interesting excuse.

The closing twenty minutes convert the card into a scheduling decision. Tomorrow's exercise may depend on the live route only when an offline pass is fresh and a live pass was recorded after the documentation check. A live skip means the assignment needs the offline fallback this workshop already is. A fail means the prompt shrinks, or the class edits by hand, because a borrowed bench can be busy.

Comparing two cards after a blip

A network blip around minute 60 is part of the lesson, not an interruption to apologize for. Pairs save the first offline card, run the live probe, and compare four fields only: mode, model status, runner status, and prompt hash. The hash is the mark that both cards describe the same prompt. A live card with a different hash is a different job, and it does not retire the offline result.

import json

def load(path):
    with open(path, encoding='utf-8') as handle:
        return json.load(handle)

def main():
    off = load('offline-card.json')
    live = load('live-card.json')
    same = off['prompt_hash'] == live['prompt_hash']
    print('same_prompt', same)
    print('offline', off['mode'], off['model_status'], off['runner']['status'])
    print('live', live['mode'], live['model_status'], live['runner']['status'])

if __name__ == '__main__':
    main()
Enter fullscreen mode Exit fullscreen mode
python3 compare_cards.py
Enter fullscreen mode Exit fullscreen mode

The comparison may be read even when the live process exited nonzero, which is why the room saved the JSON before trusting the exit code. That reading stays in the classroom. A gate in a shared pipeline should keep the nonzero exit, or a red probe will look green because someone added a shell bypass and forgot to take it out.

What the probe refuses to pretend

The substring check is a blunt gauge. It accepts the expected digits buried in a paragraph, and an exact-match variant will reject a trailing newline. Both are teaching choices, and neither certifies a patch, a migration, or a portfolio project. Pairs who want a harder second hour can require equality on a hash of the answer, then watch harmless formatting become a fail.

That frustration belongs in the workshop. It is cheaper there than in a gradebook. The runner hook splits a command the pair typed during class, and the instructor should say, while that command is still on the board, that the same pattern is a liability inside a service.

A card that executes text returned by a model is no longer a card. Adapting the request body to the documented call is part of the exercise. A mismatch should fail the card, because papering over it with a looser assertion teaches the wrong shop habit.

Who should leave this hour on the bench

An air-gapped course that must keep student source in the room should not aim prompts at a hosted model, free or otherwise. A team that needs a contract, a named region, or a budget it can forecast should rent or own the bench and write the allowance into a document it controls. Production deploys, credential handling, and any grade that cannot be repeated if the runner vanishes midweek sit outside this card. A community lab can live with a skip, and an on-call rotation cannot.

The operator describes the project as open source. The tree, when the class is actually using that route, is reading material and a place to file a mismatch. It is not a substitute for that morning's status page. Install steps belong to the tree's current instructions, and this article does not paraphrase them, because a paraphrased install is how a workshop inherits last month's flags.

Readers who already have a MonkeyCode account can aim the live probe at that route after they replace the teaching seam with the documented call. The offline fixture should stay in the repository either way. The next room can still finish the hour when the courtesy bench is closed.

Top comments (0)