A teaching lab can trust a free model swap only after three frozen prompts still score the same way. Students should be able to rerun those prompts next week and separate a wrong answer from a dead host. This session builds that separation as a fixture deck, a local scorer, and a short decision table. The deck remains useful when the remote seat is down, because stub mode never leaves the student laptop.
Recent teaching discussions keep circling the same practical question: a smaller or free host may be enough if a narrow check still passes. This workshop takes that question into a timed classroom clock instead of turning it into a product review. The check here is three frozen JSON rows, not a leaderboard and not a claim about any named model family. Students leave with a rerunnable gate, which is more useful for next week's lab than a headline about model choice.
This fixture deck grades structured answers rather than token spend, request latency, or separate tool-permission stamps. Each row names one prompt, the required JSON keys, a list of forbidden phrases, and a character ceiling. A later cohort can replace those three prompts without rewriting the scorer or the shared decision table. Transport errors stay in their own bucket so a timeout is never recorded as a content failure.
What a 65-minute pass actually proves
The lab proves that a small contract can catch obvious drift before a class depends on a new host. It does not prove that any model is generally correct, inexpensive, or stable across every course topic. Three rows are a teaching gate, not a benchmark suite and not a statement about vendor capacity. Instructors should treat every green pass as evidence about these fixtures alone, not about the wider course.
Constraints to state before minute eight
Say aloud that no product name, seat length, hardware shape, or quota number is assumed by these exercises. Students will author fixtures, run a stub, inject one deliberate drift, and only then consider a remote host. Prompts in the deck must use public sample text, never passwords, student records, or private repository contents. If the classroom cannot review a third-party terms page that morning, stop the plan after the local stub.
Session clock
Use the blocks below so the worked drift loop is not squeezed out by setup chatter. Each timed block has one visible output that the instructor can check on a nearby student screen. Do not add a fourth prompt during this sitting, because the constraint is part of the lesson. A cohort that wants more coverage should schedule a second sitting rather than stretching this clock.
Minutes 0–8: name the failure
Describe a lab where yesterday's JSON summary gained an extra field and then broke a downstream checker. Ask students which failure is content drift, which failure is transport loss, and which failure is a parser error. Write those three labels on the board before anyone opens an editor or calls a remote host. Those three labels become the only status values the scorer is allowed to emit during this session.
Minutes 8–22: author the deck
Students create a file named fixtures/deck.json that contains exactly three rows and no hidden fourth case. Each row carries an identifier, one prompt, the required keys, forbidden substrings, and a maximum character count. One row should demand a refusal to invent a version number that the source text never stated. That refusal row teaches the class that a fluent answer can still fail the written contract on purpose.
Minutes 22–40: prove the harness
Run the scorer in stub mode so the first green result does not depend on any network. The stub returns canned JSON that already satisfies every required key and avoids every forbidden phrase listed. Students should read the exit code, the status column, and the reason string before they edit anything. A green stub result only proves the harness, not the quality of any live model response.
Minutes 40–52: inject one drift
Change the stub for the endpoints row so the path field disappears while the method field remains. Rerun the same command and confirm that only that row becomes content drift while the other rows stay passing. Restore the stub and rerun the command once more so the deck returns to a clean baseline score. This short loop is the worked example that students can repeat later without a teacher in the room.
Minutes 52–65: record the swap rule
Fill the decision table before anyone discusses a remote host, a product seat, or a model name. Content drift blocks the planned swap until the fixture or the prompt is intentionally revised by the instructor. A transport failure blocks the swap for that class day, while a parser error blocks it until the expected shape is clarified. Only a full pass on these three rows allows an optional retarget of the same unchanged deck.
Fixture students commit
Save the following document as the only fixture input that this sitting is allowed to grade today. The prompts are public samples, and the ceilings are teaching limits rather than any product measurements. Keep the row identifiers stable so later diffs show prompt edits instead of merely renamed rows. If a student wants a different sample, replace the source sentence and leave the required keys in place.
{
"rows": [
{
"id": "summary",
"prompt": "Source: the patch note says retries now wait 200ms. Return a JSON object with keys summary and risk. Do not add fields.",
"required_keys": ["summary", "risk"],
"forbidden": ["password", "api_key"],
"max_chars": 240
},
{
"id": "endpoints",
"prompt": "Source: clients call GET /health. Return a JSON object with keys method and path. Do not invent other routes.",
"required_keys": ["method", "path"],
"forbidden": ["delete /", "drop table"],
"max_chars": 180
},
{
"id": "version",
"prompt": "Source text has no version number. Return a JSON object with keys status and reason. Refuse to invent a version.",
"required_keys": ["status", "reason"],
"forbidden": ["v1.2.3", "version 2"],
"max_chars": 200
}
]
}
Scorer students rerun
The script below is an unexecuted teaching artifact for the lab, not a measured client for any hosted product. Stub results are defined by the canned objects in this file, so a green run does not describe a live model. HTTP mode is a class-defined adapter with a placeholder path, and instructors must replace that path only after reading current host documentation. Do not treat the sample URL shape as a vendor interface or as evidence of uptime.
#!/usr/bin/env python3
"""Three-row fixture scorer for a classroom model-swap gate.
Label: proposed lab code. Status lines from stub mode are defined here.
They are not observations from a hosted model run.
"""
import argparse
import json
import sys
import urllib.error
import urllib.request
def load_deck(path):
with open(path, encoding="utf-8") as handle:
deck = json.load(handle)
rows = deck.get("rows")
if not isinstance(rows, list) or len(rows) != 3:
raise SystemExit("deck must contain exactly three rows")
return rows
def stub_body(row, drift_id):
canned = {
"summary": {"summary": "Patch notes add a retry hint.", "risk": "low"},
"endpoints": {"method": "GET", "path": "/health"},
"version": {"status": "refused", "reason": "source omits a version"},
}
if row["id"] not in canned:
raise SystemExit("stub has no canned body for " + row["id"])
body = dict(canned[row["id"]])
if drift_id == row["id"] and "path" in body:
body.pop("path")
return json.dumps(body)
def fetch_remote(base_url, prompt, timeout):
"""Placeholder adapter. Replace the path after reading current host docs."""
payload = json.dumps({"prompt": prompt}).encode("utf-8")
request = urllib.request.Request(
base_url.rstrip("/") + "/lab-complete",
data=payload,
headers={"Content-Type": "application/json"},
method="POST",
)
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
raw = response.read().decode("utf-8")
code = response.status
except (urllib.error.URLError, TimeoutError, OSError) as exc:
return None, "transport_failure: " + exc.__class__.__name__
if code != 200 or not raw.strip():
return None, "transport_failure: bad response"
return raw, None
def grade(row, raw):
try:
parsed = json.loads(raw)
except json.JSONDecodeError:
return "parser_error", "body is not json"
if not isinstance(parsed, dict):
return "parser_error", "body is not an object"
missing = [key for key in row["required_keys"] if key not in parsed]
if missing:
return "content_drift", "missing " + ",".join(missing)
blob = json.dumps(parsed).lower()
for phrase in row["forbidden"]:
if phrase.lower() in blob:
return "content_drift", "forbidden phrase present"
if len(blob) > int(row["max_chars"]):
return "content_drift", "over character ceiling"
return "content_pass", "keys present and ceiling held"
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--deck", required=True)
parser.add_argument("--mode", choices=["stub", "http"], default="stub")
parser.add_argument("--base-url", default="")
parser.add_argument("--timeout", type=float, default=20.0)
parser.add_argument("--drift", default="")
args = parser.parse_args()
failures = 0
for row in load_deck(args.deck):
if args.mode == "stub":
raw, err = stub_body(row, args.drift), None
else:
if not args.base_url:
raise SystemExit("http mode requires --base-url")
raw, err = fetch_remote(args.base_url, row["prompt"], args.timeout)
if err:
status, reason = "transport_failure", err
else:
status, reason = grade(row, raw)
if status != "content_pass":
failures += 1
print(f"{row['id']}|{status}|{reason}")
return 1 if failures else 0
if __name__ == "__main__":
sys.exit(main())
Expected lines from the worked example
The clean stub is defined to print three content_pass lines and to exit with code zero. The drifted stub is defined to print content_drift only for the endpoints row and to exit with code one. Students should compare their terminal text with those two patterns before they change any prompt wording. If the drifted run still exits zero, the harness is wrong and the remote step stays closed.
summary|content_pass|keys present and ceiling held
endpoints|content_pass|keys present and ceiling held
version|content_pass|keys present and ceiling held
summary|content_pass|keys present and ceiling held
endpoints|content_drift|missing path
version|content_pass|keys present and ceiling held
Commands for the two local passes
Run the clean command first and save the text beside the cohort date in the lab notes. Run the drift command second and confirm that the exit status changed from zero to one. Do not point HTTP mode at a host until both local patterns match the expected lines above. Replace the placeholder path in the adapter before any remote command is shared with the class.
python3 score_deck.py --deck fixtures/deck.json --mode stub
python3 score_deck.py --deck fixtures/deck.json --mode stub --drift endpoints
python3 score_deck.py --deck fixtures/deck.json --mode http --base-url "$LAB_BASE" --timeout 20
Decision table for the swap
| Observed status | What the row showed | Decision for this class day |
|---|---|---|
| content_pass | Required keys exist, forbidden phrases are absent, and the ceiling holds | Row may stay in the swap set |
| content_drift | A key is missing, a forbidden phrase appears, or the ceiling breaks | Block the swap until the instructor revises the row |
| parser_error | The body is missing, not JSON, or not an object | Block the swap until the expected shape is clarified |
| transport_failure | The call times out, returns a non-200, or comes back empty | Block the swap today and keep grading on the stub |
Read the table from top to bottom and stop at the first failing row rather than averaging the three results. A single drifted row is enough reason to keep the previous host in place for that class day. Parser errors should be fixed in the prompt notes before anyone retries a remote call again. Transport failures should be logged with the clock time and then left out of any accuracy discussion.
Where an optional free host fits
Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is relevant here only as one optional place to obtain free model access and a free server option for the retarget step. This plan names no models, quotas, hardware, duration, or permanence, and those facts must be read from current terms that morning. If those terms are missing, unclear, or incompatible with classroom policy, keep the lab on stub mode and still grade the deck.
After local statuses are green, a class that already has the free server option enabled can retarget this same unchanged deck. That retarget is optional, and a missing seat does not lower the local grade for the stub exercise. Students should paste only the public prompts from the deck, then compare status lines against the table above. Archive the status lines with the cohort date so the next section can see whether the host still matches the frozen rows.
Outputs the instructor collects
Collect four artifacts before the room is dismissed, and do not accept a screenshot of a chat window as a substitute. The deck file shows what was frozen, while the two transcripts show that drift is detectable. The written decision line shows whether a remote retarget is even eligible for that particular class day. Store the four artifacts together so a later cohort can rerun the same checks without reconstructing the lesson from memory.
- Students submit a three-row deck file that keeps stable identifiers and uses only public source sentences.
- Students submit a clean stub transcript that exits zero and shows three content_pass lines for every row.
- Students submit a drifted stub transcript that exits one and isolates the failure on the endpoints row.
- Students submit one decision line that either blocks the swap or records an optional retarget after a full pass.
Limitations and who should skip the lab
Exact key checks will miss a correct paraphrase that uses different field names or a nested object. Three public prompts cannot support a claim about accuracy, cost, latency, or uptime for a whole course. Do not send secrets, grades, or unpublished student work to any remote host, whether the seat is free or paid. Skip this workshop when you need a statistical evaluation, regulated medical or legal wording, or a graded exam that cannot tolerate a missing seat.
A team that already runs a large offline eval suite will gain little new procedure from a three-row gate. A classroom with no time to review third-party terms should not add the remote step at all. The scorer also ignores semantic similarity, so two answers can differ in wording and still need a human note. Those limits are intentional, because the sitting is about catching obvious contract breaks before a swap.
What to archive before the next cohort
Store the deck file, the scorer version, and the status lines from both the clean stub and the drifted stub. Add one sentence that names who may edit prompts and who may only rerun the commands. Do not archive credentials, host tokens, or private base URLs anywhere inside the shared class repository. The next cohort should diff the deck first, then decide whether a new row belongs in a later sitting.
Top comments (0)