If I cannot swap the free model and still give you the same score, the score was never about your code. It was about a lucky sentence. This lab exists to kill that kind of grade.
You already know the failure mode. Two laptops. Same fixture. Same prompt. One free model emits a tidy JSON object, and the other buries the same fields inside a cheerful paragraph. One student "passes." The other eats a zero for vibes. Fair? No.
So here is the rule I would print on the whiteboard before anyone opens a chat box. Grade the pin and the structure, not the prose. If the run cannot name the model id and the server image, it does not count. If a second model on that same image flips your pass/fail bit, you do not get credit for the prettier demo.
I have not executed this harness against a live cohort in this draft. Treat the script as a proposed lab starter, not as a benchmark, a leaderboard, or a promise about any vendor plan.
The grade I actually want
You are not proving that a model is smart. You are proving three smaller, meaner things.
- The fixture is frozen.
- The checker ignores wording.
- The run can be named, hashed, and repeated after a model swap.
That is a one-week bootcamp lab. It is not a production eval, and it is not a beauty contest between free models. Why start there? Because a cohort that cannot separate "the model spoke" from "the harness checked" will grade luck for the rest of the term.
What I will not argue about
I will not argue about tone. I will not argue about which reply "felt more senior." I will argue about missing keys, empty strings, a floating image tag, and a fixture whose hash moved after the run.
If that sounds harsh, good. Harsh and checkable beats kind and fuzzy.
Setup, before the model speaks
I want one shared box, not twelve snowflake laptops. A free server is useful here only when every student can land on the same image digest. A free model list is useful only when the manifest can name the id you actually called.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The operator behind this draft describes MonkeyCode as an open-source project and points at free model access plus a free server option. That is the whole product claim I am willing to put in a rubric. I am not pasting model names, token quotas, hardware sizes, or "how long free lasts." Those are live-plan facts. If the docs on lab morning disagree with a screenshot in Slack, the docs win, and the screenshot loses.
Make a fresh directory. Do not reuse last week's agent workspace.
mkdir -p swaplab/fixtures swaplab/runs swaplab/checker
cd swaplab
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
printf '%s\n' 'pytest==8.3.4' > requirements.txt
python -m pip install -r requirements.txt
python -c 'import sys; print(sys.version)'
Pin the interpreter in the manifest too. I do not care which patch you have. I care that the grade names it.
export LAB_MODEL_ID="${LAB_MODEL_ID:?set the model id from today's free list}"
export LAB_SERVER_IMAGE="${LAB_SERVER_IMAGE:?set an image digest, not a floating tag}"
export LAB_FIXTURE="fixtures/ticket.json"
See the :?? Forget the pin and the shell dies before the model spends a single token. That is checkpoint zero, and it is rude on purpose.
Secrets stay out of the repo. If your client needs a key, it comes from the server environment, not from a committed .env. I am not teaching a leak scanner this week. I am refusing to grade a diff that already contains one.
The fixture stays boring
One ticket. No customer data. No real company. The model must return a JSON object with owner, severity, and next_step. Everything else is trash, not signal.
{
"id": "T-14",
"title": "Checkout button ignores keyboard focus",
"notes": "Reported on the staging box. No customer data in this fixture."
}
Save that as fixtures/ticket.json. Do not "improve" it mid-lab. A moving fixture is a moving grade, and I will not chase it.
Hash it before the first run.
sha256sum fixtures/ticket.json | tee runs/fixture.sha256
Why a digest, not a tag
latest is a mood, not a pin. If I rebuild the free server overnight and your grade still says latest, I cannot tell whether you passed on Tuesday's image or Thursday's. A digest is ugly. Ugly is the point. Can you replay the run without asking the student what they "think" they used? If not, the manifest is decoration.
The checker
This is the artifact. It does not call a vendor SDK, and it does not know MonkeyCode's API. You bring a tiny adapter that returns a string. The checker only accepts a JSON object with the three required keys. Extra keys are fine. Missing keys are not. Prose is not.
Save it as checker/grade.py.
#!/usr/bin/env python3
"""Proposed lab starter. Not executed against a live model in this draft."""
import hashlib, json, os, sys
from pathlib import Path
REQUIRED = ("owner", "severity", "next_step")
SEVERITIES = {"low", "medium", "high"}
def sha256(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()
def extract_object(raw: str) -> dict:
raw = raw.strip()
start, end = raw.find("{"), raw.rfind("}")
if start < 0 or end < start:
raise ValueError("no JSON object")
obj = json.loads(raw[start:end + 1])
if not isinstance(obj, dict):
raise ValueError("JSON was not an object")
missing = [k for k in REQUIRED if k not in obj]
if missing:
raise ValueError(f"missing {missing}")
if str(obj["severity"]).lower() not in SEVERITIES:
raise ValueError("severity not in allowed set")
if not str(obj["owner"]).strip() or not str(obj["next_step"]).strip():
raise ValueError("empty owner or next_step")
return obj
def main() -> int:
model = os.environ["LAB_MODEL_ID"]
image = os.environ["LAB_SERVER_IMAGE"]
fixture = Path(os.environ["LAB_FIXTURE"])
obj = extract_object(sys.stdin.read())
manifest = {
"model_id": model,
"server_image": image,
"fixture_sha256": sha256(fixture),
"python": sys.version.split()[0],
"fields": sorted(obj),
}
out = Path("runs") / f"{model.replace('/', '_')}.json"
out.parent.mkdir(exist_ok=True)
out.write_text(json.dumps(manifest, indent=2) + "\n")
print(json.dumps({"ok": True, "manifest": str(out)}))
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except (KeyError, ValueError, json.JSONDecodeError) as exc:
print(json.dumps({"ok": False, "error": str(exc)}))
raise SystemExit(1)
Run the local cases first. No network. No model. If these fail, you are not ready to touch a free quota.
# pass: bare object
printf '%s\n' '{"owner":"sam","severity":"high","next_step":"restore focus ring"}' \
| python checker/grade.py
# pass: fenced blob, wording ignored
printf '%s\n' 'Sure, here you go:' \
'{"owner":"sam","severity":"high","next_step":"restore focus ring"}' \
'Hope that helps!' | python checker/grade.py
# fail: prose only
printf '%s\n' 'The button is bad, please fix focus.' | python checker/grade.py
# fail: empty next_step
printf '%s\n' '{"owner":"sam","severity":"high","next_step":" "}' \
| python checker/grade.py
The prose case must exit non-zero. If you "helpfully" pass any non-empty string, I grade the checker, not the model. You built a compliment machine. I do not want one.
Checkpoints
Do these in order. A shiny demo does not unlock the next box.
-
Pin or stop.
LAB_MODEL_IDandLAB_SERVER_IMAGEare set. The image is a digest, notlatest. Showenv | grep '^LAB_'. -
Fixture hash. The SHA-256 in
runs/fixture.sha256matchesfixture_sha256after a stdin run. -
Structure only. Fenced JSON passes. Prose fails. Empty
next_stepfails.severity: "urgent"fails, because it is not in the allowed set. - Swap drill. Same fixture, same image digest, two model ids from today's free list. Both manifests exist. Both pass, or both fail, for the same assertion. A pass/fail split is a lab fail.
-
Clean tree.
git grep -nE 'api_key|BEGIN PRIVATE KEY|sk-'is empty. No committed token budget either. A hardcoded quota is a smell, not a feature.
Checkpoint 4 is the whole week. A warmer voice does not earn a higher score. A pass on model A and a fail on model B means your harness was coupled to one voice. Fix the contract, or drop the assertion that only one model can satisfy. Do not "pick the winner" and hide the other manifest. I will ask for both files.
A conversation I do not want
"But model B is worse, so my code is fine." Maybe. Maybe not. This lab cannot see that. It can see that you graded a single lucky reply and called it engineering.
Want to discuss model quality? Fine, after the swap drill, in a separate note, with both manifests attached. That note is not the grade.
Stretch goals
Pick one. I would rather see one finished drill than three half demos.
- Add
blocked_byas an optional field. Do not loosenseverity. Update one fixture and the checker together. - Record elapsed milliseconds in the manifest, and do not grade on it. Free servers hiccup. A slow cold start is not a character flaw.
- Add a 15-line pytest that feeds three stdin strings into
extract_object. Those tests grade the checker. They do not call a model.
If a test needs the network, it is an integration drill. It does not belong in the default score. Label it, skip it in CI, and keep the unit path offline.
Rubric
| Checkpoint | Full marks | Zero for that row |
|---|---|---|
| Pins | Model id plus image digest in the manifest | Floating tag, missing env, or a made-up quota |
| Fixture | SHA-256 matches the file you hashed | Fixture edited after the run |
| Checker | Fenced JSON passes; prose and empty fields fail | Any non-empty string passes |
| Swap | Two ids, same pass/fail, both manifests saved | Only the pretty model was shown |
| Hygiene | No secrets and no hardcoded budget in the diff | A key, or a token number copied from a blog |
Partial credit exists, but not for charm. A broken swap with a clean checker is a 60. A perfect demo on one model, with the second manifest "lost," is a zero on that row. I am not negotiating with a screenshot of a chat window.
Limits, and who should skip this
This will not tell you which free model is best. I am not publishing latency numbers, token prices, or hardware, because I will not invent figures a live plan can contradict next week.
Do not use this lab for semantic grading. "Is next_step actually a good fix?" is a human question. The checker only asks whether the field exists, is non-empty, and sits in an allowed set. That is a floor. It is not a code review.
Do not use it as a security audit. The fixture has no customer data on purpose. If your notes include a real inbox, a real company, or a real secret, stop and strip them before any model call.
Do not park production traffic on a teaching server. A free image can sleep, reset, or change terms. That is a limitation, not a footnote you skip in the README.
Skip the lab entirely if the cohort cannot yet read a JSON error. Teach json.loads first. A swap drill on top of panic is just noise with extra steps.
Also skip it if you need a stable score across months. Free model lists move. When an id disappears, the lab pauses. You do not substitute a guess, and you do not edit yesterday's manifest to match today's menu.
What I would actually run
I would put the venv and the checker on the free server so the image digest is real, not "works on my laptop." I would copy two ids from the current free model list into the shell for checkpoint 4, then throw the list away at the end of the day. No id gets baked into the repo. No quota gets baked into the script.
Before you photocopy this rubric, open the current free-model and free-server notes, cross out anything this draft refused to invent, and run checkpoint 4 with whatever two ids are actually listed that morning. The checkpoint your cohort fails is the part worth keeping.
Top comments (0)