Last Tuesday a campus library in Halifax felt like a waiting room for dying batteries. Mine hit 11 percent just as a chat tab handed me three slick sentences about precision and recall. They sounded sharp enough to paste into my notes. I closed the lid anyway.
Thursday the tab was gone, and a classmate asked for the counterexample. All I had was a polished paragraph, with no fixture, no command, and no expected fail. Have you ever tried to rebuild a lab from the feeling that it looked right?
This is the card I should have saved before the battery died.
{"id": "hidden-stem", "term": "precision", "sentence": "Precisionism is a style, not a metric.", "expect": false}
A substring check would accept it, because "precision" is hiding inside "Precisionism". The learning question is small enough to finish on a low battery. Can a tiny term gate reject that card, write a receipt, and still run after the chat window is gone?
The night the tab vanished
I study AI and CS in Halifax, and my notes grow faster than my reruns. A portfolio page can tell the story of a project. It cannot replace the file a classmate types into a terminal. The argument that a vibe-coded site is not the same thing as a portfolio landed for me in a smaller room than a job hunt. It landed in a group lab where the pretty explanation had no command.
This case is one evening and one gate. A definition card counts only when the term shows up as its own word and the sentence actually tries to define it. I am not fine-tuning anything. I am building the boring check I should have written before I asked a model to invent examples.
What Thursday-me needed
I wanted a future Tuesday to be boring. The fixture should diff in git. The checker should depend on Python, not on whichever tab is still open. And I wanted a written rule for what a free model may draft versus what a free server may hold.
Why write that rule down when I am tired? Because last month I let a chat invent both the sentence and the expected bit, then I felt clever for an hour. A draft is a suggestion. A shelf is a copy. Neither one is the label.
Prerequisites are plain. Python 3.11 or newer, standard library only. I used json, hashlib, pathlib, and sys. There is no notebook kernel, no key in the file, and no package install. If python3 --version prints something older than 3.11, the list[str] hints are the first thing to simplify, not the gate itself.
A gate mean enough to trust
The gate is intentionally mean. It lowercases tokens, strips a short set of punctuation, and passes a card only when the term is a whole token, a marker such as is, means, or equals is present, and the sentence has at least four tokens.
"Precision matters" contains the term and still is not a definition. Should a lab accept it? I do not think so. Friendly sentences are how student checkers learn to lie.
#!/usr/bin/env python3
"""Tiny definition-card gate. Stdlib only. Python 3.11+."""
import hashlib
import json
import sys
from pathlib import Path
MARKERS = {"is", "means", "equals"}
def tokens(sentence: str) -> list[str]:
return [w.strip(".,:;!?").lower() for w in sentence.split() if w.strip(".,:;!?")]
def gate(card: dict) -> bool:
term = card["term"].strip().lower()
words = tokens(card["sentence"])
return term in words and any(m in words for m in MARKERS) and len(words) >= 4
def main() -> int:
path = Path(sys.argv[1])
fixture = json.loads(path.read_text(encoding="utf-8"))
rows = []
for card in fixture:
ok = gate(card)
rows.append({"id": card["id"], "ok": ok, "expect": bool(card["expect"])})
mismatches = [r for r in rows if r["ok"] != r["expect"]]
digest = hashlib.sha256(path.read_bytes()).hexdigest()[:12]
report = {"fixture_sha": digest, "n": len(rows), "mismatches": mismatches}
print(json.dumps(report, indent=2))
return 1 if mismatches else 0
if __name__ == "__main__":
raise SystemExit(main())
I labeled the expected bits by hand before any draft existed. That order is the experiment. Reverse it, and you are scoring a mirror.
[
{
"id": "ok-precision",
"term": "precision",
"sentence": "Precision is the share of positive calls that were right.",
"expect": true
},
{
"id": "hidden-stem",
"term": "precision",
"sentence": "Precisionism is a style, not a metric.",
"expect": false
},
{
"id": "missing-term",
"term": "recall",
"sentence": "It is the share of real positives you caught.",
"expect": false
},
{
"id": "no-marker",
"term": "recall",
"sentence": "Recall catches more of the real positives today.",
"expect": false
}
]
Save the script as gate.py and the fixture as cards.json. Then run the command you can repeat on a classmate's machine.
python3 gate.py cards.json
echo "exit=$?"
I walked the four cards through gate by hand. This is not a captured log from a remote machine, so I will not invent a hash. On this fixture you should see exit code 0, "n": 4, and "mismatches": []. The fixture_sha field should be the first twelve hex characters of the SHA-256 of cards.json. If your hash differs, the file differs. That is the receipt doing its job.
Now feed it a bad label. Flip hidden-stem so "expect" is true, and run the same command. The stem is still buried inside another word, but your label claims the gate should pass. You should get exit code 1 and one mismatch row, {"id": "hidden-stem", "ok": false, "expect": true}. A green run only says the labels and the gate agree. It does not say the English is true.
"Precision is a kind of fruit" would pass this gate. Shape is not knowledge. Did you notice how easy it is to forget that?
A wrong path is a different failure, and it should look different. python3 gate.py missing.json raises FileNotFoundError from pathlib. A trailing comma in the JSON raises JSONDecodeError. Neither of those is a failed definition. If you treat every traceback as a model problem, you will debug the wrong layer for an hour.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is the open-source project that outreach names, and the two availability claims I am willing to use are free model access and a free server option. I am not naming a model, quoting a token quota, or promising hardware, latency, or how long a free server stays up. I have not re-checked a primary page for those figures this week. Stale numbers are how student write-ups go rotten.
If I use that project for this lab at all, free model access is only the draft bench, and the free server option is only a shelf. After the four cards exist, I can ask for one more sentence that should fail, then I add it only if I hand-label it and the command agrees. The files are small enough to park where a classmate can fetch them and run python3 gate.py cards.json when my laptop is in a bag. I am not shipping an API client here. I do not have a verified base URL or request shape I am willing to freeze into a tutorial, and a guessed call would teach the wrong habit.
Four cards, no score
Hand-walking the fixture, ok-precision passes because the term is its own token and is is present. hidden-stem fails, which is the bug a term in sentence check would miss. missing-term fails because "recall" never appears as a token, even though a person can guess the topic. no-marker fails because the term is present and the sentence still is not a definition.
The result I want is smaller than a score. Thursday-me can type one command and see the same mismatches. The chat tab is allowed to vanish. The receipt is not. A fluent draft of "Precisionism is a style" can still sound like a lesson. The gate does not grade tone. It grades tokens. That is a claim I can put in a lab report without blushing.
What did I not measure? Latency, cost, and whether a hosted draft would have proposed hidden-stem on its own. I am not going to backfill those numbers. A case study that invents a benchmark is just a brochure with a code block taped on.
The shelf is not the proof
I would not point this workflow at a graded exam, a classmate's private write-up, or any text I would be embarrassed to see in a shared terminal. A free model is still a remote system. A free server is still someone else's disk. Keys, student records, and unpublished assignments stay off both.
I also would not treat free access as a promise. I have no measurement of quota, cold start, or how long a shelf remains. If the copy on a server disappears, the git commit is the lab. The URL was a convenience I can lose. Who should skip the approach? Anyone who needs a trained classifier, a private evaluation set, or a statistical claim. This gate will accept a false definition that happens to contain the word and the marker. If your course is about truth, you still have to read the sentence yourself.
Common slips, from the version of me who rushes, are using in on the raw string, letting the draft write expect, and treating the first green exit as understanding. Another quiet one is punctuation I did not strip. precision's will not match the term precision. If your course notes use possessives, the gate is too blunt, and you should say so in the report instead of hiding it.
Drafting can be cheap. Proof should be local. A borrowed shelf can hold a public fixture. The label stays mine. Would I put a hosted model inside gate()? No. The critical path is the command. The hosted draft is for the moment I am stuck inventing a counterexample and I am willing to throw it away.
What should you understand when the script exits 0? Only that these cards and this gate agree. What should you understand when it exits 1? That a label and a token rule disagreed, and you can see which id. That is the transferable idea. Receipts beat memories.
An extension that still fits in one sitting: add a card whose term is f1 and whose sentence says "The F1 is the harmonic mean of precision and recall." Predict the result before you run it. The capital letter is easy. The token f1 versus F1 is the part I would not trust my memory to settle.
If you already have a free model tab and somewhere small to park a public file, try one extra card before you trust either of them. Predict the mismatch, run the script, and keep the git copy as the one that matters. I would rather see a sentence this gate gets wrong than another polished paragraph with nowhere to rerun it.
Top comments (0)