Second run, same bytes. Otherwise I mark a zero. That is the conclusion, and I am not interested in how smooth the demo felt.
If src/ still moves when the fixer runs again, the submission fails. Why would I pass a script that only behaves on a fresh tree?
The failure I want in the room
Bootcamp agents are fluent on the first edit. They are sloppy on the second.
You ask for a return-string change, run the script, and the file looks done. Run it again. Did it append another helper? Did it flip quotes? Did it "clean up" a comment it just wrote? Then you do not own a fix. You own a generator with no stop condition.
This lab asks one question. After the change is already present, does another run stay still? If you cannot show that with a hash, do not submit.
Setup
Build a fixture so small that excuses sound silly. I would rather debug a stop condition than a fake product.
Commands I put on the board
- Create
lab-twice/and initialize git. - Add
src/greet.pythat returns"hello". - Commit it with the message
freeze base. - Add
student_fix.py. That file is what I grade. - Add
tools/grade_twice.pyfrom the scaffold below. Call it an unexecuted lab proposal, not a measured class result.
mkdir -p lab-twice/src lab-twice/tools
cd lab-twice
git init
printf 'def greet():\n return "hello"\n' > src/greet.py
git add src/greet.py
git commit -m "freeze base"
The target behavior is tiny on purpose. Make greet() return "hello, lab", and do it so a second run does not change one byte under src/.
Why tiny? I want the miss to be idempotency, not domain trivia. Enlarge the fixture next week if you must. Starting large hides the bug inside a story.
One gate before any model call: git status --porcelain is empty. A dirty start is a zero. I will not sort your scratch files from your submission.
Checkpoint 1 — hand-write the stop condition
Ten minutes. No model. Can you stop yourself before you stop a chatbot?
# student_fix.py — illustrative submission, not a class result
from pathlib import Path
PATH = Path("src/greet.py")
OLD = 'return "hello"'
NEW = 'return "hello, lab"'
def main() -> None:
text = PATH.read_text()
if NEW in text:
return
if OLD not in text:
raise SystemExit("fixture not recognized")
PATH.write_text(text.replace(OLD, NEW, 1))
if __name__ == "__main__":
main()
See the early return? That branch is the lab. A blind rewrite can still be stable, if and only if the template itself never drifts.
What happens if you append a helper whenever your check is slightly wrong? The second hash moves, and the grade falls over. Which design are you actually choosing: rewrite-from-template, or edit-if-absent? Pick one. Make run two prove it.
Checkpoint 2 — grade the rerun, not the session
Students may sketch a second draft with a hosted model. Disclosure: This article was prepared as part of MonkeyCode's product outreach. For this lab I rely on MonkeyCode only for free model access and the free server option, so a first sketch does not require a private runner or a paid seat.
I am not stating a token quota, a machine size, a speed number, or a promise that the free option lasts. If a syllabus needs those figures, read that week's docs. Do not copy a number out of a chat and into a rubric.
The session is optional. If the tab dies, write the script locally. I would not move a deadline because a draft bench blinked.
What the checker proves
The script below is a scaffold. I have not run it as a study, and its print line is not a benchmark. It assumes Python 3.9+ and a clean git worktree.
#!/usr/bin/env python3
"""Run student_fix.py twice. Unexecuted lab scaffold."""
import hashlib
import subprocess
import sys
from pathlib import Path
def src_hash() -> str:
root = Path("src")
if not root.is_dir():
sys.exit("ZERO: src/ is missing")
digest = hashlib.sha256()
files = sorted(p for p in root.rglob("*") if p.is_file())
for path in files:
digest.update(path.relative_to(root).as_posix().encode())
digest.update(b"\0")
digest.update(path.read_bytes())
return digest.hexdigest()
def porcelain() -> str:
proc = subprocess.run(
["git", "status", "--porcelain"],
check=True,
text=True,
capture_output=True,
)
return proc.stdout.strip()
def changed_paths(status: str):
paths = []
for line in status.splitlines():
# Two status characters, one space, then the path.
# Renames and quoted paths are outside this scaffold.
if len(line) < 4:
continue
paths.append(line[3:])
return paths
def run_fix() -> None:
proc = subprocess.run([sys.executable, "student_fix.py"], check=False)
if proc.returncode != 0:
sys.exit("ZERO: student_fix.py failed")
def main() -> None:
if porcelain():
sys.exit("ZERO: dirty tree at start")
before = src_hash()
run_fix()
mid = src_hash()
if mid == before:
sys.exit("ZERO: first run did not change src/")
run_fix()
after = src_hash()
if after != mid:
sys.exit("ZERO: second run changed src/")
for path in changed_paths(porcelain()):
if not path.startswith("src/"):
sys.exit("ZERO: unexpected path " + path)
print("PASS", mid)
if __name__ == "__main__":
main()
python3 tools/grade_twice.py
A pass means the start was clean, run one changed src/, and run two did not. Status also did not grow a file outside src/.
Does a pass mean the string is right? No. Hash stability is not an oracle. I still open src/greet.py and read it. Keep that read on its own rubric row, or a stable wrong answer will borrow points from a clean rerun.
Checkpoint 3 — force a zero
Sabotage the script before you trust the grader. If you cannot make it fail, you do not understand the points.
- Delete the early return and append the new string on every run. The second hash should move.
- Also write
NOTES.txt. The path check should fire. - Leave an uncommitted edit, then start the grader. The dirty-start check should fire.
Did the sabotaged script still print PASS? Then the grader is the bug. Fix that before you argue with a model.
Stretch goals
Try these only after the three checkpoints are honest.
- Run a third time. Same hash, or it is still a zero.
- Refuse the submission if
student_fix.pycontainshttp://orhttps://. A fixer that calls out to a host is not a lab artifact. - Cap the diff with
git diff --numstat. I would use 40 changed lines as a teaching cap, not as a law of nature. Over the cap, stop and ask why the draft rewrote the file. - Preload the correct string, run the grader, and confirm a no-op first run is a zero. Doing nothing is not the same as being safe to rerun.
How I would mark it
| Check | Points | Zero when |
|---|---|---|
| Clean tree before run 1 | 15 | Porcelain was not empty |
Run 1 changes only src/
|
25 | No hash change, crash, or extra path |
Run 2 keeps the src/ hash |
35 | Any byte under src/ changes |
Human read sees "hello, lab"
|
15 | Stable, but wrong |
| Sabotage notes | 10 | They never produced a real zero |
No points for a screenshot. No points for "the model agreed." Partial credit stops at the first zero in that row. I am not averaging vibes.
Oral check
I would ask three questions, out loud, with the file open.
- Where does the script stop?
- Which paths may change?
- If
src/greet.pyis missing, do you recreate it or exit?
Recreate can be idempotent. Only if run two writes the same bytes and run three writes nothing new. If you choose recreate, add that case to the checker. Do not leave it as a feeling.
Decision list when you are tired
Use this when the hour gets noisy. Do not invent a fourth outcome.
- Start dirty? Zero. Do not tidy up and pretend the start was clean.
- First hash equals the base hash? Zero. Empty work is not a pass.
- Second hash differs? Zero. That is the whole lab.
- Unexpected path? Zero, even if
src/was stable. - String wrong, hashes stable? Lose the 15 behavior points. Keep the idempotency points.
- Hosted tab died? Irrelevant. Grade the file on disk.
Who should not use this
Do not hash-grade work that should change every run. Timestamps, random IDs, and training logs fail for honest reasons. You would be punishing the assignment, not the student.
Do not use this on a one-shot migration. Those should refuse a second apply with a lock, and side files may still move. Teach the lock. Do not smash it into this rubric.
Do not point the grader at secrets. Hashing bytes is fine. Printing a credential while you debug is not. Would you paste that terminal into a model session? Then do not debug that way.
Skip it when you meant to grade taste. This rubric punishes drift, not dullness. A careful one-line edit can outscore a clever rewrite. That is the deal.
Limits, said plainly
Free model access will not insert a stop condition you forgot. It also must not sit on the grading path. The free server is a draft bench, not a witness, and uptime is not part of the score.
This scaffold does not understand renames, quoted git status paths, or binary fixtures. Say that in the README. Otherwise a filename with a space becomes a fake model failure.
I do not have a measured pass rate for this write-up. Run the scaffold on your own fixture before lab day, and change the path rule if you assigned something wider than src/.
A green second run can still be the wrong function. Schedule a later lab with a human oracle. Do not fold the two scores into one mood.
Close the tab, keep the script
If you want a sandbox for the draft, use that free model on the free server, save student_fix.py, and close the session.
Then answer the only question that counts. You ran it twice. Did src/ stay still?
Top comments (0)