DEV Community

Charlie Xu
Charlie Xu

Posted on

Second Run, Same Bytes: A Bootcamp Lab on Idempotent Fixes

Second run, same bytes. Otherwise I mark a zero. That is the conclusion, and I am not interested in how smooth the demo felt.

If src/ still moves when the fixer runs again, the submission fails. Why would I pass a script that only behaves on a fresh tree?

The failure I want in the room

Bootcamp agents are fluent on the first edit. They are sloppy on the second.

You ask for a return-string change, run the script, and the file looks done. Run it again. Did it append another helper? Did it flip quotes? Did it "clean up" a comment it just wrote? Then you do not own a fix. You own a generator with no stop condition.

This lab asks one question. After the change is already present, does another run stay still? If you cannot show that with a hash, do not submit.

Setup

Build a fixture so small that excuses sound silly. I would rather debug a stop condition than a fake product.

Commands I put on the board

  1. Create lab-twice/ and initialize git.
  2. Add src/greet.py that returns "hello".
  3. Commit it with the message freeze base.
  4. Add student_fix.py. That file is what I grade.
  5. Add tools/grade_twice.py from the scaffold below. Call it an unexecuted lab proposal, not a measured class result.
mkdir -p lab-twice/src lab-twice/tools
cd lab-twice
git init
printf 'def greet():\n    return "hello"\n' > src/greet.py
git add src/greet.py
git commit -m "freeze base"
Enter fullscreen mode Exit fullscreen mode

The target behavior is tiny on purpose. Make greet() return "hello, lab", and do it so a second run does not change one byte under src/.

Why tiny? I want the miss to be idempotency, not domain trivia. Enlarge the fixture next week if you must. Starting large hides the bug inside a story.

One gate before any model call: git status --porcelain is empty. A dirty start is a zero. I will not sort your scratch files from your submission.

Checkpoint 1 — hand-write the stop condition

Ten minutes. No model. Can you stop yourself before you stop a chatbot?

# student_fix.py — illustrative submission, not a class result
from pathlib import Path

PATH = Path("src/greet.py")
OLD = 'return "hello"'
NEW = 'return "hello, lab"'

def main() -> None:
    text = PATH.read_text()
    if NEW in text:
        return
    if OLD not in text:
        raise SystemExit("fixture not recognized")
    PATH.write_text(text.replace(OLD, NEW, 1))

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

See the early return? That branch is the lab. A blind rewrite can still be stable, if and only if the template itself never drifts.

What happens if you append a helper whenever your check is slightly wrong? The second hash moves, and the grade falls over. Which design are you actually choosing: rewrite-from-template, or edit-if-absent? Pick one. Make run two prove it.

Checkpoint 2 — grade the rerun, not the session

Students may sketch a second draft with a hosted model. Disclosure: This article was prepared as part of MonkeyCode's product outreach. For this lab I rely on MonkeyCode only for free model access and the free server option, so a first sketch does not require a private runner or a paid seat.

I am not stating a token quota, a machine size, a speed number, or a promise that the free option lasts. If a syllabus needs those figures, read that week's docs. Do not copy a number out of a chat and into a rubric.

The session is optional. If the tab dies, write the script locally. I would not move a deadline because a draft bench blinked.

What the checker proves

The script below is a scaffold. I have not run it as a study, and its print line is not a benchmark. It assumes Python 3.9+ and a clean git worktree.

#!/usr/bin/env python3
"""Run student_fix.py twice. Unexecuted lab scaffold."""
import hashlib
import subprocess
import sys
from pathlib import Path

def src_hash() -> str:
    root = Path("src")
    if not root.is_dir():
        sys.exit("ZERO: src/ is missing")
    digest = hashlib.sha256()
    files = sorted(p for p in root.rglob("*") if p.is_file())
    for path in files:
        digest.update(path.relative_to(root).as_posix().encode())
        digest.update(b"\0")
        digest.update(path.read_bytes())
    return digest.hexdigest()

def porcelain() -> str:
    proc = subprocess.run(
        ["git", "status", "--porcelain"],
        check=True,
        text=True,
        capture_output=True,
    )
    return proc.stdout.strip()

def changed_paths(status: str):
    paths = []
    for line in status.splitlines():
        # Two status characters, one space, then the path.
        # Renames and quoted paths are outside this scaffold.
        if len(line) < 4:
            continue
        paths.append(line[3:])
    return paths

def run_fix() -> None:
    proc = subprocess.run([sys.executable, "student_fix.py"], check=False)
    if proc.returncode != 0:
        sys.exit("ZERO: student_fix.py failed")

def main() -> None:
    if porcelain():
        sys.exit("ZERO: dirty tree at start")
    before = src_hash()
    run_fix()
    mid = src_hash()
    if mid == before:
        sys.exit("ZERO: first run did not change src/")
    run_fix()
    after = src_hash()
    if after != mid:
        sys.exit("ZERO: second run changed src/")
    for path in changed_paths(porcelain()):
        if not path.startswith("src/"):
            sys.exit("ZERO: unexpected path " + path)
    print("PASS", mid)

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode
python3 tools/grade_twice.py
Enter fullscreen mode Exit fullscreen mode

A pass means the start was clean, run one changed src/, and run two did not. Status also did not grow a file outside src/.

Does a pass mean the string is right? No. Hash stability is not an oracle. I still open src/greet.py and read it. Keep that read on its own rubric row, or a stable wrong answer will borrow points from a clean rerun.

Checkpoint 3 — force a zero

Sabotage the script before you trust the grader. If you cannot make it fail, you do not understand the points.

  • Delete the early return and append the new string on every run. The second hash should move.
  • Also write NOTES.txt. The path check should fire.
  • Leave an uncommitted edit, then start the grader. The dirty-start check should fire.

Did the sabotaged script still print PASS? Then the grader is the bug. Fix that before you argue with a model.

Stretch goals

Try these only after the three checkpoints are honest.

  • Run a third time. Same hash, or it is still a zero.
  • Refuse the submission if student_fix.py contains http:// or https://. A fixer that calls out to a host is not a lab artifact.
  • Cap the diff with git diff --numstat. I would use 40 changed lines as a teaching cap, not as a law of nature. Over the cap, stop and ask why the draft rewrote the file.
  • Preload the correct string, run the grader, and confirm a no-op first run is a zero. Doing nothing is not the same as being safe to rerun.

How I would mark it

Check Points Zero when
Clean tree before run 1 15 Porcelain was not empty
Run 1 changes only src/ 25 No hash change, crash, or extra path
Run 2 keeps the src/ hash 35 Any byte under src/ changes
Human read sees "hello, lab" 15 Stable, but wrong
Sabotage notes 10 They never produced a real zero

No points for a screenshot. No points for "the model agreed." Partial credit stops at the first zero in that row. I am not averaging vibes.

Oral check

I would ask three questions, out loud, with the file open.

  • Where does the script stop?
  • Which paths may change?
  • If src/greet.py is missing, do you recreate it or exit?

Recreate can be idempotent. Only if run two writes the same bytes and run three writes nothing new. If you choose recreate, add that case to the checker. Do not leave it as a feeling.

Decision list when you are tired

Use this when the hour gets noisy. Do not invent a fourth outcome.

  1. Start dirty? Zero. Do not tidy up and pretend the start was clean.
  2. First hash equals the base hash? Zero. Empty work is not a pass.
  3. Second hash differs? Zero. That is the whole lab.
  4. Unexpected path? Zero, even if src/ was stable.
  5. String wrong, hashes stable? Lose the 15 behavior points. Keep the idempotency points.
  6. Hosted tab died? Irrelevant. Grade the file on disk.

Who should not use this

Do not hash-grade work that should change every run. Timestamps, random IDs, and training logs fail for honest reasons. You would be punishing the assignment, not the student.

Do not use this on a one-shot migration. Those should refuse a second apply with a lock, and side files may still move. Teach the lock. Do not smash it into this rubric.

Do not point the grader at secrets. Hashing bytes is fine. Printing a credential while you debug is not. Would you paste that terminal into a model session? Then do not debug that way.

Skip it when you meant to grade taste. This rubric punishes drift, not dullness. A careful one-line edit can outscore a clever rewrite. That is the deal.

Limits, said plainly

Free model access will not insert a stop condition you forgot. It also must not sit on the grading path. The free server is a draft bench, not a witness, and uptime is not part of the score.

This scaffold does not understand renames, quoted git status paths, or binary fixtures. Say that in the README. Otherwise a filename with a space becomes a fake model failure.

I do not have a measured pass rate for this write-up. Run the scaffold on your own fixture before lab day, and change the path rule if you assigned something wider than src/.

A green second run can still be the wrong function. Schedule a later lab with a human oracle. Do not fold the two scores into one mood.

Close the tab, keep the script

If you want a sandbox for the draft, use that free model on the free server, save student_fix.py, and close the session.

Then answer the only question that counts. You ran it twice. Did src/ stay still?

Top comments (0)