Workshop: Cap a Shared Lab Hour With a Session Envelope in 75 Minutes
A shared lab hour stays usable only when each student run carries a written session envelope before the first model call. The envelope records a turn cap, a local token budget, a timeout, and a concurrency limit that the harness can refuse. This workshop teaches that preflight in seventy-five minutes, with a checker students can rerun on their own laptops. It does not measure a vendor, and it does not treat any live quota as a permanent classroom constant.
The failure this hour is built to catch
Early October rooms often mix a classroom exercise with a contribution sprint, and both patterns share one scarce resource. Several notebooks then retry the same prompt until a demo looks finished for the person who started first. One laptop can open parallel loops while another student is still reading the worksheet and has not called yet. Shared capacity then looks slow, even though the defect is an unbounded client rather than a mysterious network.
How this differs from after-the-fact metering
Spend ledgers and latency charts explain a run that has already finished and left a log. This session starts earlier, and the harness blocks the call when the envelope is missing or internally inconsistent. Students leave with a file they can diff in git, not with a dashboard crop they cannot reproduce next week. The same checker runs against a mock client, so the lesson does not depend on a live allowance.
Seventy-five minute timetable
Keep the order below, and do not skip the refusal exercise near the middle of the hour. Acceptance-only demos hide the bug this workshop exists to teach, which is a client that never refuses. Pairs should rotate the keyboard every time the timetable moves into a new section of the hour. The instructor collects reason codes rather than model prose, because the gradeable artifact is the gate itself.
- From minute 0 to minute 10, state the shared-hour constraint and project one unbounded retry trace on the board.
- From minute 10 to minute 25, write the envelope fields and agree on class-local numbers for today's worksheet only.
- From minute 25 to minute 45, run the checker on three fixtures, including two files that must fail closed.
- From minute 45 to minute 60, place the checker in front of a stub client, then optionally attempt one live call.
- From minute 60 to minute 75, compare two refused runs and write a three-line lab note in the shared repository.
Envelope fields the room must fill
Keep the envelope small enough that a pair can review every field in under two minutes. Four numeric limits plus an exercise id are enough for a first shared hour in this room. Extra fields tend to become unread comments that nobody checks again before the call leaves the laptop. Treat every number as class policy that the next cohort can revise in the lab note.
- The
max_turnsfield limits model round trips inside one exercise, and that count includes every automatic retry. - The
local_token_budgetfield is the student's own ceiling, not a statement about any provider allowance or grant. - The
timeout_secondsfield bounds a single call so one hung socket cannot own the rest of the hour. - The
max_in_flightfield stops one laptop from opening several loops while the rest of the room waits. - The
exercise_idfield ties the envelope to today's worksheet so yesterday's accepted file cannot be reused by accident.
Class-local starting numbers for this seventy-five minute room are only a proposal, not a measured optimum. Use eight turns, two thousand local tokens, twenty seconds, and one in-flight call as the starting envelope. Shorten the token band if the worksheet asks for a single function, and record that reason in the note. Do not copy these numbers onto a slide as if they were a published vendor limit.
Trace shape students should recognize
Before writing envelopes, show a six-line trace from a stub log so the room can see fan-out without a live model. Each line carries a timestamp, a student id, an exercise id, and a turn index for the board. The sample below is synthetic teaching data, and it is not a capture from any production host. Ask pairs to count overlapping starts rather than to judge the prompt text inside the trace.
t=0 student=s1 exercise=lab-hour-03 turn=1 state=start
t=1 student=s1 exercise=lab-hour-03 turn=1 state=start
t=1 student=s1 exercise=lab-hour-03 turn=2 state=start
t=2 student=s2 exercise=lab-hour-03 turn=1 state=waiting
t=9 student=s1 exercise=lab-hour-03 turn=1 state=retry
t=9 student=s1 exercise=lab-hour-03 turn=3 state=start
In this fixture, student s1 has three starts overlapping while student s2 is still waiting on the first turn. That pattern is the defect the max_in_flight field is meant to refuse before any live call starts. Do not treat the timestamps as latency measurements, because the file was written by hand for the board.
Worked checker students can rerun
The script below is a teaching artifact meant for local fixtures on student laptops during the hour. It has not been executed against a production fleet for this article, and it does not call a model by itself. Run it on the pass file and the two fail files before anyone in the room discusses providers. If the output disagrees with the comments beside the commands, fix the fixture before you change the checker.
#!/usr/bin/env python3
"""Session envelope checker. Local teaching artifact, not a vendor benchmark."""
import json
import sys
REQUIRED = (
"exercise_id",
"max_turns",
"local_token_budget",
"timeout_seconds",
"max_in_flight",
)
def load_envelope(path):
with open(path, encoding="utf-8") as handle:
return json.load(handle)
def check(card, expected_exercise):
errors = []
for key in REQUIRED:
if key not in card:
errors.append(f"missing:{key}")
if card.get("exercise_id") != expected_exercise:
errors.append("exercise_id_mismatch")
turns = card.get("max_turns")
if not isinstance(turns, int) or turns < 1 or turns > 12:
errors.append("max_turns_outside_class_band")
budget = card.get("local_token_budget")
if not isinstance(budget, int) or budget < 100 or budget > 4000:
errors.append("local_token_budget_outside_class_band")
timeout = card.get("timeout_seconds")
if not isinstance(timeout, int) or timeout < 5 or timeout > 30:
errors.append("timeout_seconds_outside_class_band")
if card.get("max_in_flight") != 1:
errors.append("max_in_flight_must_be_one")
return errors
def main():
if len(sys.argv) != 3:
print("usage: check_envelope.py ENVELOPE.json EXERCISE_ID", file=sys.stderr)
return 2
errors = check(load_envelope(sys.argv[1]), sys.argv[2])
if errors:
print("REFUSE")
for item in errors:
print(item)
return 1
print("ACCEPT")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Passing fixture, saved as pass_envelope.json:
{
"exercise_id": "lab-hour-03",
"max_turns": 8,
"local_token_budget": 2000,
"timeout_seconds": 20,
"max_in_flight": 1
}
Broken fixture, saved as missing_turns.json:
{
"exercise_id": "lab-hour-03",
"local_token_budget": 2000,
"timeout_seconds": 20,
"max_in_flight": 1
}
Commands for the room:
python3 check_envelope.py pass_envelope.json lab-hour-03
python3 check_envelope.py missing_turns.json lab-hour-03
python3 check_envelope.py pass_envelope.json lab-hour-04
The expected teaching outcomes are classroom labels only, and they are not measured metrics from a live service. The first command should print ACCEPT, and the other two should print REFUSE plus a reason code. A missing numeric field can emit two codes, one for absence and one for the band, and that pairing is expected. A different printout means the fixture or the Python version needs a look, not that a model misbehaved.
Exercise sequence
Exercise A, minutes 10–25
Each pair writes one valid envelope for lab-hour-03 and one intentionally broken envelope for the same checker. Broken envelopes should omit a required field, reuse yesterday's exercise id, or set max_in_flight above one. The useful result is a refuse code the other pair can predict before the command even runs. Prompt wording is out of scope for this segment, so leave the model client closed on every laptop.
Exercise B, minutes 25–45
Run the three commands and paste reason codes into the shared note without pasting secrets or prompts. Add one fixture that sets local_token_budget to 50000 and confirm the class band rejects that envelope. That rejection is local policy for this room, and it is not evidence about any provider allowance. If two pairs disagree on a reason code, diff the envelope files before anyone edits the checker.
Exercise C, minutes 45–60
Put the checker in front of a stub that sleeps only after ACCEPT and never opens a socket on REFUSE. Only after that gate works should one volunteer try a live call, and only inside the accepted timeout. Everyone else watches the reason code on the terminal, not the generated text from the optional call. If the live page cannot be confirmed that morning, stop at the stub and mark the live step skipped.
def gated_call(card, exercise_id, client):
errors = check(card, exercise_id)
if errors:
return {"status": "refused", "errors": errors}
return client.complete(
max_turns=card["max_turns"],
timeout_seconds=card["timeout_seconds"],
)
The function above is whiteboard pseudocode, and students should implement the sleep-or-refuse stub themselves in class. Copying it without a failing test recreates the unread-comment problem this hour is trying to remove from the lab. A three-line test that expects refused for a stale exercise id is enough to lock the gate. Run that small test before any volunteer is allowed to open a live client from the room.
Where a free model path fits this hour
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
Operator notes for this piece describe MonkeyCode as an open-source project with free model access and a free server option. A class may point the optional live step at those paths only after reading the current project page that morning. Students who will read or fork the code should confirm the license on that same page before they clone anything.
This article does not name models, freeze a token grant, describe hardware, or claim that free access will stay unchanged. If the live page disagrees with a slide, the live page wins and the stub exercise remains the gradeable path. Use a free server only as a shared place to run the checker when local Python installs are uneven. It is not a private production host, and it is not a reason to skip the refuse fixtures in the timetable.
Refusal decision table
| Condition | Harness result | Class action |
|---|---|---|
| Required field missing | REFUSE | Fix the envelope, and do not retry a model |
| Exercise id is stale | REFUSE | Mint today's id and discard yesterday's envelope |
| In-flight value above one | REFUSE | Close extra notebooks before any live step |
| Live access terms unknown | Skip live step | Finish on the stub and record the gap |
| Envelope accepted and stub passes | Optional live call | One volunteer, inside the written timeout |
Read the table from top to bottom, and stop at the first matching row before you take any other action. Do not average the rows into a score, because a single missing field is already a failed preflight. The last row permits one call from one volunteer, not a fan-out from the rest of the room. Record a skipped live step as a completed outcome when the access page cannot be confirmed.
Lab note template
Keep the note to six lines so review stays faster than the exercise itself for the instructor. Ask each pair to commit the note beside the fixtures before they leave the room today. The change_next_week line should name a fixture gap, not a feeling about model quality or speed. An acceptable example is to add an envelope whose timeout is four seconds so the lower band is tested.
exercise_id: lab-hour-03
class_band: turns<=12 tokens 100-4000 timeout 5-30 inflight=1
refuse_codes_seen: missing:max_turns, exercise_id_mismatch
stub_gate: pass
live_step: skipped|one_volunteer
change_next_week: Add an envelope whose timeout is four seconds.
Limitations
The two-thousand-token class band is a teaching cap, not a cost model and not a provider quota. The checker does not inspect prompt content, tool side effects, or whether data leaves the room. A student can still paste secrets into an accepted call, so the envelope is not a privacy control. Shared free capacity can still saturate for reasons this script cannot see, including other tenants on the same host.
Do not publish timings from one classroom as a benchmark, and do not compare models with this harness. The synthetic trace is hand-written, so its gaps are illustrations rather than observations from a running host. A missing fixture file should fail the command with a visible error, which the script currently leaves to the interpreter. Teach that traceback as an environment failure, not as a refused envelope, so students do not confuse the two.
Who should not use this hour
Skip this workshop if the room needs authenticated production traffic, regulated data, or a graded accuracy comparison between models. Skip it if students cannot install Python 3 and you also lack a shared server you are allowed to use. Skip the live segment when you cannot confirm today's access terms in writing before the session starts. A team that already enforces admission control at a gateway should not replace that gateway with this laptop script.
Close
Leave the three fixtures, the synthetic trace, and the checker in the class repository beside the timetable. Next week, change only the exercise id and one class-band edge, then rerun the same commands before any live client opens. If a shared optional client would help the last segment, read the current MonkeyCode access notes before you rely on them. Keep the stub path as the graded path even when that optional client is available to the room.
Top comments (0)