DEV Community

Charlie Xu
Charlie Xu

Posted on

Grade the Workspace, Not the Demo: A Bootcamp Lab on Shared Agent Boxes

A green demo is not a pass. If another student's files, notes, or session crumbs are still in your agent workspace, I grade the lab a zero.

Shared boxes feel efficient. One class server, one agent login, twelve people "just trying a fix." That is also how prompts leak and caches lie. This lab grades the handoff, not the screenshot.

Why the demo lies

You get a tiny bug. The agent edits one function. The diff looks tidy. So why am I still unhappy?

Because I cannot see whose shell you borrowed. I cannot see which checkout the agent read. And if that shared server was rebuilt at lunch, can you stand the workspace up again without opening a classmate's history?

That is the failure this lab exists to catch. Not a clever model. A messy box.

What you may use

Run the task on your laptop, or on a shared class server. One allowed runtime is MonkeyCode, an open-source coding-agent project. Disclosure: This article was prepared as part of MonkeyCode's product outreach.

The lab brief I was given says two options exist: free model access, and a free server option. I am not going to print a token count, a model list, a machine size, or a promise that the offer stays up. Those details move, and a stale number in a student write-up is worse than no number.

On lab day you open the current project docs, copy the published limit into limits.txt, and date that file. If a chat, a slide, or this article disagrees with the docs, the docs win. An invented quota is a zero.

Take the product name out and the rubric still stands. One student. One workspace. One recorded limit. One teardown I can rerun.

Setup, before the bug

Do not touch the fixture until the folder is boring and yours.

  1. Pick a lab handle, not an email. Example: s17.
  2. Make a private root: ~/lab-box/s17/.
  3. Clone only the fixture into work/. Not your home directory.
  4. Put a fake env file inside the root. Mode 600. LAB_TOKEN=dev-only is enough. No real keys. Ever.
  5. Write limits.txt from the docs you opened. Timestamp, URL, and one sentence on what is actually offered today.
  6. Write box.txt. Local runs say local plus the absolute path. Free-server runs record hostname, username, and path.
  7. Start a fresh agent session. Resuming a shared thread is an automatic miss.
id="s17"
root="$HOME/lab-box/$id"
mkdir -p "$root/work" "$root/out"
chmod 700 "$root"
umask 077
printf 'LAB_TOKEN=dev-only\n' > "$root/.env"
chmod 600 "$root/.env"
date -u +%Y-%m-%dT%H:%M:%SZ > "$root/limits.txt"
printf '\n# paste the published free-tier note and its URL under this line\n' >> "$root/limits.txt"
printf 'runtime=local\npath=%s\n' "$root" > "$root/box.txt"
Enter fullscreen mode Exit fullscreen mode

The prompt you actually send

Save this as prompt.md. If your prompt is a paragraph of vibes, I cannot grade the scope checkpoint. Would you accept a test with no assertion?

Fix the status string in work/app.py so the fixture test passes.
Read and write only under work/.
Do not read the home directory.
Do not print environment variables.
Allowed tools: read file, edit file, run the unit check.
Stop if any path outside work/ shows up in context.
Enter fullscreen mode Exit fullscreen mode

The fixture itself should be dull. A function returns "ok" when the test wants "ready". You are not shipping a product. You are proving the box was yours.

Checkpoints

I grade these in order. A later win does not repair an earlier miss.

  • C1 — Boundary. box.txt exists. The path contains your id. Nothing under your root names another student id.
  • C2 — Limit note. limits.txt has a timestamp and a pointer to docs you can open. "Unlimited," with no source, is not a note.
  • C3 — Secret mode. .env is owner-only. The diff and the report do not contain LAB_TOKEN or anything that looks like a real key.
  • C4 — Session scope. The agent was aimed at work/. prompt.md names the directory and the tools. If you allowed network, you named the host.
  • C5 — Teardown. Delete out/ and regenerate the report with the script. A screenshot is not a report.

The checker

This script is a lab artifact. It is not a benchmark, and I am not claiming a timing, a pass rate, or a result from any particular machine. Run it on your tree. If it fails, fix the workspace. Do not comment out the check.

#!/usr/bin/env python3
"""Workspace isolation checker. Lab artifact until you run it."""
import stat
import sys
from pathlib import Path

def main() -> int:
    root = Path(sys.argv[1]).resolve()
    sid = sys.argv[2]
    fails = []
    if sid not in root.name:
        fails.append("workspace folder name missing student id")
    env = root / ".env"
    if not env.is_file():
        fails.append("missing .env")
    elif env.stat().st_mode & (
        stat.S_IRGRP | stat.S_IROTH | stat.S_IWGRP | stat.S_IWOTH
    ):
        fails.append(".env is not owner-only")
    for name in ("limits.txt", "box.txt", "prompt.md"):
        if not (root / name).is_file():
            fails.append(f"missing {name}")
    limits_path = root / "limits.txt"
    limits = (
        limits_path.read_text(encoding="utf-8", errors="ignore")
        if limits_path.is_file()
        else ""
    )
    if "http://" not in limits and "https://" not in limits:
        fails.append("limits.txt has no doc URL")
    enrolled = {"s16", "s17", "s18", "s19"}
    for path in root.rglob("*"):
        if not path.is_file():
            continue
        text = path.read_text(encoding="utf-8", errors="ignore")
        if path.name != ".env" and "LAB_TOKEN=" in text:
            fails.append(f"token leaked into {path.name}")
        for other in sorted(enrolled - {sid}):
            if other in text:
                fails.append(f"foreign id {other} inside {path.name}")
    report = root / "out" / "report.txt"
    report.parent.mkdir(parents=True, exist_ok=True)
    body = "PASS\n" if not fails else "FAIL\n" + "\n".join(fails) + "\n"
    report.write_text(body, encoding="utf-8")
    print(body, end="")
    return 0 if not fails else 1

if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode
python3 check_box.py "$HOME/lab-box/s17" s17
Enter fullscreen mode Exit fullscreen mode

The enrolled-id set is a teaching stub. Replace it with the handles actually in your section. A checker that only knows your own id will smile at a copied tree. Would you trust that smile?

Local box or the free server?

Question Stay local Use the free server option
Can you install the fixture runtime without admin help? Yes, if your laptop policy allows it Only if the image already has it
Might a classmate share the host? No Assume yes until box.txt shows your own path
Will the same path exist next week? Your disk, your backups Record hostname and path, then plan to rebuild
Is today's published limit enough for one fixture? Read limits.txt and stop if the session says otherwise Same rule
Is any file real customer data? Still no No

Free access is a teaching convenience. It is not a capacity plan. If the session stops because the published limit is exhausted, write that in the report and shrink the fixture. Opening a second account to dodge the cap is a zero. It is also how a class drains a shared pool.

How I fail a submission

  1. The demo passes, and ~/lab-box/ contains s17 plus a copied s04 folder you used "for reference."
  2. limits.txt quotes a token number from an old post, and the URL is missing.
  3. The agent was pointed at $HOME because "it needed context."
  4. .env is mode 644, and the report helpfully echoes it.
  5. Teardown is a photo of a terminal. I cannot rerun a photo.

Notice what is not on that list. I do not fail you for a short diff. I do not fail you for staying local. I fail you for a boundary I cannot check.

Stretch goals

  • Create s18 beside s17 and show each tree passes alone.
  • Ask the agent to list files it read. Any path outside work/ fails C4, even if the test is green.
  • Run once locally and once on the free server option, if the docs still offer it today. The fixture diff should match. box.txt must not.
  • Timebox the base lab at 45 minutes. Stretch cannot lift a failed C2.

Rubric

Item Points Pass looks like
C1 boundary 25 Folder name has your id; no foreign id in the tree
C2 limits note 20 Dated, with a doc URL, no invented quota
C3 secrets 20 .env is mode 600; report stays clean
C4 session scope 20 prompt.md names the directory and the tools
C5 teardown 15 Script regenerates out/report.txt
Stretch +0 to 10 Bonus only if the base score is already 80 or higher

Partial credit stops at the first failed checkpoint unless you can explain the miss in two sentences. Charm is not a column.

Who should skip this

Skip it on a production fleet, a client repo, or any tree with live credentials. The fixture is fake on purpose. Skip it if your program bans third-party agents. Keep the same rubric and point the prompt at a local script.

Skip it if you need a capacity number for a budget meeting. I did not measure throughput, and a free tier is the wrong spreadsheet. Also skip the free server option if today's docs say the offer is paused, region-locked, or waitlisted. An article is not an entitlement.

The checker is crude by design. It will not catch a renamed copy of a classmate's notes, a secret split across two files, or a process still running after you delete the folder. Treat a PASS as "the obvious misses are gone," not as a security review.

Leave a report, not a vibe

The model did not make the workspace safe. You did, or you did not. Write the boundary down. Point at today's docs for whatever free model access or free server option you actually used. Leave a report I can regenerate after I delete out/.

If you want a place to confirm those two options before lab, read the current MonkeyCode project docs, copy only what they publish into limits.txt, and run one boring bug. Then show me the report.

Top comments (0)