You send the take-home on Friday. By Sunday the candidate has not touched the bug. They are still hunting a model key that will not 402. The packet you wrote was about judgment under a constraint.
The weekend measured a credit card instead. That mismatch is quiet, and it ruins the comparison. Two people can write the same harness. One already has a paid route on a laptop they use every day.
The other is rationing a trial key and hoping the fan holds. You are not looking at the same task. You are looking at two wallets wearing one prompt. Fix the packet, not the lecture you give afterward.
So you change what you ship. You stop asking them to bring a runtime. You attach one. A free model route and a free server are enough, if you freeze the contract before either of them moves.
The grade is not whether the model sounded clever. The grade is whether the harness kept its shape when you swapped the route and the box. Think of a relay baton. The runner can change, and the baton cannot.
A shared workbench is the right picture, not a stipend. A stipend still leaves one person debugging invoices. A workbench is already in the room, and both candidates stand at it. What you watch is how they clamp the wood.
Here is the prompt you paste into the packet. Keep it on one screen. If they have to scroll to find the rule, they will negotiate with the rule.
Task swap-01.
You get two environment variables: FREE_MODEL_ROUTE and FREE_SERVER.
Do not add a paid fallback. If either variable is missing, exit 2.
Read tasks/swap-01.json. Call the route. Write out/result.json.
Required keys: task_id, status, used_paid_fallback, model_route,
declared_route, server, declared_server, answer.
status must be "ok". used_paid_fallback must be false.
model_route must equal declared_route. server must equal declared_server.
answer must equal the "expect" field in the task file.
Do not write secrets, tokens, or environment dumps into the file.
We will run this twice. The second run changes both variables.
The file shape must not change. Only the declared values may change.
Timeout for each run is 60 seconds. A slow start is not a failure
unless the process is still running when the clock ends.
That prompt is the assignment. The rubric is not a paragraph underneath it. The rubric is a test you run against the repo they submit. You can read the test aloud in the debrief, and the argument stays about files.
# contract_test.py
# Rubric you run. Not a benchmark, and not a score for prose.
import json
from pathlib import Path
def load():
return json.loads(Path("out/result.json").read_text())
def test_shape():
data = load()
assert data["task_id"] == "swap-01"
assert data["status"] == "ok"
assert data["used_paid_fallback"] is False
def test_declared_matches_used():
data = load()
assert data["model_route"] == data["declared_route"]
assert data["server"] == data["declared_server"]
assert isinstance(data["model_route"], str) and data["model_route"]
assert isinstance(data["server"], str) and data["server"]
def test_answer_matches_task():
task = json.loads(Path("tasks/swap-01.json").read_text())
data = load()
assert data["answer"] == task["expect"]
def test_no_secret_spill():
raw = Path("out/result.json").read_text().lower()
for needle in ("api_key", "secret", "bearer ", "token="):
assert needle not in raw
Pair the test with a tiny task file. The expected answer is boring on purpose. You are not crowning a model. You are checking whether a person can hold a shape while the bench moves.
{
"task_id": "swap-01",
"prompt": "Return the expect string and nothing else.",
"expect": "baton-ok"
}
The sample solution below is a proposal. It has not been executed against a live route for this draft, and it does not pretend to know a vendor path. Your harness should call whatever URL you place in the variable, then write that same URL back into the file.
If the call fails, exit 1. Do not catch the failure and quietly switch providers. A second host is a different assignment. You did not send it.
# harness.py — proposal, unexecuted
import json, os, sys, urllib.request
from pathlib import Path
def main():
route = os.environ.get("FREE_MODEL_ROUTE")
server = os.environ.get("FREE_SERVER")
if not route or not server:
print("missing free runtime", file=sys.stderr)
return 2
task = json.loads(Path("tasks/swap-01.json").read_text())
payload = json.dumps({"input": task["prompt"]}).encode()
req = urllib.request.Request(
route,
data=payload,
headers={"Content-Type": "application/json"},
)
# FREE_SERVER is the box this process is allowed to run on.
# Record it. Do not open another host from this file.
with urllib.request.urlopen(req, timeout=30) as resp:
body = json.loads(resp.read().decode())
answer = str(body.get("text", ""))
status = "ok" if answer == task["expect"] else "mismatch"
out = {
"task_id": "swap-01",
"status": status,
"used_paid_fallback": False,
"model_route": route,
"declared_route": route,
"server": server,
"declared_server": server,
"answer": answer,
}
Path("out").mkdir(exist_ok=True)
Path("out/result.json").write_text(json.dumps(out))
return 0 if status == "ok" else 1
if __name__ == "__main__":
raise SystemExit(main())
You run the swap yourself before you send the packet. Two passes, same test, different variables. If your own packet cannot survive the move, you do not get to blame the candidate for a crack you shipped.
export FREE_MODEL_ROUTE="$ROUTE_A"
export FREE_SERVER="$BOX_A"
python harness.py
python -m pytest contract_test.py -q
cp out/result.json /tmp/pass-a.json
export FREE_MODEL_ROUTE="$ROUTE_B"
export FREE_SERVER="$BOX_B"
python harness.py
python -m pytest contract_test.py -q
diff -u /tmp/pass-a.json out/result.json
Save pass A before the second run. The diff should show the route and the server moving, and nothing else. That diff is what you attach to the review note. A diff is harder to argue with than a hunch about senior feel.
If the double swap fails, split it. Change only the route and rerun, then change only the server. You want to know which move bent the file. A route-only failure usually means they parsed a body they invented after one lucky response.
A server-only failure usually means a hardcoded disk path. The second machine does not owe them /tmp from the first. Treat those as different bugs. Do not average them into one vague note about flakiness.
Watch the failures that look like success. Some people put a paid key in a shell profile, then write used_paid_fallback as false because the branch you can see never mentions a card. Start their process from a shell that holds those two variables and nothing else. If it still completes, the paid path is still in the room.
Ask for the process list, not the paragraph where they promise they stayed on the free route. A promise is cheap. A process table is not. You are grading the handoff, so look at the hand.
Others pass on the first box because they wrote a scratch file under /tmp and read it back. The second box does not have that file. The free server is not a metaphor for somewhere in the cloud. If the prompt says the process runs there, a laptop-only path is a broken handoff.
You will see a missing file. You will not see a wrong adjective. Grade the missing file, and say so in the note. Adjectives can wait for a different exercise.
A third group edits the contract until the model looks right. They change expect, or they delete the equality check, or they call the rest a style nit in the README. A modified rubric is a failed submission.
You can still hold the conversation. You do not pretend they solved the task you sent. Put their edited test next to yours and let the diff talk. Silence after that diff is usually clearer than another round of feedback.
In the review, open the diff first, then the test, then the harness. Leave the chat log closed until those three agree. A fluent writeup can hide a paid fallback in one import. The file cannot hide it once the test checks the flag and the declared route together.
The approach has edges, and you should say them in the packet. A free route can slow down, refuse a burst, or disappear between Friday and Monday. Confirm the current terms the morning you send the prompt, and write that date next to the variables.
This draft does not state a token quota, a machine size, a region, or a promise that the free option remains. If the docs and this paragraph disagree, the docs win. Stale numbers are how a fair assignment becomes a trap.
Do not use this when the job needs a specific paid model's quirks. A free route will not reproduce those quirks, and a green test would be false comfort. Do not use it for a writing role either. A contract file will not show you how someone argues with a stakeholder.
Skip it when the candidate cannot reach the network you chose. Offer a recorded fixture for that case, and mark the fixture run as a separate lane. Do not mix that lane with the live swap, or you will grade a recording as if it were a machine.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. The two availability claims in this workflow, free model access and a free server option, are the claims this draft was asked to treat as available. They are not a measured benchmark, and they are not a quota.
MonkeyCode is one place you can point FREE_MODEL_ROUTE and FREE_SERVER if those options are still listed when you check the project. Keep the pytest file in your repo either way. The runtime is something you swap. It is not the grade.
If those options are already on the account you use, set them as ROUTE_A and BOX_A, run both passes, and file the diff beside the prompt. Send that packet. Leave the wallet out of the story.
Top comments (0)