Last month a candidate sent back a take-home that looked like a shipped product. Clean modules, a test suite, a README with an architecture diagram. Then I asked them to change one constraint live — run the same questions over a different folder — and the whole thing fell apart in about four minutes. Nothing in the repo showed how the model had decided anything.
That story is not rare right now. Look at what has been circulating this week: portfolios nobody has to build by hand anymore, agent writeups where the interesting detail is what a model noticed and then chose not to report. Both point at the same hiring problem. The artifact is cheap. The process is not.
So grade the process. Below is the take-home I've been handing out for tool-using agent roles, the grader that goes with it, and the ways people break it.
The prompt
Give a fixed corpus and exactly one endpoint. No framework, no scaffold — a budget and a required output format. The whole thing fits on one screen.
Build a small tool-using agent that answers questions over ./corpus
(plain text, roughly 40 files, no external index).
Constraints:
- One model endpoint, passed as MODEL_BASE_URL and MODEL_API_KEY
- At most 8 tool calls per question
- Emit trace.jsonl: one JSON object per step, in order
- No network calls outside the endpoint
- Also answer questions.jsonl (15 questions, 3 of which are not answerable)
Deliver: repo + trace.jsonl + a 300-word note on what you would fix
with another day.
Timebox: 3 hours. If you use a codegen assistant, say so in the note.
Every constraint earns its place. A single endpoint makes runs comparable across candidates. The eight-call cap forces tool selection to become visible; a brute-force grep sweep eats the budget and you can see it. The unanswerable questions check whether the agent ever says "not in the corpus" instead of inventing a clause. And the timebox plus the assistance disclosure keeps the exercise honest without pretending codegen does not exist.
One practical wrinkle: if candidates have to pay for tokens out of pocket, you are quietly ranking wallets. Hand everyone the same starting line. MonkeyCode is an open-source agent project that offers free model access and a free server option, and the operator states the free tier includes a generous token allowance — 10 million tokens as currently described. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Before you bolt any of this into a real pipeline, read the project's current terms yourself, because allowances and limits move and a number I quote today is not a contract.
The trace is the real submission
The repo shows what they built. The trace shows how they think. Ask for a flat JSONL file with three step types, and suddenly every submission is machine-readable and reviewable in one pass.
{"step":1,"type":"tool_call","tool":"grep","args":{"pattern":"refund window","path":"corpus"},"latency_ms":41}
{"step":2,"type":"tool_result","tool":"grep","bytes":812,"rows":3}
{"step":3,"type":"tool_call","tool":"read","args":{"path":"corpus/returns.txt","offset":120,"limit":40}}
{"step":4,"type":"final","answer":"30 days from delivery","citations":["corpus/returns.txt"]}
Format rules matter more than they look. Fixed step types mean a grader can run without an LLM in the loop, which means it costs nothing and gives the same answer twice. The citations field is the part candidates underestimate; a citation that does not resolve to a file, or resolves to a file that does not contain the answer, is worse than no citation at all.
The grader
This is the shape of the script I hand reviewers. Treat it as a template, not a validated benchmark — I have not run it across a cohort large enough to quote failure rates, and I would distrust anyone who did in a blog post.
import json, pathlib
KEY = {"q07": "30 days", "q12": "not in the corpus"}
def grade(trace_path, question_id, corpus):
steps = [json.loads(l) for l in pathlib.Path(trace_path).read_text().splitlines() if l.strip()]
calls = [s for s in steps if s["type"] == "tool_call"]
final = next((s for s in steps if s["type"] == "final"), None)
signatures = {(c["tool"], json.dumps(c["args"], sort_keys=True)) for c in calls}
return {
"within_budget": len(calls) <= 8,
"stopped": final is not None,
"citations_resolve": bool(final) and all((corpus / c).exists() for c in final.get("citations", [])),
"answer_matches_key": bool(final) and KEY.get(question_id, "") in (final.get("answer") or ""),
"duplicate_calls": len(calls) - len(signatures),
}
Read the output as a profile, not a score. duplicate_calls above zero on an easy question usually means the agent ignored a tool result and retried the same query — a control-flow bug, not a knowledge gap. within_budget false plus a correct answer tells you the person brute-forced their way out of designing a search step.
The rubric
| Signal | Weak | Strong |
|---|---|---|
| Budget use | Always 8 calls | Varies with question difficulty |
| Citation | Cites README or nothing | Cites the clause that contains the answer |
| Recovery | Aborts after a failed tool call | Narrows the query and continues |
| Unanswerable question | Invents a plausible answer | Says the corpus does not cover it |
| Trace honesty | Steps logged at the end, out of order | Written as it runs |
The last row is the one that catches the most people. If timestamps and ordering look retrofitted, the trace is fiction. I usually re-run question one live in the follow-up call; an agent whose trace was generated after the fact almost never reproduces its own step shapes.
Common failure modes
The most frequent submission is the grepper: the agent searches the whole corpus on every question because it never learned to narrow. Its trace is monotonous and its budget is always exactly eight.
The second is the confident summarizer. It reads one file, generalizes to all forty, and produces a fluent answer with a citation that resolves to something tangentially related. This one survives a casual review and dies the moment you open the cited line.
The third is the hardcoder. Twelve answers, memorized, with no tool calls worth mentioning. This is exactly why you hold back three questions and run them yourself during the debrief instead of accepting the batch they submitted. A memorized solution cannot answer a question it has never seen.
The fourth is structural: logging bolted on at the end, offsets missing, final emitted twice. It reads less like dishonesty and more like someone who never expected the trace to be the deliverable. Tell them in the prompt, twice, that it is.
A sample solution skeleton
Not a reference answer — a shape. The interesting decisions live in the system prompt, not the loop.
def answer(question, tools, model, budget=8):
ctx = [{"role": "system", "content": SYSTEM}] # SYSTEM: cite files, stop at 8, say "not in corpus"
for _ in range(budget):
msg = model.chat(messages=ctx + [{"role": "user", "content": question}])
if msg.tool_call:
result = tools[msg.tool_call.name](**msg.tool_call.args)
ctx += [msg.raw, {"role": "tool", "content": result}]
continue
return msg.text
return "budget exhausted"
Two lines there carry most of the signal: the system prompt's stopping rule and the citation requirement. A candidate who spends their note explaining why they chose those two lines is telling you something a passing test run never will.
Limitations, and who should skip this
This only works when answers are checkable against a key. For design-heavy roles, a written proposal graded by a human is still the better instrument, and no amount of JSONL will fix that. Setup also costs more than a LeetCode screen — you write the corpus, the key, and the grader before the first candidate ever sees it.
Free tiers have terms, and terms change. Do not build a hiring process on a quota you read in someone's article, including this one. And if you cannot spare a reviewer hour for the debrief, skip the take-home entirely; an unreviewed submission is just unpaid work.
If the reproducibility angle is what appeals to you, the free model access and free server option in MonkeyCode is the least interesting part of it. The useful part is that everyone starts from the same endpoint and every run leaves a trace you can re-read six months later. Check the current terms, then decide.
Top comments (0)