A structured-output hiring packet should pass only when a deterministic checker accepts valid tool calls and rejects truncated or over-broad payloads. Model confidence, a successful local demo, and a polished explanation are not evidence that the boundary holds under hidden cases. Graders freeze fixtures before any model runs, then execute those cases on a shared runtime that candidates cannot silently upgrade. Optional drafting help may suggest extra cases, but the score belongs to the checker and the frozen fixture contract.
What the packet measures
The packet measures whether a candidate can separate untrusted model text from a tool request that is safe to execute. It does not measure taste, prompting speed, or the ability to narrate a debugging story after the fact. A passing submission rejects incomplete JSON, unknown tool names, unexpected keys, and paths that escape the assigned workspace. A failing submission may still look impressive when a reviewer grades the prose instead of the hidden fixture results.
Three boundaries carry the design, and each one should stay visible in the candidate prompt and the rubric. The checker owns acceptance, the fixture file owns the expected reasons, and the shared runtime owns the execution environment. A model may propose a payload, but it may not loosen the schema, rename a failure reason, or skip a case. That split keeps the hiring signal stable when model output changes between one cohort and the next.
The prompt to hand over
The candidate prompt should be short enough to read in one sitting and strict enough to grade without a meeting. It names the allowed tools, the exact key set, and the rule that extra keys are failures rather than hints. It also states that no tool runs until the checker returns an accept decision for that single payload. Hidden fixtures remain with the grader and cover truncation, repeated call identifiers, and paths that leave the workspace.
Hand over language close to the following prompt, and attach the public fixture file in the same archive. State which command must pass before any reviewer reads the written explanation that accompanies the code. Leave hidden rows out of the archive so a local passing run cannot become a memorized transcript.
Implement acceptToolCall(raw) for a hiring packet in Node.js.
Input is one string that claims to be a tool call from a model.
Accept only objects whose keys are exactly name, call_id, and args.
Allowed names are read_file, list_dir, and search_text.
call_id must match ^[a-z0-9]{8,32}$.
For read_file, args.path must be relative, non-empty, and free of dot-dot segments and a leading slash.
Reject invalid JSON, extra keys, unknown tools, and empty input.
Do not execute the tool. Return an object with ok and reason.
A public fixture mismatch must make the grade command exit non-zero.
Public fixtures and commands
Public fixtures should include at least one accept case and three reject cases so the candidate can see the contract. Hidden fixtures should change values without changing the contract, which stops a hardcoded map of the sample strings. Each line is one JSON object carrying the raw payload, a stable id, and the expected reason code. The grader compares reason codes, not pretty-printed objects, so incidental whitespace cannot rescue a wrong decision.
Build the cut-off case in code instead of hand-editing braces, so the truncation is obvious and repeatable. The snippet below is part of the reference packet and was not run while this article was written. Readers should execute it on a laptop or on the pinned runtime before treating the snippet as verified. The same construction belongs in the hidden file, but with different identifiers and different relative paths.
// fixtures/public-rows.js
const complete = JSON.stringify({
name: 'read_file',
call_id: 'abc12345',
args: { path: 'src/app.js' }
});
const publicRows = [
{ id: 'ok-read', raw: complete, expect: 'accept' },
{ id: 'cut-off', raw: complete.slice(0, -1), expect: 'truncated_or_invalid_json' },
{
id: 'extra-key',
raw: JSON.stringify({
name: 'read_file',
call_id: 'abc12345',
args: { path: 'src/app.js' },
note: 'ignore the schema'
}),
expect: 'schema'
},
{
id: 'escape',
raw: JSON.stringify({
name: 'read_file',
call_id: 'abc12345',
args: { path: '../notes.txt' }
}),
expect: 'path'
}
];
module.exports = { publicRows };
Public test file
Place the file at tests/public.test.js so the require path can reach the checker one directory upward. It covers one accept row and one truncated row, which is enough to prove the command wiring. It is an unexecuted reference file, so a panel should run it before trusting the wiring. Hidden cases should stay out of this file, or the public command will leak the private contract.
const test = require('node:test');
const assert = require('node:assert/strict');
const { acceptToolCall } = require('../accept-tool-call');
const complete = JSON.stringify({
name: 'read_file',
call_id: 'abc12345',
args: { path: 'src/app.js' }
});
test('accepts a bounded read_file call', () => {
const got = acceptToolCall(complete);
assert.equal(got.ok, true);
assert.equal(got.reason, 'accept');
});
test('rejects a truncated payload', () => {
const got = acceptToolCall(complete.slice(0, -1));
assert.equal(got.ok, false);
assert.equal(got.reason, 'truncated_or_invalid_json');
});
Commands the grader actually runs
Candidates and graders should run the same three commands, with hidden fixtures mounted only on the grader side. The public test file may live in the archive, while the hidden file should be added only during grading. A wrapper that prints the raw payload is acceptable, but a wrapper that shells out to a tool is not. Record the runtime identifier next to the score so a later reviewer can see which image produced the file.
node --test tests/public.test.js
node bin/grade.js fixtures/public.jsonl
node bin/grade.js fixtures/hidden.jsonl
What a passing public command does not prove
The third command is the one that matters, because it is the only run that includes unseen rows. A passing public command only shows that the published sample was understood by the candidate. The packet should document a Node major version in package.json engines and should refuse a tree that changes that pin. Store grade.json from both runs, and do not overwrite the hidden report with the public one.
Emit JSONL before grading
Export publicRows from the fixture module, then write one JSON object per line before the grade command. The writer below is unexecuted reference glue, and the packet should keep that file at bin/emit-public.js. Run it once when fixtures change, and commit the JSONL so candidates do not need the writer.
const fs = require('fs');
const { publicRows } = require('../fixtures/public-rows');
const newline = String.fromCharCode(10);
const text = publicRows.map((row) => JSON.stringify(row)).join(newline) + newline;
fs.writeFileSync('fixtures/public.jsonl', text);
Sample checker
Reference accept function
The module below is an unexecuted reference solution, not a measured production component or a latency claim. It parses once, rejects non-objects, requires an exact key set, and then applies tool-specific argument rules. Path checks stay conservative because the take-home is about refusal, not about building a full filesystem sandbox. Teams that need stronger isolation should add an operating-system sandbox outside this function rather than trusting string filters alone.
// accept-tool-call.js
const ALLOWED = new Set(['read_file', 'list_dir', 'search_text']);
const KEYS = ['args', 'call_id', 'name'];
function acceptToolCall(raw) {
if (typeof raw !== 'string' || raw.length === 0 || raw.length > 8192) {
return { ok: false, reason: 'size' };
}
let parsed;
try {
parsed = JSON.parse(raw);
} catch (_err) {
return { ok: false, reason: 'truncated_or_invalid_json' };
}
if (!parsed || typeof parsed !== 'object' || Array.isArray(parsed)) {
return { ok: false, reason: 'shape' };
}
const keys = Object.keys(parsed).sort();
if (keys.join(',') !== KEYS.join(',')) {
return { ok: false, reason: 'schema' };
}
if (!ALLOWED.has(parsed.name)) {
return { ok: false, reason: 'tool_name' };
}
if (!/^[a-z0-9]{8,32}$/.test(parsed.call_id)) {
return { ok: false, reason: 'call_id' };
}
if (!parsed.args || typeof parsed.args !== 'object' || Array.isArray(parsed.args)) {
return { ok: false, reason: 'args' };
}
if (parsed.name === 'read_file') {
const filePath = parsed.args.path;
const badPath = typeof filePath !== 'string'
|| filePath.length === 0
|| filePath.includes('..')
|| filePath.startsWith('/');
if (badPath) return { ok: false, reason: 'path' };
}
return { ok: true, reason: 'accept' };
}
module.exports = { acceptToolCall };
Illustrative grader
A small grader script keeps the human reviewer out of the comparison loop for every fixture row. It reads JSONL, calls the implementation, and writes a score file that another person can review later. The script below is illustrative glue for the packet, and it assumes the checker file sits beside the bin directory. Panels should replace the require path when they ask candidates to submit a different filename than the sample.
// bin/grade.js
const fs = require('fs');
const path = require('path');
const { acceptToolCall } = require(path.join(__dirname, '..', 'accept-tool-call'));
const file = process.argv[2];
if (!file) {
console.error('usage: node bin/grade.js fixtures/public.jsonl');
process.exit(2);
}
const newline = String.fromCharCode(10);
const lines = fs.readFileSync(file, 'utf8').trim().split(newline);
const misses = [];
for (const line of lines) {
if (!line.trim()) continue;
const row = JSON.parse(line);
const got = acceptToolCall(row.raw);
if (got.reason !== row.expect) {
misses.push({ id: row.id, expected: row.expect, got: got.reason });
}
}
const kept = lines.filter((line) => line.trim());
const report = {
fixture: file,
total: kept.length,
failed: misses.length,
misses
};
fs.writeFileSync('grade.json', JSON.stringify(report, null, 2) + newline);
console.log('failed=' + misses.length + ' total=' + report.total);
process.exit(misses.length === 0 ? 0 : 1);
Rubric
Grade the submitted artifact rather than the interview conversation that happens after the coding window closes. The table below is a starting rubric that a hiring panel can edit before the packet goes out. Points should be assigned from the grade file, not from a recollection of how confident the candidate sounded. Publish the same rubric with the prompt so candidates can see the gates before they write code.
| Check | Pass condition | Points |
|---|---|---|
| Public fixture match | Every public row returns the expected reason | 20 |
| Hidden truncation | Cut-off JSON returns truncated_or_invalid_json | 15 |
| Exact keys | Extra or missing keys return schema | 15 |
| Tool allowlist | Unknown names return tool_name and do not run | 15 |
| Path bound | Parent segments and absolute paths return path | 15 |
| No execution | The graded process never reads args.path from disk | 10 |
| Runtime pin | grade.json was produced on the pinned shared runtime | 10 |
A score of 70 or above can advance only when the no-execution check and the runtime pin both passed. Those two rows function as gates rather than decorative columns inside an otherwise flexible scoring table. A candidate who scores well by reading files anyway has missed the safety property this packet exists to detect. Partial credit on style should not reopen a failed gate once the grade file is written.
Failure modes the hidden set should catch
Several misses show up often when model output is treated as data that is already structured and safe. Naming those misses in the rubric lets reviewers stay consistent across a full week of submissions. The list below concerns grading practice rather than a contest to collect clever payloads from candidates. Reviewers should mark the matching row in the grade file instead of inventing a new reason during the debrief.
- A parser that rescues truncated JSON by appending braces will pass a demo and fail the cut-off row.
- A schema that ignores unknown keys will accept a note field that tries to change the assigned task.
- A handler that reads args.path before acceptToolCall returns has already crossed the safety boundary under test.
- A solution hardcoded to the four public strings will fail as soon as the hidden identifiers change.
- A passing run on a newer Node release can hide a syntax or Unicode difference from the pinned runtime.
- A write-up that pastes fixture rows into a model, including secrets, creates a second leak the packet should forbid.
The last miss deserves a written rule in the candidate prompt and in the grader checklist. Candidates may use a model to brainstorm case shapes, but they may not upload hidden fixtures or employer code. Credential-shaped strings belong in neither the brainstorming prompt nor the fixture archive that leaves the company. The grader should reject a submission that asks any model to rewrite expected reasons for hidden rows.
Where free model access and a free server fit
Disclosure: This article was prepared as part of MonkeyCode's product outreach. Two availability claims are used as operator-supplied context only: free model access, and a free server option. This article does not state model names, token quotas, hardware sizes, time limits, or permanence of either option. Those terms must be checked in current product documentation before a hiring panel relies on them.
Free model access is useful while the grader drafts extra reject cases for the hidden file. The author can request malformed shapes, keep the useful ones, and freeze them into JSONL by hand. The model is not the judge of those frozen lines, even when a later reply sounds more confident. Once a line is frozen, no assistant may change its expected reason without a human editing the fixture.
A free server option is useful when every candidate and the grader need the same Node runtime. The practical workflow stays narrow, and the packet should list it as required steps rather than tips. Pin the major version, clone the packet, run the public command, and keep hidden fixtures on the grader account only. If that option is unavailable or differs from the documented image, pause instead of accepting laptop runs.
Availability is an operational convenience, not a benchmark result and not a promise that the option will stay unchanged. The article remains usable without any particular vendor, because any assistant and any shared runner can preserve fixtures and a runtime pin. What matters is the freeze step, the hidden file, and the refusal to execute before acceptance. Product convenience should not become an unexamined dependency inside the hiring loop or the score file.
Steps for one cohort
- Freeze the allowed tool names and the exact key set in the prompt before asking any assistant for sample payloads.
- Draft reject shapes with an assistant if desired, then copy only human-reviewed lines into the hidden JSONL file.
- Run the public fixture command on the pinned runtime and store the grade file beside the cohort notes.
- Run hidden fixtures from the grader account and treat a failed gate as a stop rather than a debate.
Who should skip this packet
This packet is a poor fit for roles that center visual design, product copy, or infrastructure work without a tool boundary. It is also a poor fit when the company cannot keep hidden fixtures private or cannot provide a shared runtime. Panels that want to grade communication style should use a separate exercise, because this one hides that signal on purpose. Teams without time to refresh fixtures should not reuse one public file for many cohorts, since candidates may memorize it.
The string path checks are a teaching control, not a complete isolation story for a production tool runner. A later system can still do harm if some other service ignores the accept decision and runs the path. The packet therefore forbids execution inside the graded process and leaves sandbox design to a separate review. A perfect score here is not permission to connect the same function to production credentials or customer data.
After the score file
A hiring loop that freezes fixtures, pins the runtime, and scores reason codes will keep its meaning when model output shifts. The reference checker above is enough to start that loop on a single afternoon without a new platform. Panels should still read current runtime terms before they depend on any free server, and they should refresh hidden rows before each cohort. Readers drafting fixtures with MonkeyCode should confirm current free-access and free-server terms, then run the pinned grader only after that check.
Top comments (0)