DEV Community

Riley Zhu
Riley Zhu

Posted on

Grade the Fence Before the Green Diff: A Scope-Control Take-Home

A coding-agent take-home should score scope control before it scores polish, speed, or a green local demo. Candidates who cannot keep a narrow ticket fenced will ship adjacent edits that no reviewer requested. This packet gives interviewers a prompt, a rubric, a path probe, and a list of misses that still look competent. Free model access and a disposable server can host the rehearsal, but neither one certifies production behavior.

Why this packet exists

Hiring loops already see agents rename files, upgrade dependencies, and rewrite tests while they fix a single comment. Those extra edits often compile, so a green exit code hides the defect that the reviewer actually needed to catch. The defect is an expanded change set, not a syntax error or a missing semicolon in the requested line. Interviewers need a repeatable way to separate a careful operator from a candidate who accepts every agent suggestion.

Public developer discussion often mixes portfolio demos, vibe-coded sites, and agent events into one noisy stream. None of that noise replaces a scoped review of the diff the candidate is willing to merge. Earlier exercises associated with this account already covered delivery predicates, replay safety, latency gates, and untrusted exit codes. This packet stays on a different fault, which is an agent that widens a documentation ticket into unrelated project edits.

The prompt to hand the candidate

Interviewers should hand the candidate a tiny repository and a written ticket, then forbid edits on any production system. The ticket should request one documentation correction in a single allowed path, with no implied permission to clean the tree. The planted agent transcript should also propose a dependency bump, a workflow change, and a broad test rewrite. The candidate must accept, reject, or split those proposals, and each choice needs a written reason tied to risk.

Constraints that keep grading mechanical

The prompt should state constraints so grading does not collapse into taste, seniority, or familiarity with a particular editor. Allowed work is the documentation fix plus a probe that fails when the diff leaves the named allowlist. Disallowed work includes new services, secret handling, calls toward private hosts, and claims about unverified model quotas. A free model session and a free server may serve only as scratch space, and unverified limits must be recorded as unknown.

Ask for four artifacts, and reject any packet that omits one of them without a documented blocker. The artifacts are a decision note, a diff or patch series, a probe script, and a short miss log. A time box keeps unlimited polishing from hiding a missing fence behind extra prose and screenshots. Sample code inside the interviewer guide should be labeled as an unexecuted reference, not as a scored benchmark.

How to plant the fixture

Interviewers should commit the toy repository before the exercise and record that commit hash in the prompt. The planted transcript can live in a checked-in fixture so every candidate sees the same three out-of-scope proposals. The interview must not ask the model to invent attacks or to reach systems outside the toy tree. The candidate job is classification and fencing, not a live contest against an unrestricted coding model.

Rubric

Score the packet on observable behavior, not on the confidence or polish of the accompanying write-up. Partial credit applies only when the candidate documents the miss before the interviewer discovers it. A silent extra file is a fail even when the requested documentation sentence is factually correct. A beautiful README that never runs the probe is also a fail, because the fence was never exercised.

  1. Allowlist fidelity. The accepted diff touches only paths named in the ticket, unless the note justifies a probe file that the prompt explicitly allows.
  2. Rejection quality. Each rejected agent proposal names the risk, the missing evidence, and the safer next step in a separate ticket.
  3. Fail-closed probe. The probe exits non-zero when the diff is missing, the allowlist is empty, or an unexpected path appears.
  4. Environment honesty. The note states which claims came from the operator and which limits were not verified, including quota, hardware, and duration.
  5. No secret spill. The transcript and server notes contain no credentials, customer data, or private endpoints from a real system.

Decision table

The decision table below keeps graders aligned when two interviewers read the same packet on different days. Pass rows require both a correct diff boundary and an honest statement about what the environment did not prove. Fail rows cover expansions that compile, quota claims that were never sourced, and secrets placed on a shared server. Graders should mark the row that matches the evidence, rather than averaging a vibe across the whole packet.

Evidence in the packet Grade Reason
Diff stays on allowlisted paths, and the probe exits non-zero on a planted extra file Pass The fence is specified and actually exercised
A dependency bump is accepted inside the documentation ticket Fail The ticket was widened without a separate change
The write-up states a token quota, machine size, or expiry that the packet did not source Fail Environment facts were invented
Free-server notes include a credential or a private endpoint Fail Scratch space was treated as a secret store
The probe is missing, but the decision note lists that blocker before grading Partial The miss is visible and not laundered into a pass

Sample probe

The following Python is an unexecuted proposal for interviewers, not a result from a measured or hosted run. It reads a unified diff on standard input and an allowlist file, then exits non-zero on any unexpected path. Candidates may rewrite it, provided they keep the fail-closed branches for a missing file, an empty list, and an empty diff. Interviewers should not treat a clean run of this sample as a benchmark of any hosted model.

#!/usr/bin/env python3
'''Unexecuted reference probe: fail if a diff leaves the allowlist.'''
import pathlib
import sys

def changed_paths(diff_text):
    paths = []
    for line in diff_text.splitlines():
        if line.startswith('+++ b/'):
            path = line[6:]
            if path != '/dev/null':
                paths.append(path)
    return paths

def main():
    if len(sys.argv) != 2:
        print('usage: scope_fence.py ALLOWLIST < diff', file=sys.stderr)
        return 2
    allow_file = pathlib.Path(sys.argv[1])
    if not allow_file.is_file():
        print('missing allowlist', file=sys.stderr)
        return 2
    allowed = {p.strip() for p in allow_file.read_text().splitlines() if p.strip()}
    if not allowed:
        print('empty allowlist', file=sys.stderr)
        return 2
    diff_text = sys.stdin.read()
    if not diff_text.strip():
        print('empty diff', file=sys.stderr)
        return 2
    unexpected = [p for p in changed_paths(diff_text) if p not in allowed]
    if unexpected:
        print('unexpected paths:', ', '.join(unexpected), file=sys.stderr)
        return 1
    print('scope ok')
    return 0

if __name__ == '__main__':
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Allowlist contents

A minimal allowlist for this ticket can live in a plain text file placed beside the probe. The probe file itself belongs on that list only because the prompt explicitly allowed a new tool path. Documentation edits outside the named billing note should fail even when the added prose looks helpful to a reader. Interviewers who change the ticket must change the allowlist in the same commit, or the fixture will grade the wrong boundary.

docs/billing-note.md
tools/scope_fence.py
Enter fullscreen mode Exit fullscreen mode

Commands for the interviewer

These shell illustrations show the expected branch, and they are not a published performance result from any hosted run. The first command should print a scope confirmation when the allowlist contains that documentation path and the diff is otherwise empty. The second command should exit non-zero because the package manifest was never named in the ticket or the allowlist. Interviewers should run both commands against the candidate script, not against an unseen remote repository or a production host.

printf '%s\n' '+++ b/docs/billing-note.md' | python3 tools/scope_fence.py scope.allow
printf '%s\n' '+++ b/package.json' | python3 tools/scope_fence.py scope.allow
Enter fullscreen mode Exit fullscreen mode

An unexecuted decision note can look like the fragment below, and it exists only to show the expected tone. It is a template, not a candidate submitted answer, and it must not be pasted into a score sheet as evidence. Each line maps one path or one unknown to a single decision, which keeps the later grading pass mechanical. Interviewers should replace the paths when their toy ticket uses different filenames or a different probe location.

Accept docs/billing-note.md because the ticket names that path.
Reject package.json because a dependency bump needs its own advisory and ticket.
Reject .github/workflows/ci.yml because workflow edits are outside the allowlist.
Unknown: token quota, server size, and session duration were not verified.
Enter fullscreen mode Exit fullscreen mode

Common failure modes

Several misses recur when candidates treat the agent as an author instead of a junior patch proposer. Each miss can look productive in a portfolio screenshot, which is why the probe has to be mechanical. The list is not exhaustive, but it covers the expansions that most often sneak past a friendly reading. A strong packet rejects these moves in writing and may schedule a separate ticket for real follow-up work.

  • The candidate accepts a lockfile or dependency bump because the agent called the edit a security cleanup, without a separate advisory.
  • The candidate adds continuous integration, formatting, and a workflow file, then describes that expansion as harmless drive-by quality.
  • The probe shells out to a formatter that rewrites the tree, so the fence checks a moving target instead of the submitted diff.
  • The write-up cites a free token allotment, a machine size, or an uptime promise that this packet never verified against a primary source.
  • The candidate pastes repository secrets into a free server notebook so the agent can inspect a failure that the toy repo does not contain.
  • The decision note argues from taste, and it never maps each rejected hunk to a concrete risk, a missing fact, or a safer follow-up ticket.

Where a free session fits

Disclosure: This article was prepared as part of MonkeyCode's product outreach. The operator brief describes MonkeyCode as an open-source project and supplies two availability claims for this draft. Those claims are free model access and a free server option, which can host the toy repository for a shared rehearsal. This article does not treat numeric quotas, model names, hardware sizes, durations, or permanence as verified facts.

Readers should read the current project terms before relying on either availability claim in a live hiring loop. A bounded rehearsal still has a clear operational shape that does not depend on a particular vendor dashboard. The candidate runs the agent against the toy repository, exports the diff, and feeds that diff to the probe. The free server, if used, should hold only the toy repository and the probe, with no production credentials mounted.

What the rehearsal must not assume

If the session ends early, the packet should record the stop and grade the written fence rather than invent a completed run. Unknown budget metadata is a reason to stop the session, not a reason to guess a quota or a hardware profile. The same pattern stays useful if every product name is removed from the interviewer guide and the candidate notes. Any editor with a local diff and a Python interpreter can host the probe without a remote account.

The product mention matters only when the interviewer wants a shared scratch server that candidates need not provision themselves. That convenience does not transfer test coverage, on-call ownership, or a security review from the toy tree to production. Teams should keep the hiring record in their existing system and delete the scratch server when the exercise ends. A public write-up may describe the method, as this one does, without reproducing a candidate submitted solution.

Who should not use this approach

This packet is a poor fit for several real hiring needs, and forcing it into those loops will waste interviewer time. Teams hiring for distributed systems, kernel work, or data modeling need a domain task, not a documentation fence. Roles that require supervised production access should not be simulated on an unverified free server or a shared notebook. Interviewers who cannot read a diff should not outsource the score to an agent summary or a green status badge.

Candidates must not be told that a free allotment is permanent, unlimited, or equal for every account and region. Security exercises that ask models to attack a named organization are outside this packet and should not be improvised here. The sample parser also misses renames, binary patches, and commits that never emit a standard new-file diff header. Pair the script with a name-only diff when the repository is local, and do not call the sample a supply-chain control.

Closing

Scope is the first hiring signal in an agent-assisted change, and a short allowlist probe makes that signal cheap to grade. Interviewers can hand out the prompt, require the four artifacts, and fail packets that widen the ticket or invent environment facts. Readers who want a disposable bench for the toy repository can review MonkeyCode current free model access and free server option. They should confirm the live terms before a candidate depends on either option for a timed exercise.

Top comments (0)