DEV Community

Morgan Zhou
Morgan Zhou

Posted on

Put a Meter in the Take-Home

You open the submission on a Monday and the README still has a public URL in it. The health check answers. A debug header sits in the response, cheerful and uninvited.

The candidate finished the feature. They never finished the room. That is the failure this exercise is built to catch.

You are not testing whether someone can coax a model into a longer diff. You are testing whether they can ship a small change inside a budget, then turn the lights off. Volume is cheap now. A stop is not.

Unconstrained take-homes have started to reward the cheap part. A model will happily invent auth, a database, and a frontend while the clock burns. The folder looks impressive. The judgment is missing.

You fix that by writing the constraint into the prompt before you write the feature. Send the packet below, not a vague request to build an API. The fence is the assignment, not a suggestion.

Take-home: status budget

Time box: 90 minutes. Stop when the clock ends, even if the diff is unfinished.

Add GET /status to a tiny HTTP service. Return a JSON object with ok set to true and rev set to a git short SHA, or to the string unknown. Run it only on a disposable server you can delete the same day. A model assistant is allowed. Keep status_note.md with what you asked, what you rejected, and why you stopped.

Tear the server down before you submit, and prove it. A failing curl, or a delete confirmation, is enough. Do not put secrets, tokens, or live URLs in the repo.

Submit the diff, the note, and the teardown proof. Auth, a database, a frontend, and extra endpoints are out of scope. If model access or the server option is unavailable when you start, stop. Submit the local handler and name the blocker.
Enter fullscreen mode Exit fullscreen mode

You are hiring for restraint, so the prompt has to model restraint. A blank page invites a blank check. A bounded page invites a bounded answer.

Candidates who skip the last sentences are telling you how they treat a limit. That is the signal. Do not soften it in the pairing round just because the diff was pretty.

A free model tier and a free server option belong here only as the room you lend, not as the subject of the interview. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Those two availability claims are enough to let a candidate without a cloud bill draft the handler and run it somewhere they can delete.

Re-check both before you name them in a packet. This draft assumes no quota, no hardware size, no duration, and no model name, because those details move. If the docs disagree with this article on the day you send the packet, the docs win.

Never write the word unlimited into the instructions. Say they should use what current free access allows, and stop if it does not. A take-home that depends on a number you did not verify is a fixture you cannot grade fairly.

The sample below is a proposal. It has not been executed against a live host in this article. Read it as the shape of a passing diff, not as a measured result.

# proposal: app.py — unexecuted example
from http.server import BaseHTTPRequestHandler, HTTPServer
import json
import os

REV = os.environ.get("REV", "unknown")

class Handler(BaseHTTPRequestHandler):
    def do_GET(self):
        path = self.path.split("?", 1)[0]
        if path != "/status":
            self.send_response(404)
            self.end_headers()
            return
        body = json.dumps({"ok": True, "rev": REV}).encode()
        self.send_response(200)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(body)))
        self.end_headers()
        self.wfile.write(body)

    def log_message(self, fmt, *args):
        return

if __name__ == "__main__":
    HTTPServer(("127.0.0.1", 8080), Handler).serve_forever()
Enter fullscreen mode Exit fullscreen mode

Binding to 127.0.0.1 is the point, not a style preference. A passing candidate either stays local and closes any tunnel, or binds a disposable host and deletes it before the zip file exists. A failing candidate binds 0.0.0.0, commits the URL, and calls that deployed.

The standard library server is intentional. You are not grading framework taste. You are grading whether the process can die. A framework demo that cannot be killed is a miss, however modern the imports look.

Ask them to keep the commands in the note, next to a one-line result. The local curl should print a small JSON object and exit. If it hangs, they bound the wrong interface or never started the process, and that failure belongs in the note.

# proposal: session commands — unexecuted
export REV="$(git rev-parse --short HEAD 2>/dev/null || echo unknown)"
python app.py &
PID=$!
sleep 0.3
curl -fsS http://127.0.0.1:8080/status
kill "$PID"
wait "$PID" 2>/dev/null || true
# after the remote host is deleted, this must fail closed
curl -fsS --max-time 5 "$OLD_PUBLIC_URL/status" && echo "STILL UP" || echo "GONE"
Enter fullscreen mode Exit fullscreen mode

STILL UP is a failing grade even if the JSON was perfect. The feature without the teardown is the Monday folder you already opened. GONE, plus a short note, is enough. You do not need a novel about cleanup.

You grade the stop, not the prose. On the sixth packet, every README starts to sound thoughtful. A small local checker keeps the first pass boring, which is what you want from a first pass.

It scores files on disk. It does not call a model, and it does not open a socket. Treat the script as an unexecuted proposal, the same way you treat the sample app.

# proposal: grade_stop.py — unexecuted local checker
import json
import sys
from pathlib import Path

root = Path(sys.argv[1])
note = (root / "status_note.md").read_text(encoding="utf-8", errors="replace")
diff = (root / "submission.diff").read_text(encoding="utf-8", errors="replace")
low = (note + "\n" + diff).lower()

score = 0
reasons = []

def add(points, ok, why):
    global score
    if ok:
        score += points
        reasons.append(f"+{points} {why}")
    else:
        reasons.append(f"+0 missing: {why}")

add(2, "/status" in diff and "ok" in diff, "handler mentions /status and ok")
add(2, "127.0.0.1" in diff or "localhost" in low, "local bind or a disposable-host note")
add(2, any(word in low for word in ("teardown", "deleted", "gone")), "teardown is recorded")
add(2, not any(mark in low for mark in ("api_key=", "api_key:", "secret=", "token=")), "no inline secret assignment")
add(2, any(word in low for word in ("stopped", "out of scope", "did not", "blocker")), "a stop or a blocker is named")

print(json.dumps({"score": score, "out_of": 10, "reasons": reasons}, indent=2))
Enter fullscreen mode Exit fullscreen mode

You run it from the folder that holds one candidate, then you read the reasons before the number.

python grade_stop.py ./candidate-a
Enter fullscreen mode Exit fullscreen mode

A 10 with no teardown sentence means the checker is wrong, not that the candidate is brilliant. A 6 with a clear stop is often the person you want in the pairing round. The script greps. It cannot tell a real delete from the word deleted, so the number is a crib, not a verdict.

Said in sentences, the rubric is short on purpose. Two points if the handler exists and the body carries the two fields. Two points if the process is local, or the note explains a host they can delete the same day. Two points if teardown is evidenced rather than waved at.

Two more if the tree has no pasted secret. Two more if the note names a refusal, such as an extra endpoint, a live URL, a second model pass after the clock, or a blocker when access failed. Ten is not the goal. A legible stop is the goal.

The failures rhyme, which is why the prompt is fussy. One candidate pastes a model transcript and never writes the handler, hoping the transcript counts as work. Another accepts every suggestion, so the diff grows auth and a database while the status route never appears.

A third leaves the host running and writes that they are happy to share the link. A fourth copies a dashboard key into app.py because a completion said to inline it. A fifth treats free access as infinite, spends the session on renames, and blames the tool in the note.

You fail the fifth for the same reason you fail the third. They did not notice the meter. Blame is not a teardown.

There is a softer miss you should not over-punish. A candidate who binds locally, returns rev as unknown because git was missing on the disposable host, and writes that down, has done the job. The prompt allowed unknown. Pedantry about the short SHA selects for people who match your laptop, not for people who read the constraint.

Skip this approach if the role is production on-call and you need judgment about real traffic, real data, and a real rollback. A toy status route will not show that. Skip it if you will not read status_note.md yourself.

Skip it if your company forbids external model tools, or if sending a candidate to any third-party server would break hiring policy. Skip it in any week when you cannot confirm that the free model access and the free server option still exist on the terms you are about to describe. A stale perk in an interview packet is a broken fixture, and candidates can tell.

The limits are deliberate. Ninety minutes measures a slice of taste, not seniority. Free access can be rate-limited, region-limited, or gone by the time they start, so the prompt already says what gone looks like: stop, name the blocker, submit the local handler.

Do not promise hardware, a duration, or a token budget you have not read in a current primary doc. If you mention a product at all, point candidates at that doc and keep the product out of the score.

If you need a room for this exercise, confirm MonkeyCode's current free model access and free server option in their own docs, then lend that room without grading loyalty to it. You are grading the stop. The tool is only the light they were supposed to switch off.

Top comments (0)