DEV Community

Riley Zhu
Riley Zhu

Posted on

Sent Before Ack: A Kill-Window Fixture for an Outbox Take-Home

A status write that lands before broker acknowledgement is a data-loss bug, and a happy-path print will not catch it. This packet asks a candidate to expose that window with a kill hook, then recover the row without double-applying the payload. The exercise stays small on purpose, because a hiring loop should grade one state machine rather than a platform rewrite. A free check host can run the fixture, but a green process exit is not a durability claim.

What this packet decides

The packet decides whether the candidate can separate three facts that assistant drafts often blur together. Those facts are write order inside one process, what a later process observes after a crash, and what an in-memory double may prove. Candidates who collapse those facts into a claim that the sample looked fine do not pass, even when the surrounding prose is clean.

Interviewers should hand out the buggy worker, the in-memory broker, and a one-page prompt, and they should withhold the fixed state machine. Withholding the fix matters, because the point is to see whether the candidate derives a distinct in-flight state. They should also withhold invented incident stories, since colorful production anecdotes are not evidence in this packet. The grader later compares the candidate fixture output with the expected stranded row, rather than scoring tone.

The candidate prompt

The prompt below is the artifact a coordinator can paste into a private repository for a timed exercise. It names the constraint, the forbidden shortcuts, and the required observation, and it does not ask for a slide. A blog post or a broker product comparison would measure writing taste, which is not the decision this packet makes.

You are given buggy_worker.py and an in-memory broker double.
The worker marks an outbox row sent before publish returns.
Add a kill hook that stops the process after that status write.
Show that the row is sent while the broker message list is empty.
Propose a state machine that can resume after that crash.
Publish must be idempotent on the row id.
Do not add a real network broker.
Do not treat a clean run, without the kill hook, as a pass.
Time box: 90 minutes. Submit the test file and a short note.
Enter fullscreen mode Exit fullscreen mode

The time box is a hiring constraint, not a claim about how long a careful production fix should take. Coordinators who need a deeper broker exercise should choose a different packet with a real cluster. This packet isolates the early commit, the resume path, and the idempotent key, and it stops there.

Fixture the grader actually runs

The teaching fixture uses a hook instead of a real signal, so a reviewer can run it inside one process. A later optional step can start a subprocess and send a signal, but that step is not required to see the bug. The code below is a proposal for the packet, and it is meant to be executed as a unit test. It should not be imported into a running service, because the hook raises SystemExit on purpose during the tick.

from dataclasses import dataclass

@dataclass
class Row:
    id: str
    payload: str
    status: str
    attempts: int = 0

class MemoryBroker:
    def __init__(self):
        self.messages = []
        self.seen = set()

    def publish(self, key: str, body: str) -> None:
        if key in self.seen:
            return
        self.seen.add(key)
        self.messages.append(f"{key}:{body}")

class BuggyWorker:
    def __init__(self, rows, broker, kill_after_status=False):
        self.rows = rows
        self.broker = broker
        self.kill_after_status = kill_after_status

    def tick(self) -> None:
        for row in list(self.rows.values()):
            if row.status != "pending":
                continue
            row.status = "sent"
            row.attempts += 1
            if self.kill_after_status:
                raise SystemExit("killed after status write")
            self.broker.publish(row.id, row.payload)

def show_buggy() -> None:
    rows = {"m1": Row("m1", "hello", "pending")}
    broker = MemoryBroker()
    worker = BuggyWorker(rows, broker, kill_after_status=True)
    try:
        worker.tick()
    except SystemExit as exc:
        print(exc)
    print(rows["m1"].status, broker.messages)
Enter fullscreen mode Exit fullscreen mode

The listing is an unexecuted teaching proposal for the packet, and a coordinator should run it before treating the assertions as confirmed. A local failure of the import path is a setup problem, not evidence about the outbox state machine. The expected diagnosis print is the exit text, then the word sent, then an empty list.

Diagnosis command

The matching tests make the failure observable without a log scrape or a manual console session. The first test expects a sent row and an empty message list after the hook fires. The second test belongs to the sample solution and should fail when it is pointed at the buggy worker.

def test_kill_window_strands_sent_row():
    rows = {"m1": Row("m1", "hello", "pending")}
    broker = MemoryBroker()
    worker = BuggyWorker(rows, broker, kill_after_status=True)
    try:
        worker.tick()
    except SystemExit:
        pass
    assert rows["m1"].status == "sent"
    assert broker.messages == []
Enter fullscreen mode Exit fullscreen mode

Graders can run the file with one ordinary command from a clean checkout on a machine that has Python available. The command does not depend on a particular model vendor, a private token, or a hosted console. A failing first test is the pass condition for the diagnosis half of this packet, before any patch is discussed.

python -m pytest -q test_outbox_takehome.py
python -c "from outbox_takehome import show_buggy; show_buggy()"
Enter fullscreen mode Exit fullscreen mode

A candidate who edits the assertion until the suite goes green has not found the bug. The grader should diff the submitted test against the prompt, and should not stop at the process exit code. A green exit after a deleted hook is a fail, even when the remaining assertions describe the happy path accurately.

Rubric

Scores stay binary on purpose, because partial credit invites essay padding during a short hiring exercise. Each row is something the submitted test or the short note must show without extra interpretation. A pass on one row does not borrow credit from a neighboring row, which keeps a fluent note from hiding a skipped hook.

Check Pass looks like Fail looks like
Kill hook after the status write Row is sent and the broker list is empty Only the happy path is executed
Resume state An in-flight status is distinct from sent An exception handler leaves sent in place
Second tick One message, then sent The second tick skips the row forever
Idempotent key A replay does not append a second body Duplicate publish is renamed and left in place
Scope The patch stays inside the worker and the test A new queue product or cloud account appears
Evidence note The note states what the double cannot prove The note claims disk durability or multi-host fencing

A candidate can pass the table and still be a weak hire for reasons this packet does not measure. The table does not measure system-design breadth, on-call judgment, or writing aimed at a non-engineering audience. It measures whether this early-commit bug can be made to fail in a way another engineer can rerun.

Sample solution sketch

The sample solution sets an in-flight status before the broker call and writes sent only after publish returns successfully. A crash during that in-flight status leaves a row that the next tick must claim again. The broker double ignores a repeated key, which models idempotent produce rather than exactly-once delivery across a cluster.

In-flight worker

class LeaseWorker:
    def __init__(self, rows, broker, kill_after_status=False):
        self.rows = rows
        self.broker = broker
        self.kill_after_status = kill_after_status

    def tick(self) -> None:
        for row in list(self.rows.values()):
            if row.status not in ("pending", "publishing"):
                continue
            row.status = "publishing"
            row.attempts += 1
            if self.kill_after_status and row.attempts == 1:
                raise SystemExit("killed while publishing")
            self.broker.publish(row.id, row.payload)
            row.status = "sent"
Enter fullscreen mode Exit fullscreen mode

The recovery assertion is the other half of the artifact, and it should be present in the submission. After the first tick dies, the row must not say sent, and the message list must still be empty. After the hook is cleared, one message exists and the status catches up to sent on that same row.

Assertions the note must include

def test_resume_publishes_once():
    rows = {"m1": Row("m1", "hello", "pending")}
    broker = MemoryBroker()
    worker = LeaseWorker(rows, broker, kill_after_status=True)
    try:
        worker.tick()
    except SystemExit:
        pass
    assert rows["m1"].status == "publishing"
    assert broker.messages == []
    worker.kill_after_status = False
    worker.tick()
    assert broker.messages == ["m1:hello"]
    assert rows["m1"].status == "sent"

def test_replay_does_not_duplicate():
    rows = {"m1": Row("m1", "hello", "publishing")}
    broker = MemoryBroker()
    broker.publish("m1", "hello")
    worker = LeaseWorker(rows, broker, kill_after_status=False)
    worker.tick()
    worker.tick()
    assert broker.messages == ["m1:hello"]
    assert rows["m1"].status == "sent"
Enter fullscreen mode Exit fullscreen mode

This sketch does not implement lease expiry, a fencing token, or a dead-letter cap for poison payloads. Those controls belong in a follow-up prompt when the open role actually owns a long-running worker fleet. Naming them as out of scope is a better note than silently adding untested machinery to the patch. The idempotent key covers the dual window, where publish succeeded and the status write never landed.

Failure modes that still look polished

Several submissions read well and still lose the message or append it twice under a replay. The list below is what graders should mark before they comment on sentence style or confidence. These misses show up in the diff, which is why the grader reads assertions before adjectives.

  • A try/except around publish logs the error and leaves the row sent, which hides the loss from every later tick.
  • A retry loop runs without a maximum attempt or a dead-letter status, so one poison payload can block the worker.
  • A timestamp is labeled a lease, but nothing stops an expired worker from publishing after a newer worker claimed the row.
  • The note accepts the assistant draft because a manual print matched the payload, and the kill hook was never executed.
  • The kill hook is removed so the suite passes on a clean machine, and the note describes that deletion as the fix.
  • The row is deleted before acknowledgement, which removes the only local copy of the payload if publish then fails.

Polished wording does not offset a missing assertion, and graders should quote that assertion or its absence. That habit keeps the written score about the state machine rather than about how certain the note sounded. A second reader should be able to rerun the quoted test without asking what the candidate meant.

How the hour is scored

  1. The coordinator clones the private repo and runs the diagnosis test before opening the candidate note, so the fixture result comes first.
  2. The grader records the status string and the message list, and treats a missing kill hook as a fail even if other tests pass.
  3. The grader then reads the proposed state names and checks that sent is written only after a successful publish return.
  4. The grader replays the same row id and confirms the message list length stays at one after the second successful tick.
  5. The grader marks any durability claim that the in-memory double cannot support, including disk flush and multi-host fencing.

This order keeps the review mechanical, which matters when several candidates submit notes that sound similar. Similar notes are expected if the same free model access sketched them, so the differing artifact is the test diff. A coordinator who cannot explain the stranded row should not score the packet, regardless of seniority.

Where a draft host and a check host fit

Assistant drafts are useful for the first worker file and for a candidate note, and they are a poor source of the verdict. The task brief describes MonkeyCode as an open-source project with free model access and a free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach. License terms and current limits should be confirmed in the project repository before a hiring loop relies on either option.

This article does not name models, token quotas, hardware, or a duration, because those details change and were not re-checked against a primary source. The workflow worth teaching is that split, not a preference for one vendor console over another. One step uses free model access to sketch the buggy worker and a short review note for the candidate packet. A second step uses the free server option only as a check host for the pinned fixture and its expected failure.

The score still comes from the rubric table, including the explicit note about what the double cannot prove. If either option is unavailable, the same files run on any machine with Python and pytest installed. That fallback is required, because a hiring packet must not depend on a private console staying up. Coordinators who want that split can check MonkeyCode's current free model access and free server docs, then keep the fixture in git.

Limits and who should skip it

The in-memory double does not model disk flushes, a network partition, or duplicate delivery from a real broker. A green run on a free server shows that this fixture behaved on that host, and it shows nothing about production latency. Free model access and a free server option are availability claims for drafting and checking, not a service level agreement. Numeric quotas, hardware size, and retention were left out on purpose, so this post does not freeze a limit that may already have changed.

Teams hiring for production broker ownership should not use this packet as the only exercise, because the double hides cluster failure. Loops that score confidence, slide quality, or how closely a note matches an assistant's tone should skip it. Anyone who would treat a free-tier check host as proof of durability should not adopt the workflow described here. Candidates without time to run a test should not be judged on this packet, because the artifact is the failing assertion.

The useful outcome is a stranded row that another engineer can reproduce, followed by a resume path that publishes once. Everything else written in the note remains optional color around that pair of results from the fixture. A hiring loop that can show those two results has enough evidence to score this packet and to stop.

Top comments (0)