DEV Community

Riley Zhu
Riley Zhu

Posted on

Grade the Missing Sticky Store: A Feature-Flag Review Take-Home

A hiring panel should treat a green diff from a short-context assistant as incomplete evidence, because the missing module often holds the real invariant. This packet asks a candidate to review a feature-flag patch that looked finished on a disposable two-file server. The review must name each hidden check that server never ran and must end in a reject or revise decision. Panels that already publish take-home packets can place this exercise beside a coding task without turning the interview into a product demo.

Why a short sandbox changes the review

Public portfolio repos now ship quickly, and many of those changes begin in an assistant that only saw the files an author pasted. A disposable server run can show that a snippet boots while still omitting the store, the kill switch, and the audit path production calls. Interviewers who grade only the visible diff will favor candidates who trust a narrow passing run over candidates who map constraints. This packet makes that narrowness part of the specification, so the candidate scores the boundary rather than the coding style.

The exercise stays fair when every candidate receives the same sealed fixture, the same prompt, and the same ninety-minute time box. It does not require a paid model seat, a private monorepo, or a shared cloud account for the panel. It also does not ask the candidate to attack a live system, scrape hidden endpoints, or expand the fixture with outside data. The work is a review of supplied files plus a small local probe that runs on a laptop inside one folder.

The candidate prompt

Give the candidate the brief below without edits, and attach the fixture in the next section as the only allowed corpus. The brief tells the candidate that an assistant read only flags.js and flags.test.js, then reported a passing test from a disposable server. Hidden files in the packet were not mounted for that run, and the candidate must decide whether the patch may merge. The review is capped at four hundred words, must cite file names, and must not rewrite the service or add dependencies.

State that network calls are out of scope and that tool use is allowed only inside the supplied folder. Pasting the fixture into an external chat is a process failure when the team forbids third-party retention of interview material. That rule matters once the fixture grows to include realistic identifiers, even if the first version uses only synthetic keys. Record the prohibition in the same brief so every candidate sees the data rule before opening any file.

A candidate reviews a patch that an assistant produced after reading only flags.js and flags.test.js on a disposable server. That server reported a passing test, but hidden files in this packet were never mounted for the assistant run. The review stays within four hundred words, states a merge decision, and cites each file plus the predicate that citation protects. Rewrites, new dependencies, probes outside this folder, and external pastes of the packet are all process or scope failures.

Fixture the assistant did not fully see

The visible patch replaces a static boolean with a percentage bucket, and it looks complete if the reviewer never opens the store. The assistant could hash a user id, compare the bucket with a percent, and return an enabled flag without reading any other module. The hidden file in the candidate packet keeps a sticky assignment and a kill switch that the new resolver never calls. A one-file test can stay green while both invariants fail, which is the exact trap the rubric is built to catch.

Visible resolver

// flags.js — visible to the assistant
const { createHash } = require('crypto');

function bucket(userId, percent) {
  const hex = createHash('sha256').update(String(userId)).digest('hex');
  const n = parseInt(hex.slice(0, 8), 16) % 100;
  return n < percent;
}

function resolveFlag(userId, flag, config) {
  const percent = config[flag] ?? 0;
  return { flag, enabled: bucket(userId, percent) };
}

module.exports = { resolveFlag, bucket };
Enter fullscreen mode Exit fullscreen mode
// flags.test.js — the only test the assistant ran
const test = require('node:test');
const assert = require('node:assert/strict');
const { resolveFlag } = require('./flags');

test('full percentage enables the bucketed user', () => {
  const result = resolveFlag('u1', 'beta', { beta: 100 });
  assert.equal(result.enabled, true);
  assert.equal(result.flag, 'beta');
});
Enter fullscreen mode Exit fullscreen mode

Hidden invariants

// assignmentStore.js — hidden from the assistant run
class AssignmentStore {
  constructor(rows) {
    this.rows = new Map(rows);
  }

  get(userId, flag) {
    return this.rows.get(`${userId}:${flag}`) ?? null;
  }
}

function killSwitchOn(config, flag) {
  return Boolean(config.killed && config.killed[flag]);
}

module.exports = { AssignmentStore, killSwitchOn };
Enter fullscreen mode Exit fullscreen mode

Local probe

The probe is a proposed local check for this packet, not a production harness and not a benchmark. It loads the sticky row and the kill switch, then compares both with the patch output. Interviewers should treat a printed miss list as a fixture self-check, not as the candidate score. Candidates may extend the probe, but the written review still has to name the predicates in prose.

// probe.js — run locally; do not present the output as a benchmark
const { resolveFlag } = require('./flags');
const { AssignmentStore, killSwitchOn } = require('./assignmentStore');

const store = new AssignmentStore([['u1:beta', false]]);
const config = { beta: 100, killed: { beta: true } };

const sticky = store.get('u1', 'beta');
const fromPatch = resolveFlag('u1', 'beta', config).enabled;
const killed = killSwitchOn(config, 'beta');

const misses = [];
if (sticky === false && fromPatch === true) misses.push('sticky-override');
if (killed && fromPatch === true) misses.push('kill-switch');

const decision = misses.length ? 'reject' : 'accept';
console.log(JSON.stringify({ decision, misses }));
Enter fullscreen mode Exit fullscreen mode

Running node probe.js in that folder prints a reject decision with both miss names when the patch ignores the store and the switch. The command is a reproducibility check for the interviewer, not a score, and it is not a latency benchmark. A candidate who only pastes that JSON without explaining the predicates has not finished the written review. Interviewers should run the probe once on a clean machine before the first candidate and should fix relative paths if the layout differs.

Rubric

Score the written review, not the assistant that drafted the patch, and use the same four rows for every candidate. Sticky assignment is worth up to three points and requires a clear statement that an existing false row wins over a new hash bucket. A missing row may fall through to the percentage, and the review should say that fall-through is allowed only after the sticky lookup. Kill switch is worth up to three points and requires a killed flag to stay disabled even at full percentage.

The review must also note that the visible patch never reads the killed map, so the helper in the hidden file is currently dead. Evidence boundary is worth up to two points and requires an explicit refusal to treat the two-file passing run as merge proof. Decision hygiene is worth up to two points and requires a reject or revise call with file citations and no drive-by rewrite. A total of seven or higher passes, and a six may proceed only when the miss is wording rather than a dropped invariant.

Do not average this score with a live-coding score from a different exercise, because the constructs are not interchangeable. The packet measures review judgment under a stated context limit, which is a separate hiring signal from implementation speed. Panels should publish the four rows with the prompt so candidates are not graded against a hidden taste rule. Keep the point values stable across a hiring cycle, and change them only between cycles with a written note.

Required resolution order

  1. Read the kill switch and force the flag disabled when the killed map contains the flag name.
  2. Read the sticky store and return the stored boolean when a row exists for that user and flag.
  3. Fall through to the percentage bucket only when the kill switch is off and no sticky row exists.
  4. Treat a two-file passing test as non-evidence until a probe imports the hidden module and fails on both misses.
Check If true If false or missing
Kill switch for the flag Return disabled and stop Continue to the sticky lookup
Sticky row for user and flag Return the stored boolean and stop Continue to the percentage bucket
Percentage present Return whether the hash bucket falls inside that percent Return disabled

Sample review that should pass

The following review is a proposed answer key for interviewer calibration, not a transcript from a real candidate. It should not be published as evidence that any person sat this test or that any team adopted the packet. The sample cites both files, names both predicates, and refuses the green run without claiming a timing result or a customer incident. Interviewers should accept equivalent structure when the prose differs, provided the decision and the citations stay intact.

The patch in flags.js computes a stable bucket and returns the enabled bit from that bucket alone. The assignment store already holds a sticky false value for the synthetic user and the beta flag, and the new resolver never calls get. A user held at disabled would flip on whenever the percentage is high, which breaks the sticky contract the store exists to protect. The kill-switch helper reads the killed map for beta, but resolveFlag never consults that helper at any branch.

A killed flag therefore still enables when the percentage is one hundred, even though the hidden module already encodes the stop. The assistant reported a green flags test after a two-file run, and that run cannot import the store or the switch. That report is not merge evidence, so the decision is reject until the kill switch, the sticky row, and the bucket resolve in that order. Add a local probe that fails when either hidden predicate flips, and keep that probe inside the supplied folder.

Common failure modes

Candidates miss this packet in a small set of repeating ways, and naming those ways in the debrief keeps the panel off taste debates. The first cluster is a competency miss, where the candidate argues about the hash and never mentions the unread store. The second cluster is a process miss, where the candidate treats the disposable server as production or pastes the packet into a forbidden chat. Mark competency misses and process misses separately, and fail the packet on a process miss when a written data rule exists.

  • Treating the hash function as the defect, and proposing a different hash, while the sticky row and the kill switch remain unread.
  • Accepting the patch because the visible test exits zero, which repeats the assistant evidence boundary instead of challenging it.
  • Rewriting the resolver inside the review, which spends the time box on unrequested code and hides whether the invariants were seen.
  • Calling the disposable server production-like without listing mounted files, which turns a convenience into a false environment claim.
  • Pasting the packet, including any later realistic identifiers, into a hosted chat after the brief forbade third-party retention.

Reproducing the narrow run

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

MonkeyCode is an open-source coding assistant, and the only availability claims used here are free model access and a free server option. An interviewer can reproduce the short-context run on that disposable server, mount only the two visible files, and keep the hidden files off the run. Free access lets a panel try the same fixture without buying a seat for every interviewer who wants to see the narrow session. Those claims do not name a model, a quota, hardware, or duration, and current documentation must be checked first.

Clone the fixture into a fresh directory, and start the free server option only when the panel wants the original two-file assistant session. Do not mount the assignment store for that session, and keep the candidate review on the laptop rather than in the hosted chat. Record the date the terms were checked, because free access is an operator-supplied current option and not a permanence promise. If the documented offer has changed, drop the hosted step and run the probe locally, since the rubric does not depend on the host.

Readers who want a calibration pass can open the MonkeyCode project and confirm the current free access terms before mounting any files. Compare the assistant note with the local probe output, and treat that comparison as a panel aid rather than a candidate requirement. That single check is enough, and this article does not ask readers to sign up under a deadline or to prefer one assistant over another.

Limitations and who should skip this

This packet does not measure system design, people management, or security research, and it should not be the only signal in a hire. Teams that ban hosted models should skip the assistant step and still use the local probe, the prompt, and the rubric. Teams that need every candidate to hit an identical model build should not depend on a free tier whose routing may change. Roles that never touch feature flags will find the domain distracting, and a neighboring packet in their actual stack would be fairer.

Do not load real user identifiers, customer code, or secrets into the fixture or into any hosted run used to reproduce the assistant. The probe is a proposed example, and this article does not present it as an executed benchmark or as a measured pass rate. No hiring outcomes, view counts, or customer stories are claimed, because none were supplied for this account. Interviewers should run the probe on their own machines before the first scheduled candidate and should retire the fixture if a path breaks.

What the hiring file should keep

A short-context patch can be locally consistent and still wrong, and the hiring packet should score that gap directly. The prompt, the two-file boundary, the hidden store, the probe, and the ten-point rubric are enough to run the exercise in one sitting. Free model access and a free server can reproduce the narrow run when the panel verifies current terms and keeps hidden files unmounted. The candidate review, not the assistant green note, is the artifact that belongs in the hiring file.

Top comments (0)