The take-home was a small service: read a CSV of transactions, apply some rules, expose two endpoints, write tests. It had been the first stage of our hiring process for four years, and it did its job well. Submissions spread out. Some were rough and fast, some careful and slow, some over-engineered, some missed a rule, and that spread told us things.
By last autumn the spread had gone. Twenty submissions in a row were clean, idiomatic and fully tested, with a README, a Dockerfile and a CI config we hadn't asked for. They were also close to interchangeable, right down to how the test file was laid out. Nobody was cheating in any sense we'd defined. They'd used the tools every working engineer uses now, and those tools turn a well-specified problem into a good answer in an hour.
What our test measured was whether a candidate could drive a coding agent through an easy task, and everyone can, so it told us nothing.
What we were actually hiring for
Before changing anything we had to answer a question we'd been dodging: what does an engineer on this team do that the agent doesn't?
After some frank discussion we landed on four things. They decide what to build, at the level of a ticket, when the ticket is vague. They notice something is wrong in code they didn't write, before it ships. They can tell when the agent's output is right and when it's plausible but wrong, and that second part is the whole skill. And they carry the system's context across weeks, so what gets built this week fits what was built last month.
You won't see any of that in a take-home with a clear spec. You see all of it when you watch someone work with an agent on a task that's slightly wrong.
The exercise
The new first stage is a ninety minute video call. The candidate shares their screen and uses whatever agent and editor they normally use. We tell them so in as many words: use your tools, we want to see how you work, not how you work with one hand tied.
The task is a small existing codebase, about 2,000 lines, that we hand over at the start. It comes with a feature request written the way real tickets get written, meaning incompletely. And there are three things wrong with it that we don't mention.
One is a bug in the existing code that the feature request will bring to the surface. One is a place where the obvious way to build the feature is wrong, for a reason you only see by reading a different part of the codebase. The third is a test that passes when it shouldn't, because it asserts the buggy behaviour.
We ask the candidate to implement the feature, and that's the whole brief.
What we watch
Finishing isn't what we're looking at. Most people more or less finish, because the agent writes an implementation in a few minutes. We watch what happens around that.
Do they read the codebase before prompting, or prompt first and read the output? Either can work. Prompting first and then reading the output carefully is fine. Prompting first and accepting the output isn't, and the exercise is built so that accepting it ships the bug.
When the agent's implementation is plausible but wrong, because of that thing in the other part of the codebase, do they notice? This is the middle of the whole exercise. The strongest candidates notice within a few minutes, usually because they read the module the feature touches and see the constraint. The middle group notices when a test fails, if they wrote a test that covers it, which the agent won't do unprompted since it doesn't know the constraint exists. The rest ship it.
When they find the existing bug, what do they do with it? They can fix it silently, fix it and mention it, note it and leave it, or ask. Any of the last three is fine. The first is a small flag, because silently changing behaviour outside the ticket is the habit that causes incidents.
When the agent's test suite goes green, do they believe it? The passing test that should fail is there to see who reads tests as claims and who reads them as proof. We want the ones who read it and say "this is asserting the wrong thing".
And how do they talk to the agent? This one surprised us. Some candidates give the agent context, "this module has a constraint that X, implement the feature so that it holds". Others paste in the ticket as it is and iterate on what comes back. The first group finishes faster and ships fewer bugs, and the difference comes down almost entirely to whether they'd read the code before asking.
What the exercise found
We ran it with about forty candidates over the winter and spring, and a few results changed how I think about the job.
Years of experience predicted almost nothing. Some of the best sessions came from people with three years and some of the worst from people with fifteen. What separated them was whether they'd changed how they work to fit an agent, or were using it as a faster autocomplete while working the way they did in 2019.
Reading predicted the most, much more than writing. Candidates who spent the first ten minutes reading the codebase, before touching the agent, found all three problems far more often. Reading code was always the underrated skill, and the agent made it the main one.
The second predictor was being comfortable saying "I don't know if this is right". There's a moment in the exercise where the agent's output looks right and isn't. Candidates who stopped there and said "I want to check this against the other module" did well. The ones who said "looks good" didn't. That's as much temperament as skill, and I'd now hire for it over almost anything else.
The rest of the loop
After the session there's one more technical stage, a conversation about a system the candidate built, where we ask about their decisions and what went wrong. We didn't change it, because it never measured typing.
The take-home is gone and we haven't missed it. The ninety minutes takes more of our time per candidate than reviewing a submission did, and it tells us ten times as much.
We also changed what the offer says about tools. It used to say nothing. Now it says the agent is part of the job, that we expect people to use one, and that we expect them to read what it produces. That second sentence is there because the failure we saw in the exercise is the same one we see in the team, and naming it on day one is cheaper than finding it in a post-mortem.
The exercise, in enough detail to steal
People ask for the codebase. I won't share it, because it is the exercise, but the recipe is more useful than the artefact anyway, so here it is.
Start from a real, small service. Ours is a cut down version of a scheduling API we used to run: a Postgres schema with five tables, an HTTP layer with eight endpoints, a job that sends reminders, and about 60 tests. Two thousand lines is the right size. Any smaller and there's nowhere to hide the problems. Any larger and ninety minutes isn't enough to read it.
Write the feature request the way your product manager writes them. Ours is four sentences long and asks for recurring events. It doesn't say what happens to a recurring event's reminders, or whether editing one occurrence edits the series, or what the API should return for a series.1
Then plant the three problems. The existing bug is in the reminder job. It uses the event's start time in UTC and the user's timezone offset from when the event was created, so an event created before a daylight saving change sends its reminder an hour off once the clocks move. Recurring events expose it because a series spans the change. The constraint in the other module is that the reminder job assumes one row per event. The obvious way to build recurring events, one row per occurrence, would send one reminder per occurrence, and for a daily event over a year that's 365 reminders on the day the series is created. The wrong test asserts that a reminder goes out at the stored offset, so it passes against the bug.
None of the three is a trick. Problems like these exist in every real codebase, and a strong engineer who reads the reminder job before building the feature runs into all of them.
Send the candidate the repository fifteen minutes before the call so the first ten minutes of the session don't go on setup. Tell them the session is recorded for the panel, that we'll be watching their screen and their agent's conversation, and that both are fine to show us.
During the session the interviewer says almost nothing. Two prompts are allowed: "how is it going" at the halfway mark, and "what would you want to check before merging this" at seventy-five minutes. That second one gives candidates who noticed something and kept quiet their chance to say it.
Score four lines, yes or no each, written down before any discussion with the panel. Did they read before prompting? Did they catch the constraint? Did they question the passing test? Did they deal with the existing bug in a way that wasn't silent? Two yeses gets you to the next stage.2
What changed for the team
The exercise turned out to show us ourselves. The first four times we ran it, the panel disagreed on scores. When we dug into why, it was because the panel members worked differently with agents themselves. Two of us read first and prompt second, the other two prompt first and review after, and each pair was scoring candidates who worked like them higher.
That was useful to learn about ourselves, and it led to a team session where we ran the exercise on each other. Everyone on the panel now does it once a year with a fresh planted bug, and the scoring calibration meeting comes after.3
The other change is that the ninety minutes became the template for onboarding. A new engineer's first task is a small feature on a real service with a real, unannounced bug next to it, and their onboarding buddy watches the same four things the interview panel watched. The interview and the first week now run into each other, which never happened when the interview measured typing and the job measured reading.
The candidates' side
Several candidates told us afterwards it was the first interview in a year where the process matched the job. A couple said it was the first where they'd been allowed to use their tools at all.4
One candidate, who we hired, said something that stayed with me. She said the exercise was the first time an interviewer had seemed to care whether she could tell when the machine was wrong. That's the job now. Checking your own work and other people's was always part of it, but in two years the amount of plausible work that needs checking went up by an order of magnitude, and the hiring process had to follow the job.
Originally published at zeybek.dev.
-
A candidate who asks about those in the first ten minutes has already told you a lot. ↩
-
Four is rare, and every one of those people has turned out to be a strong hire. ↩
-
It's the only training we run that people ask to do again. ↩
-
I found that remarkable, since the alternative is an interview that measures a way of working nobody uses any more. ↩
Top comments (0)