A North Star Metric is one number that represents value delivered. That is the whole definition.
The difficulty is never understanding it. The difficulty is that when you pick one honestly for something you actually do, it is almost never the number you have been keeping.
Try it on a job search
Suppose you are looking for work, and you want one number to tell you whether the search is going well.
The first candidate that comes to mind is applications sent. It is easy to count, it goes up every week, and it feels like effort. Most people searching for a job are tracking it whether they call it a metric or not.
It is also close to useless. Fifty applications sent tells you nothing about whether any of them were read. It is perfectly possible to send fifty and be no closer to a job than when you started, and the number will still look like a productive month.
So try the next one: hours spent job searching. This is worse. It measures input rather than output, and it rewards inefficiency directly. Spending six hours rewriting a cover letter for a role you were never going to get scores higher than spending twenty minutes on one you were.
Try again: responses received. Better, because it requires someone on the other side to have done something. But it still counts rejections, and a rejection is not value delivered to anyone.
What you are actually after is something like interviews earned from roles you would genuinely accept. That number only moves when a real employer, for a job you actually want, has decided you are worth their time.
| Candidate metric | What it measures | Why it fails |
|---|---|---|
| Applications sent | Effort | Says nothing about whether anyone read them |
| Hours spent searching | Input | Rewards inefficiency |
| Responses received | Someone reacted | Counts rejections too |
| Interviews earned (roles you'd accept) | Value delivered | — |
Notice what happened across those four attempts. Each one moved further from what you did and closer to what resulted. That direction of travel is the whole idea.
Why the wrong metric is so tempting
Activity is easy to count. Applications sent is a number you already have. Interviews earned requires you to define what counts as a role you would accept, then wait, then attribute the outcome to something you did weeks ago.
Activity also rises reliably. Send ten applications, the number goes up by ten. Outcome metrics sit still for long stretches and then move in steps, and a flat week on a number you care about feels much worse than a rising week on a number you do not.
That asymmetry is the reason people pick badly, and it has nothing to do with not understanding the concept. Choosing the honest metric means accepting a number that will often refuse to move while you are working hard. Most people, reasonably, would rather watch the one that rewards them daily.
The same trade shows up at work. Tickets closed, commits pushed, messages answered, items reviewed. All real work. None of them say whether anything was better afterwards. Compare what each one could be instead:
- Tickets closed → problems that stayed fixed
- Commits pushed → features users actually used
- Items reviewed → items usable afterwards
One number is not enough
A North Star on its own is gameable, usually without anyone intending to game it.
Optimise interviews earned and the obvious move is to apply only to roles where your odds are highest, which over a few months quietly narrows your search to jobs slightly below what you could get. The metric rises. The search gets worse.
So it needs a counterweight (often called a guardrail metric), and one is normally enough. Here it might be the share of applications going to roles that would be a genuine step up. That number stops you buying interviews with ambition.
The pairing is the point.
One metric tells you how to grow, the other tells you what growth is not allowed to cost.
Counting your own work honestly
Once you have tried to pick an honest metric for one mundane thing, it becomes hard not to do it for everything else.
It changes what you report about your own work most of all. It is easy to say you reviewed two hundred items last week, and much harder to say how many were usable afterwards. The first number is always available and always flattering. The second one is the one that was worth doing.
The question that gets you there is short:
If this number goes up and nothing else changes, has anything actually improved?
If the answer is no, you are counting activity.
Top comments (0)