The problem
Two cycles ago you shipped the migration that cut p99 latency in half. Everyone said it was the best work of your career. This cycle you did it again — same rigor, same impact, a different system, a bigger number. And your rating went down.
Nobody can point to a mistake. Your manager says the work was "great, consistent with last cycle." That sentence is the whole problem, and almost no one clocks it as one.
Why it happens
Calibration committees don't score absolute output. They score trajectory — the story of change since last time. A committee sits in a room with fifteen minutes per person and a forced curve to fill. The fastest signal they can extract isn't "was this good," it's "was this different from what we already believe about this person."
If your story last cycle was "shipped the big win," this cycle's story needs to be a new big win, ideally a bigger one, or the room reads you as flat. Repeat excellence doesn't update anyone's prior, so it doesn't generate the kind of narrative that wins the argument in the room. Growth reads as signal. Consistency reads as noise.
There's a second, quieter mechanism: advocacy fatigue. The person who fought for you last cycle spent their political capital telling that story once. They don't have a fresh anecdote to spend capital on this time, and reusing the old one in the room sounds like padding, not evidence. So the strongest voice in the room for you is quieter the second time, even though the work was just as good — or better.
This is the same failure mode as recency-weighted metrics in any noisy system: the committee is measuring the derivative, not the value. A flat line at a high number gets treated like nothing happened, even though "nothing happened" at that altitude is the hard part.
What to do about it
You can't fix a broken measurement system from below, but you can stop feeding it the wrong inputs.
First, manufacture the delta yourself. If the work's absolute value is flat, change the framing before the room does it for you: scope, blast radius, who now depends on it, what broke without it. "I did the same kind of thing" and "I did the thing that is now load-bearing for four other teams" are the same work described at different resolutions — pick the one that shows change.
Second, get a new advocate into the room, not just a returning one. A different manager, skip-level, or partner team lead telling your story for the first time reads as fresh evidence, even if the underlying work is a continuation.
Third, name the trap out loud in your self-review. Write the sentence "this is the second consecutive cycle where I delivered at this level, and consistency at this altitude is the achievement" — you're pre-empting the "flat" read before the committee invents it themselves.
Fourth, if you're the one calibrating other people's cases: catch yourself doing this to someone else. Ask "would this same output, from someone I hadn't seen do it before, read as a big deal?" If yes, you're penalizing them for not surprising you, which is not a performance measure — it's a novelty tax.
Key takeaways
- Calibration committees implicitly score trajectory (change), not absolute output — repeat excellence gets under-weighted because it doesn't update anyone's belief.
- Advocacy fatigue compounds it: the same champion has a weaker story to tell the second time, even for equally strong work.
- Counter it by reframing scope/impact to surface real delta, bringing in a fresh advocate, and naming the "flat" narrative before the room invents it.
- If you sit on a calibration committee, watch for this bias in how you evaluate others — "not surprising" isn't the same as "not valuable."
Top comments (0)