DEV Community

speed engineer
speed engineer

Posted on

Your Promo Case Wasn't Judged on Its Own Merits. It Was Judged Against Whoever Went Before You.

The problem

Two engineers, same level, same tenure, comparable years of scope. One gets "exceeds expectations" in calibration. The other gets "meets." Same manager, same evidence packet quality, same quarter. The only real difference: the order they were discussed in a three-hour room with forty cases on the docket.

I've sat in enough calibration meetings — as the person presenting cases and as one of the raters — to know this isn't an edge case. It's the default failure mode of how leveling and performance calibration actually happens, and almost nobody names it out loud.

Why it happens

Calibration meetings are a textbook setup for anchoring bias. The first case discussed in the room sets an implicit reference point for "what exceeds looks like" and "what meets looks like" — and every case after it gets judged relative to that anchor, not against a fixed bar.

Run the math on a typical session: eight raters, forty people to calibrate, three hours. That's under 4.5 minutes per person once you subtract the inevitable tangents. Nobody has time to re-derive a rubric from first principles for case #23. They pattern-match to case #1 or #2, because that's what's fresh and vivid in working memory.

It compounds in a specific direction, too. If the first case presented is a strong, well-documented "exceeds," the bar for everyone after is dragged up — good "meets" performers start looking merely adequate by contrast. If the first case is a middling "meets," the opposite happens: the bar sags, and genuinely strong later cases don't stand out because the room's calibration is already loose. The order of presentation isn't neutral. It's load-bearing.

There's a second-order effect that makes this worse: managers who present early in their careers at a company learn (correctly, if cynically) that going first with your strongest case is a real lever. It's not gaming the system maliciously — it's responding rationally to an unstated rule nobody wrote down.

What to do about it

The fix isn't "try to be more objective." Anchoring survives good intentions; it's a property of how comparative judgment works under time pressure, not a character flaw in the raters. You have to change the structure, not the willpower.

Three things that actually work, in order of how much they cost to implement:

Independent pre-scores before the room opens. Every rater submits a written score against the written rubric — not a gut read, an evidence-backed score — before anyone talks. The discussion becomes about resolving disagreement between pre-scores, not building consensus from a blank slate live in the room. This alone kills most of the anchoring effect, because the anchor gets set individually, forty separate times, instead of once for the whole room.

Randomize presentation order every session. If order is going to have an effect no matter what, at least make the effect random instead of systematic. Don't let managers self-select who goes first. Draw it.

Write the rubric's evidence bar down before the meeting, with real examples. Not "exceeds = significant impact." A concrete example of what shipped, what scope it touched, what the blast radius of failure would have been. Cases get compared to a written example instead of to whichever case is loudest in short-term memory.

None of these are exotic. They're the same fixes structured interviewing uses to fight interviewer anchoring — write the rubric first, score independently, discuss after. Performance calibration is just structured interviewing with worse incentives to fix it, because the "customer" of a bad calibration outcome is an employee who usually never finds out why.

Key takeaways

  • Calibration meetings anchor hard on whichever case is discussed first — the effect is structural, not a rater character flaw.
  • Time pressure (minutes per case) is the mechanism: nobody has bandwidth to re-derive a bar from scratch forty times in a row.
  • Fix the process, not the people: independent pre-scores, randomized order, and a written evidence-based rubric before the room opens.
  • If your org calibrates on live-discussion consensus with no pre-scoring, the outcome for any given person depends more on scheduling than on their work — and that's worth saying out loud to whoever owns the process.

Top comments (0)