<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Atticus</title>
    <description>The latest articles on DEV Community by Atticus (@hey_atticus).</description>
    <link>https://dev.to/hey_atticus</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3843629%2F747485fa-94c0-4ae6-b16e-9007d766937f.jpg</url>
      <title>DEV Community: Atticus</title>
      <link>https://dev.to/hey_atticus</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hey_atticus"/>
    <language>en</language>
    <item>
      <title>Decide What Counts as a Win Before You Test</title>
      <dc:creator>Atticus</dc:creator>
      <pubDate>Thu, 23 Jul 2026 19:14:22 +0000</pubDate>
      <link>https://dev.to/hey_atticus/decide-what-counts-as-a-win-before-you-test-1pnb</link>
      <guid>https://dev.to/hey_atticus/decide-what-counts-as-a-win-before-you-test-1pnb</guid>
      <description>&lt;h1&gt;
  
  
  Decide What Counts as a Win Before You Test
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Meta description:&lt;/strong&gt; Medicine proved that picking your primary metric after seeing the data is a structural bias. The five-minute fix most experimentation programs skip.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A test finishes. The metric everyone agreed to watch is flat. But average order value moved, or a secondary funnel step improved, and the room quietly reframes what "the win" was — after seeing which number happened to move.&lt;/li&gt;
&lt;li&gt;The cleanest large-scale evidence that this specific sequence is a real bias, not a harmless judgment call, comes from medicine: large NIH-funded trials that hadn't registered a primary outcome in advance reported a significant benefit far more often than the same kind of trials after registration became mandatory.&lt;/li&gt;
&lt;li&gt;The fix isn't fewer metrics. It's naming, in writing, which one metric decides ship or kill, what bar it has to clear, and what happens for each outcome — before the test starts, when no result yet exists to make one candidate metric look more flattering than another.&lt;/li&gt;
&lt;li&gt;Every other metric still matters. It just can't retroactively become the proof for a test it wasn't named to decide — it can only generate the next hypothesis.&lt;/li&gt;
&lt;li&gt;The actual skill isn't precommitting itself. It's choosing, in advance, the metric that's genuinely closest to the business decision at stake — not the one that happens to be easiest to move.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A test wraps. The metric everyone agreed mattered when the test was planned — checkout conversion, say — is flat, within noise, a clean null. In the same readout, average order value ticked up. Or time-to-first-value improved. Or a downstream retention cohort looks slightly better, though it's early. Someone reasonably says: "isn't that actually what we cared about?" The room nods. The test gets written up as a win, on the metric that happened to move, and nobody in the room did anything that feels like misconduct. That's exactly what makes this failure mode durable — it doesn't require anyone to act in bad faith, only to make a completely natural decision in the wrong order.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cleanest evidence this is a real bias, not a judgment call
&lt;/h2&gt;

&lt;p&gt;Medicine ran, unintentionally, one of the largest natural experiments available on this exact question. Before the year 2000, researchers running large clinical trials were not required to publicly register which outcome they considered primary before the trial began. After 2000, &lt;a href="https://en.wikipedia.org/wiki/Preregistration_%28science%29" rel="noopener noreferrer"&gt;prospective registration&lt;/a&gt; on a public registry — &lt;a href="https://en.wikipedia.org/wiki/ClinicalTrials.gov" rel="noopener noreferrer"&gt;ClinicalTrials.gov&lt;/a&gt; — became standard practice for major trials. Kaplan and Irvin's 2015 analysis in PLOS ONE, "Likelihood of Null Effects of Large NHLBI Clinical Trials Has Increased over Time," found that 17 of 30 trials (57%) published before 2000 reported a statistically significant benefit on their primary outcome — compared to just 2 of 25 (8%) published after 2000.&lt;/p&gt;

&lt;p&gt;The biology being studied didn't change in 2000. The treatments, the patient populations, the underlying science — none of it shifted because a registry requirement went into effect. What changed was narrower and more specific: researchers lost the ability to look at everything they'd measured and, after the fact, designate whichever outcome looked best as "the" primary result. Prospective registration doesn't make anyone more honest in the moment they're analyzing data. It removes the moment of choice from after the data arrives and forces it to before, which is the entire mechanism. The authors are careful to note this is a strong association rather than an experimentally proven cause — other things about trial conduct changed across those decades too — but the size and direction of the shift, on exactly the outcome you'd predict from the mechanism, makes this some of the best real-world evidence available for a bias that's otherwise mostly discussed in the abstract.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same mechanism, running quietly in a normal analytics stack
&lt;/h2&gt;

&lt;p&gt;Almost no CRO or growth program operates with anything as formal as the pre-2000 clinical-trial process — but most operate with its structural equivalent. A test launches against a loosely stated goal ("improve checkout"), and a modern analytics stack surfaces a dozen or more numbers that moved by the time it's read out: conversion, AOV, time on page, scroll depth, a funnel step two clicks downstream, a retention cohort that's too early to trust but is right there in the dashboard anyway. Nothing was pre-registered as &lt;em&gt;the&lt;/em&gt; metric. So when the obvious one is flat, there's no structural obstacle to the room settling on whichever of the other dozen moved in the direction everyone was hoping for. That's not a smaller, informal version of what pre-2000 trials did. It's the identical mechanism, running by default because nobody removed the moment of choice from after the results arrived.&lt;/p&gt;

&lt;p&gt;It's worth being precise about how this differs from a related bias already worth watching for in any testing program: an already-selected "winning" metric being inflated relative to its true effect, simply because it was chosen as the best-looking result out of several noisy candidates. That's a real and separate phenomenon. This one happens earlier and is arguably more consequential — it's not about the size of the metric that won being overstated, it's about &lt;em&gt;which&lt;/em&gt; metric gets to be called the win in the first place, chosen after the fact from whichever one happened to move. In practice the two compound: outcome-switching selects a flattering metric to call primary, and the standard selection-inflation problem then overstates how large that metric's true effect actually is. Two separate biases, stacking in the same direction, from two different moments in the same process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: name it before you look
&lt;/h2&gt;

&lt;p&gt;The fix costs nothing to run and doesn't require new tooling — it requires deciding three things in writing before the test launches, not after the readout:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The one metric that decides ship or kill.&lt;/strong&gt; Not the one metric you'll look at — the one metric whose result determines the decision. Everything else is still worth watching; it just doesn't get a vote on this test's verdict.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The bar it has to clear&lt;/strong&gt;, stated as a specific direction and magnitude, not "it should go up."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happens for each outcome&lt;/strong&gt;, decided for all three realistic cases: it clearly clears the bar, it clearly doesn't, or it lands in the ambiguous middle. Naming the action for the ambiguous case in advance matters most, because that's the case most likely to trigger exactly the after-the-fact metric search this whole practice exists to prevent.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this requires ignoring the other numbers on the dashboard. A secondary metric moving in an interesting direction is a legitimate, valuable thing to notice — it's a candidate hypothesis for the &lt;em&gt;next&lt;/em&gt; test, where it can earn the same precommitment treatment before being trusted. What it can't do is retroactively become the proof for the test that was actually designed and powered around a different question. Keeping that boundary explicit is the entire practice. Blurring it is how a null result quietly becomes a win without anyone deciding, on purpose, that it should.&lt;/p&gt;

&lt;h2&gt;
  
  
  The judgment call underneath the mechanics
&lt;/h2&gt;

&lt;p&gt;Precommitting a metric is not, by itself, a hard skill — anyone can lock in a number in a planning doc. The actual skill is choosing, in advance, which metric is genuinely closest to the business decision this test is supposed to inform, rather than whichever one is easiest to move or most likely to look good.&lt;/p&gt;

&lt;p&gt;A team optimizing for a good-looking readout will gravitate toward a shallow, upstream metric — a click, a scroll, a micro-conversion — because those move easily and often, and an easy win is more pleasant to report than an honest null. A team optimizing for a decision that will actually hold up gravitates toward the metric closest to the outcome the business cares about, even when that metric is noisier, slower to move, and more likely to come back flat. That choice has to be made before the test runs, precisely because it's much harder to make honestly once a shallow metric has already moved in a flattering direction and a deeper one hasn't. The senior version of this practice isn't the discipline of precommitting. It's the judgment of knowing which metric deserved the commitment in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Doesn't locking in a metric early just mean I might miss something important?
&lt;/h3&gt;

&lt;p&gt;It doesn't mean missing it — it means categorizing it correctly. Anything unexpected that moves is still visible, still worth discussing, and still a legitimate seed for the next test. Precommitment only determines what's allowed to close out &lt;em&gt;this&lt;/em&gt; test's verdict. An interesting secondary movement earns its own precommitted test before it gets to be called a finding rather than an observation.&lt;/p&gt;

&lt;h3&gt;
  
  
  What if the primary metric is flat but something big and unexpected shows up elsewhere?
&lt;/h3&gt;

&lt;p&gt;Treat it exactly as valuable as it is: a strong new hypothesis, not a result. Write it down, and if it's promising enough, design a test where it's the named primary metric next. What it shouldn't do is get reported as this test's outcome — that's the precise substitution this practice exists to prevent, and it's also usually the moment where a program's reported win rate quietly stops meaning what people assume it means.&lt;/p&gt;

&lt;h3&gt;
  
  
  How is this different from choosing when to stop a test?
&lt;/h3&gt;

&lt;p&gt;They're separate decisions about separate questions. Stopping-rule methodology — &lt;a href="https://dev.to/blog/sequential-testing-sprt-stop-test-early"&gt;covered in depth elsewhere in this series&lt;/a&gt; — governs &lt;em&gt;when&lt;/em&gt; you're allowed to look at a result without inflating your false-positive rate. This is about &lt;em&gt;what you're allowed to call the result once you do look&lt;/em&gt; — which of possibly many measured numbers gets to be the verdict. A test can have a perfectly rigorous stopping rule and still fall into outcome-switching if the primary metric wasn't named until after the data arrived.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does this relate to deciding how much evidence a bet needs?
&lt;/h3&gt;

&lt;p&gt;They answer different questions at different points in the process. &lt;a href="https://dev.to/blog/confidence-tier-model-deciding-with-insufficient-data"&gt;The Confidence Tier Model&lt;/a&gt; helps size a bet against however much evidence you actually have, once you know what the evidence says. This practice determines what's allowed to count as "what the evidence says" in the first place. Skipping this step doesn't just risk a wrong tier assignment — it risks tiering a metric that was never the right one to be measuring the decision against.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does this only apply to formal A/B tests?
&lt;/h3&gt;

&lt;p&gt;No — it applies to any decision where more than one plausible metric could be used to declare success, which is most consequential business decisions, not just controlled experiments. A new hire's first-quarter review, a feature launch, a pricing change rolled out without a holdout — all of them have multiple candidate metrics available after the fact, and all of them benefit from the same fix: name the one that decides the verdict before you have a result that could bias the choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Related reading:&lt;/strong&gt; &lt;a href="https://dev.to/blog/confidence-tier-model-deciding-with-insufficient-data"&gt;The Confidence Tier Model&lt;/a&gt;, &lt;a href="https://dev.to/blog/sequential-testing-sprt-stop-test-early"&gt;Sequential Testing and the SPRT&lt;/a&gt;, &lt;a href="https://dev.to/blog/winners-curse-growth-teams-wins-dont-replicate"&gt;Why Most "Wins" Don't Replicate: The Winner's Curse&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Medicine didn't fix its outcome-switching problem by asking researchers to be more honest. It fixed it by moving one decision — which metric counts — to before the data existed to bias it, and the reported result changed by a wide enough margin, at a large enough scale, to make the mechanism hard to argue with. Nothing about a growth or product team's version of this problem is different in kind, only in formality. Naming the metric, the bar, and the action for every outcome, in writing, before the test starts, is a five-minute habit with the same effect: it doesn't make anyone more honest in the room. It just removes the moment where honesty would have been the only thing standing between a flat result and a win that was chosen, not found.&lt;/p&gt;

&lt;p&gt;If you're building or auditing an experimentation program and want an outside read on this, &lt;a href="https://dev.to/contact"&gt;get in touch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>analysis</category>
      <category>productivity</category>
      <category>testing</category>
    </item>
    <item>
      <title>How to Tell If a Growth Hire Understands Risk</title>
      <dc:creator>Atticus</dc:creator>
      <pubDate>Thu, 23 Jul 2026 19:14:21 +0000</pubDate>
      <link>https://dev.to/hey_atticus/how-to-tell-if-a-growth-hire-understands-risk-19nk</link>
      <guid>https://dev.to/hey_atticus/how-to-tell-if-a-growth-hire-understands-risk-19nk</guid>
      <description>&lt;h1&gt;
  
  
  How to Tell If a Growth Hire Understands Risk
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Meta description:&lt;/strong&gt; A great win story tells you almost nothing about judgment. Two borrowed interview probes — from forecasting research and intelligence tradecraft — do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A confident story about a big win is one of the weakest signals available in a growth-leadership interview — it's vivid, it's memorable, and it tells you almost nothing about the judgment that produced it versus the luck that carried it.&lt;/li&gt;
&lt;li&gt;Most interviews test for articulateness under a friendly, low-stakes question. Almost none test for the specific skill that actually separates a good experimentation leader from a confident one: how someone reasons under uncertainty, in real time.&lt;/li&gt;
&lt;li&gt;Two techniques, borrowed from forecasting research and intelligence tradecraft, convert that skill from a vibe into something you can actually probe for in forty-five minutes.&lt;/li&gt;
&lt;li&gt;Neither probe can be prepared for with a rehearsed answer, because neither asks the candidate to tell a story — both ask them to do something, live, that a good story can't substitute for.&lt;/li&gt;
&lt;li&gt;The candidate with the thinner resume who passes both probes is very often the better long-term hire. The skill you're actually buying is how their next ten ambiguous calls go, not how their best past one turned out.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A hiring manager sits across from a growth-lead candidate who tells a great story: a redesign that lifted conversion, a pricing change that held up, a launch that hit its number. It's specific, it's confident, it's told with the easy fluency of someone who's told it many times before. Everyone in the room nods. It's also almost worthless as evidence, and not because the candidate is lying — because a vivid, well-rehearsed story about a past win is exactly the kind of input that's easiest to be fooled by and hardest to verify in the room. It could be the record of genuinely excellent judgment. It could be one lucky call, retold with the confidence of a pattern. From the interview seat, the two are indistinguishable, because the story format itself doesn't carry the information that would let you tell them apart.&lt;/p&gt;

&lt;p&gt;That's a real gap, not a minor one. Growth and product leadership is a job that consists almost entirely of making calls under incomplete evidence — and most interview processes have no mechanism for testing that skill directly. They test for confidence, for fluency, for a track record that may or may not have been earned by the judgment being evaluated. What follows are two techniques that test the actual thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the resume story is the wrong evidence to weigh heavily
&lt;/h2&gt;

&lt;p&gt;The reason a compelling war story is a weak signal isn't specific to hiring — it's a well-documented feature of how people process vivid information generally. A single, detailed, emotionally satisfying account is disproportionately persuasive regardless of whether it's been corroborated, because coherence feels like truth even when it isn't evidence of it. Intelligence analysts are trained specifically to resist this, because the same trap — a single vivid, uncorroborated source outweighing several duller, more reliable ones — is one of the best-documented causes of analytical failure. A hiring interview has the identical structure: one polished anecdote, delivered with total conviction, against no corroborating information at all.&lt;/p&gt;

&lt;p&gt;The practical implication isn't to distrust every story a candidate tells. It's to stop treating story quality as a proxy for judgment quality, because the two are only loosely correlated — a genuinely lucky call and a genuinely well-reasoned one get told with the same confident cadence after the fact, especially by someone who's had years to polish the retelling. If the story is the only evidence on the table, you're not evaluating judgment. You're evaluating narrative skill, which is a real skill, just not the one the role actually requires.&lt;/p&gt;

&lt;h2&gt;
  
  
  Probe one: ask for a number, not a story
&lt;/h2&gt;

&lt;p&gt;Forecasting research — most famously &lt;a href="https://en.wikipedia.org/wiki/The_Good_Judgment_Project" rel="noopener noreferrer"&gt;the Good Judgment Project&lt;/a&gt;'s multi-year tournaments — has spent years measuring what actually separates good judgment from confident judgment, and the finding that matters most here is almost embarrassingly simple: people with genuinely good judgment under uncertainty routinely assign real probabilities to their beliefs and check those probabilities against what actually happened. Almost everyone else just remembers whether they were right, in the vague, self-flattering way memory tends to work.&lt;/p&gt;

&lt;p&gt;The interview version: when a candidate describes a call they made with real conviction, stop them and ask two follow-up questions. First — "What percentage would you have put on that, at the time, before you knew the answer?" Not "were you confident," a number. Second — "How did you find out whether you were right, and when?"&lt;/p&gt;

&lt;p&gt;Watch for three things. Whether they can produce an actual number instead of reaching for "pretty confident" — most candidates, asked for the first time, visibly struggle with this, and that struggle is itself informative. Whether the number sounds calibrated rather than reflexively extreme — someone who says "95%" about a genuinely uncertain call before they had evidence is telling you something different than someone who says "65%." And whether they have an actual answer to the second question, or whether the honest answer is that they moved on to the next project and never checked. That last gap is the most common finding, and it's the more important one — a candidate who has never once gone back and scored their own past confidence against reality has never built the one habit most reliably associated with good judgment in this specific line of work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Probe two: give them an ambiguous drop and listen for the alternatives
&lt;/h2&gt;

&lt;p&gt;The second technique comes from intelligence tradecraft rather than forecasting research, and it targets a different failure mode: not overconfidence, but premature certainty about &lt;em&gt;why&lt;/em&gt; something happened.&lt;/p&gt;

&lt;p&gt;Give the candidate a short, realistic scenario: "A core conversion metric dropped by a meaningful margin the week after a launch. Walk me through how you'd figure out what happened." Then just listen, and count how many genuinely distinct explanations they generate before they commit to one. The specific, structured version of this — &lt;a href="https://en.wikipedia.org/wiki/Analysis_of_competing_hypotheses" rel="noopener noreferrer"&gt;Analysis of Competing Hypotheses&lt;/a&gt;: listing every plausible explanation before looking closely at any of them, then actively hunting for evidence that would rule each one out rather than evidence that confirms the first one — is a real technique taught throughout the intelligence community precisely because unaided judgment reliably jumps to the first plausible story and then spends its remaining effort defending it rather than testing it.&lt;/p&gt;

&lt;p&gt;A weak answer commits early: "sounds like the new checkout flow broke something," followed by a description of how they'd confirm that specific theory. A strong answer holds off: the launch, yes, but also a tracking regression, a traffic-mix shift from a paid campaign change that landed the same week, a seasonal pattern, a competitor promotion — several candidate explanations, named before any of them gets investigated, with a specific idea of what evidence would rule each one out rather than just support the favorite. The gap between those two answers is not intelligence or experience. It's a discipline, and it's one of the more reliable predictors of whether someone's future root-cause calls will be trustworthy or just confidently plausible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each source of evidence actually tells you
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence source&lt;/th&gt;
&lt;th&gt;What it reveals&lt;/th&gt;
&lt;th&gt;What it can't tell you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A resume win story&lt;/td&gt;
&lt;td&gt;That something good happened once, and that the candidate can narrate it persuasively&lt;/td&gt;
&lt;td&gt;Whether the outcome reflected their judgment or favorable noise they got credit for&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The calibration probe&lt;/td&gt;
&lt;td&gt;Whether they've ever quantified their own confidence and checked it against reality&lt;/td&gt;
&lt;td&gt;Domain expertise, or how they'll perform on a problem type they haven't faced before&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The competing-hypotheses probe&lt;/td&gt;
&lt;td&gt;Whether their default mode under ambiguity is to search for the truth or to defend the first plausible story&lt;/td&gt;
&lt;td&gt;Whether they'll actually apply that discipline under real time pressure, not just in an interview&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of the three is sufficient alone. But the first one is the only one most interview processes actually collect, which is precisely backwards — it's the one that's easiest for a good storyteller to pass regardless of the judgment underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The diagnostic catch this surfaces
&lt;/h2&gt;

&lt;p&gt;Here's the finding that makes this worth doing, stated plainly: a candidate with an excellent-sounding track record can fail both probes, and a candidate with a noticeably thinner one can pass both. That's not a paradox — it's exactly what you'd expect once you separate outcome from process. A strong track record can be a small number of calls that happened to land well, retold with the confidence of a pattern, by someone who's never once gone back to check their own calibration or practiced distinguishing a favorite explanation from a verified one. A thinner track record can belong to someone earlier in their career who nonetheless already has both habits — and habits, unlike a specific past win, are the thing that transfers to your business, under your constraints, on problems that look nothing like the ones on their resume.&lt;/p&gt;

&lt;p&gt;That's the actual hire you're making. Not their best past call. How their next ten ambiguous ones are likely to go, somewhere you can't yet see, on evidence you haven't yet collected. A win story can't tell you that. A number and a list of ruled-out alternatives, produced live under a question they couldn't have rehearsed, comes a lot closer.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Isn't this too abstract to actually use in a normal interview?
&lt;/h3&gt;

&lt;p&gt;No — both probes are two follow-up questions, not a separate exercise. Ask a candidate for a story you'd ask for anyway, then push with "what percentage, at the time" and "how did you find out." Ask a candidate a scenario question you'd ask for anyway, then just count the alternatives before they commit. The discipline is in what you listen for, not in redesigning the interview.&lt;/p&gt;

&lt;h3&gt;
  
  
  What if the candidate has never explicitly tracked their calibration? Is that disqualifying?
&lt;/h3&gt;

&lt;p&gt;Not by itself — almost nobody has, which is exactly the point of asking. What matters is what happens next: do they recognize the gap immediately and reason about why it matters, or do they deflect with a version of "I just have good instincts." The first response is a promising sign in someone who hasn't built the habit yet. The second is the more useful warning, regardless of how strong their resume looks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does this work for evaluating an existing team, not just a candidate?
&lt;/h3&gt;

&lt;p&gt;Yes, and arguably it's more useful there, because you have real, current decisions to probe instead of a rehearsed history. Pick a live ambiguous result from the last quarter and ask the same two questions of whoever owned the call. The same gaps show up, and unlike an interview, you can act on what you find.&lt;/p&gt;

&lt;h3&gt;
  
  
  How is this different from a standard behavioral interview question?
&lt;/h3&gt;

&lt;p&gt;A standard behavioral question ("tell me about a time you...") asks for a story, which is exactly the format that's easiest to prepare and hardest to verify. Both probes here ask for something a rehearsed story doesn't contain — a specific number produced in the moment, or a live count of alternative explanations — which is much harder to fake convincingly on the spot.&lt;/p&gt;

&lt;h3&gt;
  
  
  What's the actual pass/fail signal here?
&lt;/h3&gt;

&lt;p&gt;There isn't a hard pass/fail — the value is comparative. A candidate who produces a real number and checks it, and who generates several genuine alternatives before committing to one, is demonstrating a specific, transferable discipline. A candidate who can only offer conviction and a single confident theory is demonstrating storytelling. Weigh the difference the same way you'd weigh any other job-relevant skill you tested for directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Related reading:&lt;/strong&gt; &lt;a href="https://dev.to/blog/calibration-training-forecasting-tournaments-trusting-your-gut"&gt;Calibration Training&lt;/a&gt;, &lt;a href="https://dev.to/blog/what-intelligence-analysts-know-about-evidence"&gt;What Intelligence Analysts Know About Evidence&lt;/a&gt;, &lt;a href="https://dev.to/blog/confidence-tier-model-deciding-with-insufficient-data"&gt;The Confidence Tier Model&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;A great story is the easiest thing for a strong candidate to prepare and the hardest thing for an interviewer to verify — which makes it close to the worst evidence available for the judgment a growth-leadership role actually requires. Asking for a number instead of a feeling, and counting alternatives instead of accepting the first plausible one, tests something a rehearsed answer can't fake. What you're really buying isn't the candidate's best past call. It's how the next ambiguous one, on your business, is likely to go — and that's a specific, probeable habit, not a vibe you pick up from a good story.&lt;/p&gt;

&lt;p&gt;If you're trying to tell whether a growth hire — or a fractional leader you're evaluating — actually has this judgment, not just a good story about it, &lt;a href="https://dev.to/contact"&gt;get in touch&lt;/a&gt;.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
