The improvement you measured is a covariance, not a cause
In 2013, MD Anderson Cancer Center began building a clinical decision support tool with IBM, using Watson technology. It was called the Oncology Expert Advisor. It was supposed to recommend treatments, match patients to trials, and help evaluate cases. Over roughly four years the project consumed at least $62 million.
By September 2016 it was not in clinical use. IBM ended support for the pilot and demo systems effective September 1 of that year. A University of Texas System audit followed, and the press coverage arrived the next February.
The number everyone quotes is the $62 million. The fact that actually matters sits one line below it in the audit: the system had never been piloted anywhere outside MD Anderson.
That is not a statement about whether the software was any good. It is a statement about what could possibly have been known. A clinical decision support system trained on one institution's practice, evaluated only inside that institution, is being measured against the people who taught it. Agreement is the expected result. It would have been the expected result whether the recommendations were medically excellent or medically wrong, because the thing being measured and the thing doing the measuring share a source.
There was no out-of-sample test to fail. That is what the $62 million bought.
Keep two findings apart
The audit is easy to misread, and misreading it is the fastest way to get this story wrong.
What the auditors found was a procurement problem. Contracts, approvals, compliance. Nearly all of the spending went ahead without board approval. They also noted that the tool could not exchange data with the Epic electronic health record system the hospital had moved to, which left oncologists consulting protocol and trial data from the system Epic replaced.
What the auditors explicitly did not do was evaluate the model. The review's own scope statement limits it to contracting and compliance and says it did not cover project management or system development. The 48-page document is not a verdict on whether the Oncology Expert Advisor gave good advice. Nobody produced that verdict, because producing it would have required the external pilot that never happened.
So there are two failures here and they are not the same failure. One is that money moved without the approvals it needed. The other is that four years of work generated no evidence that could travel outside the building. The first is a governance problem. The second is an epistemics problem, and it is the one that repeats everywhere.
What external validation looks like when someone actually runs it
Watson for Oncology, a related but distinct product, was trained with Memorial Sloan Kettering and did get deployed to other institutions. So we can see what happens when the test finally runs.
Concordance rates, meaning how often the system's recommendation matched the local tumor board, vary enormously by site and by cancer type. A double-blind study of 638 breast cancer patients in India reported 93% concordance for recommended or for-consideration treatment. Work presented at ASCO in 2017 covering 525 patients in Korea reported 73% for colon cancer and 49% for gastric cancer. A Korean study at Gil Medical Center put absolute concordance for colon cancer at 48.9%, rising to 65.8% when cases judged acceptable were included. For gastric cancer the same kind of analysis reported 41.5% at the recommended level and 87.7% at the for-consideration level.
Read those numbers as a group rather than individually. The same system, asked the same kind of question, agrees with local practice more than nine times in ten in one setting and fewer than half the time in another.
The interesting reading is not that the system was bad. It is that concordance was never measuring medical correctness in the first place. It was measuring the distance between one institution's encoded practice and another's. High agreement where local practice resembles the training institution, low agreement where it does not. That is a covariance between two things that share a cause, and it produces a number that looks like an accuracy score and behaves like a similarity score.
An organization that only ever measured this metric at home would see a high number every time and would learn nothing at all.
Why the home number is so convincing
It would be comforting to file this under carelessness, but the internal number is persuasive for reasons that survive being smart and careful.
The first is that it is not imprecise. You can measure concordance at your own institution to as many decimal places as you like, and every one of them will be accurate. The precision is real. It is attached to the wrong quantity. High precision on the wrong estimand feels exactly like rigor from the inside, and it produces the confident charts.
The second is that the people doing the evaluating are the people who supplied the training signal. When the system disagrees with them, the natural reading in the room is that the system made a mistake, because in that room it did. Disagreement gets logged as error and corrected away. The evaluation loop is therefore not neutral about which direction the system moves. It rewards convergence on local practice, which is the same operation as destroying the system's ability to tell you anything you did not already believe.
The third is that nobody in the building has an incentive to produce the number that would embarrass everyone in it. An external pilot is the only party to the process that has no stake in the result. That is not a cynical point about human character. It is a structural point about where disconfirming evidence comes from, and it is why "we validated it internally" is a description of a procedure rather than of evidence.
The same bug, in your own field
If oncology feels far away, here is the version with a coefficient attached.
In 2024 a team at Scale AI built GSM1k, a fresh benchmark designed to mirror the style and difficulty of GSM8k, the widely used grade school math benchmark. The point was to ask a question that GSM8k can no longer answer about itself: when a model scores well, how much of that is arithmetic reasoning and how much is having seen the test?
Their paper, "A Careful Examination of Large Language Model Performance on Grade School Arithmetic," reports accuracy drops of up to 8% when models move from GSM8k to the fresh set. Several model families show systematic overfitting across nearly all sizes. And the finding that makes it more than a suspicion: a positive relationship, Spearman's r² = 0.36, between how likely a model is to generate GSM8k examples and how much its score falls on the held-out set. Models that can recite the test do worse when the test changes.
That coefficient is what turns an argument about principle into a measurement. A benchmark score is a joint measurement of capability and exposure, and nothing inside the score tells you the ratio.
Three numbers that drifted on the way here
While assembling this piece we ran into the thesis three separate times, in our own sources.
The GSM1k accuracy drop is reported in a good deal of secondary coverage as 13%. The paper's abstract says up to 8%. We took the primary.
The University of Texas audit is described in several retellings as posted on January 31, 2016, which cannot be right, because the events it audits run through September 2016. The posting date is January 31, 2017. A year fell off in transit.
The split of the $62 million between IBM and the consulting firm that supported the project is given slightly differently across accounts, roughly $39 to $40 million to IBM and roughly $21 to $23 million to the consultancy, while the total holds steady. This is why the figure here is "at least $62 million," which is the reporting's own hedge, rather than a precise sum.
None of these drifts is scandalous. Each write-up was doing its honest best to report what a study or an audit found. That is the point. A number does not need anyone to lie about it in order to arrive somewhere false. It only needs to be copied a few times by people who did not go back to the source, and each copy is individually reasonable.
The counterweight, because the lazy version of this argument is wrong
The GSM1k paper also found that frontier models showed minimal signs of overfitting. They generalized to problems guaranteed to be absent from their training data.
That matters, and leaving it out would be the same sin this essay is about. The honest claim is not that benchmarks are theater or that measured progress is fake. Real capability gains exist and the same study that quantifies contamination also demonstrates them.
The claim is narrower and more useful: a measured improvement licenses a causal reading only when the thing you changed is the only thing that moved. When your evaluation shares a source with your training, whether that source is an institution, a market regime, or a scraped corpus, some fraction of the agreement you observe is an echo, and the score alone will not tell you how much.
The zero that was a base rate
Here is the version of this that nearly got past us this morning.
Working a research question on July 27, 2026, we wanted a data-grounded read on whether a particular kind of skilled work had been touched by AI adoption. The Anthropic Economic Index publishes a labor market impacts release with task penetration and job exposure data, so we pulled the task-level file and looked at every calibration and metrology task in it.
Every one read 0.0.
That is a striking result. Measured AI penetration into this work, zero. It would have made an excellent line, and we were most of the way to writing it.
Then we counted the rest of the file. By our count of that file, of 17,998 task statements only 1,354 carry a nonzero value. About 92.5% of the file reads 0.0. Our finding was the modal value of the dataset. We had measured the sparsity of a file and were about to report it as a fact about an occupation.
The fix was to move to the occupation-level file, where the distribution actually discriminates: by our count, a mean of 0.0770 across 756 occupations with 54.4% sitting at exactly zero. There the relevant proxies read 0.0324 and 0.0. Still low, but now meaningfully low, because there was a spread to be low against.
The zero did not change. What changed was knowing how ordinary a zero was in the population it came from.
What to compute before you believe a delta
The check that catches all four of these cases is short, and it is worth running before the number goes in a slide.
Ask what else moved during the window you measured. If the answer is anything other than "only the thing I changed," the improvement is a covariance and you should say so out loud rather than let the reader assume otherwise.
Ask what the base rate of your value is in the population it came from. An extreme reading means nothing until you know how common that reading is. A zero drawn from a file that is mostly zeros is a description of the file.
Ask what a fresh sample would say. Not a held-out split of the same collection, which shares every bias the training data has, but a matched sample sourced independently, the way GSM1k was built to mirror GSM8k without inheriting it. If getting one is impractical, that is a real constraint and worth stating. What is not acceptable is spending four years and $62 million without ever noticing that you never had one.
Three questions. None of them require new tooling, and any of them would have raised a hand somewhere in Houston in 2014.
Sources
University of Texas System Administration, special review of MD Anderson's Oncology Expert Advisor procurement (48 pages; report November 2016, results posted January 31, 2017). Scope limited to contracting, procurement and compliance; explicitly excludes project management and system development.
Matthew Herper, "MD Anderson Benches IBM Watson In Setback For Artificial Intelligence In Medicine," Forbes, February 19, 2017.
The Register, coverage of the MD Anderson audit, February 20, 2017.
Hugh Zhang, Jeff Da, Dean Lee, Vaughn Robinson, Catherine Wu, Will Song, Tiffany Zhao, Pranav Raja, Charlotte Zhuang, Dylan Slack, Qin Lyu, Sean Hendryx, Russell Kaplan, Michele Lunati and Summer Yue (Scale AI), "A Careful Examination of Large Language Model Performance on Grade School Arithmetic," arXiv:2405.00332, submitted May 1, 2024, revised November 22, 2024. Figures quoted from the abstract: accuracy drops of up to 8%, Spearman's r² = 0.36, frontier models showing minimal signs of overfitting.
S. P. Somashekhar et al., "Watson for Oncology and breast cancer treatment recommendations: agreement with an expert multidisciplinary tumor board," Annals of Oncology, 2018 (PMID 29324970). 638 breast cancer cases, Manipal Comprehensive Cancer Center, 2014–2016. Concordance 93%, where a recommendation counts as concordant if the tumor board's choice was designated "recommended" or "for consideration" by the system. The study included a blinded second review by the tumor board in 2016 of the cases where the two disagreed; earlier conference reporting of the same work quotes a lower pre-review figure.
ASCO 2017, Gachon University Gil Medical Center, Incheon: 525 patients treated 2012–2016 (340 colon, stage II–IV; 185 chemotherapy-naïve advanced gastric). Concordance 73% colon, 49% gastric.
Won-Suk Lee et al., "Assessing Concordance With Watson for Oncology, a Cognitive Computing Decision Support System for Colon Cancer Treatment in Korea," JCO Clinical Cancer Informatics, 2018 (doi:10.1200/CCI.17.00109, PMID 30652564). 656 patients, stage II–IV colon cancer, 2009–2016. Absolute concordance 48.9%; 65.8% (432 of 656) when cases judged acceptable are included.
Concordance Rate between Clinicians and Watson for Oncology among Patients with Advanced Gastric Cancer: Early, Real-World Experience in Korea (PMC6377977). 65 advanced gastric cancer patients, Gachon Gil Medical Center, 2016–2017. Concordance 41.5% (27 of 65) at the recommended level, 87.7% (57 of 65) at the for-consideration level.
Anthropic Economic Index, labor market impacts release:
task_penetration.csv(task level) andjob_exposure.csv(occupation level). The task-level and occupation-level counts given above are our own counts of those two files as accessed on July 27, 2026, not figures published by the index.
Related reading from us
Your Agent Eval Is One-Factor-at-a-Time, and Fisher Proved That's Blind. That piece is about experimental design, changing one factor at a time versus factorial designs. This one is the observational counterpart: what you can and cannot infer when you did not run an experiment at all and only have a measured delta.
The Answer Key Was in the Training Data. Direct contamination, which is the sharpest special case of the mechanism described here.
Zillow Disabled Its Human Pricing Override. Then It Wrote Down $407.9 Million. A governance failure rather than a validation failure. Different bug, comparable bill, and worth reading alongside this one precisely because the two are so often confused.
Top comments (1)
"Some fraction of the agreement you observe is an echo" is the sentence, and it generalises well below the $62M tier. I build internal tools for a hospital as a non-developer, so the scale is nothing like this, but I ran the same structure last month and it cost me weeks instead of millions.
My version: AI writes most of my code, so I had AI review it — separate agents, one implementing, one running tests, one reading for security flaws. Different roles, same model underneath. For months the reviewer flagged almost nothing and I read that as evidence the code was clean. Then over three days, strangers on this site found eight real defects in those same systems, including a guard that had been silently mis-scoring six clean files for weeks. My reviewer found zero of them. Not because it was worse than the strangers — because it shared a prior with the thing it was reviewing. Its agreement was an echo, exactly as you put it, and I'd been booking the echo as validation.
The part of your framing I'd underline for anyone doing this at small scale: "no out-of-sample test to fail" is not an absence of testing, it's a test that cannot produce a negative result. I had five layers of automated checking by the end, and they genuinely fail for different reasons — a rule regression, a dead gate, a stopped scheduler. But every layer was designed by me and written by the same model, so diversity of failure mode was masking a single point of origin. The only thing that has reliably been out-of-sample for me is other people, which doesn't scale, doesn't run nightly, and remains the only check I have that isn't downstream of my own assumptions. Cheaper way to learn it than an audit, at least.