Google DeepMind’s Cognitive Taxonomy splits general intelligence into ten faculties: perception, attention, memory, reasoning, metacognition, executive function, and so on. Kaggle ran it as a $200,000 hackathon, paying contestants to build a benchmark for each one. The pitch is that this borrows real rigor from psychology and neuroscience instead of relying on vibes. Follow that borrowing back to where it actually comes from, though, and every piece of it turns out to depend on something that doesn’t survive the trip.
The categories aren’t natural kinds
Cognitive psychology’s categories are useful theoretical constructs rather than discrete natural kinds. They are factor-analytic groupings, and even inside clinical psychology, where they are used for real diagnoses, the boundaries get settled by committee rather than by nature. For instance, DSM-5’s own work group spent a decade building a scientifically better dimensional model to replace the old categorical thresholds, and the APA’s Board of Trustees overruled them anyway, because clinicians found the new version less familiar and less billable. If the field’s own diagnostic manual resolves disputes by vote and workflow convenience, the categories were never facts waiting to be found. They were tools kept around because someone needed to make a call by Friday, and that’s the thread that unravels everything downstream.
The reason those tools were worth keeping was never that they carved nature at the joints. A neuropsych score is inert on its own. It only becomes meaningful once it’s attached to a specific decision someone is accountable for: does this patient meet criteria for dementia, are they safe living alone, are they responding to treatment. Strip that referral question out and you are left holding a number with nothing to be true or false about. DeepMind’s taxonomy never had a referral question to begin with. Nobody is deciding anything based on a model’s “metacognition” score. So the categories arrive already unmoored twice over, arbitrary by the field’s own admission, and then stripped of the one thing that gave the arbitrary label a job to do.
Human cognitive tests derive validity from the decisions they support. Once detached from those decision contexts, the burden shifts to demonstrating a new source of validity rather than assuming the construct transfers unchanged.
A map without the territory
The whole mapping from clinical cognitive profiles to underlying neurobiology is messy, heterogeneous, and often non-isomorphic, and that messiness is exactly what the DeepMind/Kaggle story glosses over.
In people, the syndromes we call “memory-predominant,” “executive-predominant,” or “language-predominant” emerge from a long list of etiologies: neurodegenerative (Alzheimer’s, FTD, PSP, Lewy body), vascular (large infarcts, small-vessel disease, strategic lacunes), neoplastic (slow-growing tumors, meningiomas), traumatic (repetitive TBI, focal contusions), inflammatory/autoimmune (MS, encephalitis), infectious (HIV, neurosyphilis), metabolic/endocrine (thyroid, B12, liver/renal failure), toxic, hydrocephalus, and more. Two things follow:
- Same profile, many causes: Amnesia, for example, can come from Alzheimer’s pathology, but also from hippocampal infarcts, medial temporal tumors, hypoxic injury, herpes encephalitis, or severe B12 deficiency. The surface-level “memory deficit” label doesn’t tell you which network, which cell type, or which time scale is at fault.
- Same cause, many profiles: Vascular disease can look like pure executive dysfunction, pure slowing, pure memory problems, or a mixed picture depending on where the infarcts, microbleeds, or white-matter lesions land. A single etiology such as small-vessel disease doesn’t map neatly onto one of the ten faculties. It scrambles several at once in patterns that reflect anatomy and connectivity, not a pre-labeled taxonomy. In humans, the “dissociations” we use to motivate separate faculties are already etiologically heterogeneous, anatomically distributed, and developmentally shaped. Many different pathologies can produce the same cognitive label. Damage to different nodes and connections in large-scale networks yields overlapping syndromes. Lifelong wiring, plasticity, and compensation mean the same lesion can look different in different brains. And even then, clinicians don’t treat “executive function equals 1.2 SD below mean” as a standalone truth. They tie it to a concrete question: capacity, safety, diagnosis, treatment response. Strip that context away and you are left with a number that can’t be right or wrong on its own.
But the LLMs don’t have vascular territories, tumor mass effects, developmental trajectories, or amyloid plaques creeping up from the olfactory system. They have uniform layers and training objectives, and the kinds of “damage” they take such as pruning, quantization, or fine-tuning do not respect any human-like cognitive map. Even if these labels are to be interpreted purely functionally rather than biologically, the evidence used to motivate the taxonomy draws substantial motivation from human cognitive dissociations whose interpretation depends on developmental, anatomical, and etiological structure absent from transformers.
The hackathon isn’t built to fix this
None of this is a problem a hackathon can fix, because a hackathon isn’t built to fix it. It’s built to reward whoever makes the problem look solved to the people judging it, not whoever actually solves it. Real instrument validation takes large samples and years of iteration against outcomes nobody can spin. A contest judged on a submission deadline optimizes for “the judges found this persuasive,” which is a different target than “this task has demonstrated construct validity,” and no amount of clean plotting closes that gap.
The competition itself cannot supply the external validation required to establish construct validity, because the taxonomy only has to survive contact with a leaderboard judged by the people who wrote it.
It’s Goodhart’s Law with a $200,000 check stapled to it, running inside a closed loop where the paper builds on DeepMind’s own new taxonomy framework, the hackathon builds on the paper, and the winners, now $200,000 richer and invested in the taxonomy’s success, are incentivized to keep citing and building on it in follow-up work. The citations will keep on feeding each other, even if none of them ever come from outside the loop.
This isn’t a new problem. Interdisciplinary reviews of AI benchmarking practices find that roughly half of benchmarks test abstract phenomena like “reasoning” or “harmlessness” without clear, uncontested definitions. Only about half justify that they actually measure what they claim to measure. Very few report statistical significance or allow easy replication. Many reuse existing datasets or exams, increasing risks of contamination and memorization.
A recent NeurIPS paper puts it bluntly: AI evaluations, just like cognitive assessments for humans, suffer from poor construct validity. The systematic review of around 445 benchmarks found that only 16 percent used statistical tests in their comparisons, and that most benchmarks aimed at abstract concepts lacked rigorous definitions and justification. The DeepMind/Kaggle initiative inherits all of these issues and then wraps them in a cognitive-science veneer that makes them look more settled than they are.
So what is this, really?
The winning submissions were never going to be anything but polished mumbo jumbo. Not because anyone cheated, but because the premise did the cheating already: arbitrary categories, stripped of the purpose that once justified the arbitrariness, wrapped in an excuse for any result that contradicts them, borrowing the credibility of a biological process that is networked, developmental, and decades in the building, but has no counterpart in the thing being measured. Score that with a leaderboard and a deadline and you don’t get science. You get very confident-sounding noise with citations.
That’s not an accident of execution, it’s the business model. It’s not so different from buying followers or GitHub stars: pay for the appearance of organic adoption and let the volume do the credibility work. The hackathon was used to rapidly pull in a thousand teams who’d spend their own months of unpaid labor building on top of the pseudoscience, citing the authors’ paper as the reason their task exists, and handing them a leaderboard’s worth of apparent validation they didn’t have to produce themselves. The money isn’t funding research. It’s funding adoption of a framework that hadn’t earned adoption yet.
The central challenge is not inventing tasks but demonstrating construct validity, discriminant validity, predictive validity, and robustness across models and prompting conditions. Human cognitive categories are pragmatic, not Platonic. Their meaning comes from the decisions they support, not from the labels themselves. The mapping from “cognitive profile” to “mechanism” in humans is already messy, heterogeneous, and non-isomorphic. Transformers have none of the developmental, anatomical, or etiological structure that makes those messy mappings even partially intelligible in people. AI benchmarks in general already struggle with construct validity, statistical rigor, and clear definitions. A hackathon optimized for “persuasive by the deadline” is the wrong institutional form to fix any of that.
Demonstrating construct validity would require showing that tasks intended to measure the same faculty converge across diverse implementations, remain distinct from other faculties, predict meaningful downstream behavior, and continue to do so across model families, prompting strategies, and future architectures. Until then, the taxonomy should be treated as a working hypothesis rather than a validated decomposition of intelligence. DeepMind has presented a speculative cognitive decomposition before demonstrating the construct validity needed to justify it.
If you want a research agenda that actually moves the field forward, it starts by admitting that the ten-faculty story is a narrative device, not a scientific discovery. Then you build narrow, task-specific evals with explicit decision contexts, statistical rigor, and external validation, and you stop pretending that slapping “memory” or “executive function” on a leaderboard makes it psychology.
Top comments (0)