A handful of AI-designed drugs have reached human trials, but the grand promise of curing cancer remains a cliche, according to Anthropic CEO Dario Amodei. A biotech startup called Vivodyne says it knows exactly why: the AI drug-discovery industry has a data problem, as reported by TechCrunch. It argues the most advanced models are missing the causal, biological data on living human tissue they need to graduate from curing cancer in mice to curing it in people.
Vivodyne’s alternative is a machine called HIVE, a modular robotic lab that grows human tissues and runs experiments autonomously. It’s a bet that solving medicine’s hardest problems requires a new kind of infrastructure, not just a more powerful model.
The AI Cure-All Myth Hits a Cold, Hard Data Wall
The concept is simple but daunting. Ninety percent of drugs effective in animal testing fail to gain regulatory approval for humans. That statistic underscores a fundamental mismatch: models trained on mouse biology or static cellular snapshots are learning the wrong lessons. As Vivodyne CEO Andrei Georgescu puts it, "Absent human testing, what are these [AI] models going to do? They're going to cure cancer in mice."
This is the core disconnect fueling recent skepticism. While leaders like Sam Altman and Demis Hassabis have long cited curing disease as a primary justification for pursuing artificial general intelligence (AGI), the practical track record is thin. Nobel-prize winning Alphafold revolutionized protein folding but has not yet produced a new drug. Isomorphic Labs, built on its technology, expects its first trials by the end of this year, a delay from the original 2025 target.
The emerging consensus, which Vivodyne embodies, is that the bottleneck isn't raw compute or algorithmic brilliance. It's the quality and nature of the training data. Georgescu calls for "a sanity check," arguing existing models lack the data to capture human biology's complexity. The hype has outpaced the foundational work required to validate it.
The "Human Data Center" and a New Kind of Lab Rat
Vivodyne’s plan is to generate that missing data at scale. Spun out of the University of Pennsylvania in 2021, the company has raised just under $80 million, led by Khosla Ventures. Last week, it opened what it calls the world's largest "human data center" near San Francisco.
Its HIVE systems can grow 20 kinds of human tissue. The company claims impressive predictive accuracy: its liver cells show 94% accuracy against human toxicity trials, its airway tissue matches real behavior 96% of the time, and its bone marrow tests achieved 100% concordance across 20 chemotherapy drugs. Georgescu says the lab's throughput is already double that of all animal trials conducted in the United States.
The goal is not to replace human clinical trials, but to make them far more efficient. Clinical trials typically cost tens of millions of dollars, and most candidates fail. Vivodyne's approach is akin to automotive crash testing. "An automaker is typically confident its car will pass NHTSA requirements before testing it," Georgescu told TechCrunch, "but drugmakers rarely have that same confidence going into a clinical trial."
The company is working with multiple major pharma partners, though it won't name them publicly. The value proposition is clear: de-risk the most expensive phase of drug development by providing higher-fidelity human data much earlier.
Beyond Screening: Building Models That Understand Cause and Effect
The immediate application is improving drug candidate selection. However, Georgescu’s larger vision is about rebooting AI model training for biology itself. He points to a recent study in Nature Methods that found no clear data scaling laws when training generative AI on existing cellular data. The problem, he argues, is a lack of causality.
"All the training is done on static snapshots of these cells, and the models are not conditioned at all by the how a cell got to that state," Georgescu said. "In other words, the model learns 'this is cell state A,' 'this is cell state B,' but never 'cell state B is the effect of inflaming cell state A.'"
Current models see pictures. Vivodyne aims to provide a movie, with a script.
The HIVE machines track hundreds of thousands of ongoing experiments where diseased tissue is exposed to stimuli. This generates sequences of events, cause and effect, that could train models through a form of biological reinforcement learning. Georgescu believes this is essential for the next frontier of medicine: combination therapies that target multiple disease pathways.
"If we want combination therapies, the space that has to be searched explodes, it can't be an experimental approach," he said. "You have to say, 'I want this effect to happen, so what cause should I invoke?' Establishing causality in human biology is the basis of all of this."
The Long Road From Algorithm to Actual Cure
Vivodyne’s thesis exposes a gap in the AI-pharma narrative. The field has celebrated algorithmic milestones, like a model identifying a potential drug molecule, but has struggled to navigate the messy, causal reality of human physiology. This is similar to the challenges faced when applying AI to other complex systems, where data quality dictates real-world utility, as seen when Amazon bulldozes rare books for AI-fueled data wars.
A sober look at the evidence ladder for medical AI shows why. Success in a retrospective benchmark or a controlled lab setting is merely step one. It must then prove itself in prospective studies and, ultimately, randomized clinical trials that show improved patient outcomes. The Swedish MASAI mammography study, which showed AI could improve screening performance in a real-world workflow, is a rare and meaningful example of this progression.
XOOMAR Analysis: Vivodyne’s argument implicitly critiques a "build it and they will come" approach to medical AI. The assumption has been that enough compute and clever algorithms would overcome data limitations. Vivodyne asserts the opposite: without a fundamental upgrade in the biological data itself, the most powerful models will remain intellectually constrained. They may optimize within the known paradigm of mouse studies and single-protein targets, but they won't unlock fundamentally new biological insights.
What Winning the Data Race Would Actually Mean
The implications are vast. If Vivodyne’s approach gains traction, the competitive moat in AI-driven drug discovery shifts decisively.
It becomes less about who has the best AI team and more about who controls the best human biological datasets. The "data center" it built is a physical manifestation of that future moat. This could spark a new arms race for high-fidelity, causal human data, potentially creating tension between proprietary platforms like Vivodyne and open-science initiatives.
Furthermore, it reframes the problem from a pure software challenge to a bioengineering and systems integration challenge. The hard part isn't just writing the code. It's building the robotic lab that can keep human tissues alive, dosing them accurately, monitoring outcomes, and processing the resulting data flood. Success depends as much on mechanical and fluidic engineering as on machine learning.
Finally, it places a harsh spotlight on the entire drug development pipeline. A system where 9 out of 10 drugs fail after animal testing is not just expensive, it's inefficient. Tools that can compress that failure curve earlier, saving years and hundreds of millions of dollars per drug, would be transformative. The question is whether the data Vivodyne generates can reliably predict human outcomes at the scale and speed it promises.
The watch item now is validation. Can Vivodyne’s partners, the unnamed major pharma companies, point to a drug candidate that successfully navigated trials because of this human tissue data? Can the causal datasets they produce actually train a new generation of AI models that outperform the current state-of-the-art? The promises of leaders like Altman and Hassabis set a high bar. The work of startups like Vivodyne reveals just how much foundational work remains to be done on the factory floor of biology before those promises can be kept.
Why This Changes Everything
- It confronts the fundamental data gap causing 90% of drugs that work in animals to fail in humans.
- It shifts focus from just building more powerful AI models to creating the biological infrastructure needed for accurate human drug discovery.
- It suggests the path to curing human diseases like cancer requires a new, human-tissue-based approach, not just faster computation.
Originally published on XOOMAR. For more news and analysis, visit XOOMAR.
Top comments (0)