DEV Community

Cover image for Medical Data as Commodity: Why Your Health Records Are the Most Valuable Training Data
VelocityAI
VelocityAI

Posted on

Medical Data as Commodity: Why Your Health Records Are the Most Valuable Training Data

In 2021, a company you've probably never heard of quietly acquired the de-identified medical records of roughly 50 million Americans. The price was reported at around $1 billion. The records included diagnoses, prescriptions, lab results, and visit histories, stripped of names and Social Security numbers. The sellers were hospitals and health systems that had collected this data as a byproduct of care. The buyers were building products nobody had consented to be part of.

That's not a scandal. It's a business model. And it's the most important thing happening in medical AI that almost nobody is talking about.

We keep debating whether AI will replace doctors or whether models will hallucinate a dosage. Meanwhile, the actual ground truth of the entire medical AI industry is being extracted, packaged, and sold, one EHR at a time. If you want to understand where medical AI is going, stop reading model cards. Start reading data licensing agreements.

What follows is a teardown of how this market actually works, why your health record is worth more than your genome, and one contrarian claim that will make most AI builders uncomfortable. You'll walk away with a mental model you can apply the next time someone says "we train on de-identified clinical data."

Why Health Data Is Worth More Than Anything Else You Produce
Let's start with the economics, because the ethics only make sense once you understand the incentives.

Your social media posts are worth fractions of a cent. Your browsing history, maybe a few dollars a year. Your genome, in raw form, has been commoditized to the point where sequencing companies will give you the test for free to get the data. But your longitudinal medical record, the years of diagnoses, prescriptions, lab trends, imaging notes, and outcomes, is genuinely rare.

Here's why:

It's longitudinal. A snapshot of your health is nearly worthless. A decade-long record with outcomes is a goldmine, because it lets models learn what happens after a diagnosis, not just what the diagnosis was.

It's labeled. Unlike scraped web text, clinical records come with structured codes, standard vocabularies, and physician annotations. That's the difference between an unsupervised pile of text and a supervised training set.

It's scarce. The entire open web is a firehose. Clean, longitudinal, outcome-linked clinical data is a trickle. Scarcity drives price.

It's multimodal. Imaging, labs, notes, genomics, vitals. Most training data is single-modality. Clinical records are naturally multi-modal, which is exactly what frontier medical models need.

The result is a market where de-identified patient records trade in the hundreds of millions to low billions, and where the buyers are not just hospitals but insurers, pharma companies, and increasingly, AI labs that need clinical data to fine-tune medical models.

The Consent Illusion
Here's the part that should bother you, even if you're a pragmatist.

Most of this data is collected under the banner of "de-identification," governed by HIPAA in the US and GDPR in Europe. The pitch is simple: strip the identifiers, and the data is no longer personal. Problem solved.

Except de-identification is not the same as anonymization. A 2019 study in Nature Communications showed that re-identification of de-identified health data is often trivially easy with auxiliary datasets. A 2022 paper demonstrated that a model trained on "de-identified" records could recover patient names from clinical notes at rates that should concern anyone who has ever signed a HIPAA form.

And the consent that patients give when they enter a hospital is not consent to train models. It's consent to treatment. The two are legally distinct, ethically worlds apart, and practically conflated all the time.

Consider the actual mechanics:

You sign a Notice of Privacy Practices. It's 10-15 pages. It says your data may be used for "treatment, payment, and operations."

Somewhere in there, buried, is language about "research" and "business associates."

You don't read it. Nobody reads it. That's the point.

Your data is then de-identified and sold to a data broker, who sells it to a model developer, who uses it to train a system that may eventually be sold back to your hospital.

You were never asked. You were never told. And you were never compensated. The transaction happened in a language you weren't fluent in, on a timeline you didn't control.

The Sacred Cow: "De-Identified Data Is Ethically Safe"
I want to slaughter a sacred cow that the AI community has been grazing on for years: the belief that de-identified data is ethically equivalent to non-personal data.

This belief is load-bearing. It's the legal and moral foundation of an entire industry. And it's mostly wrong.

Here's why. De-identification removes direct identifiers. It does not remove distinguishability. A de-identified record with a rare diagnosis, a specific age, a ZIP code, and a procedure date is often unique to one person in a population. That uniqueness is not a bug; it's the feature that makes the data valuable for training. The very properties that make clinical data useful for models are the properties that make it re-identifiable.

The tech industry has spent a decade telling itself that if the name is gone, the person is gone. That's not true. A person is a distribution, not a field. Remove the name field, and the distribution remains. A model can learn your patterns without ever knowing your name, and that's enough to make inferences about you, sell to you, or discriminate against you.

The ethical upshot is uncomfortable: de-identification is a legal shield, not a moral one. The industry relies on it not because it's right, but because it's defensible. And defensibility is not the same thing as consent.

Where the Value Actually Flows (Spoiler: Not to You)
Follow the money for a moment.

The patient produces the data as a byproduct of care. Gets no compensation, no transparency, no opt-in.

The hospital collects the data and sells it to a data broker for a lump sum or a per-record fee. The revenue is often small relative to the hospital's budget, which is why the sale happens quietly.

The data broker aggregates and cleans the data, then licenses it to multiple buyers. This is the layer where most of the margin lives.

The pharma company or AI lab uses the data to train models, validate drugs, or develop diagnostics. This is where the value compounds.

The hospital may eventually buy back a model trained on its own patients' data, at a price it can't negotiate because it's locked into a vendor.

The patient is the source. The patient is the subject. The patient is the last person in the chain to see any benefit.

There's a word for this in other industries: extraction. The tech community has spent years critiquing "data extraction" from users by social media platforms. It's time to apply the same lens to medicine.

What Good Looks Like (and Why It's Rare)
I don't want to leave you with nothing but critique. There are models that work, and they're worth studying.

Synthetic data. GANs and diffusion models can generate synthetic patient records that preserve statistical properties without exposing any real individual. Early results in oncology and cardiology are promising, though the fidelity gap is real and often undersold.

Federated learning. Training across hospitals without moving data. Google's work with hospital networks on mammography is the canonical example. It's slower, more expensive, and harder to deploy, which is exactly why it's rare.

Patient consent frameworks. Some cohorts, like UK Biobank and All of Us, build explicit, revocable consent into the design. The tradeoff is scale: these cohorts are orders of magnitude smaller than the aggregated data broker market.

Data trusts and cooperatives. Models where patients hold governance rights and share in revenue. Early pilots exist. None have reached the scale of the commercial market.

Notice the pattern. The ethical approaches are slower, smaller, and less profitable. That's not a coincidence. It's the entire reason the commercial market exists in its current form.

What To Actually Do About It
If you're a builder working in medical AI, you have more leverage than you think. Three concrete moves.

  1. Ask for provenance, not just de-identification. When you acquire or license clinical data, ask where it came from, how consent was obtained, and whether patients have any recourse. If the vendor can't answer, you're building on sand. Document the answer in your model card.

  2. Instrument for re-identification risk. Before you train, run a re-identification audit. Use k-anonymity metrics, membership inference tests, and auxiliary dataset attacks. If your "de-identified" dataset fails these tests, it's not de-identified. It's just unidentified until someone tries.

  3. Prefer federated and synthetic pipelines where they fit. They're harder and slower. They're also the only defensible long-term path if regulation tightens. Build the muscle now, before you're forced to.

If you're a patient, here's the harder truth: you have almost no leverage in the current system. The best you can do is read the Notice of Privacy Practices at your next visit, ask whether your data is shared with third parties, and opt out where allowed. It won't stop the market. But it will teach you how the market works, and that knowledge compounds.

The Mental Model
Here's the lens I want you to keep.

Medical data is not a resource. It's a relationship. It's a record of a person's body, their fears, their diagnoses, and their worst days, compressed into structured fields and sold to the highest bidder. The industry treats it as a commodity because that's the only way the economics work. But the economics are a choice, not a law of nature.

Every time you build a model on clinical data, you are borrowing someone's worst day to build something better for the next person. That's a legitimate thing to do. It's also a debt. The question is not whether we'll use medical data to train models. We already are. The question is whether the people who produced that data will ever be treated as parties to the transaction rather than inputs to it.

The market has an answer. It's not a good one.

What's your read on this? If you've worked with clinical data, how did you handle consent and provenance? If you haven't, would you want to know if your records were sold? Drop it in the comments. I'm genuinely curious how the builders in this space are thinking about it.

Top comments (1)

Collapse
 
wrobeltomasz profile image
Tomasz •

Does the article address any legal requirements for obtaining explicit patient consent before using their de‑identified health records to train AI models?