Before a single line of extraction logic existed for LungTrace, I spent hours writing fictional patients. That felt backwards at the time. It turned out to be the highest-leverage work in the whole project.
The demo is only as convincing as the data
LungTrace tracks a patient's lung scan findings over time using Hindsight as the memory layer, and flags follow-up scans that were recommended but never happened. The idea is easy to explain. Proving it works is not, unless the data is built to prove it.
If I'd generated 20 random synthetic reports, the odds that any of them would show a clean growth pattern, a genuinely missed follow-up, and a stable control case were low. So I stopped treating the data as filler and started treating it as the test suite.
Designing patients like test cases
Every patient I wrote had a job to do:
- A clear growth case: A nodule at 3mm, then 4.5mm, then 6mm across three reports spanning eight months. Read alone, each report says "small nodule, likely benign." Read together, it's doubled in size.
- A missed follow-up case: One report, a recommendation for a follow-up scan, and a due date that's now well in the past with nothing after it.
- A stable, boring case: Same size across three annual scans. This one mattered as much as the alarming ones, because a system that flags everything is as useless as one that flags nothing.
- A shrinking, reassuring case: A nodule getting smaller, correctly marked as needing no further follow-up.
- A different-organ case: A liver lesion instead of a lung nodule, to check the grouping logic didn't secretly assume "lung" everywhere.
Each report also needed to read like something a radiologist actually wrote, not a data table in paragraph form, since the whole point was testing whether an LLM could pull structure out of prose:
Top comments (0)