Generative Simulation Benchmarking for precision oncology clinical workflows during mission-critical recovery windows
Introduction: When a Downtime Drill Taught Me That Accuracy Isn't Enough
Six months ago, I was part of a small research group tasked with a deceptively simple question: if our clinical decision-support models went dark for four hours, could we still trust the workflows that depended on them? We were studying precision oncology pipelines — the kind that ingest genomic variants, match them to targeted therapies, and surface recommendations to tumor boards. The models themselves were solid. What broke during our simulation was everything around them.
That experiment sent me down a rabbit hole I didn't expect. I started exploring how we could generate synthetic clinical scenarios at scale, benchmark entire workflows (not just models) against them, and specifically stress-test those workflows during what I came to call mission-critical recovery windows — the fragile periods after a system outage, a data pipeline failure, or a model rollback when clinicians are trying to catch up on delayed cases.
While learning about generative simulation for healthcare, I realized that most benchmarking literature treats model accuracy as the terminal metric. But in oncology, a 94% accurate variant classifier means very little if the surrounding workflow collapses when it's unavailable for 90 minutes during a molecular tumor board. This article is my attempt to share what I discovered building a generative simulation harness for exactly this problem, and why I now believe recovery-window benchmarking deserves to be a first-class discipline in clinical AI.
Why Precision Oncology Is a Brutal Testbed
Precision oncology is the practice of matching a patient's tumor genomics to targeted therapies. A typical workflow looks like this:
- Sample acquisition — tissue or liquid biopsy enters the lab.
- Sequencing — NGS panel produces raw reads.
- Variant calling — somatic variants are identified and filtered.
- Annotation — variants mapped to databases (OncoKB, CIViC, ClinVar).
- Therapeutic matching — candidates ranked against approved trials and drugs.
- Molecular tumor board (MTB) — human review and decision.
- Ordering and follow-up — prescription, monitoring, escalation.
Each stage has different latency tolerances. Variant calling can tolerate a few hours. MTB preparation cannot — the board meets on a schedule, and a delayed case means a patient waits another two weeks. This asymmetry is exactly what makes recovery windows so dangerous: a system that comes back online at 80% capacity may look healthy while silently starving the highest-priority queue.
In my research of clinical workflow resilience, I found that most institutions measure recovery in terms of "system up / system down." That binary framing hides the real failure mode: degraded throughput during the catch-up period.
The Idea: Generative Simulation as a Benchmarking Substrate
The core insight from my experimentation was this — instead of waiting for real outages to test workflows, we can generate plausible clinical scenarios, inject controlled failures, and measure how the workflow recovers. This is generative simulation benchmarking.
There are three ingredients:
- A scenario generator — produces realistic patient cases with genomic profiles, urgency levels, and dependencies.
- A fault injector — simulates outages, latency spikes, partial data loss, and model rollbacks.
- A recovery scorer — measures how quickly and how fairly the workflow returns to steady state.
Let me walk through each.
Ingredient 1: Generating Realistic Oncology Scenarios
The temptation is to use an LLM to hallucinate cases. That's fine for prose, but oncology cases need structured genomic realism — variant frequencies, co-occurring mutations, and clinically plausible pairings. I ended up combining a structured prior (from public cohorts like TCGA-style distributions) with a generative model for the narrative layer.
import numpy as np
from dataclasses import dataclass
from typing import List
@dataclass
class OncologyCase:
case_id: str
urgency: int # 1 (routine) .. 5 (critical)
variants: List[str]
tumor_type: str
arrival_time: float # hours since simulation start
# Empirical prior: how often genes co-occur in lung adenocarcinoma
GENE_PRIORS = {
"EGFR": 0.15, "KRAS": 0.30, "ALK": 0.05,
"TP53": 0.50, "STK11": 0.15, "KEAP1": 0.12,
}
def sample_variants(rng: np.random.Generator, tumor_type: str):
# Sample a plausible variant set with co-occurrence structure
genes = list(GENE_PRIORS.keys())
probs = np.array([GENE_PRIORS[g] for g in genes])
mask = rng.random(len(genes)) < probs
selected = [g for g, m in zip(genes, mask) if m]
return selected or ["TP53"] # ensure at least one driver
def generate_case(rng, case_id, tumor_type="LUAD"):
urgency = rng.choice([1, 2, 3, 4, 5], p=[0.3, 0.3, 0.2, 0.15, 0.05])
return OncologyCase(
case_id=case_id,
urgency=urgency,
variants=sample_variants(rng, tumor_type),
tumor_type=tumor_type,
arrival_time=rng.exponential(scale=2.0),
)
The key design choice: urgency is sampled independently of variant complexity. This is critical because real outages reveal whether a workflow prioritizes the right cases, not just the easy ones. A naive scheduler will process fast, low-urgency cases first and let critical ones rot in the queue.
Ingredient 2: Fault Injection During Recovery Windows
A recovery window isn't just "the system was down." It's the period after the system returns when the backlog is being processed. I modeled it as a state machine with a degradation profile.
from enum import Enum
class SystemState(Enum):
HEALTHY = "healthy"
DEGRADED = "degraded"
DOWN = "down"
RECOVERING = "recovering"
class FaultInjector:
def __init__(self, schedule):
# schedule: list of (time, new_state, capacity_fraction)
self.schedule = sorted(schedule)
def state_at(self, t):
current = (SystemState.HEALTHY, 1.0)
for (start, state, cap) in self.schedule:
if t >= start:
current = (state, cap)
else:
break
return current
# Example: outage at t=8h, partial recovery at t=12h, full by t=14h
injector = FaultInjector([
(0, SystemState.HEALTHY, 1.0),
(8, SystemState.DOWN, 0.0),
(12, SystemState.RECOVERING, 0.4),
(14, SystemState.HEALTHY, 1.0),
])
While experimenting with fault schedules, I found that the shape of the recovery curve matters more than the outage duration. A system that comes back at 40% capacity for two hours is often worse than a system that stays down for two hours and returns at 100% — because the degraded mode creates partial results, inconsistent state, and rework.
Ingredient 3: Scoring Recovery Fairness and Throughput
The scoring function is where most benchmarking efforts go wrong. Accuracy and throughput are necessary but insufficient. I settled on three metrics:
- Time-to-steady-state (TSS): hours until the queue length returns to baseline.
- Critical-case delay (CCD): average additional latency for urgency-5 cases.
- Fairness ratio (FR): correlation between urgency and processing order.
def score_recovery(cases, completion_times, baseline_latency=2.0):
latencies = {c.case_id: completion_times[c.case_id] - c.arrival_time
for c in cases}
critical = [c for c in cases if c.urgency == 5]
ccd = np.mean([latencies[c.case_id] - baseline_latency for c in critical])
# Fairness: does higher urgency correlate with faster processing?
urgencies = np.array([c.urgency for c in cases])
lat = np.array([latencies[c.case_id] for c in cases])
fr = np.corrcoef(urgencies, -lat)[0, 1] # negative latency = faster
return {"critical_case_delay": ccd, "fairness_ratio": fr}
In my investigation of scheduler behavior under degraded capacity, I discovered that most default schedulers scored negative on fairness ratio during recovery — they actively prioritized the wrong cases. This was invisible without the simulation.
Building the Benchmarking Harness
Putting it together, the harness runs a discrete-event simulation. Here's a compact version using simpy:
import simpy
def run_simulation(cases, injector, service_rate=1.0):
env = simpy.Environment()
queue = simpy.PriorityResource(env, capacity=10)
completion_times = {}
def process(case):
with queue.request(priority=-case.urgency) as req:
yield req
state, capacity = injector.state_at(env.now)
# Effective service time scales inversely with capacity
service_time = 1.0 / max(capacity, 0.01) / service_rate
yield env.timeout(service_time)
completion_times[case.case_id] = env.now
for case in cases:
env.process(process(case))
env.run(until=48)
return completion_times
Two things I learned while wiring this up:
- Priority queues alone aren't enough. Under degraded capacity, even a priority queue can starve low-urgency cases indefinitely. You need a bounded priority scheme — urgency slots with reserved capacity.
- Simulation time must exceed the recovery window. I initially ran 24-hour simulations and missed the long tail of rework. Extending to 48 hours revealed that some workflows never fully recovered within the observation period.
Where Generative Models Actually Help
The structured prior above is fine for baseline scenarios, but it can't produce the weird cases that break workflows — the ones with contradictory annotations, missing tumor type, or rare co-occurring variants. This is where a language model earns its keep.
I used a small open-weights model with a constrained decoding schema to generate narrative case notes that matched the structured variant data. The trick was to keep the LLM downstream of the structured generator — never let it invent the genomics.
from pydantic import BaseModel
from typing import List
class CaseNarrative(BaseModel):
case_id: str
clinical_summary: str
prior_therapies: List[str]
complicating_factors: List[str]
def generate_narrative(case: OncologyCase, llm_client) -> CaseNarrative:
prompt = (
f"Write a brief clinical summary for a {case.tumor_type} patient "
f"with variants {case.variants}. Do NOT invent additional variants. "
f"Return JSON matching the CaseNarrative schema."
)
response = llm_client.generate(prompt, schema=CaseNarrative)
return response
One interesting finding from my experimentation with constrained decoding: the failure modes of narrative generation were mostly inconsistencies with the structured data — the LLM would invent a prior therapy that contradicted the variant profile. Validating the narrative against the structured case before admitting it to the simulation caught 90%+ of these.
Quantum-Inspired Sampling for Rare Failure Modes
Rare failure modes are, by definition, rare. Standard Monte Carlo wastes most of its samples on boring scenarios. I spent a few weeks exploring whether quantum-inspired sampling — specifically amplitude amplification analogues like Grover-style boosting — could concentrate samples on interesting regions.
I won't oversell this: real quantum advantage isn't available yet for this problem. But the classical analogue — importance sampling with an amplitude-inspired reweighting — was surprisingly effective.
def importance_sample(rng, n_samples, rarity_fn, boost=4.0):
"""Sample scenarios with emphasis on rare failure modes."""
samples = []
while len(samples) < n_samples:
case = generate_case(rng, case_id=f"c{len(samples)}")
rarity = rarity_fn(case) # 0..1, higher = rarer
accept_prob = min(1.0, (1 + boost * rarity) / (1 + boost))
if rng.random() < accept_prob:
samples.append(case)
return samples
During my investigation of this approach, I found that a modest boost factor (2–4x) gave the best tradeoff between coverage and realism. Push it too high and you're benchmarking against a synthetic world that doesn't resemble reality.
Agentic Workflows: The Next Frontier
The most recent thread of my exploration has been agentic AI systems — LLM-driven agents that autonomously triage cases, fetch annotations, and prepare MTB summaries. These are exactly the systems where recovery-window benchmarking matters most, because agents make decisions about what to work on next.
An agent that naively processes cases in arrival order will fail catastrophically during recovery. An agent that over-prioritizes urgency will starve routine cases. The right behavior is somewhere in between, and it has to be learned or specified explicitly.
class TriageAgent:
def __init__(self, policy):
self.policy = policy # maps (queue_state, system_state) -> case_id
def select_next(self, queue, system_state):
# Policy must account for degraded capacity
if system_state == SystemState.RECOVERING:
# Reserve 30% of slots for critical cases
critical = [c for c in queue if c.urgency >= 4]
if critical and len(queue) > 10:
return max(critical, key=lambda c: c.urgency)
return self.policy(queue, system_state)
While learning about agentic scheduling, I observed that the biggest failure mode wasn't choosing the wrong case — it was the agent re-choosing the same case repeatedly when downstream services were unavailable. Idempotency and backoff, boring as they sound, were the difference between a recovering workflow and a thrashing one.
Challenges I Ran Into
Challenge 1: Validation without ground truth. You can't validate a synthetic outage against a real one you didn't have. I ended up using historical incidents (anonymized) as calibration anchors — matching the distribution of real recovery times to my simulated ones.
Challenge 2: Computational cost. Full simulations with LLM-generated narratives are expensive. I split the pipeline: structured simulation runs fast and cheap, and narratives are generated only for cases that pass a rarity threshold.
Challenge 3: The metric problem. Every metric I added made the benchmark more realistic and less comparable to prior work. I settled on reporting a small core set (TSS, CCD, FR) plus optional extensions, so results remain comparable across studies.
Challenge 4: Human factors. The most important variable — how a clinician behaves under recovery pressure — is the hardest to simulate. I ended up modeling clinician behavior as a separate stochastic process with its own latency distribution, and treating it as a tunable parameter rather than a fixed constant.
Future Directions
Three directions I'm actively watching:
Agentic recovery policies learned via RL. Instead of hand-coding triage rules, train agents in the simulator and transfer them to real workflows. Early results are promising but the sim-to-real gap is real.
Federated simulation. Different institutions have different workflows. A federated simulation harness could benchmark recovery behavior across sites without sharing patient data.
Quantum-assisted scenario search. As quantum hardware matures, amplitude amplification for rare failure modes could become genuinely advantageous over classical importance sampling. I'm skeptical of near-term timelines, but the mathematical framing is already useful.
Conclusion: Benchmark the Recovery, Not Just the Model
The single biggest lesson from my learning journey is this: the moment a clinical AI system fails is not when it goes down — it's when it comes back up. Recovery windows are where workflows are most fragile, where priorities are most likely to invert, and where patients are most likely to be silently harmed.
Generative simulation benchmarking gives us a way to study this without waiting for real incidents. It lets us generate scenarios at scale, inject controlled faults, and measure not just throughput but fairness — whether the workflow prioritizes the patients who need it most.
If you take one thing from this article, let it be the fairness ratio. Track it. Measure it. Because a workflow that recovers fast but prioritizes the wrong cases isn't recovering at all — it's just failing quietly. And in precision oncology, quiet failures have names and faces attached to them.
The code in this article is a starting point, not a finished harness. I'd encourage you to fork it, inject your own fault schedules, and see what your workflows actually do when they're under pressure. You might be surprised — I certainly was.
Top comments (0)