DEV Community

Sritha Reddy
Sritha Reddy

Posted on

Making a Clinical AI Admit What It Doesn't Know: Grounded Generation With Evidence

In a clinical setting, a fluent wrong answer is worse than no answer. A co-pilot for therapists should only recommend what a child's records actually support.

The failure to avoid

A stateless model told a therapist to ask an autistic child to verbally describe his feelings during a meltdown. Reasonable on the surface, and the wrong move for that child at that moment.

Constrain the model to retrieved records

Ask for structured output where every recommendation must cite a memory:

SYSTEM = """You are a clinical co-pilot. Use ONLY the memories provided.
Return JSON: {"recommendations": [{"text": str, "evidence_ids": [str]}],
              "gaps": [str]}
If the memories do not support a recommendation, put it in "gaps"."""
Enter fullscreen mode Exit fullscreen mode

Validate the output

Don't trust the prompt alone. Check citations in code:

def validate(output, memory_ids):
    for rec in output["recommendations"]:
        if not rec["evidence_ids"]:
            return False
        if any(e not in memory_ids for e in rec["evidence_ids"]):
            return False
    return True
Enter fullscreen mode Exit fullscreen mode

If validation fails, retry once, then fall back to "insufficient records."

Test it

Build a small evaluation set of realistic queries, including ones where the records have no answer. Compare a stateless baseline against the grounded agent and track how often each invents an unsupported intervention.

Takeaways

  1. Force citations so every claim traces to a record.
  2. Give the model an explicit "gaps" channel so it can decline.
  3. Validate in code, and evaluate with cases that have no right answer.

Top comments (1)

Collapse
 
aifrontierpost profile image
AI Frontier Post •

The retry-once-then-fallback rule is the load-bearing detail here, but there's a subtle hole: if you validate evidence_ids against memory_ids, a retried model can still pass by citing IDs from its own first draft, so the memory_ids set has to exclude anything the model generated this session. The other underrated risk is measuring the wrong thing — an eval that only counts hallucination rate will teach the model that stuffing everything into "gaps" is the safest way to never get flagged.