Healthcare data is notorious for being "messy." If you've ever worked in medical data engineering, you know the pain of staring at inconsistent hospital reports, shorthand notes, and non-standard lab results. Achieving healthcare interoperability has historically required thousands of lines of regex and manual mappingβuntil now.
In this tutorial, we are going to build an automated pipeline that transforms "dirty" clinical text into the FHIR Standard (Fast Healthcare Interoperability Resources) using LangChain and Pydantic. By the end, you'll see how to leverage Large Language Models (LLMs) to turn unstructured medical gibberish into high-quality, queryable data for your PostgreSQL database. π
Why FHIR and Why Now?
The FHIR standard is the backbone of modern health tech. However, most legacy systems still output unstructured text. Using an automated data pipeline powered by LLMs allows us to bridge this gap without losing context or clinical nuance.
The Architecture: From Raw Text to Structured JSON
Before we dive into the code, let's look at the data flow. We ingest raw text, pass it through an LLM guided by a strict schema, validate it, and sink it into a relational database.
graph TD
A[Dirty Medical Text] --> B{LLM + Pydantic Parser}
B --> C[Structured FHIR JSON]
C --> D{Validation Check}
D -- Pass --> E[PostgreSQL Storage]
D -- Fail --> F[Human-in-the-loop Review]
E --> G[Analytics / EHR Integration]
Prerequisites
To follow along, you'll need:
- Python 3.9+
- An OpenAI API Key (or equivalent LLM provider)
pip install langchain langchain-openai pydantic sqlalchemy psycopg2
Step 1: Defining the FHIR Schema with Pydantic
The secret sauce to getting structured output from an LLM is the Pydantic Output Parser. We define exactly what we want the output to look like using Python classes.
from pydantic import BaseModel, Field
from typing import List, Optional
class FHIRObservation(BaseModel):
"""Represents a single clinical observation in FHIR format."""
resourceType: str = Field("Observation", description="The type of FHIR resource")
status: str = Field(..., description="Status: final, amended, corrected, etc.")
code_coding_code: str = Field(..., description="LOINC code for the test")
code_coding_display: str = Field(..., description="Human-readable name of the test")
valueQuantity_value: float = Field(..., description="The numerical result of the test")
valueQuantity_unit: str = Field(..., description="Unit of measurement (e.g., mg/dL, mmol/L)")
effectiveDateTime: str = Field(..., description="ISO 8601 timestamp of the observation")
class FHIRBundle(BaseModel):
"""A collection of FHIR resources."""
observations: List[FHIRObservation]
Step 2: The LangChain Extraction Chain
Now, let's build the prompt and the chain. We want to tell the LLM to act as a medical coder.
from langchain_openai import ChatOpenAI
from langchain.prompts import PromptTemplate
from langchain.output_parsers import PydanticOutputParser
# Initialize the LLM
llm = ChatOpenAI(model="gpt-4-turbo", temperature=0)
# Set up the parser
parser = PydanticOutputParser(pydantic_object=FHIRBundle)
# Create the prompt template
template = """
You are an expert medical data engineer. Extract clinical observations from the following raw text
and format them strictly according to the FHIR standard.
Text: {raw_text}
{format_instructions}
"""
prompt = PromptTemplate(
template=template,
input_variables=["raw_text"],
partial_variables={"format_instructions": parser.get_format_instructions()},
)
chain = prompt | llm | parser
Step 3: Running the Pipeline
Let's feed it some messy data. Imagine a doctor's note that looks like this:
"Patient visited on 2023-10-12. Glucose levels were 110 mg/dL, slightly high. Blood pressure was 120/80. Potassium 4.5 mmol/L."
raw_clinical_note = "Patient visited on 2023-10-12. Glucose levels were 110 mg/dL. Potassium 4.5 mmol/L."
try:
fhir_data = chain.invoke({"raw_text": raw_clinical_note})
print(fhir_data.json(indent=2))
except Exception as e:
print(f"Error parsing data: {e}")
Advanced Patterns for Production π₯
While the example above works for simple cases, production-grade medical systems require handling complex edge cases, multi-step validation, and HIPAA-compliant data handling.
For a deeper dive into production-ready architectures and how to handle high-throughput medical data streams, check out the specialized guides at WellAlly Blog. They offer excellent insights into advanced LLM orchestration and data reliability patterns that go beyond basic tutorials.
Step 4: Persisting to PostgreSQL
Once we have the validated Pydantic object, saving it to PostgreSQL is straightforward using an ORM like SQLAlchemy.
from sqlalchemy import create_engine, Column, Integer, JSON
from sqlalchemy.ext.declarative import declarative_base
from sqlalchemy.orm import sessionmaker
Base = declarative_base()
class MedicalRecord(Base):
__tablename__ = 'fhir_records'
id = Column(Integer, primary_key=True)
content = Column(JSON) # Store the FHIR JSON directly
# Setup DB connection
engine = create_engine('postgresql://user:pass@localhost/medical_db')
Base.metadata.create_all(engine)
Session = sessionmaker(bind=engine)
session = Session()
# Insert the record
new_record = MedicalRecord(content=fhir_data.dict())
session.add(new_record)
session.commit()
print("β
Record successfully saved to PostgreSQL!")
Conclusion
Standardizing healthcare data doesn't have to be a manual nightmare. By combining LangChain's orchestration, Pydantic's validation, and the FHIR standard, we can build resilient, automated pipelines that turn "dirty" pixels and text into actionable medical insights.
What's next for your pipeline?
- π₯ Adding HIPAA-compliant de-identification layers.
- π§ͺ Integrating with real-time lab feeds via Webhooks.
- π€ Using RAG (Retrieval-Augmented Generation) to verify clinical codes against the LOINC database.
If you enjoyed this tutorial, don't forget to subscribe for more "Learning in Public" content! Let me know in the comments: what's the messiest data format you've ever had to deal with? π¬
Top comments (0)