If a patient reports a pain score of 6 out of 10 on their third morning post-knee replacement, should their surgical team be worried?
If you ask an LLM in a stateless prompt containing only that morning's intake form, it will give you a generic, hedged answer. It might suggest resting, icing the joint, and taking paracetamol. But in surgical recovery, a pain score of 6 is completely meaningless in isolation.
If that patient was an 8 yesterday and a 9 on Day 1, a 6 is great news. The acute inflammatory response is subsiding, analgesics are working, and the recovery trajectory is positive.
If that same patient was a 2 yesterday, walking comfortably, and suddenly jumps to a 6 alongside localized warmth and shivering, they are in trouble. They may be developing a deep surgical site infection or a hematoma.
Early on, we treated check-ins as independent point-in-time events. We passed the daily intake into an LLM, asked for an assessment, and hoped for clinical insight. It was an architectural dead end. The model either caused alert fatigue by treating every expected ache as a crisis, or worse, missed gradual, multi-day deterioration because no single day looked catastrophic on its own.
To monitor recovery safely, an agent cannot simply read today's status. It must understand trajectory, baseline, and past clinical decisions. That is why we integrated Vectorize agent memory
through Hindsight into CareTraceAI.
Here is how the system works in production, how we wired episodic memory into our backend, and what we learned about designing longitudinal systems that clinicians can actually trust.
What CareTraceAI Does in Production
CareTraceAI is an automated post-operative recovery monitoring engine. When a patient is discharged following a major procedure—such as a knee replacement, appendectomy, or cardiac bypass—our backend enrolls them into a scheduled monitoring protocol.
Every morning at 9:00 AM, a background scheduler triggers a WhatsApp or SMS check-in tailored to the patient's procedure:
Pain level on a 0–10 scale
Mobility relative to yesterday (Better, Same, Worse)
Localized swelling (None, Mild, Moderate, Severe)
Temperature flags (fever > 100.4°F or chills)
Free-text symptom notes (drainage, odor, breathlessness)
When the patient replies, CareTraceAI evaluates their response, assigns a risk tier (Low, High, Critical), computes trajectory trends (improving, stable, worsening), and alerts hospital staff if deterioration occurs. Clinicians review these alerts on an operational dashboard, record their determinations, and close the loop.
The entire system relies on one premise: every check-in must be evaluated against the patient's longitudinal history.
Why Naive Vector RAG Collapses in Longitudinal Care
When engineers first realize an agent needs context, the standard reflex is traditional Retrieval-Augmented Generation (RAG). You chunk past notes, embed them into a vector database, and run cosine similarity searches against the latest message.
In post-operative care, naive vector RAG fails for three concrete reasons:
Similarity Is Not Relevance: If a patient writes, "My knee feels stiff and swollen today," semantic search retrieves every prior check-in mentioning "knee" or "swelling"—pulling Day 1, Day 3, and pre-op notes while ignoring yesterday's Day 5 baseline that showed sudden mobility loss.
Loss of Temporal Sequence: Standard vector databases treat documents as static coordinates in embedding space. In recovery, temporal sequence is the entire signal. A mild fever on Day 1 is common atelectasis; a new fever on Day 7 is a probable infection.
No Delta Computation: Vector search cannot calculate numerical deltas. It cannot compute that pain shifted from 3 to 7 (+4 points) or that swelling worsened by two categorical ranks.
Instead of treating patient history as an unstructured document pile, we needed structured episodic memory designed for agentic workflows. That led us to the Hindsight GitHub
repository and the @vectorize-io/hindsight-client
SDK.
Hindsight provides bank-isolated episodic memory that allows an agent to retain structured experiences and recall relevant temporal context when analyzing new events.
The Core Technical Story: Building the Memory Loop
In CareTraceAI, we implemented a four-phase memory pipeline for every check-in:
Temporal Recall: Fetch baseline and past recovery events from Hindsight.
Delta Math: Compute deterministic physiological deltas against the prior checkpoint.
Guardrailed Assessment: Merge recalled history and deltas through our clinical safety engine.
Retention: Commit the check-in and clinician review back into persistent Hindsight memory.
Here is how our dual-layer memory service is implemented in backend/services/memory.js:
javascript
const { HindsightClient } = require("@vectorize-io/hindsight-client");
const db = require("../database");
const client = new HindsightClient({
baseUrl: process.env.HINDSIGHT_URL || "http://localhost:8888",
apiKey: process.env.HINDSIGHT_API_KEY || undefined,
});
const BANK = process.env.HINDSIGHT_BANK || "caretrace";
async function retain(patient, text, options = {}) {
const { event_type = "daily_checkin", day_number = 1, context = {} } = options;
// Local SQLite store for auditability and compliance
await run(
INSERT INTO hindsight_memories (patient_id, day_number, event_type, memory_text, context_json),
VALUES (?, ?, ?, ?, ?)
[patient.id, day_number, event_type, text, JSON.stringify(context || {})]
);
// Transmit to external Vectorize Hindsight Client
try {
await client.retain(BANK, Patient ${patient.name} (ID ${patient.id}): ${text});
return true;
} catch (e) {
return true; // Fallback gracefully if client is offline
}
}
We maintain a local SQLite hindsight_memories table alongside the Hindsight client. This gives us zero-latency local fallback, audit logging for medical compliance, and rapid structured queries, while Hindsight handles long-horizon episodic indexing.
When a check-in arrives in backend/routes/checkins.js, we run deterministic delta math between the previous day's snapshot and today's input:
javascript
const previous = await get(
SELECT * FROM checkins WHERE patient_id = ? AND status = 'completed' ORDER BY day DESC, id DESC LIMIT 1,
[patient_id]
);
const comparison = compareWithPrevious(current, previous);
let memories = [];
if (memory !== "off") {
memories = await recall(
patient,
Post-op recovery Day ${day}. Current update: ${message}. Pain: ${current.pain}/10. Mobility: ${current.mobility}. Swelling: ${current.swelling}. Fever: ${current.fever ? "yes" : "no"}. Symptoms: ${current.symptoms}. Trend: ${comparison.trend}.
);
}
The compareWithPrevious helper calculates concrete differentials: pain delta, mobility shift, swelling progression, and fever appearance. These deltas and recalled memories feed directly into our assessment service.
Never Let an LLM Downgrade a Clinical Floor
Generative models must never have unchecked authority over patient risk classification. LLMs are prone to sycophancy. If a patient reports severe pain or a +5 delta spike, an LLM might produce a calming classification if the patient writes in a polite tone.
In backend/services/risk.js, we implemented hard clinical rules that compute a deterministic safety floor. If red flags or critical delta spikes occur, the floor is set to Critical or High.
We then pass the prompt to an LLM to synthesize clinical explanations, but strictly forbid it from downgrading below the floor:
javascript
const clinicalEvaluation = evaluateClinicalRisk({
patient, day, pain, mobility, swelling, fever, symptoms, message, hindsight, previous
});
const floorRisk = clinicalEvaluation.risk_level;
let finalResult = { ...clinicalEvaluation };
if (process.env.GROQ_API_KEY) {
const parsed = await callLlmAssessment({ patient, day, history, clinicalEvaluation });
if (parsed?.risk_level) {
const parsedIdx = ORDER.indexOf(parsed.risk_level);
const floorIdx = ORDER.indexOf(floorRisk);
// Safety floor: LLM can escalate risk, but NEVER downgrade
finalResult.risk_level = parsedIdx >= floorIdx ? parsed.risk_level : floorRisk;
finalResult.explanation = parsed.explanation || clinicalEvaluation.explanation;
}
}
This hybrid pattern delivers physiological safety combined with fluent, context-aware reasoning grounded in Hindsight memories.
Retaining Clinician Feedback: Closing the Memory Loop
Most agentic systems only store what the user says. In healthcare, what the doctor decides is far more important.
When CareTraceAI generates an alert, a clinician reviews it on the dashboard, selects an outcome (Confirmed concern, False alarm, or Resolved), and adds clinical notes.
If this decision is not retained in agent memory, the next check-in fails: the patient reports elevated pain or medication side effects, and the agent panics again, unaware that an intervention is already underway.
In backend/routes/checkins.js, we explicitly retain the clinician's review back into Hindsight:
javascript
await retainClinicianReview(patient, {
outcome,
note: clinician_name ? [${clinician_name}] ${note} : note,
alert_title: alertTitle,
day_number: dayNumber
});
By retaining clinician determinations as first-class memory events, the agent transforms from a passive symptom counter into an informed participant in the care plan.
Example Interaction: A Four-Day Recovery Trajectory
Consider the recovery trajectory of Rajesh, 54, who underwent total knee replacement:
Day 1: Baseline Established
Patient Input: Pain 4/10, Mobility Same, Swelling Mild, Fever No. Notes: "Resting with leg elevated."
Hindsight Recall: No prior records. Baseline created.
Output: Risk: Low, Trend: baseline. Baseline recorded in Hindsight.
Day 2: Normal Progression
Patient Input: Pain 3/10, Mobility Better, Swelling None, Fever No. Notes: "Walking with walker, pain manageable."
Hindsight Recall: Day 1 baseline retrieved.
Delta Engine: Pain -1 pt, Swelling down. Trend: improving.
Output: Risk: Low, Trend: improving.
Day 3: Acute Deterioration (The Spike)
Patient Input: Pain 8/10, Mobility Worse, Swelling Severe, Fever Yes (101.2°F). Notes: "Knee is throbbing and hot. Shivering."
Hindsight Recall: Days 1 & 2 retrieved (Pain 3/10, improving).
Delta Engine: Pain spike: 3 → 8 (+5 pts). Swelling: None → Severe. Fever: No → Yes. Trend: worsening.
Output: Risk: Critical, Trend: worsening.
Alert Triggered: 🚨 Critical Recovery Alert — Day 3: Acute pain spike (+5 pts) with severe swelling and new fever.
Clinician Intervention: Dr. Mehta reviews the alert, calls Rajesh, diagnoses suspected cellulitis, and prescribes oral antibiotics. Dr. Mehta logs: "Prescribed oral Amoxicillin-Clavulanate 625mg BID."
Memory Retained: Clinician determination stored in Hindsight.
Day 4: Post-Intervention Stabilization
Patient Input: Pain 5/10, Mobility Same, Swelling Moderate, Fever No. Notes: "Started antibiotics last night. Fever broke."
Hindsight Recall: Recalls Day 3 spike (+5 pain, fever) AND Dr. Mehta's antibiotic prescription.
Delta Engine: Pain: 8 → 5 (-3 pts). Fever: Yes → No. Trend: improving.
Output: Risk: Low, Trend: improving.
Explanation: "Patient shows positive response to initiated antibiotic therapy. Pain reduced by 3 points from Day 3 spike; fever resolved."
Without Hindsight, Day 4 would trigger another alert: pain 5/10 is still elevated. Because the agent recalled the antibiotic prescription and Day 3 spike, it recognized a stabilizing trajectory, avoiding an erroneous false alarm.
Four Lessons Learned from Shipping Agent Memory
Building CareTraceAI changed how we approach agent memory in high-stakes domains:
Longitudinal Delta Beats Absolute Value: In physiological monitoring, absolute metrics are noisy. The true clinical signal lies in the first derivative—the delta between yesterday and today. If agent memory cannot compute and contextualize deltas, it cannot monitor patients safely.
Close the Loop with Clinician Decisions: An agent that only remembers user statements is half-blind. In our system, the most valuable memories retained in Hindsight are clinician determinations. Retaining care team interventions keeps the agent aligned with the active treatment plan.
Build a Dual-Store Memory Architecture: Never rely exclusively on an external service for operational state. We maintain local SQLite tables (hindsight_memories, checkins) for transactional safety and compliance, and layer Hindsight on top for episodic indexing and semantic recall.
Never Trust an LLM with Safety Floors: Generative models excel at narrative synthesis and multi-day reasoning, but they should never decide whether acute symptoms warrant emergency triage. Use deterministic rules to set an unbreakable floor, and let the LLM operate above it.
Conclusion
Stateless prompts work well for one-off tasks like drafting emails or code completion. But the moment an agent monitors a human process over days or weeks, statelessness becomes an architectural liability.
By integrating structured episodic memory with Hindsight, we turned an isolated check-in script into a production-grade recovery intelligence system. CareTraceAI tracks the entire recovery journey, detects complications early, and gives surgical teams the longitudinal context they need to intervene when it counts.
To learn more about implementing agent memory, check out the Hindsight docs
and the Hindsight GitHub
.



Top comments (0)