DEV Community

slgayatri praharshitha
slgayatri praharshitha

Posted on

I Gave an AI Agent a Memory with Hindsight. Then I Taught It When to Ignore It.

Most AI agents have a simple memory story: remember something, retrieve something similar, and use it again.

That sounds useful until the world changes.

I built ED Resolve, an operational decision engine for a recurring Emergency Department problem: a patient is ready for ICU/HDU transfer, but something in the hospital workflow is blocking it. ED Resolve uses Hindsight to remember previous operational experiences — but it does not assume that an old solution is still correct.

The hard part was not making the agent remember.

It was making it answer:
Does this old experience actually apply to what is happening right now?
The problem is bigger than “find a bed”

An ICU transfer can depend on bed availability, bed cleaning, transport, coordination, escalation paths, and competing transfers. Two situations can look similar at a high level while requiring different actions.

That is why I did not want to build only a prediction model. I wanted a system that can reason over the current operational state, recall previous experiences, reject impossible interventions, simulate feasible alternatives, and learn from the outcome.

A 2026 observational study from a tertiary-care centre in India examined 510 adults who remained in an Emergency Department for at least 24 hours without specialty transfer and reported system-related causes including ICU/HDU bed waits and interdepartmental disputes. That work motivated the operational setting; ED Resolve itself uses synthetic scenarios and does not claim clinical effectiveness.

https://pubmed.ncbi.nlm.nih.gov/42326213/ · https://pmc.ncbi.nlm.nih.gov/articles/PMC13275199/

What happens inside ED Resolve?

CURRENT ED STATE
↓
WORLD MODEL + BOTTLENECK
↓
HINDSIGHT RECALL
↓
APPLICABILITY CHECK
↓
GPT-OSS CANDIDATE ACTIONS
↓
CP-SAT HARD CONSTRAINTS
↓
SIMPY COUNTERFACTUAL SIMULATION
↓
DETERMINISTIC SELECTION
↓
OUTCOME + PREDICTION ERROR
↓
HINDSIGHT RETAIN

Every stage has a different responsibility. That separation is deliberate: I did not want one LLM call to quietly become the entire control system.

Hindsight is the experience layer

I use Hindsight as the persistent experiential memory layer. Instead of keeping a few previous messages in a prompt, ED Resolve retains completed operational episodes: the state, bottleneck, candidate actions, constraints, selected action, outcome, and prediction error.

Later, another episode can recall those experiences.

The key distinction is:

Memory says:
“This happened before.”

ED Resolve asks:
“Was it the same kind of world?”

That is why Hindsight is central to the architecture rather than just an attached history store. For background, I also used Vectorize's guide to agent memory.

The key layer: applicability

A naive memory loop is:

Recall something similar → Repeat what worked

Our loop is different. Retrieved experiences pass through an a*pplicability gate* before they can influence the decision.

The prototype combines semantic relevance with structural signals:

relevance = (
0.20 * bottleneck_match
+ 0.15 * scenario_match
+ 0.10 * graph_match
+ 0.55 * semantic_relevance
)

The operational state also has a lightweight dependency graph. In a familiar case, it can identify a path such as:

ED → BED CLEANING → ICU

and identify BED_CLEANING as the dominant dependency.

A memory that mentions an ICU transfer is not enough. The current bottleneck or scenario structure should align with the old experience. The gate maps applicability to HIGH, MEDIUM, or LOW; in the current implementation, HIGH begins at 0.75 and MEDIUM at 0.45.

The lesson was simple: retrieval and applicability are different engineering problems.

The LLM proposes. It does not get the final vote.

GPT-OSS generates candidate interventions and reasoning context. The action space is deliberately bounded:

WAIT
EXPEDITE_BED_CLEANING
REQUEST_ALTERNATIVE_UNIT
PRIORITIZE_TRANSPORT
ESCALATE_TO_BED_MANAGER

Those candidates go through OR-Tools CP-SAT. If no alternative HDU capacity exists, REQUEST_ALTERNATIVE_UNIT is rejected before selection.

The solver is also deterministic:

solver.parameters.max_time_in_seconds = 0.25
solver.parameters.num_search_workers = 1
solver.parameters.random_seed = 0

So the architecture is easy to explain:

LLM proposes. Constraints filter. Simulation evaluates. Deterministic logic selects. Hindsight remembers.

Why simulate before deciding?

After impossible actions are removed, the remaining interventions are tested with SimPy, a discrete-event simulation library.

For a layman, think of this as a small virtual hospital workflow. Instead of immediately saying “expedite bed cleaning,” the system asks what would happen under the current conditions if we did that — and separately what would happen with the other feasible actions.

The simulator tracks synthetic consequences such as resolution time, boarding time, escalation count, resource conflicts, and whether the transfer resolves.

One verified familiar-regime run produced an expected resolution time of about 48.7 minutes **and an observed synthetic outcome of about **48.1 minutes, giving a prediction error of roughly 0.6 minutes.

Those are not hospital performance claims. They are measurements from our synthetic environment.

The learning signal is prediction error

I did not want to store only “what action won.” I wanted to store the gap between expectation and reality.

EXPECTED CONSEQUENCE
↓
ACTION
↓
OBSERVED CONSEQUENCE
↓
PREDICTION ERROR
↓
RETAIN EPISODE IN HINDSIGHT

In a controlled retain → recall test, Episode A was retained successfully, and a later Episode B recalled matching experience from Hindsight and used it as experience support.

That is the behavior I wanted: Episode B has access to institutional experience that did not exist in its own prompt.

The result I almost did not want to publish

In one memory ablation, I disabled experiential memory.

With memory enabled, evidence scores changed. But the final selected action** did not change** in that particular controlled case.

That was useful.

It showed that “the agent has memory” is not a meaningful evaluation by itself. The stronger questions are whether memory changes decisions when the current state is aligned with past experience, reduces repeated mistakes across sequences of episodes, or improves how the system reacts to prediction error.

I would rather show that limitation than pretend every memory retrieval automatically makes an agent better.

The strongest test: an unseen regime

I then created a deliberately unfamiliar synthetic scenario:

CASCADE_RESOURCE_COLLAPSE_X53

Critical resources are exhausted, normal escalation paths are disabled, and the operational structure no longer matches the familiar regimes.

Hindsight can still retrieve memories. But the applicability gate marks them LOW, the experiential prior falls to zero, and the simulation finds no executable resolution pathway.

So ED Resolve does not invent a recommendation:

AUTONOMOUS ACTION WITHHELD
HUMAN REVIEW REQUIRED

For me, this is more interesting than simply proving that memory retrieval works. A useful agent should know when the past is the wrong reference point.

What is actually doing what?

GPT-OSS → candidate generation and reasoning context.

Hindsight → persistent experience storage and recall.

NetworkX → operational dependency graph and bottleneck structure.

OR-Tools CP-SAT → hard feasibility constraints.

*SimPy *→ counterfactual operational consequences.

LangGraph → orchestration of the decision state.

FastAPI → the service layer behind the dashboard.

The important part is the contract between them:

LLM → suggestions
Hindsight → experience
Graph → structure
CP-SAT → feasibility
SimPy → consequences
Selector → authority
Outcome → learning signal

What building this taught me
**

  1. Similar does not mean applicable**

Semantic similarity is useful for finding experiences. Operational decisions need state alignment.

2. Bounded action spaces make agents easier to inspect

Five explicit actions are less flashy than unrestricted autonomy, but constraints, simulation, and failure cases become much easier to test.

3. Simulation is the bridge between reasoning and reality

The model can suggest something that sounds sensible. The current world can still make it impossible or ineffective.

4. Negative results are valuable

The memory ablation did not always change the final action. That became an evaluation lesson instead of a result to hide.

5. Safe Mode is part of the product

“I do not have a validated resolution pathway here” is a real operational outcome, not a crash.

Where I want to take it next

The current system is intentionally narrow: one workflow, synthetic scenarios, and a bounded action space.

The next step is a larger sequential evaluation where the same agent accumulates experience across many episodes. That would let us test the question I care about most:

Can institution-specific experience reduce repeated decision mistakes as the environment changes?

I would also add durable PostgreSQL audit storage, process-mining over event logs, richer uncertainty calibration, and a larger memory-on versus memory-off benchmark.

The core loop would remain:

REMEMBER → RETRIEVE → VALIDATE → SIMULATE → DECIDE → OBSERVE → LEARN

I did not want to build an agent that simply remembers the past.

I wanted to build one that can remember the past, compare it with the present, test what would happen now, and know when the past should be ignored.

Top comments (0)