Hindsight Remembered the Cases; Code Chose the Evidence
The hardest part of giving an agent memory was not storing more information. It was deciding what that memory was actually allowed to prove.
I built Echo, an accounts-payable exception investigation system that combines persistent institutional memory with deterministic evidence processing. When an invoice hits an exception, Echo compares the written policy with what happened in similar historical cases, shows the conditions behind those decisions, and gives a human the evidence needed to make the final decision.
The project is open on GitHub:
Echo — AP Exception Intelligence Agent
The important architectural decision is simple:
Hindsight stores memory. Deterministic code decides what counts as evidence. The LLM explains that evidence.
That separation changed the way I designed the entire system.
The problem: policy is not always practice
Accounts payable looks straightforward until an invoice does not fit the normal process.
An invoice may have:
- an amount mismatch
- a missing purchase order
- a GST mismatch
- a suspected duplicate
- a policy threshold violation
- a special condition that changes how the exception is handled
The written policy can tell an AP operator what should happen.
But organizations accumulate another layer of knowledge over time:
- which exceptions were actually approved
- which approver handled them
- which workaround was used
- under which condition the approval happened
- whether a particular exception pattern repeatedly appeared
- whether historical practice differed from the written route
That knowledge is usually scattered across old cases, spreadsheets, emails, and human memory.
I wanted Echo to make that institutional memory queryable.
But there was an immediate problem.
If an agent retrieves five “similar” historical cases, are all five really evidence?
Not necessarily.
A case can be semantically similar while belonging to a different policy regime, amount band, condition, or exception pattern.
So Echo does not allow retrieval alone to determine the conclusion.
What Echo does
At a high level, Echo sits between the written AP policy and the human decision-maker.
The workflow is:
The architecture intentionally gives each component a narrow responsibility.
The Hindsight integration
The application keeps Hindsight behind a backend memory layer rather than spreading memory calls throughout the UI.
For example, the case expansion path verifies that Hindsight is configured before writing historical cases:
const cfg = hindsightConfig();
if (!cfg) {
console.error(
"Hindsight is not configured. Set HINDSIGHT_BANK_ID and HINDSIGHT_BASE_URL in .env.local"
);
process.exitCode = 1;
return;
}
const cases = JSON.parse(
readFileSync(EXPANSION_PATH, "utf8")
) as ExpansionCase[];
The important design choice is that the frontend never becomes the memory system itself.
The runtime retrieval path follows the same separation:
User / Invoice
↓
Identify exception family
↓
Recall relevant Hindsight pattern
↓
Read supporting case IDs
↓
Fetch exact historical cases
↓
Apply deterministic comparability
↓
Compute evidence + confidence
↓
LLM explains the evidence
↓
Human decision
↓
Retain new resolution
This gives the system a clear boundary between memory, evidence, explanation, and decision-making.
Two kinds of memory
Echo separates case memory from pattern memory.
A case represents a concrete historical event:
case_id: C12
invoice: INV-NW-1193
issue_family: amount_mismatch
approver: Priya Nair
condition: vendor annexure attached
outcome: approved
A pattern is an aggregate observation supported by multiple cases:
pattern:
fuel-surcharge exception
support_count:
5
approved_count:
5
rejected_count:
0
condition:
fuel-surcharge annexure attached
supporting_case_ids:
C02, C04, C08, C10, C12
The distinction matters because the pattern can point back to the exact historical cases that support it.
Hindsight provides persistent memory.
Echo's deterministic layer decides how that memory becomes evidence.
The retrieval boundary
A simple memory implementation could look like this:
Query
↓
Retrieve similar cases
↓
Give everything to the LLM
I did not want Echo to work that way.
Instead, the runtime path is:
Query
↓
Identify exception family
↓
Recall relevant Hindsight pattern
↓
Extract supporting case IDs
↓
Read those exact cases
↓
Apply comparability rules
↓
Exclude the current invoice
↓
Compute evidence
↓
LLM explains the result
The distinction matters.
Hindsight can find memories that are relevant to the query. That does not automatically mean every retrieved memory is valid evidence for the current invoice.
Echo therefore evaluates historical cases against deterministic rules.
The comparison considers the issue family, relevant pattern, applicable policy regime, amount or difference information when available, the side of a policy threshold, and whether the case contains enough information to be compared.
A historical case that cannot be compared reliably is not silently treated as evidence.
That became one of the central rules of the project:
Similarity finds candidates. Deterministic logic decides evidence.
Confidence is computed, not generated
I also wanted to keep confidence outside the language model.
Echo calculates the confidence level from the actual evidence.
For example:
5 comparable historical cases
5 / 5 approved
MEDIUM — SMALL SAMPLE
The language model can explain what that evidence means, but it does not invent the approval count or confidence level.
That creates a clear provenance boundary:
COMPUTED
↓
Deterministic application logic
EVIDENCE
↓
Historical Hindsight cases
SYNTHESIZED
↓
LLM explanation
This is important because an explanation can sound convincing even when the underlying evidence is weak.
Echo therefore keeps the evidence and the explanation visibly separate.
A concrete investigation
Consider:
INV-NW-1188
Northwind Freight
Amount mismatch
With memory enabled, Echo retrieves five comparable historical cases:
C02
C04
C08
C10
C12
All five were approved.
The recurring condition is:
fuel-surcharge annexure attached
Instead of simply telling the user that five invoices were “similar,” Echo exposes the evidence chain:
Exception
↓
Fuel-surcharge pattern
↓
Annexure attached
↓
5 comparable historical cases
↓
5 approvals
At the same time, Echo shows the written policy separately.
The user can therefore see:
WHAT THE POLICY SAYS
vs.
WHAT HISTORICAL PRACTICE SHOWS
The point is not to silently replace written policy with history.
The point is to expose the gap between the written route and historical practice so a human can make the final decision with better context.
Before and after persistent memory
The same invoice provides a useful comparison.
Before: written policy only
Without historical memory, Echo can still evaluate the invoice against the written rule.
The system can identify the exception and show the applicable policy route.
But it has no institutional precedent to show.
The result is effectively:
Written policy:
CFO approval required for mismatch
above INR 50,000.
Historical precedent:
None available in memory mode.
After: persistent institutional memory
With Hindsight enabled, the same investigation can include:
5 comparable historical cases
C02
C04
C08
C10
C12
5 / 5 approved
Condition:
fuel-surcharge annexure attached
That changes the investigation from a pure policy lookup into a policy-plus-practice analysis.
The user can see both sources of information rather than receiving a single opaque AI conclusion.
Memory ON vs Memory OFF
Echo also makes the memory boundary explicit.
With memory enabled:
MEMORY ACTIVE
Historical cases:
5
Historical outcome:
5 / 5 approved
With memory disabled:
WRITTEN POLICY ONLY
The system must not use historical precedent when memory is off.
This matters because a visual toggle is not enough by itself. The actual data flow has to respect the boundary.
The same invoice therefore becomes a useful test:
MEMORY ON
Policy + Historical Evidence
MEMORY OFF
Written Policy Only
The no-history case
A memory system also needs to handle the situation where there is no useful precedent.
Echo has an explicit no-history state:
NO INSTITUTIONAL PRECEDENT
In that situation, the system does not invent:
- historical cases
- approval counts
- conditions
- confidence
- past outcomes
Instead, it falls back to the written policy.
That gives the system a useful failure mode:
No sufficient historical precedent was found.
Written policy remains the available guidance.
I consider that an important property of the architecture.
An agent does not need to answer every historical question.
Sometimes the correct answer is that the available memory is insufficient.
Learning from human decisions
Memory is useful only if it can grow.
When a human resolves an exception, Echo retains the new case and updates the corresponding pattern.
The update flow is:
Human resolution
↓
Retain new case
↓
Read current pattern
↓
Read supporting cases
↓
Add new case
↓
Recompute aggregate
↓
Store updated pattern
↓
Read back and verify
The pattern stores structured information such as:
support count
approved count
rejected count
approvers
condition
policy route
last seen
supporting case IDs
The language model can explain those facts, but it does not become the source of those facts.
That keeps the learning loop inspectable.
Echo Replay
Echo Replay uses the same evidence model to compare historical precedent under different conditions.
For example:
Current condition:
fuel-surcharge annexure attached
Historical evidence:
5 comparable cases
5 approved
The user can then inspect an alternative condition.
If historical precedent does not exist for that condition, Echo shows:
No comparable historical precedent found.
It does not manufacture an approval or rejection merely because the condition changed.
Absence of evidence remains absence of evidence.
Reflection is deliberately separated
Hindsight also provides reflection capabilities.
Echo has an opt-in reflection path for producing a concise synthesis of retrieved evidence.
However, reflection is not treated as the source of structured truth.
It should not determine:
case identity
approval count
confidence
policy truth
historical outcome
Those remain grounded in the historical memory and deterministic application logic.
The responsibility boundary is:
Hindsight
↓
Persistent memory
Deterministic Echo logic
↓
Evidence + confidence + comparability
LLM
↓
Explanation
Human
↓
Final decision
This separation made the system much easier to reason about.
What changed as the memory corpus grew
The system started with a smaller verified set of historical cases.
I then expanded the corpus with 150 additional structured cases covering:
- amount mismatches
- missing POs
- GST mismatches
- duplicate suspects
The expansion was validated before being seeded into Hindsight.
Adding more memory did not simply make retrieval “better.”
It also introduced a new problem.
The larger the memory corpus became, the more likely it was that a new query could retrieve something that looked similar but belonged to a different policy regime or condition.
That made the evidence boundary more important.
The solution was not to give the language model more historical context.
The solution was to give it a smaller and better-defined evidence window.
An honest engineering lesson
One of the biggest lessons from building Echo was that persistent memory creates a second problem after solving the first one.
The first problem is:
How do I remember what happened?
The second problem is:
How do I know whether what I remembered actually applies here?
Those are different engineering problems.
The first needs persistent memory.
The second needs deterministic rules.
That distinction changed the architecture of Echo.
Instead of:
Retrieve everything
↓
Ask the LLM to decide
the system became:
Retrieve
↓
Validate
↓
Compare
↓
Compute
↓
Explain
That extra boundary is what makes the memory useful without allowing memory itself to become an uncontrolled source of truth.
What I would reuse in another agent
1. Separate retrieval from evidence selection
Semantic retrieval is useful for finding candidates.
It should not automatically determine which candidates are valid evidence.
2. Keep confidence outside the LLM
Counts, sample sizes, and evidence strength are easier to verify when they come from deterministic logic.
3. Store conditions alongside outcomes
“Approved” is less useful than:
Approved when the vendor annexure was attached.
The condition explains when historical behaviour occurred.
4. Treat no precedent as a valid result
An agent should be able to say:
No sufficient historical precedent was found.
rather than manufacturing a conclusion from unrelated memories.
5. Let human decisions become future memory
The system becomes more useful when resolved cases can become part of its future institutional memory.
The architecture I would reuse
The most reusable part of Echo is not the AP domain itself.
It is the separation of responsibilities.
This pattern can extend beyond accounts payable.
The domain can change.
The responsibility boundaries can remain the same:
Memory preserves experience. Deterministic logic establishes evidence. Models explain the evidence. Humans make the final decision.
Screenshots
The screenshots are part of the technical story, not just decoration.
Investigation workspace
Memory ON vs Memory OFF
Final thoughts
I started Echo thinking that the interesting engineering problem would be giving an agent long-term memory.
The harder problem turned out to be defining what that memory was actually allowed to prove.
Hindsight gave Echo persistent institutional memory.
The rest of the architecture determines how responsibly that memory is used.
The result is not a database of answers.
It is a system that remembers what happened, shows the conditions under which it happened, compares that history with written policy, and gives a human a clearer basis for the next decision.
That is the part of agent memory I find most interesting:
Not remembering everything, but remembering enough to make the next decision better grounded.
Project
Echo — AP Exception Intelligence Agent













Top comments (0)