DEV Community

Cover image for Why Evidence Matters When AI Agents Use Memory

Why Evidence Matters When AI Agents Use Memory

Why Evidence Matters When AI Agents Use Memory

Giving an AI agent memory creates a new problem:

What happens when the agent remembers correctly but reasons beyond the evidence?

Persistent memory makes agents more capable because they can use information from earlier interactions.

It can also make unsupported conclusions feel more convincing.

PactTrace is a vendor-memory application designed around this exact tension.

It uses Hindsight to remember historical vendor commitments, recalls them when a new quote arrives, and then compares the new evidence against that history.

But the comparison is not accepted blindly.

PactTrace validates the evidence before presenting the result.

Why memory alone is not enough

Suppose a vendor previously says:

We will provide a 15% renewal discount if your account exceeds 100 seats.

Later, the company receives a renewal quote.

A memory-enabled agent can recall the earlier promise.

That solves one problem.

But several new questions appear:

  • Does the current quote actually prove the account exceeds 100 seats?
  • Does the quote explicitly include or exclude the discount?
  • Is the recalled memory from the correct vendor?
  • Did the model copy the historical condition correctly?
  • Is the model interpreting silence as contradiction?

These are evidence problems, not memory problems.

That distinction is important.

Hindsight provides historical context

PactTrace uses Hindsight to retrieve relevant vendor memories.

The recall layer asks for earlier commitments, discounts, pricing promises, fee waivers, conditions, and negotiation history.

But PactTrace does not treat the recall query itself as an isolation boundary.

It verifies vendor metadata after retrieval:

ts
const owner = memory.metadata?.vendor;

if (
!owner ||
normalizeVendor(owner) !== normalizedVendor ||
!memory.text.trim()
) {
return [];
}

This reduces the risk of unrelated vendor information entering the comparison.

The memory system finds candidates.

The application still has responsibility for deciding what evidence it will accept.

Preserve exact historical text

PactTrace keeps the original memory ID and text when selecting relevant memories.

It also avoids fuzzy deduplication for details where small differences matter.

For example:

  • 10% discount vs 15% discount
  • more than 100 seats vs at least 100 seats
  • renewal discount vs onboarding discount
  • May 1 vs May 31

Those are not cosmetic differences.

The code deliberately avoids merging memories when numbers, dates, negation, or conditions change.

ts
return memories
.filter((memory) => {
const key = memory.text
.split(/ | (?:When|Involving):/)[0]
.normalize("NFKC")
.toLowerCase()
.replace(/\s+/g, " ")
.trim()
.replace(/[.!]+$/, "");

if (seen.has(key)) return false;

seen.add(key);
return true;
Enter fullscreen mode Exit fullscreen mode

})
.slice(0, 5);

That conservatism protects the meaning of historical commitments.

The model has to cite its evidence

The comparison engine returns structured findings.

Each finding contains:

  • memoryId
  • historical commitment
  • current evidence
  • condition
  • condition status
  • status
  • severity
  • explanation

But generating those fields is not enough.

PactTrace checks them.

ts
if (
!memory ||
!conflict.historicalCommitment.trim() ||
!memory.text.includes(conflict.historicalCommitment) ||
!conflict.currentEvidence.trim() ||
!input.quote.includes(conflict.currentEvidence)
) {
throw new Error(
"Comparison evidence was not grounded in the supplied sources"
);
}

The historical excerpt must literally exist inside the recalled memory.

The current evidence must literally exist inside the new quote.

This is a very practical guardrail.

A fluent explanation cannot replace missing evidence.

Three outcomes are better than two

Many systems want a binary answer:

conflict / no conflict

PactTrace uses three states:

Potential conflict

Explicit evidence suggests that an applicable historical commitment may be contradicted.

Honored

Explicit evidence shows the applicable promise is being fulfilled.

Insufficient evidence

The system does not have enough information to support either conclusion.

That third option is essential.

An AI agent should be allowed to say:

I do not have enough evidence.

The CloudNova example

The clearest example is CloudNova.

Historical memory:

CloudNova agreed to waive the onboarding fee and promised a 15% renewal discount if the account exceeds 100 seats.

Current quote:

CloudNova renewal quote for 130 seats is ₹460000 annually. The quote does not include any renewal discount.

For the renewal discount:

Historical condition: account exceeds 100 seats.

Current evidence: 130 seats.

Because 130 is greater than 100, the condition is satisfied.

The current quote explicitly states that the renewal discount is not included.

That supports a potential conflict.

But the onboarding-fee promise is different.

The new document is a renewal quote.

If it simply says nothing about onboarding, that silence is not enough to establish a contradiction.

PactTrace explicitly instructs the model:

text
A quote's silence about a discount or fee alone is
insufficient evidence.

Therefore, that comparison remains insufficient evidence.

This is exactly the kind of distinction that prevents memory-enabled systems from becoming overconfident.

Conditions have to be proven

A remembered promise may only apply under certain circumstances.

PactTrace tracks four condition states:

  • met
  • not_met
  • unknown
  • not_applicable

And then validates consistency.

ts
if (
["unknown", "not_met"].includes(conflict.conditionStatus) &&
(
conflict.status !== "insufficient_evidence" ||
conflict.severity !== "low"
)
) {
throw new Error("Inconsistent comparison status");
}

Unknown conditions cannot magically become high-confidence conflicts.

That rule is simple, but it gives the system a much stronger evidence model.

Why PactTrace says “potential conflict”

Another good design choice is language.

PactTrace does not claim:

The vendor violated the contract.

The system calls the result a potential conflict.

That leaves the final interpretation to a human.

A remembered promise might lack legal context.

There may be contract amendments.

The quote might not be final.

The historical conversation might require additional interpretation.

AI should surface evidence, not pretend to be the final legal authority.

Memory plus validation is more useful than memory alone

Hindsight gives PactTrace the ability to preserve and recall history across interactions.

But the broader lesson is that persistent memory needs an evidence discipline around it.

Useful agent systems need both:

Memory

What happened before?

Grounding

Can I prove that my current conclusion follows from the supplied sources?

The second question becomes even more important as the first capability becomes more powerful.

For anyone exploring the distinction between stored context and real agent memory, Vectorize's article on agent memory versus RAG is worth reading.

The takeaway I would keep is this:

An AI agent should never become more confident just because it remembers more.

It should become more useful because it can connect the right history to the right evidence.

Project: https://github.com/v9vek26/pacttrace

Live application: https://pacttrace.vercel.app/

Top comments (0)