Why I Separated Hindsight Evidence From Model Reasoning
By Bhupathi Mahesh Varun Kumar · Hindsight / AI / memory
When I inspect an AI troubleshooting answer, I want to know which part came from the machine report, which part came from historical experience, and which part is a model's inference. A paragraph that simply says “this happened before” is not enough. It needs a source that the next stage can carry forward. ForgeMind's Hindsight integration makes that problem concrete: retrieve candidate memories, shape them into a bounded evidence set, and pass their identifiers alongside the current incident.
The problem: retrieval results are not yet usable evidence
A memory system can return several records that overlap, vary in relevance, or describe different machines. Passing every result directly into a model increases prompt size and makes it harder to inspect why a record was selected. Passing only a summary loses identity: later output cannot point back to the record that supported a claim.
ForgeMind builds the recall query from the current machine, title, description, and symptoms. In m1-runtime/m1/hindsight/recall.js, the request is explicit:
const query = [
`Machine: ${incident.machineId || "unknown"}`,
`Title: ${incident.title || ""}`,
`Description: ${incident.description || ""}`,
`Symptoms: ${(incident.symptoms || []).join(", ")}`,
"Find historically relevant incidents, root causes, actions, and outcomes."
].join("\n");
return hindsight.recall(bankId, query);
That function delegates retrieval to Hindsight; it does not decide whether a result is a diagnosis. The distinction matters. A similar phrase in a memory can be useful context without proving that the same cause applies to the current machine.
Normalize, preserve identity, then bound
The orchestration layer in m1-runtime/m1/services/incidentMemoryService.js maps Hindsight's results into a shape the analysis step can inspect. It keeps the returned ID, final score, summary text, tags, and a source label. It also sorts by relevance, groups duplicate summaries, keeps the highest-scoring item within each duplicate group, then returns at most three records.
normalized.push({
memoryId: item.id,
relevance:
typeof item.scores?.final === "number"
? item.scores.final
: 0,
type: inferEvidenceType(summary),
summary,
source: "hindsight",
tags: Array.isArray(item.tags) ? item.tags : []
});
The final slice(0, MAX_HISTORICAL_EVIDENCE) is a small but meaningful boundary. It gives the later prompt a bounded set and keeps each candidate's identity attached to its content. ForgeMind then places evidence in a named prompt section and tells the model to use only supplied memory IDs for historical claims. The returned recommendation schema includes evidenceMemoryIds, making provenance visible in the response shape.
I think of this as an evidence ledger, not a confidence oracle. The ID says which record was supplied; it does not certify that the memory is correct or that the model used it correctly.
Before and after
Before this mapping, imagine passing raw recall text into analysis. The model might describe a repair as precedent, while the application has no reliable way to show which record it meant. Similar summaries can also appear more than once and consume prompt space.
After mapping, the analysis input contains a short list of structured records. For a hypothetical later report of reduced coolant flow, a recalled outcome could travel with its Hindsight memory ID, score, tags, and summary. The model can cite that ID in a cause or recommendation. If recall returns no records, ForgeMind passes an empty evidence list and the prompt says that no historical Hindsight evidence was retrieved. This is a description of the code path, not a claim that a live query returned a particular record or improved diagnosis accuracy.
The screen above is deliberately not presented as a Hindsight recall result. Factory Memory currently searches local JSON incidents and shows sample totals; Hindsight recall occurs in M1's analysis path. That mismatch is worth making visible because a UI label can otherwise suggest the wrong source of truth.
What I would keep and what I would change
The useful design choice is carrying identifiers all the way into model output. The code also keeps reflection separate from the raw evidence list, so a consumer can inspect both rather than treating one generated paragraph as the whole history. Hindsight supplies retain, recall, and reflect operations through its client; ForgeMind chooses the query, record shape, cap, and response contract.
There are limits. The evidence type is assigned by keyword checks such as “caused,” “replaced,” and “lesson.” That label is display metadata, not verified causality. The code requests only the top three after deduplication, so a relevant fourth record is excluded. And while the prompt instructs the model to reference supplied memory IDs, the application does not validate generated IDs against the actual evidence set. A schema can require an array of strings without proving those strings are valid references.
My engineering takeaway is to treat retrieval as the beginning of provenance, not the end. Preserve IDs and source metadata, make the selected set small enough to inspect, and validate references after generation if downstream users rely on them. Hindsight makes memory operations available; the application still owns the evidence contract.

Top comments (0)