

# I Let Hindsight Find the Incidents. My Code Decides Which Ones Matter.
When I first wired memory into an incident workflow, I expected the nearest retrieved story to be the most useful one. On a machine floor, "nearest" is only a starting point: the same symptom can have a different cause when the tool, temperature, or coolant has changed.
I built RECALL-X around that gap. It uses Hindsight to retain and retrieve operational experience, then applies an explicit, domain-specific ranking step before asking a reasoning agent to explain what the history does, and does not, support.
The incident is more than a prompt
RECALL-X helps operators investigate recurring CNC problems such as rough surface finish, spindle vibration, and overheating. An incident includes machine and material identifiers, tool state, spindle speed, feed rate, temperature, pressure, vibration, and the operator's observation. The system finds relevant past incidents and their outcomes, compares conditions, and returns an evidence-backed recommendation for operator review.
The backend is a small FastAPI service. POST /incident accepts a typed Pydantic model, stores the incident, retrieves history, runs the "What Changed?" comparison, and passes the evidence to the reasoning layer. POST /outcome records what the operator tried and whether it worked. The Next.js interface shows the recommendation alongside the evidence and parameter differences, and operators can then record the outcome and browse the growing history.
That loop matters. If I saved only a symptom and a suggested action, the next investigation could not distinguish advice from a verified result. So RECALL-X stores both the incident narrative and the eventual outcome as operational experience. Hindsight provides the memory layer with retain, recall, and reflect operations; the Hindsight GitHub repository and documentation describe them in detail.
Hindsight finds stories; the ranker weighs cases
Shop-floor notes are not tidy database queries. "Surface waviness after the morning batch" and "rough finish after contour pass" may describe related events without sharing a phrase. A narrative memory system can retrieve across that vocabulary and connect facts, entities, and time in ways a handful of SQL filters cannot.
But semantic relevance is not operational equivalence. A memory can describe the right symptom and still be a poor precedent if it came from another machine, a different alloy, or a tool in a different condition. I do not want the order of a semantic search to silently decide which intervention an operator sees.
So I split the problem into two questions:
- Which memories might be relevant to this incident?
- Which structured incidents are close enough to be useful evidence?
Hindsight answers the first. The application resolves those memories to structured incident and outcome records, then scores them with an explicit rubric. That second step is intentionally ordinary code: engineers can inspect the weights, change them as experience warrants, and explain why a case made the final list.
Here is the Hindsight lookup in the memory manager:
recalled = _run_in_isolated_thread(
self.client.recall,
bank_id=self.bank_id,
query=semantic_query,
)
The query starts with the current machine and symptom, then adds material, tool, RPM, feed rate, temperature, and the operator's observation when present. In production, I match the returned memories to canonical incident records before scoring. That join is important: Hindsight recovers narrative context, while the structured record gives the comparison code reliable fields and a stable incident ID.
The scorer also has a rule that sounds mundane but changes the quality of the answer: it ranks only incidents with recorded outcomes. An unresolved event can serve as context, but it should never be presented as evidence that an action succeeded or failed.
outcome_data = hist.get("outcome") or {}
if not outcome_data.get("result"):
continue
if hist.get("machine", "").strip().lower() == incident.machine.strip().lower():
score += 30.0
The full rubric awards up to 100 points: exact machine match (30), symptom match (35), material (15), tool (10), and parameter proximity (10). A machine-category match earns less than an exact machine match. RPM and feed rate count as close within ten percent of the historical value; temperature counts as close within four degrees. These are policy choices, not universal truths, which is exactly why I keep them visible in code rather than buried in a prompt.
"What changed?" is a separate calculation
Ranking identifies plausible precedents; it does not tell us whether their conditions still apply. So I compare the current incident with a successful historical baseline and calculate parameter differences separately. For temperature, even a small absolute change can matter when the relative change is large:
delta = curr_temp - prev_temp
pct = round(((curr_temp - prev_temp) / prev_temp) * 100, 1) if prev_temp != 0 else 0.0
if abs(delta) >= 1.0:
severity = "important" if abs(delta) >= 5.0 or abs(pct) >= 18.0 else "moderate"
The comparison engine applies similar rules to RPM, feed rate, vibration, pressure, coolant, tooling, and material. Its output is a list of typed differences, each with a severity and an explanation. Arithmetic and thresholds stay deterministic, and the reasoning agent receives the differences as evidence instead of inferring them from prose.
The agent uses Groq when configured, validates the response against a Pydantic model, and falls back to a deterministic reasoning path when a model call is unavailable or invalid. Either path must separate historical evidence from inference and keep the operator in control. The API response explicitly sets human_review_required, and the system never applies a CNC parameter change itself.
A concrete case: same finish problem, changed conditions
Consider a rough-finish incident on CNC-07. The current record reports tool T-19 and a temperature of 34°C. The successful historical baseline used tool T-17 at roughly 27°C with a feed rate of 450 mm/min. In the stored history, reducing feed to 410 mm/min improved surface roughness from Ra 3.2 µm to 0.8 µm. That is useful evidence, but not a command to repeat the feed reduction blindly.
The comparison step flags the tool change and the seven-degree temperature increase. The reasoning layer can then say what the evidence supports: inspect the current tool and coolant condition before relying on the old feed-rate fix. It can also surface a failed alternative from the history: an earlier attempt to raise spindle speed from 3,200 to 3,800 RPM excited fixture resonance and worsened the finish to Ra 4.8 µm.
This is more useful than "try the successful fix" because it preserves the conditions around that success and keeps the failed intervention visible. The operator can inspect the cited incidents, decide whether to act, and later submit the actual action and result. The next analysis then has another recorded outcome to retrieve.
Here Hindsight's reflect operation adds a different kind of value. Recall answers a case-level question; reflection can synthesize across accumulated memories, such as what tends to precede spindle vibration on a particular machine. That difference between lookup and synthesis is part of why I chose an agent-memory system over a static pile of documents. Vectorize's guide to agent memory offers useful background.
What I learned
Separate retrieval from decision policy. Semantic recall is good at finding relevant language and relationships. It does not replace the application's responsibility to decide which cases are comparable enough to inform a recommendation.
Store outcomes as first-class memory. An intervention without a result is only an attempted action, not a lesson. Capturing success, failure, root cause, and notes turns later recall into evidence instead of anecdote.
Keep critical arithmetic out of free-form generation. Temperature deltas, severity thresholds, and outcome filters are simple enough to calculate deterministically. Let the reasoning layer explain those facts, not invent them.
Make every recommendation inspectable. A useful answer shows the historical cases, the changed conditions, and the reason for its confidence. In operational software, "why this advice?" is part of the feature.
I began with the idea that memory would make the system better at finding old incidents. The more important decision was what to do after finding them. Hindsight lets RECALL-X remember and synthesize experience; explicit scoring and condition comparison decide when that experience is relevant to the machine in front of the operator.
Top comments (0)