My README said the agent used Hindsight's Reflect operation to reason over a machine's failure history. My code called Recall, stuffed the results into a prompt, and asked an LLM to do the reasoning itself. Nobody caught it until I went looking for a bug in something else entirely.
What this thing actually does
The idea is simple and, I think, underused: every industrial machine should have its own memory. A technician fixes a pump today; six months from now, a different technician sees the same vibration pattern on the same pump and has no idea it happened before. The service report is in a binder somewhere, or in the first technician's head, and neither is searchable at 2am when the line is down.
So I built a small system — a JSON file of synthetic maintenance records, a Flask UI, and Hindsight doing the memory work — that lets a technician describe what they're seeing and get back an answer grounded in what actually happened last time, including the fixes that didn't work. Not "here's a generic troubleshooting guide," but "R. Bhatt tried re-greasing this exact bearing in February, it didn't help, here's what did."
The architecture is intentionally small:
Technician input (symptoms / sensor readings)
│
▼
demo_agent.py
│
┌───────┴───────┐
│ │
RECALL REFLECT
│ │
└───────┬───────┘
▼
Hindsight Memory
│
▼
Groq LLM
Two modes, two different jobs: a side-by-side "with vs. without memory" comparison uses Recall — grab relevant fragments, hand them to an LLM, let it reason. A pre-repair check uses Reflect, which is supposed to do the reasoning itself, grounded directly in memory. That distinction is the whole story of this post.
The moment I found out my "Reflect" wasn't reflecting
I was doing an honest self-review before treating this as done, and I asked myself a simple question: does the code actually call reflect() anywhere? I grepped. It didn't.
Here's what check_before_repair looked like:
def check_before_repair(symptoms, lang="English"):
...
texts, err = recall(symptoms, machine)
...
a = llm("You are the Chief Reliability Engineer for a plant. " + SCHEMA, user, json_mode=True)
Recall, then a generic chat completion doing the actual judgment call. It worked fine in a demo. It was also not what I'd told anyone it was. The difference matters more than it sounds: Recall's job is retrieval — hand back relevant fragments and get out of the way. Reflect's job is synthesis — decide whether this matches a known failure pattern, weigh conflicting evidence, and answer the actual judgment question. I was doing the second thing by hand, with a hand-rolled prompt, instead of asking the memory system to do it.
The fix was to call Hindsight's reflect() directly, with structured output:
def reflect(query, context=None, tags=None, response_schema=None, budget="mid"):
r, e = call_with_retry(lambda: hs().reflect(
bank_id=BANK_ID, query=query, context=context, tags=tags,
tags_match="any", budget=budget, max_tokens=1600,
response_schema=response_schema, include_facts=True))
if e: return None, friendly_error(e)
return r, None
response_schema gives me a JSON Schema and Hindsight returns structured_output matching it — no more asking an LLM to "please respond only with valid JSON" and hoping. include_facts=True is the part that actually changed the architecture: it returns a based_on field naming the exact memories, mental models, and directives Hindsight used to produce the answer. My "why does the agent think this" panel in the UI now renders that — not a list of things I recalled and hoped the model used, but a list of things the model reports it actually used.
The gotcha that cost me an afternoon: IDs live in context, not text
Once Reflect was wired in, my citation-checking code broke. I'd built a guard that only lets the model cite a record ID if that ID appears somewhere in the retrieved text — a cheap, effective hallucination check. It started rejecting every citation.
Turned out Hindsight's fact extraction paraphrases content into atomic statements, and the [MAINT-2025-001]-style tag I'd embedded in my retained text doesn't survive that paraphrasing. But the context string I'd also passed at retain time — "maintenance record MAINT-2025-001 | machine MCH-017 | tags: ..." — comes back untouched on every fact:
id_blob = texts + [f.context for f in facts if f.context]
conf = confidence(symptoms, id_blob)
srcs = sources_from(id_blob)
Lesson, in one line: if you need something to survive Hindsight's extraction pass verbatim, put it in context, not in the content you're asking it to summarize.
A bug that only exists because of async
The other real bug: a cached Hindsight client, reused across requests, would occasionally throw Timeout context manager should be used inside a task — a classic symptom of an aiohttp session created in one asyncio event loop being reused in another. Flask's dev server spins up a new thread per request; the sync wrapper spins up a new event loop per call. Cache the client across both and eventually they collide.
The fix was cheap once I understood it: stop caching, build a fresh client per call, and retry once on anything that looks transient:
def hs():
from hindsight_client import Hindsight
return Hindsight(base_url=os.environ["HINDSIGHT_BASE_URL"],
api_key=os.environ["HINDSIGHT_API_KEY"])
def call_with_retry(fn, attempts=2, delay=1.2):
for i in range(attempts):
try:
return fn(), None
except Exception as e:
if i < attempts - 1 and _is_transient(e):
time.sleep(delay); continue
return None, e
I only found this because I tested concurrent requests on purpose. It would have shown up live, in front of the people I most wanted it not to.
What memory that grows actually looks like
The part that convinced me this wasn't just retrieval-with-extra-steps was the closed loop. Log a failed fix, and the very next diagnosis for that machine changes — not because I re-ran a script, but because the memory itself grew:
Before logging feedback:
"Extending lubrication intervals from monthly to quarterly failed twice, once due to a manual policy change and once due to a CMMS software reset."
After logging "re-greased the bearing housing — didn't work":
"Re-greasing the bearing housing has been attempted and resulted in failure, as noted in log LIVE-001. Discontinue all attempts to re-grease; replace the drive-end bearing with SKF 6309 directly."
Same machine, same underlying question, a genuinely different answer because a technician's outcome became a memory a few seconds earlier. I also gave each machine a "digital twin" — a Hindsight mental model whose source_query asks it to synthesize that machine's full history into a living profile, with trigger={"refresh_after_consolidation": True} so it updates itself as new incidents come in, rather than a summary I'd have to regenerate by hand. And I moved a couple of standing rules — lockout-tagout before touching rotating equipment, never invent a technician name or a date — out of my prompt text and into Hindsight directives, so they're enforced by the memory system on every request instead of by convention in a string I control.
What I'd tell someone starting this tomorrow
- Grep your own claims. If your docs say you use an operation, verify the code path actually calls it. I didn't, for weeks.
-
Decide what needs to survive verbatim, and put it somewhere that isn't summarized.
contextis notcontent; treat them differently on purpose. - Test concurrency before your users do. The scariest bug I found only existed under simultaneous requests, and I only found it by deliberately generating them.
- A system's memory should get smarter without you touching code. If the only way your agent "learns" is by editing a seed file, it's not really learning — it's re-authoring.
- Retrieval and reasoning are different jobs. Recall for "what's relevant." Reflect for "what does this mean." Conflating them is an easy trap, and I fell into it first.
None of this needed a bigger model. It needed the memory system to actually be doing the job I'd told everyone it was doing.
Top comments (0)