When the AI gets the answer wrong
When a production service starts falling apart, getting a diagnosis is only half the problem. The harder
question is usually: what should we do next?
We already have logs, metrics, deployments, runbooks, tickets, dashboards and incident histories. The
frustrating part is that those things tend to tell us what happened. They don't always preserve the
reasoning behind the decision: what an engineer tried, why they rejected another option, what failed,
and what eventually worked.
That was the idea behind FixLoop. I wanted to build an incident decision system that treats previous
incidents as experience rather than just documents.
The rule I kept coming back to
It is surprisingly easy to make an incident assistant look intelligent. Give an LLM some logs and
symptoms and ask it for a diagnosis. You will probably get a reasonable-looking answer.
But a reasonable-looking answer isn't necessarily a useful operational decision. It might completely
ignore what the team already learned from previous incidents.
Grok reasons. Python calculates. Hindsight remembers.
Where Hindsight fits
Hindsight is not just a history panel in FixLoop. It is the memory layer that connects one operational
decision to the next.
Retain saves the complete experience after an incident decision is verified.
Recall retrieves comparable incidents, actions, decisions and outcomes.
Reflect reasons across those memories to understand what the team learned and what matters now.
I don't really care about remembering an incident just for the sake of remembering it. I care about
remembering the consequences of decisions.
An incident comes in
Imagine a payment API with a P1 incident, a 91% error rate, timeouts, 500 errors, and a recent
deployment.
Grok turns that information into a structured understanding of the incident: a summary, likely causes,
important signals and candidate actions. I keep the output structured because the next stages need
predictable fields instead of a paragraph of generated text.
Then comes the part I actually cared about
Suppose the system has three possible actions:
Restart — historical success: 2 / 5
Rollback — historical success: 4 / 6
Restore configuration — historical success: 4 / 4
Those numbers are calculated from incident history. They are not numbers the LLM gets to make up.
This is what I call Decision Rehearsal. Before recommending an action, FixLoop looks at what
happened when those actions were tried in comparable situations. It is basically a historical
counterfactual: if we choose this action now, what happened the last time we chose something similar?
The part I really wanted to keep: human overrides
Real operational decisions aren't always going to match the AI recommendation. And I don't think an
override should simply disappear once the incident is closed.
Imagine FixLoop recommends a rollback. An engineer looks at the deployment and says the
deployment itself looks healthy and configuration drift is a better hypothesis. So they override the
recommendation and choose configuration repair.
That disagreement is valuable data.
FixLoop records the recommendation, the human decision, the reason for the override, and the
eventual outcome. That entire experience is retained in Hindsight.
Now a similar incident can recall that human correction alongside the original recommendation and
final outcome. The system isn't just remembering the incident. It is remembering how the team
corrected its own decision process.
The second incident is the real test
I think the first incident is actually the least interesting part of the system.
The first run can show that the application can analyze an incident, find historical cases and make a
recommendation. Fine.
The second similar incident is where the memory actually has to prove itself.
Incident 1 → AI recommendation → human override → verified outcome → Hindsight retain
Incident 2 → similar signals → Hindsight recall → recommendation informed by the new
experience
If the second incident is handled exactly as if the first one never happened, then the memory layer isn't
doing much. The useful part is when the newly retained experience becomes evidence for the next
decision.
Keeping the architecture deliberately boring
I deliberately kept the architecture narrow. I didn't want a giant collection of autonomous agents just
because the project involved an LLM.
The frontend is Next.js, React and Tailwind. FastAPI handles orchestration and deterministic decision
logic. Grok handles incident interpretation and explanation. Hindsight handles long-term memory.
The result is intentionally boring in a good way:
Frontend → FastAPI → Grok + Hindsight → Decision Engine → Recommendation → Human
Decision → Hindsight
That makes the system easier to debug because each component has a clear job.
What I learned building around memory
- Memory is more useful when it stores consequences. An incident summary plus the decision, override and outcome is much more useful than the summary alone.
- Keep deterministic evidence around the LLM. I want the model to explain evidence, not manufacture it.
- Human overrides are valuable data. The reason an engineer disagrees with an AI recommendation can become part of the future context.
- Test the second incident. A memory system is much easier to evaluate when new experience visibly affects what happens later.
- Narrow systems are easier to trust. I'd rather build one decision loop I can explain end to end than hide the important logic behind layers of abstraction. Why I think this is more interesting than another incident chatbot I don't think the interesting part of FixLoop is that an AI can summarize an outage. We already know language models can do that. The interesting part is turning an operational decision into durable experience and carrying that experience into the next incident. That's the role Hindsight plays in the architecture. It isn't a side database and it isn't just a history pane. It is the layer that lets the system remember what the team learned. “What did we learn last time, and should that change what we do now?” Links Hindsight GitHub: github.com/vectorize-io/hindsight Hindsight documentation: hindsight.vectorize.io
Top comments (0)