Software teams rarely encounter every failure for the first time.
A login failure today can look different from a random logout reported last month. A session disappearing after inactivity may appear to be another unrelated bug. But sometimes several incidents are symptoms of the same underlying problem.
The useful information is already there. It is just spread across historical incidents.
I built Hindsight around this idea: instead of treating every incident as an isolated ticket, use historical incidents as engineering memory and look for recurring failure patterns.
The central concept is Failure DNA: a recurring root-cause pattern discovered from multiple historical incidents.
From Incident History to Engineering Memory
The first part of Hindsight is its Incident Memory.
For the prototype, historical incidents are stored in a SQLite database. Each incident contains information such as:
*Incident ID
Title
Description
Root cause
Affected module
Previous resolution
*
For example, the prototype contains incidents such as:

*INC001
Login failure
Root cause: Token expiration
Module: authentication
Resolution: Automatic token refresh
*
Another incident contains:
INC002
Random logout
Root cause: Token expiration
Module: authentication
Resolution: Token refresh mechanism
And another:

*INC003
Session lost
Root cause: Token expiration
Module: authentication
Resolution: Session handling and token refresh
*
Individually, these look like separate incidents.
When viewed together, a pattern becomes visible.
All three involve the same root cause and the same affected module.
That recurring relationship is what Hindsight calls Failure DNA.
Building the Failure Pattern Library
Hindsight doesn't simply display the incidents. It analyzes them to identify recurring root causes.
The pattern-generation logic counts how often each root cause appears in the historical data.
A simplified part of the implementation is:
**root_causes = []
for incident in incidents:
root_causes.append(incident[4])
cause_counts = Counter(root_causes)
**
The system then groups incidents belonging to the same root cause.
For each pattern, Hindsight also looks at the affected modules and determines how consistently the pattern appears in the same area.
Token expiration
│
├── INC001 → authentication
├── INC002 → authentication
└── INC003 → authentication
The result is a stronger pattern than simply finding one similar ticket.
Hindsight therefore records several pieces of evidence for a pattern:
Frequency
Common affected module
Module consistency
Previous resolution
Evidence strength
Pattern status
A pattern appearing multiple times is classified as Recurring, while a pattern with only one historical incident is treated as Emerging.
Failure DNA
This is the part I wanted to make different from simple incident search.
Suppose a developer enters:
Users are being logged out after keeping the application open for a long time.
A traditional search could return an old login failure because the words are similar.
Hindsight goes one step further.
It asks:
What recurring failure pattern is represented by the matching incidents?
In our example, the historical evidence points toward:
Failure DNA
Pattern:
Token expiration
Module:
authentication
Historical incidents:
3
Status:
Recurring
This distinction matters because the goal isn't just to find an old ticket.
The goal is to understand the repeated failure behind several tickets.
Matching a New Issue Against History
Once the historical memory exists, Hindsight can analyze a new issue.
The prototype uses TF-IDF to convert historical incident text into numerical vectors.
For each historical incident, Hindsight combines information from the description, root cause, affected module and resolution.
The new issue is then added to the same vector space.
The similarity between the new issue and historical incidents is calculated using cosine similarity.
The core operation is:
similarities = cosine_similarity(
vectors[-1],
vectors[:-1]
)[0]
Each historical incident receives a similarity score.
The results are sorted from the strongest match to the weakest match.
The prototype currently uses a similarity threshold of 0.15 to determine which historical incidents become candidate matches.
This gives Hindsight two different types of evidence:
New issue
│
├── Text similarity
│
└── Historical Failure DNA
The first tells us whether the new issue resembles previous incidents.
The second tells us whether those incidents form a meaningful recurring pattern.
Evidence Confidence
A similarity score alone isn't enough.
If only one weakly similar incident exists, the evidence should not be treated the same way as several historical incidents pointing toward the same failure.
Hindsight therefore calculates an evidence confidence using both:
Number of matching incidents
Average text similarity
The prototype weights historical evidence more heavily than similarity:
confidence = (
evidence_score * 0.6 +
similarity_score * 0.4
)
This produces a confidence percentage and classifies it as:
High
Medium
Low
For example, if several historical incidents support the same pattern and the new issue is reasonably similar to them, Hindsight can show a stronger confidence level.
Hindsight Risk Score
I also wanted the interface to communicate the difference between textual similarity and historical evidence.
So Hindsight combines the text confidence with the Failure DNA evidence strength.
The prototype uses:
risk_score = (
text_confidence * 0.5 +
dna_strength * 0.5
)
The resulting score is classified into three levels:
75–100 → High
50–74 → Medium
0–49 → Low
The dashboard can then communicate the result as a Hindsight Alert rather than simply displaying a list of similar tickets.
For example:
HINDSIGHT ALERT — FAILURE DNA MATCHED
The important part is the evidence shown underneath the alert.
Hindsight explains the historical pattern, affected module, previous resolution and related incidents instead of presenting the alert as an unexplained prediction.
Looking Beyond Incidents: Code Change Analyzer
The second workflow in the prototype looks at code changes.
A developer can enter something such as:
**Changed file:
auth/session.py
and:
Updated token refresh logic and session handling.
**
Hindsight analyzes the change description and detects the likely module.
The prototype currently uses module aliases to map terms such as:
auth
login
session
token
to:
authentication
It then checks historical incidents associated with that module.
If the authentication module has a recurring Failure DNA involving token expiration, Hindsight produces a pre-deployment alert.
The idea is simple:

Developer changes code
↓
Changed module detected
↓
Historical incidents checked
↓
Failure DNA found
↓
Developer receives warning
This changes the point at which historical knowledge can be useful.
Instead of only consulting incident history after something breaks, the history can also be considered while modifying an area that has previously experienced failures.
A Complete Example
Consider the three authentication incidents in the prototype:
**
INC001 → Login failure
Token expiration
authentication
INC002 → Random logout
Token expiration
authentication
INC003 → Session lost
Token expiration
authentication
**
Hindsight groups these incidents around the recurring root cause:
Failure DNA:
Token expiration
Now a new issue arrives:
Users are being logged out after keeping the application open for a long time.
The text-matching layer compares the issue with the historical incidents.
The matching incidents are then grouped by root cause.
Because multiple matching incidents point toward Token expiration, Hindsight can produce:
Failure DNA:
Token expiration
Affected module:
authentication
Historical evidence:
3 incidents
Status:
Recurring
The previous resolution is also surfaced:
Implemented automatic token refresh.
This doesn't automatically prove that the new issue has the same root cause. Instead, it gives the developer a structured piece of historical evidence to investigate.
That distinction is important.
Hindsight is an evidence and memory layer, not a replacement for engineering diagnosis.
What I Learned While Building It
- *Similarity Alone Isn't Enough * A search system can find similar words without understanding whether those incidents represent the same recurring engineering problem.
Grouping historical incidents around their root causes gives the similarity result more context.
**
- Repetition Itself Is Useful Evidence ** One incident can be noise.
Several incidents pointing toward the same root cause and module provide a stronger signal.
That is why Hindsight tracks recurrence and module consistency rather than only displaying similarity scores.
**
- Explainability Matters ** An alert is much more useful when a developer can answer:
Why am I seeing this?
Hindsight therefore shows the historical incidents, common root cause, affected module and previous resolution behind an alert.
**
- Simple Models Can Make a Useful Prototype ** The current prototype uses SQLite, Python, Streamlit and Scikit-learn.
There is no need for a large infrastructure stack to demonstrate the basic idea of turning incident history into structured engineering memory.
**
- Historical Memory Should Become Part of Development ** The most interesting direction for Hindsight is connecting incident memory with the development workflow itself.
If a developer changes an area that has historically been associated with failures, that historical context can be surfaced before deployment rather than rediscovered after another incident.
Where Hindsight Can Go Next
The current prototype is deliberately focused on the core workflow:
Historical Incidents
↓
Incident Memory
↓
Failure Pattern Detection
↓
Failure DNA
↓
New Issue / Code Change
↓
Historical Matching
↓
Hindsight Alert
There are several ways this could be extended.
The similarity layer could become more semantic instead of relying primarily on lexical TF-IDF similarity. Incident ingestion could be connected to real issue trackers. Code analysis could become more sophisticated than keyword-based module detection. Hindsight could also integrate with CI/CD workflows so that historical failure evidence becomes available during development and deployment.
But the underlying idea would remain the same:
Don't let previous failures become forgotten tickets. Turn them into engineering memory.
That is what I built Hindsight to explore.

Top comments (1)
Weighting historical evidence over raw similarity (0.6/0.4) is smart. It stops one loosely worded new ticket from getting matched to a shaky old one. Have you tried swapping TF-IDF for embeddings yet? I'd guess that catches paraphrased incidents, like "token expiration" vs "auth session invalidated early," that plain word matching would miss.