Faultline: An AI Incident-Memory Agent for Smarter Incident Response
Introduction
When a software service fails, engineers usually investigate the issue, fix it, and create an incident report.
But there is a common problem: important lessons and promised follow-up tasks can be forgotten.
Faultline is an AI incident-memory agent designed to solve this problem by connecting new alerts with information from previous incidents.
The Problem
A company may have many software services. When one service fails, the incident report may contain a recommendation such as:
"Check the other services for the same problem."
If this task is never completed, another service can experience the same problem later.
The information already exists, but it is difficult for engineers to manually connect information across many old incident reports.
Our Solution: Faultline
Faultline remembers previous incidents and uses that information when a new alert arrives.
It answers five important questions:
- Have we seen this problem before?
- Which previous incidents had the same underlying cause?
- What fixed the problem previously?
- Which other services are still exposed to the same problem?
- Did someone previously promise to fix it but never complete the task?
This gives engineers useful historical context while investigating a new incident.
How It Works
When a new alert arrives, Faultline searches its incident memory and connects related information.
For example, an IAM authorization failure can be connected to previous incidents that had the same underlying cause.
Faultline can then show:
- Previous related incidents
- Common underlying cause
- Previous runbook or solution
- Other services with the same exposure
- Unfinished remediation tasks
Instead of starting the investigation from zero, engineers receive relevant context from previous incidents.
Uses
Faultline can be used to:
- Connect new alerts with previous incidents.
- Identify repeated underlying problems.
- Retrieve previous fixes and runbooks.
- Find other services exposed to the same problem.
- Track unfinished remediation tasks.
- Provide historical context during incident investigation.
- Help engineering teams learn from previous outages.
Benefits
- Reduces the need to manually read many old incident reports.
- Helps engineers identify repeated problems faster.
- Makes forgotten remediation tasks visible.
- Helps identify services that may be affected before they fail.
- Preserves organizational knowledge.
- Goes beyond simple keyword-based search by connecting information across incidents.
Example
Suppose two previous incidents were caused by the same IAM configuration problem.
When a similar alert appears, Faultline can connect the new alert with those incidents and show the previous solution.
It can also identify other services that currently have the same problematic configuration.
This gives the engineer a broader view of the problem instead of only showing the current error.
Memory ON vs Memory OFF
We also created a simple demonstration to show the importance of memory.
With memory enabled, Faultline can retrieve:
- Previous incidents
- Previous runbooks
- Unfinished remediation tasks
- Related services
With memory disabled, the same AI does not have access to those historical connections.
This demonstrates that the main value comes from giving the AI persistent incident memory.
Evaluation
Our demonstration contains 19 synthetic incidents.
We created a manually written answer key and tested whether Faultline could identify related incidents.
It correctly identified 5 out of 6 tested relationships, giving an 83% result.
We also show the case that was missed instead of hiding it.
Why This Is Different From Search
A normal search system can find documents containing similar words.
Faultline is designed to connect information across multiple incidents and services.
For example, two incidents may describe different symptoms but have the same underlying configuration problem.
The goal is not only to find similar documents, but to connect the information and provide useful incident context.
Limitations
The current demonstration uses prepared data.
Faultline does not yet directly check live Terraform or Kubernetes infrastructure.
The current system reads a prepared file describing service configurations.
Connecting Faultline to live infrastructure is an important next step toward a production-ready system.
Outcome
Faultline demonstrates how AI memory can help engineering teams preserve knowledge from previous incidents and use that knowledge when new problems occur.
The project achieved an 83% result in the demonstrated incident-matching evaluation and showed a clear difference between memory-enabled and memory-disabled responses.
Future Scope
Future improvements could include:
- Connecting to live infrastructure.
- Automatically detecting configuration changes.
- Tracking remediation tasks until completion.
- Integrating with incident-management systems.
- Continuously updating service exposure information.
- Expanding the incident memory with real production data.
Conclusion
Every incident contains information that can help prevent future incidents.
The challenge is remembering and connecting those lessons.
Faultline is an attempt to give software incidents a memory — helping engineers connect new alerts with previous incidents, previous fixes, exposed services, and unfinished remediation.
Top comments (0)