The Incident-Response Agent I Built Needed to Remember What Happened Last Time
The hardest part of incident response is not recognizing the second SSH brute-force attack. It is remembering what we learned from the first one before we repeat the same investigation.
I built Sentinel Memory around that problem: an incident-response agent for SOC analysts that combines current incident analysis with persistent, outcome-oriented memory. The core loop is deliberately simple: Detect → Analyze → Recall → Recommend → Resolve → Retain → Improve.
What I built
Sentinel Memory is a React and TypeScript SOC interface backed by a FastAPI service. The backend uses SQLAlchemy's async stack for incident state, while the agent layer separates incident analysis, response planning, and memory operations. The memory integration sits behind an adapter boundary so the rest of the application does not need to know how Hindsight is deployed.
The frontend exposes the workflow through dashboard, incident detail, memory, and learning views. The backend exposes incident APIs and coordinates the agent. The interesting part is not any individual component. It is the path an incident takes through them.
An incoming incident is analyzed for tactics, severity, and indicators. The resulting context is then used to recall relevant previous experiences from a Hindsight memory bank. Those experiences are passed to the response planner. The analyst reviews the resulting recommendation and controls the response. Once the incident is resolved, the post-mortem becomes new memory.
That last step is what closes the loop.
For the memory layer, I used Hindsight persistent agent memory, with the Hindsight documentation as the reference for the underlying memory model and API. The broader distinction between storing information and giving an agent durable experience is also well described in agent memory architecture.
The design decision that mattered: remember outcomes, not transcripts
My first instinct was to treat previous incidents like documents: store them and retrieve the closest one later.
That is useful, but it misses the most important part of incident response.
A previous incident is valuable because we know what happened after the investigation.
For Sentinel Memory, I therefore made the memory representation an experience capsule. It contains the incident identity and type, source and target, indicators, analysis summary, root cause, response actions, outcome, and lessons learned.
The code makes that decision explicit:
backend/app/hindsight/service.py
capsule = (
f"INCIDENT: {incident_data.get('id', 'UNKNOWN')} - {incident_data.get('title', '')}\n"
f"TYPE: {incident_data.get('incident_type', 'general')} | SEVERITY: {incident_data.get('severity', 'UNKNOWN')}\n"
f"SOURCE: {incident_data.get('source', 'N/A')} -> TARGET: {incident_data.get('target', 'N/A')}\n"
f"INDICATORS: {', '.join(incident_data.get('indicators', []))}\n"
f"ANALYSIS SUMMARY: {analysis.get('summary', incident_data.get('description', ''))}\n"
f"ROOT CAUSE: {postmortem.get('root_cause', 'Under investigation')}\n"
f"RESPONSE ACTIONS TAKEN: {actions_str}\n"
f"OUTCOME: {resolution.get('outcome', 'Resolved')}\n"
f"LESSONS LEARNED: {postmortem.get('lessons_learned', 'Standard runbook applied')}\n"
)
That looks almost too straightforward, but it captures an important architectural constraint: memory should preserve the relationship between context → cause → action → outcome.
"Block the attacker IP" is an action.
"Blocking the attacker IP contained the attack, but the post-mortem identified configuration drift as the actual root cause" is an experience.
The second is what I want the next investigation to see.
Keeping Hindsight behind an adapter
I did not want Hindsight-specific HTTP calls scattered throughout the incident services.
Instead, I defined a small memory contract:
backend/app/hindsight/base.py
class BaseHindsightAdapter(ABC):
@abstractmethod
async def retain(
self,
bank_id: str,
content: str,
metadata: Optional[Dict[str, Any]] = None
) -> Dict[str, Any]:
pass
@abstractmethod
async def recall(
self,
bank_id: str,
query: str,
limit: int = 5
) -> List[Dict[str, Any]]:
pass
The application talks to that interface. The concrete client adapter handles communication with Hindsight.
This gave me two useful properties.
First, the incident workflow remains independent of the memory vendor's transport details.
Second, I can test the incident workflow without requiring an external memory service for every local run.
The service selects the client adapter when Hindsight is configured for client mode and otherwise uses the local deterministic adapter:
backend/app/hindsight/service.py
if adapter is not None:
self.adapter = adapter
elif settings.HINDSIGHT_MODE == "client":
self.adapter = HindsightClientAdapter(
base_url=settings.HINDSIGHT_BASE_URL,
api_key=settings.HINDSIGHT_API_KEY
)
else:
self.adapter = MockHindsightAdapter()
For a production deployment, the important boundary is the client adapter. The application should not need to change its incident logic just because the memory service moves from local infrastructure to a hosted or self-managed Hindsight deployment.
The interesting part happens on incident two
The first incident creates memory.
The second incident tells me whether that memory is actually useful.
The primary example in the system is an SSH brute-force sequence.
INC-2026-001 targets the perimeter bastion. The investigation identifies the attack and the analyst blocks the attacker. The post-mortem then records the less obvious finding: an automated package upgrade had enabled password authentication in sshd_config.
That finding is retained with the response and outcome.
Later, INC-2026-002 arrives from a different source against another server. The source IP is different, so a naive lookup based only on the attacker identifier would not help.
Instead, the orchestrator builds a query from the current incident's title, description, indicators, incident type, and target:
backend/app/agents/incident_agent/orchestrator.py
query = f"{title} {desc} {indicators}"
recalled_experiences = await self.hindsight_service.recall_similar_incidents(
incident_context=(
f"type: {incident_data.get('incident_type')} "
f"target: {incident_data.get('target')}"
),
query=query,
limit=3
)
The important detail is that the current incident becomes the retrieval context. The agent is not simply asking, "Have I seen this exact IP before?"
It is asking whether the current combination of incident characteristics resembles something that was previously investigated.
The orchestrator also removes the current incident from the recalled results before planning the response:
filtered_memories = [
m for m in recalled_experiences
if m.source_incident_id != inc_id
]
recommendation = await self.recommender.plan_response(
incident_data=incident_data,
recalled_experiences=filtered_memories
)
That separation keeps historical evidence distinct from the incident currently being processed.
Turning memory into an actual recommendation
Retrieval by itself is not the goal.
The useful behavior is what happens after retrieval.
The response planner receives the current incident together with the recalled experiences. In the SSH example, that means the previous root cause and lessons can influence what the analyst is asked to investigate next.
For INC-2026-002, the recalled INC-2026-001 experience provides a reason to inspect the target server's SSH configuration instead of treating the new attack as an isolated brute-force event.
The system's recommendation can therefore surface a critical audit around sshd_config and a prevention runbook derived from the earlier post-mortem.
That is a materially different interaction from a generic response such as "block the source IP."
The previous incident does not automatically execute the same response. It changes the evidence available to the analyst.
That distinction is important in security systems because similar symptoms do not guarantee identical causes.
Why I did not make remediation autonomous
There is a temptation to build an incident agent that sees an alert and immediately starts changing infrastructure.
I intentionally did not make that the core model.
Sentinel Memory is recommendation-first. The analyst reviews, approves, and executes response actions.
That makes memory useful without treating it as authority.
A historical experience can be wrong for the current environment. A runbook can be outdated. A similar attack can have a different root cause. The system therefore needs to make historical precedent visible while leaving the final operational decision with the analyst.
For me, that also makes the memory system easier to reason about: memory supplies evidence; the response workflow decides what to do with it.
Retention is part of the incident lifecycle
Another design choice was to make learning an explicit service operation instead of an incidental side effect.
The learning service delegates the completed incident and post-mortem to the Hindsight service:
backend/app/services/learning_service.py
async def retain_experience(
self,
incident_data: Dict[str, Any],
postmortem_data: Optional[Dict[str, Any]] = None
) -> Dict[str, Any]:
return await self.hindsight_service.retain_incident_experience(
incident_data=incident_data,
postmortem_data=postmortem_data
)
That means the lifecycle is explicit:
Analyze the incident.
Resolve it.
Write the post-mortem.
Retain the experience.
Recall it when a later incident provides relevant context.
The system's memory is therefore not a separate knowledge-management project. It is a consequence of completing an incident properly.
The architecture is intentionally boring
The repository is split into a React frontend, a FastAPI backend, an incident agent layer, services, database models, and a dedicated Hindsight integration.
The core intelligence path is roughly:
SOC Analyst
|
v
React / TypeScript UI
|
v
FastAPI
|
+--> Incident Analyzer
|
+--> Hindsight
| |
| +--> Recall previous experiences
|
+--> Response Planner
|
v
Analyst Recommendation
|
v
Resolution / Post-Mortem
|
v
Hindsight Retain
I prefer this architecture to a large collection of autonomous agents because there is one clear question at every stage.
The analyzer answers: what is happening?
Memory answers: what have we learned before?
The planner answers: given both, what should the analyst consider?
The analyst answers: what are we actually going to do?
That division is easier to test and easier to explain during an incident.
What I learned
- Memory quality starts at the retention boundary
The quality of future recall is largely determined by what gets written today.
If I retain raw logs without the investigation outcome, I preserve noise. If I retain a structured relationship between root cause, response, outcome, and lesson, future retrieval has something operationally useful to work with.
- The second incident is the real memory test
A successful retain() call proves storage works.
It does not prove memory is useful.
The meaningful test is a later incident where the recalled experience changes what the system recommends or what the analyst investigates.
That is why the INC-2026-001 → INC-2026-002 flow matters more to me than simply showing a memory record in a dashboard.
- Provenance matters as much as retrieval
If an analyst sees a recommendation without knowing why it appeared, the system becomes another opaque AI layer.
Showing the historical incident and the relevant root cause gives the analyst something concrete to verify.
- Memory should inform, not override
Past experience is evidence, not truth.
That is especially important in security. A previous successful response can be highly relevant while still being wrong for a different asset, environment, or root cause.
- The memory boundary is an architectural boundary
Keeping Hindsight behind BaseHindsightAdapter turned out to be useful beyond testing.
It means incident logic, memory transport, deployment mode, and persistence concerns can evolve independently. That is exactly the kind of boundary I want around infrastructure that will eventually become critical to the system.
Where this goes next
The long-term value of an incident-response agent is not that it can explain one alert slightly faster.
It is that the system can accumulate operational experience without turning every post-mortem into another document someone has to remember to search.
A resolved incident becomes a structured experience. A later incident becomes a retrieval opportunity. The response planner gets both current evidence and historical precedent. The analyst remains in control. Then the new outcome becomes the next piece of memory.
That gives me a much more useful definition of an incident-response agent:
It should not just know what is happening. It should remember what happened the last time, what we did about it, what actually worked, and what we learned afterward.
The loop is simple enough to explain:
Detect → Analyze → Recall → Recommend → Resolve → Retain → Improve.
The engineering challenge is making every arrow real
Top comments (0)