DEV Community

Cover image for Designing the AI Investigation Layer for WorkMemory AI
Rithika Gudla
Rithika Gudla

Posted on

Designing the AI Investigation Layer for WorkMemory AI

An AI assistant can generate an answer in seconds, but that answer is much more useful when it has access to the team's own previous experiences.

While working on WorkMemory AI, I focused on the AI and investigation side of the system: how an LLM can use relevant historical incident knowledge to help an engineer investigate a new problem.

The idea was not to build another chatbot that answers technical questions from general knowledge. We wanted the investigation process to be grounded in what the engineering team had actually experienced before.

The Problem With Starting Every Investigation From Zero

Software incidents are rarely completely new.

A Payment API may fail after a deployment. An authentication service may start timing out. A database connection may become unstable.

The exact cause can be different every time, but engineering teams often encounter similar patterns.

The problem is that the useful details from previous incidents can be difficult to find.

An engineer might need to search:

  • old incident reports
  • tickets
  • documentation
  • team conversations
  • troubleshooting notes
  • deployment records

Even after finding an old incident, the engineer still has to determine whether it is relevant to the current problem.

This is where we designed the AI investigation layer of WorkMemory AI.

Figure 1: WorkMemory AI interface for working with engineering incidents.

What WorkMemory AI Is Designed to Do

WorkMemory AI follows a simple workflow:

New Incident
↓
Backend
↓
Recall Relevant Previous Experience
↓
AI Investigation
↓
Recommended Context
↓
Engineer Verifies and Resolves

The important part is the separation between memory retrieval and AI reasoning.

The memory layer is responsible for finding relevant previous experiences.

The LLM is responsible for interpreting those experiences in the context of the current incident.

That gives each component a clear responsibility.

Why Memory Comes Before the LLM

One design principle became clear while working on the project:

«The LLM should not be expected to know an organization's private incident history automatically.»

A general-purpose model may know that deployment configuration can cause an API failure.

But it does not automatically know:

  • what happened in our previous deployment,
  • what our team tried,
  • which solution worked,
  • what side effect occurred,
  • or what warning the previous engineer recorded.

That information needs to come from the application's memory layer.

The intended workflow therefore looks like this:

Current Incident
↓
Hindsight Memory
↓
Relevant Past Incidents
↓
Current Incident + Past Context
↓
LLM
↓
Investigation Response

Hindsight is intended to provide the historical context, while the LLM interprets that context.

Designing the Investigation Endpoint

The backend provides a dedicated endpoint for investigation:

POST /api/incidents/investigate

The frontend sends the current incident to the backend.

The backend then passes the incident to the investigation service.

The route is intentionally simple:

router.post("/incidents/investigate", async (req, res) => {
try {
const incident = req.body;

if (!incident || Object.keys(incident).length === 0) {
  return res.status(400).json({
    error: "Incident data is required"
  });
}

const result = await investigateWithMemory(incident);

return res.status(200).json(result);
Enter fullscreen mode Exit fullscreen mode

} catch (error) {
console.error("Error investigating incident:", error);

return res.status(500).json({
  error: "Failed to investigate incident",
  details: error.message
});
Enter fullscreen mode Exit fullscreen mode

}
});

The important thing here is that the route does not contain the AI reasoning logic itself.

The investigation service becomes the boundary between the API and the future memory/LLM pipeline.

What the LLM Should Receive

For an investigation, sending only the current error message is not enough.

The intended AI input contains two types of information:

Current incident

For example:

Problem:
Payment API is returning 500 errors.

Service:
Payment API

Environment:
Production

Description:
The errors started after the latest deployment.

Historical context

The memory layer can provide relevant information from previous incidents.

For example:

Previous incident:
Payment API became unstable after a deployment.

Action:
Deployment environment variables were checked and corrected.

Result:
The service recovered after redeployment.

Consequence:
The configuration mismatch was added to the team's troubleshooting notes.

The LLM can then reason over both pieces of information.

From Retrieved Memory to Investigation

The intended AI pipeline is:

Current Incident
+
Relevant Historical Incidents
↓
LLM
↓
Structured Investigation

Rather than asking the model:

«“What should I do?”»

we wanted the prompt to provide a more structured task.

For example, the investigation response can be designed around:

  1. What happened previously?
  2. What action was taken?
  3. What result did that action produce?
  4. Was there a later consequence?
  5. What should the engineer check next?

This structure makes the output easier for an engineer to read and verify.

Figure 3: Investigation results presented to the engineer.

A Concrete Example

Suppose an engineer reports:

Payment API is returning HTTP 500 errors
after a production deployment.

The memory system finds a previous incident involving the same service.

The previous incident says that an incorrect environment configuration caused a similar failure.

The AI investigation layer can then produce context such as:

Previous experience:
A similar Payment API failure occurred after deployment.

Previous action:
The deployment environment configuration was checked.

Result:
The configuration was corrected and the service was redeployed.

Recommended next check:
Compare the current deployment environment variables
with the last known working configuration.

This does not mean that the previous incident proves the current root cause.

That distinction is important.

The engineer still needs to check the current system.

The AI is providing historical context and a starting point, not replacing engineering judgment.

Keeping Retrieval and Reasoning Separate

One of the most important architectural ideas in WorkMemory AI is:

Hindsight → Retrieve relevant experience

LLM → Interpret that experience

Backend → Connect the workflow

Engineer → Make the final decision

This separation makes the system easier to reason about.

If the wrong historical incident is retrieved, that is primarily a memory/retrieval problem.

If the retrieved information is correct but the explanation is poor, that is an AI reasoning problem.

If the API fails to connect the two layers, that is a backend integration problem.

Separating these responsibilities gives each part of the system a clearer debugging boundary.

The Current Development Stage

During development, we first created the backend API and a memory-service adapter.

The adapter allowed us to test the complete request flow before connecting the full production memory and LLM implementations.

For example:

async function investigateWithMemory(incident) {
console.log("\n[HINDSIGHT] Investigation requested:");
console.log(JSON.stringify(incident, null, 2));

return {
message: "Investigation endpoint is working",
incident,
similarIncidents: [],
analysis: {
status: "pending",
message: "Connect Hindsight Recall and the LLM layer next."
}
};
}

This was useful because it allowed the frontend and backend workflow to be developed independently.

The next stage is connecting real Hindsight recall and the LLM analysis layer to this interface.

What I Learned

  1. Context is as important as generation

An LLM can generate an impressive answer, but the answer becomes more useful when it is grounded in relevant organizational knowledge.

  1. Retrieval and reasoning are different problems

Finding a relevant previous incident and explaining what it means for the current incident are two separate tasks.

  1. Structured prompts produce clearer results

Instead of asking an AI for a generic answer, defining the information we want—previous action, result, consequence, warning, and next step—makes the intended output much clearer.

  1. AI should not replace the engineer

A previous incident may look similar without having the same root cause.

The engineer still needs to verify the AI's suggestions against the current system.

  1. Build the interfaces before the integrations

Creating the backend investigation endpoint and service boundary first allowed the rest of the application to progress while the memory and LLM integrations were still being developed.

Future AI Workflow

The complete intended workflow is:

Engineer submits incident
↓
Backend validates incident
↓
Hindsight retains the experience
↓
New similar incident appears
↓
Hindsight recalls relevant experiences
↓
Backend combines current + historical context
↓
LLM analyzes the context
↓
Structured investigation response
↓
Engineer verifies the recommendation
↓
New learning becomes future memory

This creates a continuous learning loop.

Every resolved incident can potentially make the system more useful for the next investigation.

Conclusion

The AI part of WorkMemory AI is not simply about asking an LLM to answer technical questions.

The more interesting idea is connecting an LLM to the engineering team's own history.

Hindsight is intended to provide the memory of previous experiences, while the LLM provides reasoning over that context.

The backend connects these components and the engineer remains responsible for validating the result.

For me, the main lesson was simple:

«An AI assistant becomes much more useful when it can learn from the experiences of the team using it.»

Project

GitHub: https://github.com/gudlarithika60-tech/workmemory-ai

Demo: https://photos.app.goo.gl/h1fkWCBvU2xqh2dD9

Hindsight: https://github.com/vectorize-io/hindsight

Hindsight Documentation: https://hindsight.vectorize.io/

Top comments (0)