DEV Community

Shravya Kulkarni
Shravya Kulkarni

Posted on

When an AI Audit Assistant Can Remember What Happened Before

Most AI assistants are good at answering questions based on the information you give them at that moment.

While working on my latest project, I started wondering about something slightly different:

What if an audit assistant could remember what happened in previous audits?

In internal auditing, the current finding is not always the whole story. An auditor might want to know whether a similar control issue appeared a few months ago, what remediation was applied at the time, and whether the same problem appeared again.

That is what led me to build AuditMind, an AI-powered internal IT audit assistant that uses Hindsight as its long-term memory layer.

The idea was not just to store old audit reports somewhere and search them later. I wanted the historical information to actually become part of the reasoning process.

[https://github.com/vectorize-io/hindsight]
[https://github.com/kulkarnishravya132/AuditMind]


AuditMind uses historical audit memory to provide evidence-backed context for a new audit question.


The problem I was trying to solve

Let's say an auditor finds that an employee's access was not removed on time after they left the organization.

An audit assistant could suggest something like:

Disable the access immediately and improve the offboarding process.

That's a reasonable recommendation.

But I was more interested in another question:

Has this happened before?

If the organization has already had the same issue in previous audits, then the auditor probably needs more than a generic remediation suggestion.

They might want to know:

  • Which previous findings were related?
  • What remediation was applied?
  • Did the issue happen again?
  • What evidence supports that?
  • Is there enough history to actually call it a recurring issue?

That is where persistent memory becomes useful.


How I built AuditMind

The overall workflow is fairly simple:

Audit findings → Hindsight memory → historical analysis → auditor-facing result

I built the application using Python and Streamlit, with Hindsight Cloud providing the persistent memory layer.

For the initial version, I used a structured JSON dataset containing historical audit findings. Each finding has information such as the finding ID, control theme, evidence, remediation, and other audit context.

The important part for me was how this information would be used.

Hindsight provides memory operations around concepts such as retain, recall, and reflect. I used these as part of the AuditMind workflow rather than treating Hindsight as just another storage layer.

Historical findings are retained in a Hindsight memory bank, and later the application can use that accumulated history during the analysis of a new audit question.


The first problem: getting the right findings

One of the first things I had to think about was retrieval.

My initial thought was to let the system search the complete memory and decide which findings were relevant.

It sounds convenient, but I wasn't comfortable with that for an audit use case.

For example, if the auditor is asking about Employee Offboarding, I don't want a Network Monitoring finding getting into the analysis just because the two documents happen to contain some semantically similar words.

So I added a control-theme matching step before the main analysis.

When an auditor asks about a recognized control theme, AuditMind identifies that theme and retrieves the findings belonging to that exact theme.

Those findings are then used as the relevant evidence for the Hindsight analysis.

This created a separation that I found useful:

The application decides which evidence is relevant, while Hindsight handles the memory-based reasoning over that evidence.

That made the system easier to understand and debug.


Making sure one finding doesn't become a "pattern"

Another thing I wanted to avoid was overconfidence.

Suppose the system finds only one historical finding for a particular control theme.

I don't think that's enough to say that the organization has a recurring issue.

So AuditMind has an explicit "insufficient history" state.

if len(matches) < 2:
    return (
        "insufficient history"
    )
Enter fullscreen mode Exit fullscreen mode

I also added rules to the Hindsight reflection prompt telling it to cite finding IDs, avoid inventing dates or remediation details, avoid mixing unrelated control themes, and distinguish documented facts from interpretation.

This was important because an AI system can produce a very convincing explanation even when there isn't enough evidence behind it.

For an audit assistant, that distinction matters.


Seeing the difference when memory is available

This is probably the part of the project I found most interesting.

Without historical memory, an assistant can look at a new finding and provide recommendations based on the information in front of it.

With historical memory, the same question can be connected to previous findings.

For example, the AuditMind dataset contains multiple Employee Offboarding findings:

AUD-001, AUD-004, AUD-008, and AUD-014.

When the application identifies Employee Offboarding as the control theme, these historical findings can become part of the analysis.

So instead of only getting:

"Here is what you should do."

the auditor can get something closer to:

"Here is what happened before, here is the evidence supporting it, and here is what deserves attention now."

That difference is basically the reason I wanted to add memory to the project in the first place.


Using Hindsight for the reasoning

Another thing I wanted to understand while building this was how Hindsight would actually fit into an audit workflow.

I use its reflect operation for the analysis step:

The application provides the auditor's question, the requested control theme, and the historical findings that were selected as relevant.

Hindsight then performs the memory-backed reasoning used to generate the analysis shown to the auditor.

I also added a small feedback loop.

After an analysis, an auditor can mark the result as useful or needing review. That feedback is retained in Hindsight as well.

So the memory layer isn't limited to the original audit findings. It can also retain feedback about how the analysis was received.


Building a pre-audit risk view

After getting the historical analysis working, I started thinking about another use case.

Instead of waiting for an auditor to ask about one specific control, what if the system could provide some historical context before an audit starts?

That led to the idea of a pre-audit risk brief.

The idea is to look across multiple control themes and surface historical issues that may deserve attention.

Normally, an auditor would have to go through previous reports and try to identify those patterns manually.

With the accumulated memory, AuditMind can generate a broader historical view instead.

The risk brief uses Hindsight's reflection capability with a larger context window than the individual analysis.

So memory becomes useful not only when someone has already found a problem, but also during audit preparation.


The debugging problem I wasn't expecting

Interestingly, the biggest engineering issue I ran into wasn't with the AI reasoning itself.

It was async resource management.

The Hindsight Python client uses asynchronous HTTP resources. My initial implementation combined a shared client, thread execution, and different event-loop contexts.

Eventually, I started getting errors involving asyncio event loops and futures attached to different loops.

The solution was to simplify the client lifecycle.

Each Hindsight operation now creates the client, performs the operation, and closes the client within the same event loop:

async def _execute():
    client = Hindsight(
        base_url="https://api.hindsight.vectorize.io",
        api_key=os.getenv("HINDSIGHT_API_KEY")
    )

    try:
        return await operation(client)
    finally:
        await client.aclose()

return asyncio.run(_execute())
Enter fullscreen mode Exit fullscreen mode

That fixed the issue.

It was also a good reminder that adding an AI memory component doesn't remove the normal software-engineering problems.

You still have to think about things like resource lifecycles, async code, data flow, and how different parts of the application interact.


What I learned from building it

The biggest thing I learned is that having memory isn't enough.

The system also needs rules around how that memory is used.

Three things ended up being particularly important for AuditMind:

1. Relevant evidence needs to be isolated

I use exact control-theme matching so unrelated findings don't get pulled into the analysis.

2. There needs to be an insufficient-history case

One historical finding shouldn't automatically become a recurring pattern.

3. The evidence should be visible

Using finding IDs makes it possible to see where the historical reasoning came from instead of treating the AI's answer as an unexplained conclusion.


What is still missing?

AuditMind is currently a demonstration of the workflow rather than a production audit platform.

The current dataset is mainly designed to demonstrate how the system works.

A production implementation would need things such as:

  • stronger data ingestion
  • authentication and authorization
  • more audit-source integrations
  • proper evaluation datasets
  • observability
  • more extensive testing

So there is still a lot that could be built on top of this.


Where I want to take the idea

The core experiment worked.

I wanted to see whether an audit assistant could use previous audit findings instead of treating every new question as completely independent.

The answer was yes, but with an important condition:

the memory has to be constrained by relevant evidence.

That changes the kind of questions the assistant can help with.

Instead of only asking:

"What should I do about this finding?"

an auditor can start asking:

"Have we seen this before?"

"What happened last time?"

"Did the issue come back?"

"What evidence supports that?"

That's what I wanted AuditMind to explore: not just an AI assistant that can answer audit questions, but one that can use the history of previous audits when looking at a new one.

Top comments (0)