DEV Community

Ashwini Ravirala
Ashwini Ravirala

Posted on

What Changes When an AI Code Reviewer Remembers? Building ReviewMind with Hindsight

What Changes When an AI Code Reviewer Remembers? Building ReviewMind with Hindsight

Most AI code reviewers are good at understanding the code in front of them.

The harder problem is remembering what a particular team has already learned.

A team might decide to avoid print() in production code, prefer structured logging, use a particular error-handling pattern, or follow specific database-access rules. Those decisions often appear gradually through pull requests, review comments, and debugging sessions.

I wanted to explore what would happen if an AI code reviewer could actually carry some of that context from one review into the next.

That idea became ReviewMind, a memory-driven code review agent built around Hindsight.

The core workflow is simple:

Recall → Review → Feedback → Retain

The interesting part is not simply generating another AI code review. It is making previously learned team context available at the moment a new review is being generated.

The problem isn't only finding bugs

Consider a small function:

def archive_user(user):
print("Archiving user:", user)
return user

A reviewer might reasonably flag the print() statement and recommend structured logging instead.

That's useful.

But imagine that the same convention appears again in another file a week later.

If the reviewer has no access to the team's previous decision, the developer may receive essentially the same recommendation again without any continuity between the two reviews.

There is an important difference between knowing a general programming practice and remembering a team's specific convention.

Teams accumulate decisions over time.

The problem is that those decisions can remain scattered across pull requests, conversations, documentation, and individual developer knowledge.

I wanted ReviewMind to make that accumulated context part of the review process itself.

The idea: put memory in the review loop

ReviewMind uses a Next.js and TypeScript frontend, a Python FastAPI backend, Groq for LLM-based review, and Hindsight for memory.

The backend coordinates the workflow.

At a high level, every review follows four stages:

Code
↓
RECALL
↓
AI Review
↓
Developer Feedback
↓
RETAIN
↓
Future Review

  1. Recall

Before the LLM reviews the code, ReviewMind builds a query using information such as the programming language, framework, and coding-convention context.

It then asks Hindsight for relevant memories from the team's memory bank.

  1. Review

The submitted code and recalled memories are passed to the LLM.

The model can therefore consider both:

the code being reviewed
relevant knowledge from previous interactions

  1. Feedback

The developer can respond to individual findings.

The available decisions are:

Accepted
Rejected
Not relevant

This matters because an AI reviewer shouldn't assume that every suggestion is correct for every codebase.

  1. Retain

Meaningful feedback can be turned into a memory that may be useful during later reviews.

That creates the learning loop:

RECALL → REVIEW → FEEDBACK → RETAIN
↑ |
└─────────────┘
Why Hindsight is more than a storage layer

One of the design questions I had was:

Why not just store review feedback in a database?

A database can certainly store comments.

But storing feedback is only part of the problem.

For a new review, I need to find the information that is actually relevant to the current code.

That's where Hindsight's recall capability becomes important.

ReviewMind uses a team identifier as the Hindsight bank_id, creating a boundary around the memories associated with that team.

The memory integration is kept inside its own service rather than spreading Hindsight-specific logic throughout the API.

The recall operation looks like this:

results = await self.client.arecall(
bank_id=bank_id,
query=query,
budget="mid",
max_tokens=4096,
)

The service converts the returned results into a compact structure containing the memory text, type, and optional score.

Those memories are then passed to the LLM service.

This separation keeps the responsibilities fairly clear:

FastAPI
│
├── Hindsight Service
│ └── Recall / Retain
│
└── LLM Service
└── Generate Review

For more information about the underlying memory approach, see the Hindsight documentation and Vectorize's explanation of agent memory.

Giving the reviewer explicit team context

I didn't want the model to blur together its general programming knowledge and the team's retrieved knowledge.

So the LLM prompt explicitly separates the two.

The memory section is constructed roughly like this:

memory_text = ""

if memory_context:
memory_text = "\nTEAM MEMORY:\n"

for i, memory in enumerate(memory_context, 1):
    memory_text += f"{i}. {memory.get('text', '')}\n"
Enter fullscreen mode Exit fullscreen mode

else:
memory_text = "\nTEAM MEMORY:\nNo team memories available for this review.\n"

The reviewer receives the submitted code along with the retrieved team memory.

The model is also instructed to refer to a memory only when that memory was actually supplied.

ReviewMind then checks the memory_used values returned by the model against the memories that were actually provided.

That distinction is important.

A reviewer shouldn't claim that it remembered a team rule if that rule was never actually retrieved.

Feedback is where the learning loop starts

A review agent shouldn't blindly assume that its suggestions are correct.

Developers need a way to disagree.

That's why ReviewMind treats feedback as part of the memory workflow rather than just a UI state.

For example:

Finding:
Avoid print() in production code.

Developer:
Accepted

The system can retain a convention such as:

Team convention:
Avoid print statements in production code.
Use structured logging.

The retention call is kept behind the Hindsight service:

self.client.retain(
bank_id=bank_id,
content=content,
context=context,
metadata=metadata or {},
retain_async=False,
)

The retained item can also include metadata such as:

review ID
issue ID
decision
programming language
framework
team ID

This gives the memory some context about where the learning came from.

A concrete before-and-after example

Let's say the team accepts a recommendation about structured logging.

The first review contains:

def archive_user(user):
print("Archiving user:", user)
return user

The developer accepts the finding.

ReviewMind retains the team convention.

Later, another developer submits:

def delete_user(user):
print("Deleting user:", user)
return user

This is a different function.

The important part is that Hindsight may now retrieve the earlier convention:

Team convention:
Avoid print statements in production code.
Use structured logging.

The LLM receives both the new function and that retrieved context.

The resulting finding can therefore be connected to an explicit team decision rather than being presented only as generic programming advice.

The exact memories returned depend on what has previously been retained and what Hindsight retrieves for the query.

Memory doesn't guarantee that every relevant rule will be found, and it doesn't guarantee that every generated finding will be correct.

It gives the reviewer additional context.

That's the distinction I wanted to explore.

What I learned while building it

  1. Retrieval needs to happen before the decision

This sounds obvious after building it, but it changes the architecture.

If memory is retrieved after the LLM has already generated its review, that memory cannot influence the decision.

So the sequence needs to be:

Code
↓
Recall relevant context
↓
Give context to LLM
↓
Generate review

Memory has to be close to the decision point.

  1. Feedback needs meaning

A simple thumbs-up or thumbs-down doesn't tell the system very much.

A decision connected to a specific issue and explanation is much more useful.

For example:

Accepted:
"Yes, our production services use structured logging."

Rejected:
"We intentionally allow print() in CLI scripts."

The second piece of feedback can be just as useful as the first because it describes an exception to a general recommendation.

  1. Memory claims should be traceable

An AI reviewer shouldn't casually say:

"Your team previously decided..."

unless the relevant team memory was actually supplied to it.

That's why ReviewMind keeps the recalled memories in the review response and validates memory references.

The goal is to make the boundary between model knowledge and retrieved team knowledge visible.

  1. Graceful degradation matters

Memory systems can fail.

If Hindsight recall fails, ReviewMind can continue the review with an empty memory list.

That means the core code review doesn't necessarily have to stop just because the memory layer isn't available.

But this creates another requirement: the interface and logs need to make memory availability clear.

Otherwise, a developer could mistake a normal review for a memory-informed review.

Current limitations

ReviewMind is still an MVP, and there are several limitations I wouldn't hide.

The current review store is in memory.

That means review records used for feedback are lost when the backend restarts.

The project also doesn't automatically review GitHub pull requests yet.

The recall query is primarily constructed from programming language, framework, and convention-related information rather than performing a sophisticated semantic analysis of the submitted code before constructing the query.

These limitations matter because memory-based systems introduce their own problems.

For a real codebase, I'd want to think carefully about:

outdated memories
incorrect memories
conflicting team conventions
repository-level versus team-level memory
access control
sensitive code
secrets accidentally included in submissions
memory correction and deletion

A system that remembers the wrong thing can be just as problematic as one that forgets everything.

What's next?

There are several directions I'd explore next.

Persistent review storage

Reviews and feedback should survive backend restarts.

GitHub pull-request integration

Instead of manually submitting code, ReviewMind could work directly with pull requests and provide memory-aware review comments.

Repository-level memory

Different repositories can have different conventions.

A future version could combine team-level knowledge with repository-specific context.

Better memory controls

Developers should be able to understand, correct, and manage important team conventions.

Better observability

When memory retrieval fails or produces no relevant results, the system should make that visible.

These changes would move the project closer to being useful in a real development workflow.

From isolated answers to accumulated context

The central idea behind ReviewMind is fairly modest.

I don't think memory replaces code review expertise, testing, or human judgment.

Instead, it gives an AI reviewer another source of context: what the team has already learned.

Hindsight provides the retain-and-recall layer.

FastAPI coordinates the workflow.

Groq generates the structured review.

And the developer remains in the loop by deciding which suggestions are useful and which aren't.

The interesting shift is from:

Review → Forget → Review → Forget

to:

Review
↓
Feedback
↓
Remember
↓
Recall
↓
Review with context

That's what I wanted to explore with ReviewMind.

Not whether an AI can review code.

But what changes when the reviewer can remember the people and decisions behind the code.

Project






Top comments (0)