DEV Community

BOPPUDI SANTHOSH (So143)
BOPPUDI SANTHOSH (So143)

Posted on

I Built a Code Review Agent That Remembers What It Found

I wanted CodeMind to work differently: remember useful lessons from previous reviews and use them when reviewing code again.

That idea became CodeMind, an AI code review agent that combines static code analysis, LLM reasoning, evidence verification, deterministic scoring, and Hindsight persistent memory.

The interesting part isn't simply asking an LLM whether code is good or bad. The interesting part is creating a review loop where useful experience from one review can become context for the next one.

The Problem With Stateless Code Review

A typical LLM-based code review starts with the current repository and the current prompt.

It can identify a SQL injection, recommend parameterized queries, and explain why the vulnerability matters. But when another review happens later, that previous experience is gone.

The same problem appears with project-specific conventions.

A team might repeatedly use a particular error-handling pattern, prefer a specific coding style, or have already fixed a recurring security problem. A stateless reviewer sees the next piece of code without knowing that history.

This creates several problems:

  • Previous reviews are forgotten.
  • Similar mistakes are treated as completely new problems.
  • Project-specific conventions are difficult to preserve.
  • Developers can receive repetitive recommendations.
  • The reviewer has no persistent engineering experience.

I wanted CodeMind to address that problem.

Instead of treating every review as an isolated event, CodeMind gives the reviewer a persistent memory layer using Hindsight.

What Is CodeMind?

CodeMind is an AI code review agent designed to analyze source code and produce an engineering-focused review.

A developer can submit a public GitHub repository or upload a ZIP file. CodeMind then processes the source code through several stages.

The high-level pipeline looks like this:

Code Repository
      |
      v
Static AST Analysis
      |
      v
Hindsight Recall
      |
      v
Groq LLM Review
      |
      v
Evidence Verification
      |
      v
Deterministic Scoring
      |
      v
Hindsight Retain
      |
      v
Interactive Review Report




Enter fullscreen mode Exit fullscreen mode

This architecture creates a continuous feedback loop.

Before each review, CodeMind recalls relevant experience from previous reviews through Hindsight. The current code is then analyzed using static analysis and LLM reasoning.

After the review, useful findings and lessons are retained back into Hindsight.

This means that each review can contribute useful knowledge to future reviews.

Why Hindsight?

The main question I had while building CodeMind was:

What if a code reviewer could remember what it had already learned?

Hindsight became the memory layer for that idea.

It is not simply used as a place to store review results. CodeMind uses Hindsight to recall relevant experience before a review and retain useful learning after a review.

That gives the system a memory cycle:

RECALL
  |
  v
Use Previous Experience
  |
  v
Analyze Current Code
  |
  v
Learn From Findings
  |
  v
RETAIN
  |
  v
Future Review


Enter fullscreen mode Exit fullscreen mode

How Hindsight Recall Works

Before CodeMind sends the current code for LLM analysis, it retrieves relevant memories from previous reviews.

The recall process uses information from the current review to find useful previous experience.

A simplified version of the request looks like this:

payload = {
    "query": query,
    "budget": budget
}

res = await client.post(
    recall_url,
    json=payload,
    headers=self.headers
)
Enter fullscreen mode Exit fullscreen mode

The returned memories are then prepared as context for the current review.

This gives the LLM two important sources of information:

  1. The code it is currently reviewing.
  2. Relevant experience from previous reviews.

The previous memory does not replace the analysis of the current code. It provides additional context that can help the reviewer recognize recurring patterns and make more informed recommendations.

Static Analysis Still Matters

I didn't want the LLM to be the only analysis layer in CodeMind.

CodeMind performs static analysis before the LLM review. The system uses AST-based analysis and security checks to identify potentially dangerous or problematic patterns.

The review process combines deterministic analysis with LLM reasoning:

Static Analysis
      +
Current Source Code
      +
Hindsight Memory
      |
      v
LLM Reasoning

Enter fullscreen mode Exit fullscreen mode

Static analysis provides concrete signals.

Hindsight provides relevant historical context.

The LLM uses both to reason about the current code and explain the findings.

The Review Does Not End With the LLM

An AI model can produce a convincing explanation for a problem that does not actually exist in the source code.

That is why CodeMind includes an evidence verification stage.

Candidate findings are checked against the source code before they become part of the final review.

The final findings can contain information such as:

  • Severity
  • Category
  • File
  • Line number
  • Source snippet
  • Problem description
  • Recommended fix
  • Evidence

This helps make the review more grounded in the actual code instead of relying only on generated explanations.

Hindsight Retain: Turning Findings Into Learning

After the review is completed, CodeMind can turn useful findings into reusable learning.

For example, a security finding can become a learning pattern like this:

CodeMind Learned Pattern

Category:
Security

Issue:
Unsanitized user input in SQL query

Best Practice:
Use parameterized query bindings

Enter fullscreen mode Exit fullscreen mode

The important part is that CodeMind does not have to remember every line of every previous review.

Instead, it can retain reusable engineering knowledge such as:

  • recurring security problems
  • project-specific conventions
  • useful fixes
  • patterns that previously caused problems
  • lessons that can help with future reviews

This makes the memory layer more useful than simply storing old review reports.

Evidence Verification

One of the goals of CodeMind is to make review findings grounded in the actual source code.

An LLM can produce a convincing explanation even when the evidence in the repository is weak or missing. I wanted CodeMind to reduce that problem by checking findings against the code that was actually analyzed.

Each finding can be associated with evidence such as:

  • File
  • Line number
  • Source snippet
  • Problem description
  • Recommended fix
  • Evidence

This helps make the review more actionable for a developer.

Instead of receiving only a general statement such as:

"This code may contain a security vulnerability."

the reviewer can connect the finding to the actual source location and understand why it was reported.

The goal is simple:

A review finding should be connected to evidence in the code.

Deterministic Scoring

Another part of CodeMind is deterministic scoring.

I did not want the final score to depend entirely on an LLM saying that a repository "looks good."

The review findings can be converted into a consistent score using deterministic rules.

A simplified model looks like this:

Review Findings
      |
      v
Severity Classification
      |
      v
Deterministic Score
      |
      v
Final Review Result

Enter fullscreen mode Exit fullscreen mode

Interactive Review Report

The final result is presented as an interactive review report rather than a raw LLM response.

The report brings together the different stages of the analysis into one place.

A developer can see:

  • Overall review score
  • Severity of findings
  • Security issues
  • Code quality issues
  • File and line references
  • Source snippets
  • Evidence supporting each finding
  • Recommended fixes
  • Relevant previous experience recalled from memory

The goal is to make the result useful for an engineer who needs to understand not only what is wrong, but also where it is wrong and what to do next.

The review therefore follows a simple path:



Source Code
     |
     v
Analysis
     |
     v
Findings
     |
     v
Evidence
     |
     v
Score
     |
     v
Interactive Review Report
Enter fullscreen mode Exit fullscreen mode

Top comments (0)