DEV Community

Cover image for I Built a Code Reviewer That Remembers Team Rules
K Venkata Pavan Kumar
K Venkata Pavan Kumar

Posted on

I Built a Code Reviewer That Remembers Team Rules

A technical article about CodeMind, persistent agent memory, and Hindsight.

Most code review tools can tell me that something is wrong. The harder problem is getting them to remember why my team considers it wrong.

That distinction became the interesting part of building CodeMind, a code review agent that combines an LLM with persistent engineering knowledge. Instead of treating every review as an isolated request, I wanted the reviewer to carry context from previous reviews: coding standards, architectural decisions, security expectations, developer preferences, and lessons learned from feedback.

The key piece that made this possible was Hindsight.

The problem with stateless code review

A typical AI code review starts with a simple request: here is some code, find the problems. The model can inspect the code and produce useful feedback, but it has no inherent reason to know that a particular team prefers service classes over controllers, has a specific validation convention, or previously rejected a certain recommendation.

A code review is not only about whether code is technically valid. It is also about whether the code fits the engineering system around it.

If a team has established that business logic belongs in service classes rather than controllers, the reviewer should be able to use that knowledge in future reviews without requiring the developer to repeat it every time.

Putting memory before the review

The backend is built with Node.js and Express. The review endpoint validates the incoming code, language, and file type before passing the request into the review pipeline.

The important part happens inside the review pipeline. Instead of immediately sending source code to the model, CodeMind first asks Hindsight for relevant engineering knowledge.

const memoryQuery = `
Find team knowledge relevant to reviewing this code.

Consider:

  • Coding standards
  • Security policies
  • Input validation requirements
  • Architecture decisions
  • Controller/service boundaries
  • Database usage
  • Error handling
  • Maintainability
  • Developer preferences
  • Lessons learned from previous reviews

Review context:
Programming language: ${language}
File type: ${fileType}

IMPORTANT:
Only return knowledge that can reasonably influence
the current review.

CURRENT CODE:
${code}
`;

const memories = await recall(memoryQuery);

The query is deliberately contextual. The reviewer should not receive every piece of knowledge the agent has accumulated. It should receive knowledge that has a reasonable relationship with the code being reviewed.

Hindsight provides the underlying agent-memory system.

Hindsight GitHub:-

GitHub logo vectorize-io / hindsight

Hindsight: Agent Memory That Learns



What is Hindsight?

Hindsight™ is an agent memory system built to create smarter agents that learn over time. Most agent memory systems focus on recalling conversation history. Hindsight is focused on making agents that learn, not just remember.

hindsight-learning-demo.mp4

It eliminates the shortcomings of alternative techniques such as RAG and knowledge graph and delivers state-of-the-art performance on long term memory tasks.

Contents


Memory Performance & Accuracy

Hindsight is the most accurate agent memory system ever tested according to benchmark performance. It has…




Hindsight Documentation:-

Overview | Hindsight

Why Hindsight?

favicon hindsight.vectorize.io

Hindsight becomes the team's engineering memory

I kept the Hindsight integration behind a small wrapper. The application exposes three core operations: remember engineering knowledge, recall relevant knowledge, and reflect on accumulated knowledge.

export async function remember(text, options = {}) {
if (!text || !text.trim()) {
throw new Error("Memory content cannot be empty");
}

return await hindsight.retain(bankId, text, {
...options,
});
}

export async function recall(query, options = {}) {
if (!query || !query.trim()) {
throw new Error("Recall query cannot be empty");
}

const result = await hindsight.recall(bankId, query, {
...options,
});

return result;
}

export async function reflect(query, options = {}) {
if (!query || !query.trim()) {
throw new Error("Reflect query cannot be empty");
}

const result = await hindsight.reflect(bankId, query, {
includeFacts: true,
...options,
});

return result;
}

This separation matters. The rest of the application does not need to know how Hindsight works internally. It only needs to know that it can remember engineering knowledge, retrieve relevant knowledge, and reflect on accumulated knowledge.

The interesting part: the reviewer can learn

Retrieval is only half of the problem. If the agent can remember old information but cannot learn from new developer feedback, the memory quickly becomes stale.

export async function learnFromFeedback(feedback) {
console.log("Learning from developer feedback...");

await remember(`
Developer feedback from a CodeMind code review:

${feedback}

This feedback represents a team coding preference,
engineering standard, architectural decision,
or lesson learned.

Use this knowledge when reviewing future code.
`);
}

The important idea is that feedback is not treated only as UI state. It becomes knowledge that can influence later reviews.

The application also records feedback events in PostgreSQL, giving the system a separate historical record of what developers accepted, dismissed, or explicitly taught.

I wanted memory to be inspectable

One thing I learned while working on agent memory is that "the agent remembers" is not a very useful engineering explanation by itself. I wanted to know what influenced a decision.

That is why the project includes a reflection path. Reflection can ask what coding standards have been established, which architectural preferences appear repeatedly, what developer feedback has influenced those standards, and what the reviewer should prioritize in future reviews.

This matters because persistent memory introduces a new debugging question: which previous experiences influenced this response? With a normal LLM request, I can inspect the prompt and response. With an agent that has memory, I also need to inspect the retrieved knowledge.

The review pipeline

+-----------------+
| Developer |
| submits code |
+--------+--------+
|
v
+-----------------+
| Express API |
| validation |
+--------+--------+
|
v
+-----------------+
| Hindsight |
| recall |
+--------+--------+
|
relevant memories
|
v
+-----------------+
| AI Reviewer |
| code + context |
+--------+--------+
|
v
+-----------------+
| Deterministic |
| score calculation|
+--------+--------+
|
v
+-----------------+
| PostgreSQL |
| review history |
+-----------------+

There is another loop running in the opposite direction: developer feedback is retained in Hindsight and can influence future code reviews.

I do not want those responsibilities mixed together. Hindsight answers, "What should the reviewer remember?" PostgreSQL answers, "What happened in the application?"

I kept the scoring deterministic

Another design decision was to avoid letting the model completely control the final score. The reviewer calculates a score from issue severity.

const deductions = {
critical: 30,
high: 20,
medium: 10,
low: 3,
};

for (const issue of issues) {
const severity = String(issue.severity || "").toLowerCase();
score -= deductions[severity] || 0;
}

return Math.max(0, Math.min(100, score));

The model can identify issues, but the numeric score follows a deterministic rule. I prefer this separation because LLMs are useful for interpreting code and explaining problems, while deterministic application logic is easier to test when it is responsible for a calculation.

What a review looks like

Consider a developer submitting a controller containing substantial business logic. The first review can identify the architectural issue. The developer then teaches the agent that business logic should be placed in service classes instead of controllers.

Business logic should be placed in service classes instead of controllers.

That information is retained in Hindsight. During a later review, CodeMind constructs a new memory query using the language, file type, and current source code. Hindsight can return the previously learned architectural preference, which is then supplied to the reviewer as relevant context.

The important behavior is not that the second review somehow "knows" the first review. It has access to a persistent representation of engineering knowledge extracted from previous interactions.

What I learned

  1. Memory should be selective

More memory is not automatically better. Sending every historical interaction to an LLM would create noise and consume context. The review pipeline instead asks Hindsight for knowledge relevant to the current code.

  1. Learning needs an explicit path

If developers can correct the agent but those corrections disappear after the request ends, the system is not really learning. The feedback path makes learning an explicit application operation.

  1. Memory and database history solve different problems

PostgreSQL is useful for review records, scores, issues, timestamps, and feedback events. Agent memory is useful for knowledge that should influence future behavior. Keeping those responsibilities separate makes the architecture clearer.

  1. Agent debugging needs memory visibility

When a response is influenced by previous experiences, normal request/response logs are not enough. Reflection and source-memory reporting give the system a way to inspect what accumulated knowledge is influencing the agent.

  1. The interesting part is not the model

The LLM is only one component of the system. The more interesting engineering problem is everything around it: input, context, memory retrieval, model reasoning, deterministic logic, persistence, feedback, and new memory.

Where I would take it next

The next step is not simply adding more prompts. I would make the memory lifecycle more explicit.

A mature version could track when a rule was introduced, how often it is used, whether developers consistently accept recommendations based on it, and whether newer feedback contradicts an older rule.

The goal is not to build a reviewer that remembers everything. The goal is to build a reviewer that remembers the right things.

That is the reason Hindsight became such an important part of CodeMind. The code review itself is only one interaction. The accumulated engineering knowledge surrounding those interactions is what can make future reviews more relevant.

For more background, see Vectorize's explanation of agent memory:-

What Is Agent Memory? A Complete Guide | Vectorize

Agent memory lets AI agents retain, recall, and reflect on experience across sessions. Learn how it works, the key memory types, and how to implement it.

favicon vectorize.io

Project visuals

CodeMind AI Code Review Agent Interface

AI-Generated Code Review with Hindsight Memories

Developer Feedback on the Generated Review

Hindsight Recall and Memory Influence on the Review

Architecture of the Hindsight Memory Layer

Project

Top comments (0)