How I Built a Code Reviewer That Remembers Team Decisions
Introduction
AI code review tools can identify syntax problems, missing validation, possible bugs, and common coding issues.
However, there is one problem: a normal AI reviewer may not remember what a particular development team decided in previous reviews.
For example, a team may decide that inline CSS should not be used and that Tailwind classes should be preferred. The same team may also have a rule that API parameters must always be validated before they are processed.
If these decisions are not remembered, the reviewer may give generic feedback every time.
This led us to the idea of building a code review agent that does not only review the current pull request, but also remembers the team's previous decisions.
The Problem With Stateless Code Review
Software teams develop their own coding conventions over time.
These conventions can include:
- Preferred coding styles
- Architecture decisions
- Design patterns
- API validation rules
- Common bug fixes
- Frontend styling conventions
- Backend development practices
Many of these decisions are discussed during pull request reviews.
The problem is that useful review knowledge can easily get lost in previous pull requests, conversations, or discussions.
A new pull request may contain the same mistake that was already discussed several times.
A normal reviewer can identify the technical problem, but it may not know that the team has already discussed the same issue and established a specific rule.
This creates repeated review comments and additional review cycles.
The Idea
The idea behind our project is to connect a code review agent with Hindsight memory.
The agent receives a pull request from GitHub or GitLab and analyzes the changes.
Instead of looking only at the current code, the agent can also use relevant information from previous team discussions and reviews.
This allows the review to consider both:
- What is happening in the current pull request
- What the team has decided in the past
The result is a code review that is more specific to the development team's working style.
Caption: Figure 1: Architecture of the Hindsight-powered code review agent.
System Architecture
The system consists of several main components.
GitHub or GitLab provides the pull request and the latest code changes.
The Code Review Agent analyzes the pull request and communicates with Hindsight to retrieve relevant previous information.
Hindsight acts as the memory layer. It stores useful information such as team conventions, architecture decisions, and recurring bugs.
The relevant memory is then used together with the current code to generate review feedback.
The overall flow is:
GitHub/GitLab → Code Review Agent → Hindsight Memory → Review Analysis → Review Feedback
The agent can also store new useful decisions so that they can be used during future reviews.
What the Agent Remembers
The important part of this system is deciding what information should become memory.
The agent can remember team conventions such as using Tailwind classes instead of inline CSS.
It can remember architecture decisions, such as where specific business logic should be placed.
It can also remember recurring bugs that have appeared in previous pull requests.
For example, suppose a team previously identified a problem where an API parameter was used without validation.
If the team decides that all API parameters must be validated before processing, this decision becomes useful information for future reviews.
When a similar pull request appears later, the agent can retrieve this information.
Caption: Figure 2: Relevant team conventions and previous review decisions recalled for a code review.
How Hindsight Changes Code Review
Consider a developer submitting a pull request that contains an API parameter.
A basic AI reviewer may simply say that the parameter should be validated.
That feedback is technically useful, but it does not explain the team's previous decision.
With memory, the feedback can become more specific.
The agent can recognize that the team previously discussed API parameter validation and can connect the current issue with that decision.
Instead of giving only a generic recommendation, the reviewer can explain that the current implementation does not follow an existing team convention.
The same idea applies to frontend development.
Suppose the team previously decided to use Tailwind classes instead of inline CSS.
A future pull request contains inline CSS.
A memory-aware reviewer can identify the issue and connect it with the team's earlier decision.
This makes the feedback more meaningful to the developers.
Caption: Figure 3: The code review agent analyzing a pull request and providing context-aware feedback.
From the First Review to Later Reviews
The first interaction with the system can be relatively simple.
The agent receives a pull request and performs a normal code review. It can identify syntax issues, possible bugs, missing validation, and other common problems.
As the team continues using the system, more useful information becomes available.
Previous review comments and team decisions can become part of the agent's memory.
Later, when a new pull request arrives, the agent can recall information that is relevant to the current changes.
This creates a continuous feedback cycle.
The system is not simply reviewing one pull request at a time. It is building context from previous engineering decisions.
Before and After
Without memory, the review process can look like this:
Before
The reviewer identifies a missing validation check and gives a general recommendation to add validation.
The developer receives technically correct feedback, but the review does not contain information about the team's previous decisions.
With Hindsight memory, the review can provide more context.
After
The reviewer identifies the missing validation and connects it with the team's existing convention that API parameters must be validated before processing.
The developer can understand not only what should be changed, but also why the change is important for the team's codebase.
This difference is important because good code review is not only about finding problems. It is also about maintaining consistency across a project.
Caption: Figure 4: Comparison between generic review feedback and feedback using previous team decisions.
Automatically Detecting Recurring Problems
Another useful part of the idea is identifying patterns that appear repeatedly.
If the same type of architectural problem appears in multiple pull requests, it can become a recurring issue.
For example, if developers repeatedly place business logic in the wrong layer, the agent can use previous review knowledge to identify similar situations in future pull requests.
This can help teams detect architectural violations earlier.
Instead of discovering the same problem during multiple reviews, the agent can use accumulated knowledge to provide more relevant feedback.
Why Memory Is Different From a Long Prompt
One possible approach would be to put every team rule into a very large prompt.
However, that approach becomes difficult to maintain as the project grows.
Teams can have many decisions, and not every decision is relevant to every pull request.
A memory system provides a way to retrieve information that is relevant to the current situation.
For a frontend change, frontend conventions may be useful.
For an API change, API validation rules and backend architecture decisions may be more relevant.
This makes the review context more focused.
Lessons Learned
1. Not everything needs to become memory
A useful memory should help future decisions.
Temporary information or irrelevant comments do not necessarily need to be remembered.
The quality of the memory is important because irrelevant information can make future reviews less useful.
2. Context makes feedback more useful
A generic recommendation can identify a problem, but context can explain why that problem matters to a particular team.
Remembering previous decisions makes the feedback more connected to the project.
3. Code reviews contain valuable engineering knowledge
Pull request discussions are not only about fixing individual lines of code.
They often contain decisions about architecture, coding standards, design patterns, and recurring problems.
This information can be useful beyond the original pull request.
4. Repeated review comments are good candidates for memory
When the same issue appears repeatedly, it is a sign that the team may have an established convention or recurring problem.
Remembering these patterns can reduce repeated explanations.
5. Human decisions still matter
The goal of the system is not to replace developers or technical leads.
The team's decisions are still made by people.
The agent uses those decisions as context to make future reviews more consistent.
Limitations
A memory-based code review system also has limitations.
Not every previous decision should automatically be applied to every new situation.
A rule that was appropriate for one part of a project may not always apply somewhere else.
The system therefore needs relevant memory retrieval and appropriate context.
It is also important for developers to review the agent's suggestions rather than blindly accepting every recommendation.
The final decision about whether a change should be made remains with the development team.
Conclusion
The main idea behind this project is simple: a code review agent should not only understand the code it is reviewing today, but should also be able to use relevant knowledge from previous team decisions.
By connecting the review agent with Hindsight memory, previous conventions, architecture decisions, and recurring bugs can become useful context for future reviews.
This changes the review process from a stateless interaction into a more continuous learning workflow.
Instead of repeatedly explaining the same rules, the team can build a shared memory that helps the agent provide more context-aware feedback over time.
The combination of current code + team memory creates a code review process that is more connected to how the team actually develops software.
Caption: Figure 5: Running the code review agent and retrieving relevant team memory.
Hindsight Resources
Hindsight GitHub https://github.com/vectorize-io/hindsight
- Hindsight Documentation https://hindsight.vectorize.io/
- Hindsight Agent Memory Documentation https://vectorize.io/what-is-agent-memory






Top comments (0)