DEV Community

Anitha Alli
Anitha Alli

Posted on

How I Built a Code Reviewer That Remembers Every PR It's Ever Seen

I've reviewed enough pull requests to know the pattern: someone forgets to wrap an API call in a try/except, I flag it, they fix it, and three weeks later someone else on the same team makes the exact same mistake. The review tools we use don't remember any of this. Every diff is reviewed in a vacuum, as if the codebase has no history and the team has no conventions.

That bothered me enough to build something about it: a code review agent that actually remembers what it's told a team before, and uses that history to write sharper, more specific reviews over time.

What it does

At its core, the system is simple. You give it a code diff. It looks up similar diffs and reviews from its memory, uses that context to write a new review, and then stores the new review back into memory so the next diff benefits from it too. Over time, instead of generic advice like "consider adding error handling," it starts saying things like "this team has flagged missing error handling on external API calls repeatedly — matches that pattern here."

The memory layer is handled by Hindsight, an agent memory system built by Vectorize. I chose it specifically because it separates the concerns I cared about: I didn't want to build my own vector store, retrieval logic, and ranking system just to get an agent that remembers things. Hindsight's Python client exposes exactly two operations I actually needed — retain to store a memory, and recall to retrieve relevant ones — and it runs multiple retrieval strategies (semantic, keyword, graph, temporal) under the hood so I didn't have to think about which one to use.

The core loop

Here's the actual retrieval step from the agent:

python
recalled = hindsight.recall(bank_id=BANK_ID, query=diff_text)
memory_context = "\n".join(
r.text for r in recalled.results
) if recalled.results else "No prior history yet."

num_matches = len(recalled.results) if recalled.results else 0
print(f"[Memory: based on {num_matches} similar past review(s)]")

That last line ended up mattering more than I expected. Early on, the agent's memory use was invisible — it was working, but there was no way to tell from the output whether it had actually recalled anything relevant or was just generating plausible-sounding text on its own. Printing the match count turned an invisible mechanism into something I could actually verify while testing.

Once I have the recalled context, I pass it to the LLM alongside the new diff:

python
prompt = f"""You are a senior code reviewer for a software team.

Here is a summary of similar past reviews this team has given:
{memory_context}

Now review this new code diff, and where relevant, reference the team's
established patterns from the history above:

Diff: {diff_text}

Give a concise, specific review comment (2-4 sentences)."""

And after generating the review, I write it straight back into memory:

python
hindsight.retain(
bank_id=BANK_ID,
content=f"Diff: {diff_text} Review: {review}",
)

That third step is the part that makes this an agent with memory rather than just a wrapper around an LLM with a prompt template. Every review the system generates becomes a data point that shapes the next one. The loop compounds.

What surprised me: the agent started citing itself

I didn't explicitly instruct the agent to reference specific past incidents — I just told it to use the recalled context "where relevant." But once I'd stored a handful of memories, the generated reviews started producing lines I hadn't written into any prompt template, like referencing that a pattern had been "flagged repeatedly" or citing a specific prior review by description. The model was doing the work of connecting the new diff to the retrieved history on its own, because the retrieved text was concrete enough to reason over.

That's a genuinely different experience than a static linter. A linter can tell you a rule was violated. This agent can tell you the rule exists because of a specific, real precedent it has seen before — which is a much more persuasive thing to read in a code review.

Before and after, side by side

With an empty memory bank, a diff that adds an API call with no error handling gets a review like this:

"The new function calls the inventory API but lacks any error handling, which could cause unhandled exceptions or silent failures in production. Please wrap the API call in a try/catch block."

That's correct, but it's the kind of comment a linter could produce. After seeding just a handful of past reviews establishing the same convention, the same category of diff produces:

"This has been flagged repeatedly (e.g., missing try/catch and logging in prior reviews). Please wrap the API call in a try/catch block, log failures with sufficient context, and add unit tests covering both the successful response and error paths to meet our production-stability standards."

The second version doesn't just restate a rule — it treats the rule as something the team has already established and is now enforcing consistently. That's the difference memory makes.

Adding a narrow chat interface

I also added a small chat feature that lets you ask the agent questions like "what conventions have you learned for this team?" I was careful to scope this tightly — the agent only answers from what's actually stored in Hindsight, and is instructed to say so honestly if the memory doesn't cover something, rather than making up an answer. This wasn't meant to turn the tool into a general assistant. It was meant to make the memory itself inspectable, since a memory system that's invisible is hard to trust.

Lessons learned

Memory needs to be printed, not just used. The single highest-value change I made wasn't a modeling improvement — it was printing the recall count next to every output. If you're building something memory-driven, make the memory visible somewhere, even if it's just a debug line, or you'll have no way to tell if it's actually working.

Seed data quality matters more than volume. A handful of consistent, specific past reviews (same team style, same recurring issues) produced noticeably better results than a larger set of generic ones. The retrieval only helps if what's being retrieved is actually coherent.

Recall result shapes aren't always obvious. I initially assumed Hindsight's recall() would return plain dictionaries, and it actually returns typed result objects with a .text attribute. Small integration detail, but worth checking the actual response shape early rather than assuming.

A narrow scope beats a broad one. I was tempted to let the chat feature answer anything. Keeping it strictly grounded in memory — willing to say "I don't know" — made it more trustworthy, not less useful.

The loop is the feature. It would have been easy to build just the review-generation half and skip writing new reviews back into memory. That one line of retain() after every review is what turns this from a one-shot tool into something that actually improves the longer it runs.

If you're building anything where an agent needs to get better at a repetitive task over time rather than starting fresh every time, agent memory is worth looking at as a building block rather than something you roll yourself.

Top comments (1)

Collapse
 
maddy30445r profile image
Madhur Mittal •

The part I'd be curious about is forgetting. When a team fixes a pattern repo-wide, does the reviewer keep flagging it because of old PRs? Building something similar for my own agent context, being able to see and delete what's remembered turned out to matter as much as the recall itself. Also, are you keying memories by file path or by the kind of mistake? The second seems like it would survive refactors better.