DEV Community

Cover image for Building a Code Review Agent That Actually Remembers!
Bharath sai Ganipineni
Bharath sai Ganipineni

Posted on

Building a Code Review Agent That Actually Remembers!






Why code review keeps repeating itself

Picture a familiar scene. A teammate opens a pull request, and a reviewer leaves the same
comment they left last month: "Please add a timeout to this HTTP call." The fix is easy. The real
problem is that nothing carried the lesson forward.
Linters and automated reviewers have the same weakness. Each run starts fresh, so knowledge
that one person paid for is lost for the next. In this article I walk through how to design a review
system with long-term memory, and the design choices that matter more than the code

What we are building

The system has two screens. The developer screen is where a review happens. The manager
screen shows patterns over time. Behind them sits a simple three-step pipeline:
• Recall: look up what the team already knows about this file and this kind of error.
• Decide: check whether this exact problem has appeared before.
• Save or reuse: if it is new, store it. If it is a repeat, reuse the earlier fix and bump a counter

Design choice 1: use two stores, not one __
The tempting approach is a single database table: team, file, error, fix. It is easy to build and easy
to chart. But a table only matches exactly. Two developers will describe the same bug in different
words, and a table will treat them as unrelated.
So I split the work in two. A memory engine (Hindsight, using its retain and recall calls)
handles fuzzy, meaning-based lookup. A small SQLite ledger handles hard numbers. Neither is a
copy of the other, because they answer different questions

Design choice 2: one memory bank per team
You could use one global bank, one per developer, or one per team. A global bank mixes teams
whose rules conflict. A per-developer bank hides knowledge from the people who need it. Coding
standards are agreed by teams, so the team is the right boundary

def build_note(team, file_name, problem, rule, fix):
return (
)
f"In {file_name}, team {team} had this problem: {problem}. "
f"Rule broken: {rule}. Fix that worked: {fix}."
note = build_note(
"Alpha", "http_client.py",
"an HTTP request had no timeout and hung the worker",
"all outbound calls must set a timeout",
"pass timeout=5 and handle the timeout error",
)
client.retain(bank_id=bank_for("Alpha"), content=note)

Design choice 4: search by the signal, not the noise
__For recall, I query with the file name plus a short description of the error. I leave the raw code out.
Variable names and whitespace pull results toward code that merely looks similar, while the file
and the failure description describe the actual problemdef find_related(client, team, file_name, error):
query = f"{file_name}: {error}"
result = client.recall(bank_id=bank_for(team), query=query)
return [m.text for m in result.results]

Design choice 5: an exact fingerprint for counting**
__Meaning-based search is great for context, but a number on a manager's dashboard should be
exact. If the dashboard says "this issue happened three times", you must be able to defend that.
So I hash the details into a fingerprint:import hashlib
def fingerprint(team, file_name, code, error):
raw = "|".join([team.strip().lower(), file_name.strip().lower(),
code.strip(), error.strip()])
return hashlib.sha256(raw.encode()).hexdigest(

import sqlite3
db = sqlite3.connect("ledger.db")
db.execute("""CREATE TABLE IF NOT EXISTS incidents (
fingerprint TEXT PRIMARY KEY,
team TEXT, file_name TEXT, repeats INTEGER DEFAULT 1,
first_seen TEXT DEFAULT CURRENT_TIMESTAMP)""")
def is_repeat(fp):
row = db.execute("SELECT repeats FROM incidents WHERE fingerprint=?",
(fp,)).fetchone()
return row is not None

--Plan for failure from day one
Memory is a network call, and network calls fail. Notice the try/except around the save. If the
memory service is down, the review still finishes and the incident is still written to the local ledger.
A reviewer that says "my memory is unavailable right now" is far better than one that crashes.
Keep credentials in environment variables, never in code, and fail early with a clear message if
one is missing.
The manager view
Because the ledger stores counts, a simple query answers the question managers care about: are
we teaching the same thing over and over?
SELECT file_name, SUM(repeats) AS total
FROM incidents WHERE team = ?
GROUP BY file_name ORDER BY total DESC

 **

Make the memory visible
If you ask engineers to trust a system that remembers things, show your work. Display which
search was run and how many memories came back. Turning "trust the agent" into "check the
agent" makes adoption much easier.

The bigger takeaway is that the interesting part of an agent that remembers is rarely the remembering. It's deciding what deserves to be remembered, who gets to retrieve it, and how you'll explain it when it's wrong. If you're building something similar, the Hindsight repository is the place to start.

A Special thanks to pragnasri yellanki, mithin sai mojjada, and Sreevallika Balagonda.

Top comments (1)