DEV Community

Dasari Abhiram
Dasari Abhiram

Posted on

My Code Review Agent Stopped Repeating Itself With Hindsight

I replayed 21 real Flask pull requests through my review agent twice. With no memory it wrote 96 comments, and 62 of them were kinds of feedback my simulated team had already rejected. With Hindsight memory on, it wrote 47 comments and repeated none of those rejected kinds.

This post covers how I built that, the one bug that taught me the most, and what the numbers do and do not show.

The problem I wanted to fix

Anyone who has done code review on a small team knows the loop. A tool or a new reviewer suggests adding a docstring, someone says "we don't do that here", and three weeks later the same suggestion shows up again. The conventions live in people's heads, and most AI review tools start from zero on every pull request.

I wanted an agent that gets told "no" once and remembers it. I called it the Review Desk.

What the Review Desk does

It reads a pull request diff and posts comments inline, under the lines they refer to. A person clicks Accept or Reject on each comment, with an optional reason. Every decision is stored in Hindsight. On the next review, the agent recalls those decisions before it writes anything, so it stops repeating rejected kinds of feedback and keeps raising what the team values.

The stack is small: Python, FastAPI, Groq running openai/gpt-oss-120b, the GitHub API, and a single-file React frontend loaded from a CDN. I only use two Hindsight operations, retain and recall.

The Review page with inline comments and Accept / Reject buttons

Where Hindsight sits

The whole memory layer is two functions. This is the real code from replay_real.py:

def retain_memory(bank_id: str, text: str) -> None:
    _mem.retain(bank_id=bank_id, content=text, context="code-review-feedback")

def recall_memory(bank_id: str, query: str) -> list[str]:
    for attempt in range(6):
        try:
            res = _mem.recall(bank_id=bank_id, query=query)
            return [r.text for r in (res.results or [])]
        except Exception as e:
            msg = str(e).lower()
            # Bank/tables not created yet, or transient server error: wait and retry.
            transient = ("not found" in msg or "404" in msg or "does not exist" in msg
                         or "undefinedtable" in msg or "500" in msg or "503" in msg)
            if not transient:
                raise
            time.sleep(8)
    print("    (recall unavailable after retries; continuing with no memory for this PR)")
    return []
Enter fullscreen mode Exit fullscreen mode

Every Accept or Reject becomes a short piece of text, and retain_memory writes it to a bank. Before each review, recall_memory pulls the relevant notes back, and they go into the prompt. The retry loop exists because a fresh bank is not always ready on the first call, so the code waits eight seconds and tries again, up to six times. If recall still fails, the review continues without memory instead of crashing. The Hindsight docs cover the retain and recall calls in more detail.

One thing I learned the practical way: Hindsight needs about ten seconds to process a retain. In the UI I show a "Memory settling" bar after a rejection so nobody reviews again too early and thinks memory failed.

A rejected comment marked as remembered, with the ten-second settling notice

How I measured it

I fetched merged pull requests from pallets/flask, kept the ones with 4 to 200 changed diff lines, and skipped release, version bump and typo-style titles. I replayed them in merge order, once with memory off and once with memory on, with a fresh memory bank for each run.

The team is simulated. A script plays the reviewer: it accepts comments about security, validation, error handling and bugs, and rejects everything else. That gives me a fixed ground truth, so I can count how often the agent repeats something the script already rejected.

The headline run used 21 pull requests. I removed four that only touched CI or the logo.

Memory off Memory on
Comments generated 96 47
Repeat-rejected 62 0
Accepted 29 39
Acceptance rate 30% 83%

The running totals of repeat rejections show the shape of it: after 5 pull requests it was 9 against 0, after 10 it was 24 against 0, and after 20 it was 59 against 0. Without memory the agent kept making the same mistakes. With memory, it made none of the rejected ones after being told.

The Replay page: memory off on top, memory on below, with a scripted reviewer

Running total of repeated rejections: 62 without memory, 0 with memory

The bug that taught me the most

Hindsight stores what it is given as facts. When I rejected a comment, the stored fact was narrow, something like "rejected a docstring on the load function". The agent read that back correctly, but it did not generalize it. A docstring comment on a different function slipped through, and so did other documentation comments.

The fix in the live app is a decision log. Each decision carries a category, and the app builds category rules from the log (for example, "the team rejects docs comments") and puts them first in the prompt, ahead of the recalled Hindsight notes. The notes add detail, and the rules cover the broad cases.

I want to be plain about one thing: the 62 to 0 result came from replay_real.py, which uses Hindsight recall only. The decision-log rules were added to the live app afterward, so they were not part of the measured replay. I have not measured the combined version.

What the numbers do not show

  • One run per arm. Model output varies, and with a single run I cannot give a range.
  • Restarts. Comment ids live in server memory, so a restart invalidates old ids. Saved decisions survive.
  • The team was a script. Real reviewers are inconsistent, and I have not tested that.
  • Accepted counts are not a clean signal. In my earlier 10-pull-request run, accepted comments went from 18 to 14 with memory on. In the 21-pull-request run they went from 29 to 39. They moved in opposite directions, so I do not claim memory raises or lowers accepted comments. The acceptance rate is high partly because memory cut total comments roughly in half.
  • Comment quality is unverified. The comments are model output. Inline placement is approximate, and at least one security comment about the Host header is a stretch. None of them are confirmed Flask bugs.
  • Category labels drift. The same kind of feedback can get a slightly different label between runs.
  • Privacy. Decisions are stored in a plain text file, data/decisions.json, with no encryption at rest.
  • Not built yet. Comments show up in the UI. They are not posted back to the pull request on GitHub. A GitHub App that posts comments and treats resolved or dismissed threads as accept and reject signals is the obvious next step, and I have not built it.

What I would take from this

  1. Store decisions as rules, not only as events. A raw rejection is a narrow fact. If you want broad behavior, keep a category and state the rule explicitly.
  2. Put the strongest instructions first in the prompt. Recalled notes are useful context, but they should not be the only thing standing between the model and a repeated mistake.
  3. Measure repeats, not only acceptance. Counting repeat-rejected comments gave me a number I could trust more than an acceptance rate.
  4. Plan for the delay. If your memory layer takes seconds to process a write, tell the user, or they will read a normal delay as a failure.
  5. Say what you did not measure. It made the results easier to defend.

Try it

The code is at github.com/abhiram0411/review-desk. A read-only copy with the measured results is at review-desk-z659.onrender.com. It is on a free host, so the first load can take up to a minute or two. Live reviews and the paste-a-PR-link box run locally.

If you are new to the idea behind all this, the agent memory page on Vectorize is a good place to start, and the Hindsight repository has what you need to run it yourself.

Thanks to Code.in for the coding help.
Tagging Code.in.

Top comments (0)