DEV Community

Vinay Damarsing
Vinay Damarsing

Posted on

Why I Replayed Real Flask Pull Requests Through Hindsight

The fastest way to make an agent demo look fake is to test it on invented data. For my code review agent I wanted pull requests that a real maintainer team had merged, so the memory test would face real diffs. This article covers how I picked them, one concrete pull request as a before and after, and what real data does not fix.

Why real pull requests

My agent reads a diff, posts comments under the lines they refer to, and stores each Accept or Reject in Hindsight. To see whether memory changes behavior, I needed a stream of different diffs in a fixed order, where the same kinds of feedback could come up again and again. Merged pull requests from a popular open source project give exactly that, and they come with a real merge order.

Picking the set

I pulled merged pull requests from pallets/flask with the GitHub API, using a fine-grained token that has read-only access to public repositories. The fetch script saves each pull request under data/prs. Its rules are short:

  • Keep pull requests whose diff is 4 to 200 lines. Smaller ones give the agent nothing to say, and larger ones are slow and noisy.
  • Skip titles that look like a release, a version bump or a typo fix.

This is the real code for that filter:

TARGET_COUNT = 25
MAX_DIFF_LINES = 200      # keep diffs small enough to be readable + cheap for the LLM
MIN_DIFF_LINES = 4        # skip trivial one-line diffs, not interesting for a demo
SKIP_WORDS = ("release", "bump", "codespell", "flit_core", "pre-commit", "typo")
# ...
            if not pr.get("merged_at"):
                continue  # only want PRs that were actually merged, not just closed
            if any(w in pr["title"].lower() for w in SKIP_WORDS):
                continue
# ...
            diff_lines = diff.count("\n")
            if diff_lines < MIN_DIFF_LINES or diff_lines > MAX_DIFF_LINES:
                continue  # too trivial or too big for a clean demo
Enter fullscreen mode Exit fullscreen mode

After that I removed pull requests that only touched CI or the logo. In the 21-pull-request run, I dropped #5945, #5795, #5754 and #5757. In the earlier 10-pull-request run, I dropped #5924, #5865 and #5844. I list them so that anyone can check my choices. It is a judgment call, and it does change the set.

The number 21 is not a round choice. The script's target was 25, and after I dropped those four, 21 were left. I kept 21 instead of fetching more to reach a round number. I also noticed the filter works on titles only, so it would skip a genuine fix that happens to mention a typo. It is a rough filter, and I would tighten it before trusting it on a bigger set.

The Replay page: 21 real Flask pull requests in merge order

The replay

I replayed the set in merge order, once with memory off and once with memory on, using a fresh memory bank for each run. A script played the reviewer. It accepts comments about security, validation, error handling and bugs, and rejects everything else. That gives a fixed rule, so I can count how often the agent repeats something that was already rejected.

The Replay page in the app shows the run. Each block is one review comment, and the legend has three states: accepted, rejected for the first time, and rejected again.

One pull request, before and after

The pull request that shows the effect best is #6133, "add app.query route decorator". Without memory, the agent wrote 6 comments and 3 of them were repeated rejections. With memory it wrote 2 comments and none were repeated. Same diff, same model. The difference is that the memory-on run had already been told what this team rejects.

Across all 21 pull requests the totals were:

Memory off Memory on
Comments generated 96 47
Repeat-rejected 62 0
Accepted 29 39
Acceptance rate 30% 83%

Running total of repeated rejections: 62 without memory, 0 with memory

Real data in the live app too

The same pull requests are served by the app. /demo-prs lists them and /demo-prs/{n} returns one, so anyone running the app locally can review a real diff in one click. There is also a box to paste any public GitHub pull request link. It accepts only links shaped like https://github.com/owner/repo/pull/N, and it works on localhost only.

What real data does not fix

  • One project. All 21 pull requests come from Flask. Its maintainers have a style, and a different project could behave differently. I have not tested another repository.
  • The team is still simulated. The diffs are real, but the reviewer is a script. Real reviewers are inconsistent.
  • One run per arm. Model output varies, and with a single run I cannot give a range.
  • Accepted counts are not a clean signal. In my earlier run, accepted went from 18 to 14 with memory on. In this run it went from 29 to 39. They moved in opposite directions, and the acceptance rate is high partly because memory cut the total comments roughly in half.
  • The comments are unverified model output. Inline placement is approximate, at least one security comment about the Host header is a stretch, and none of them are confirmed Flask bugs.
  • The measured replay used Hindsight recall only. The live app also puts category rules from a decision log first in the prompt, and those rules were not part of the measurement.
  • Comments show in the UI. They are not posted back to the pull request on GitHub.

Takeaways

  1. Use data that a stranger can look up. It makes the test easier to check.
  2. Write down what you removed from the set, and why.
  3. Show one concrete example next to the totals. #6133 explained the totals better than the totals did.
  4. Real diffs do not make the reviewer real. Say which part is simulated.

Try it

The code is at github.com/abhiram0411/review-desk, and a read-only copy with the results is at review-desk-z659.onrender.com. It runs on a free host, so allow a minute or two for the first load. The Hindsight docs explain retain and recall, and the agent memory page on Vectorize is a good place to start on the idea.

Tagging Code.in.

Top comments (0)