<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: jyothisri laveti</title>
    <description>The latest articles on DEV Community by jyothisri laveti (@jyothisri_laveti_3ba6a948).</description>
    <link>https://dev.to/jyothisri_laveti_3ba6a948</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147012%2F7bea85d3-e17c-4339-a7f8-00b1bb9b712e.png</url>
      <title>DEV Community: jyothisri laveti</title>
      <link>https://dev.to/jyothisri_laveti_3ba6a948</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jyothisri_laveti_3ba6a948"/>
    <language>en</language>
    <item>
      <title>I taught my code reviewer with Hindsight—and it remembered too well</title>
      <dc:creator>jyothisri laveti</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:03:32 +0000</pubDate>
      <link>https://dev.to/jyothisri_laveti_3ba6a948/my-hindsight-code-reviewer-kept-enforcing-a-dead-rule-i0j</link>
      <guid>https://dev.to/jyothisri_laveti_3ba6a948/my-hindsight-code-reviewer-kept-enforcing-a-dead-rule-i0j</guid>
      <description>&lt;p&gt;Our team switched from Redux to Zustand four months ago. Last week, my AI code reviewer told a new teammate to "use Redux, as decided in PR #12."&lt;/p&gt;

&lt;p&gt;The memory worked perfectly. That was the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;I built a code review agent that connects to GitHub pull requests. When a human reviewer leaves feedback such as "use Tailwind, not inline CSS" or "validate API parameters for null," the agent stores it. On the next PR, it recalls the relevant team decisions and uses them in its review.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdre4xr4o8mlbrcl176wp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdre4xr4o8mlbrcl176wp.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The system has three parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A GitHub webhook that receives new and updated pull requests&lt;/li&gt;
&lt;li&gt;An LLM that reads the diff and drafts the review&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight, the open-source agent memory system&lt;/a&gt;, which stores team decisions (retain) and brings them back when they're relevant (recall)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without memory, the reviewer behaves like a generic linter with opinions. With memory, it sounds like a teammate who's been on the project for a year. Most of this article is about that second half — and about what went wrong once it actually worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem: teams repeat the same review comments
&lt;/h2&gt;

&lt;p&gt;Every team I've worked with leaves the same pull request comments again and again. "Don't use inline CSS, we use Tailwind." "Add a null check on this API parameter." "Follow the folder structure." The rule was decided once, but people forget it, new teammates never hear it, and a senior reviewer ends up typing it out for the fourth time.&lt;/p&gt;

&lt;p&gt;Generic AI reviewers don't fix this. They catch syntax and style issues, but they know nothing about the decisions your team made in earlier reviews. Every review starts from zero.&lt;/p&gt;

&lt;p&gt;What I wanted was a reviewer that gets better with every pull request: generic on day one, aware of the team's conventions after a few reviews, and eventually able to flag the mistakes that keep recurring.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I stopped using SQLite for memory
&lt;/h2&gt;

&lt;p&gt;My first prototype was a SQLite table of past review comments with keyword search. It broke within a day, because reviewers rarely repeat themselves in the same words. One writes "don't style inline," another writes "use utility classes." Both describe the same rule, but keyword search sees two unrelated sentences.&lt;/p&gt;

&lt;p&gt;I needed retrieval by meaning, and I didn't want to spend the project building a retrieval layer from scratch. That's what &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;agent memory&lt;/a&gt; is for. Moving to Hindsight gave me three things I would otherwise have had to build by hand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Recall by meaning.&lt;/strong&gt; A query about "styling conventions" finds the Tailwind decision even when the original comment never used those exact words.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Separate memory per team.&lt;/strong&gt; Each team gets its own memory bank, so one team's conventions never leak into another team's reviews.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A small, predictable API.&lt;/strong&gt; The &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; got me to a working retain-and-recall loop in an afternoon. The whole integration is really two function calls.
## The core loop&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg52izml5yb28g21ax1z3.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg52izml5yb28g21ax1z3.jpeg" alt=" " width="799" height="245"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When a human reviewer leaves a comment on a PR, the agent retains it along with its source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;team-rules&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PR #42 (reviewer: priya): Team decided to use Tailwind &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classes instead of inline CSS.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When a new PR comes in, the agent recalls before it writes anything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memories&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;team-rules&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;styling conventions relevant to this diff&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Review this diff.
Team decisions you must respect:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;format_memories&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memories&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Diff:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;diff&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The recalled memories go straight into the prompt, next to the diff. Every memory carries its source PR number, so the review can point to exactly where a rule came from. That detail mattered more than I expected — developers accept "decided in PR #42" far more readily than "the AI suggests."&lt;/p&gt;

&lt;h2&gt;
  
  
  When the same mistake shows up a fifth time
&lt;/h2&gt;

&lt;p&gt;Recall on its own answers "what did we decide." It doesn't answer "is this happening again and again." For that, MemReview counts how often each recalled rule actually matches a new diff, and once a rule crosses a threshold, the review cites every violated convention at once instead of just one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;This violates the convention from &lt;strong&gt;PR #51&lt;/strong&gt;: route handlers must not call the database directly; use the service layer.&lt;/li&gt;
&lt;li&gt;This violates the convention from &lt;strong&gt;PR #37&lt;/strong&gt;: all API handler functions must null-check parameters before use.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;By the time a pattern has come up this many times, the agent isn't just linting — it's pointing at a real, recurring gap in the codebase, sourced back to the exact PR where the team decided the rule.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F903u2z047y556gmavvuz.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F903u2z047y556gmavvuz.jpeg" alt=" " width="800" height="278"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Letting developers grade the agent's memory
&lt;/h2&gt;

&lt;p&gt;Not every recalled memory is still useful, and I didn't want MemReview to just assume its own suggestions were right. Every review comment ships with Accept and Reject buttons.&lt;/p&gt;

&lt;p&gt;A Reject doesn't delete the underlying memory — it's a signal. A rule that gets rejected repeatedly is probably outdated, worded badly, or no longer matches how the team actually works. An Accept reinforces that the memory is still pulling its weight.&lt;/p&gt;

&lt;p&gt;This turned out to matter more than the recall logic itself. An agent that repeats a stale rule with total confidence is worse than one with no memory, because people stop reading its comments at all — including the correct ones.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faijq7tflmmite1045pni.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faijq7tflmmite1045pni.jpeg" alt=" " width="799" height="226"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Right now, a rejected suggestion lowers that memory's trust weight, so a rule that gets rejected repeatedly shows up less often in future recalls instead of continuing to interrupt reviews.&lt;/p&gt;

&lt;h2&gt;
  
  
  What didn't work
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Recall was too broad early on.&lt;/strong&gt; Before I narrowed queries to the diff's actual topic, unrelated conventions crept into reviews and made them noisy instead of helpful.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not every rejection means the rule is wrong.&lt;/strong&gt; Sometimes a developer rejects a correct suggestion just because it's inconvenient that week. Trust weight helps, but it isn't a perfect signal on its own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agent only knows what got written down.&lt;/strong&gt; A decision made in a call or a Slack thread and never mentioned in a PR simply doesn't exist for it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Citing multiple PRs at once can overwhelm a small diff.&lt;/strong&gt; When three conventions fire on one file, the review reads more like a checklist than feedback. I'm still tuning how many memories to surface per comment.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Lessons learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Memory needs feedback, not just storage.&lt;/strong&gt; Recall alone doesn't tell you if a rule is still good — Accept/Reject does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store the source with every memory.&lt;/strong&gt; A PR number turns a suggestion into something a developer can go verify, and it's the single detail that built the most trust.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recurring violations are more convincing than single ones.&lt;/strong&gt; Citing PR #51 and PR #37 together on one diff makes the pattern undeniable in a way one citation doesn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let humans push back cheaply.&lt;/strong&gt; A one-click Reject is enough friction to catch bad memories without turning code review into a chore.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieve by meaning, and don't build it yourself.&lt;/strong&gt; For messy human feedback, semantic recall beats keyword matching every time, and using &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; meant my time went into the review logic instead of a retrieval layer.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're building agents that need to remember, start with the &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight docs&lt;/a&gt; and design a way for humans to correct it from day one — not as an afterthought once the memory has already gone stale.&lt;/p&gt;

&lt;p&gt;The Hindsight GitHub repository: &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;https://github.com/vectorize-io/hindsight&lt;/a&gt;&lt;br&gt;
The documentation for Hindsight: &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;https://hindsight.vectorize.io/&lt;/a&gt;&lt;br&gt;
The agent memory page on Vectorize: &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;https://vectorize.io/what-is-agent-memory&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>python</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
