<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Chandana Bonbon</title>
    <description>The latest articles on DEV Community by Chandana Bonbon (@chandanabonbon).</description>
    <link>https://dev.to/chandanabonbon</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147735%2F95a8bbe3-7253-4edd-ae77-81c01b83a23d.png</url>
      <title>DEV Community: Chandana Bonbon</title>
      <link>https://dev.to/chandanabonbon</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chandanabonbon"/>
    <language>en</language>
    <item>
      <title>Dependabot gives pass rates. Our Hindsight agent remembers what broke.</title>
      <dc:creator>Chandana Bonbon</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:39:37 +0000</pubDate>
      <link>https://dev.to/chandanabonbon/dependabot-gives-pass-rates-our-hindsight-agent-remembers-what-broke-728</link>
      <guid>https://dev.to/chandanabonbon/dependabot-gives-pass-rates-our-hindsight-agent-remembers-what-broke-728</guid>
      <description>&lt;h1&gt;
  
  
  Dependabot gives pass rates. Our Hindsight agent remembers what broke.
&lt;/h1&gt;

&lt;p&gt;On 15 September 2020, an engineer at Netlify deployed a dependency upgrade to their database client library. The new version shipped with a default setting far higher than their infrastructure could handle, and about 37 minutes of customer-facing trouble followed. It's exactly the kind of change that's easy to miss in a release.&lt;/p&gt;

&lt;p&gt;I spent a large part of this project reading incidents like that one, and the pattern behind them is the reason Regression Radar exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem is real, and it's documented
&lt;/h2&gt;

&lt;p&gt;The Netlify report is the clearest example I found, because it's first-party and it names the cause: a new default in a client library. It isn't the only one. In April 2023, GitLab's OIDC and OAuth sign-in broke for about 38 hours after a dependency upgrade introduced a breaking change to a key format. It arrived in a &lt;em&gt;patch&lt;/em&gt; release, the kind of version bump where nobody expects breaking changes.&lt;/p&gt;

&lt;p&gt;The cost isn't only outages. Sonatype describes a team that was spending days of development time every four weeks keeping around 30 repositories up to date, before they automated it. The same write-up makes a point I kept coming back to: breaking changes are hardest to catch when they aren't mentioned in the release notes. And when an upgrade does cause an outage, the stakes are high. In ITIC's 2024 survey, 97% of large enterprises put the cost of a single hour of downtime above $100,000.&lt;/p&gt;

&lt;p&gt;The information about what breaks usually already exists. Someone hit the problem first and wrote it up in an issue. It's just scattered across hundreds of threads that nobody reads before bumping a version.&lt;/p&gt;

&lt;h2&gt;
  
  
  What existing tools already do
&lt;/h2&gt;

&lt;p&gt;I expected to find that nobody was working on this. That turned out to be wrong, and it's worth being precise about.&lt;/p&gt;

&lt;p&gt;Dependabot's security updates can include a compatibility score: the share of CI runs that passed when other public repositories made the same update. Renovate's Merge Confidence does something similar, using test results across the repositories where its app runs. So both of them already learn from other teams' upgrades. Socket goes in another direction and can test an upgrade against your own project before you accept it.&lt;/p&gt;

&lt;p&gt;What none of them gives you is an explanation. A pass rate tells you &lt;em&gt;how often&lt;/em&gt; an upgrade went fine elsewhere. It doesn't tell you what broke, why, whether it has been fixed since, or what happened to the developers who went ahead anyway. That's the gap we built for.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Regression Radar fills it
&lt;/h2&gt;

&lt;p&gt;You describe an upgrade (we deliberately focused on one path, Next.js 14.1 to 14.2 with the App Router and Prisma) and it answers with the specific problems other developers reported, each linked to the real GitHub issue and marked fixed or still open. After you upgrade, you mark each warning "this hit us" or "didn't affect us", and the next answer changes.&lt;/p&gt;

&lt;p&gt;The memory underneath is &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight, Vectorize's open-source agent memory system&lt;/a&gt;, loaded with 98 real issues from &lt;code&gt;vercel/next.js&lt;/code&gt; and &lt;code&gt;prisma/orm&lt;/code&gt;. A typical question pulls back about 80 memories: 48 world facts, and 32 observations that Hindsight consolidated across separate reports on its own. It's live at &lt;a href="https://regression-radar.vercel.app" rel="noopener noreferrer"&gt;https://regression-radar.vercel.app&lt;/a&gt;, and the code is at &lt;a href="https://github.com/ekupekuAI/regression-radar" rel="noopener noreferrer"&gt;https://github.com/ekupekuAI/regression-radar&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring the difference instead of describing it
&lt;/h2&gt;

&lt;p&gt;Having spent so long on sources, I cared most about one thing: can the answer be checked? So we measured it. The same question goes to a capable general model with no memory, and every issue number in its answer is checked against our data.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cited&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;matchAll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;(?:&lt;/span&gt;&lt;span class="sr"&gt;#|issues&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;|issue&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+#&lt;/span&gt;&lt;span class="se"&gt;?)(\d{4,6})\b&lt;/span&gt;&lt;span class="sr"&gt;/gi&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
      &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
  &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;verified&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;cited&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;lookup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With memory off, the model cited anywhere from zero to five issue numbers across our runs, and none of them could ever be verified against our dataset. With memory on, it cited seven to twelve, and every one shown to the user is verified, because anything that can't be checked is dropped before it reaches the screen.&lt;/p&gt;

&lt;p&gt;The wording there matters. "Couldn't be verified" is what we actually know. Some of those numbers may well be real issues that simply aren't among our 98. Claiming more than the evidence shows is exactly the habit a project like this should avoid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Protecting the memory
&lt;/h2&gt;

&lt;p&gt;A public feedback button is also a way to write into shared memory, so the learning endpoint only accepts issue numbers that exist in our snapshot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;lookup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;rejected&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reports are still anonymous, and each one counts the same, which is a real limitation. Weighting reports by identity and corroboration is the next step.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Start with a named incident.&lt;/strong&gt; One first-party incident report persuades people more than any survey number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Be fair to the competition.&lt;/strong&gt; Dependabot and Renovate already learn from other teams, as pass rates. Pretending otherwise would have been wrong, and the honest version of the gap is sharper anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Measure the before and after as a number.&lt;/strong&gt; "None of the cited issues could be verified" against "every one shown is verified" settles the argument faster than adjectives do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Say exactly what you checked.&lt;/strong&gt; "Couldn't be verified" is a claim you can defend. Anything stronger usually isn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep the scope narrow.&lt;/strong&gt; One upgrade path, done properly, is more convincing than a promise to cover everything.&lt;/p&gt;

&lt;p&gt;If you want to go further, &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight's documentation&lt;/a&gt; covers how retain, recall and reflect fit together, and Vectorize's piece on &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;what agent memory is&lt;/a&gt; explains why remembering outcomes is different from searching documents.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Built with Ekansh, Shreya and Meghana.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg31oj341d43bgb9qhfj7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg31oj341d43bgb9qhfj7.png" alt=" " width="799" height="359"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
