DEV Community

Cover image for Dependabot gives pass rates. Our Hindsight agent remembers what broke.
Chandana Bonbon
Chandana Bonbon

Posted on

Dependabot gives pass rates. Our Hindsight agent remembers what broke.

Dependabot gives pass rates. Our Hindsight agent remembers what broke.

On 15 September 2020, an engineer at Netlify deployed a dependency upgrade to their database client library. The new version shipped with a default setting far higher than their infrastructure could handle, and about 37 minutes of customer-facing trouble followed. It's exactly the kind of change that's easy to miss in a release.

I spent a large part of this project reading incidents like that one, and the pattern behind them is the reason Regression Radar exists.

The problem is real, and it's documented

The Netlify report is the clearest example I found, because it's first-party and it names the cause: a new default in a client library. It isn't the only one. In April 2023, GitLab's OIDC and OAuth sign-in broke for about 38 hours after a dependency upgrade introduced a breaking change to a key format. It arrived in a patch release, the kind of version bump where nobody expects breaking changes.

The cost isn't only outages. Sonatype describes a team that was spending days of development time every four weeks keeping around 30 repositories up to date, before they automated it. The same write-up makes a point I kept coming back to: breaking changes are hardest to catch when they aren't mentioned in the release notes. And when an upgrade does cause an outage, the stakes are high. In ITIC's 2024 survey, 97% of large enterprises put the cost of a single hour of downtime above $100,000.

The information about what breaks usually already exists. Someone hit the problem first and wrote it up in an issue. It's just scattered across hundreds of threads that nobody reads before bumping a version.

What existing tools already do

I expected to find that nobody was working on this. That turned out to be wrong, and it's worth being precise about.

Dependabot's security updates can include a compatibility score: the share of CI runs that passed when other public repositories made the same update. Renovate's Merge Confidence does something similar, using test results across the repositories where its app runs. So both of them already learn from other teams' upgrades. Socket goes in another direction and can test an upgrade against your own project before you accept it.

What none of them gives you is an explanation. A pass rate tells you how often an upgrade went fine elsewhere. It doesn't tell you what broke, why, whether it has been fixed since, or what happened to the developers who went ahead anyway. That's the gap we built for.

How Regression Radar fills it

You describe an upgrade (we deliberately focused on one path, Next.js 14.1 to 14.2 with the App Router and Prisma) and it answers with the specific problems other developers reported, each linked to the real GitHub issue and marked fixed or still open. After you upgrade, you mark each warning "this hit us" or "didn't affect us", and the next answer changes.

The memory underneath is Hindsight, Vectorize's open-source agent memory system, loaded with 98 real issues from vercel/next.js and prisma/orm. A typical question pulls back about 80 memories: 48 world facts, and 32 observations that Hindsight consolidated across separate reports on its own. It's live at https://regression-radar.vercel.app, and the code is at https://github.com/ekupekuAI/regression-radar.

Measuring the difference instead of describing it

Having spent so long on sources, I cared most about one thing: can the answer be checked? So we measured it. The same question goes to a capable general model with no memory, and every issue number in its answer is checked against our data.

const cited = Array.from(
  new Set(
    Array.from(answer.matchAll(/(?:#|issues\/|issue\s+#?)(\d{4,6})\b/gi))
      .map((m) => Number(m[1]))
  )
);
const verified = cited.filter((n) => lookup(n));
Enter fullscreen mode Exit fullscreen mode

With memory off, the model cited anywhere from zero to five issue numbers across our runs, and none of them could ever be verified against our dataset. With memory on, it cited seven to twelve, and every one shown to the user is verified, because anything that can't be checked is dropped before it reaches the screen.

The wording there matters. "Couldn't be verified" is what we actually know. Some of those numbers may well be real issues that simply aren't among our 98. Claiming more than the evidence shows is exactly the habit a project like this should avoid.

Protecting the memory

A public feedback button is also a way to write into shared memory, so the learning endpoint only accepts issue numbers that exist in our snapshot:

if (!lookup(n)) {
  rejected.push(n);
  continue;
}
Enter fullscreen mode Exit fullscreen mode

Reports are still anonymous, and each one counts the same, which is a real limitation. Weighting reports by identity and corroboration is the next step.

Lessons

Start with a named incident. One first-party incident report persuades people more than any survey number.

Be fair to the competition. Dependabot and Renovate already learn from other teams, as pass rates. Pretending otherwise would have been wrong, and the honest version of the gap is sharper anyway.

Measure the before and after as a number. "None of the cited issues could be verified" against "every one shown is verified" settles the argument faster than adjectives do.

Say exactly what you checked. "Couldn't be verified" is a claim you can defend. Anything stronger usually isn't.

Keep the scope narrow. One upgrade path, done properly, is more convincing than a promise to cover everything.

If you want to go further, Hindsight's documentation covers how retain, recall and reflect fit together, and Vectorize's piece on what agent memory is explains why remembering outcomes is different from searching documents.

Built with Ekansh, Shreya and Meghana.

Top comments (0)