DEV Community

Cover image for Hindsight finds the bug reports. My code decides what's true.
G. Ekansh
G. Ekansh

Posted on AI-assisted

Hindsight finds the bug reports. My code decides what's true.

Hindsight finds the bug reports. My code decides what's true.

The first real answer our upgrade agent gave me looked perfect: a numbered list of things that break when you move a Next.js app from 14.1 to 14.2, each tagged FIXED or OPEN, each citing a GitHub issue. Then I checked the citations against our own data. Three of the ten warnings said OPEN for issues that had been closed for months.

That one check ended up shaping the whole architecture of Regression Radar.

What it does

Regression Radar tells you what is likely to break before you upgrade a dependency. You describe the upgrade, for example "Next.js 14.1 to 14.2, app router + Prisma", and it answers with the specific problems other developers hit on that stack. Every warning links to the real GitHub issue and says whether it is fixed or still open. After you upgrade, you mark each warning "this hit us" or "didn't affect us", and the next developer gets an answer shaped by what actually happened to you.

The memory behind it is Hindsight, the open-source agent memory system from Vectorize. We loaded 98 real issues from vercel/next.js and prisma/orm into a Hindsight bank: 56 still open, 42 closed, all from 2024 onwards. A Next.js API on Vercel sits in front of it with three endpoints. One answers the question, one records what happened after the upgrade, and one asks the same question with no memory so you can compare. It's live at https://regression-radar.vercel.app and the code is at https://github.com/ekupekuAI/regression-radar.

The bug that shaped everything

My first version asked Hindsight's reflect() to do everything at once: find the relevant bug reports, decide whether each one was fixed, and cite them. The output read well. It was also wrong in a way that would have stayed invisible if I hadn't checked.

Issues #64603, #66248 and #64609 all came back as OPEN. All three were closed. The reason was obvious once I read the threads. People write "still broken on 14.2.1" in comments weeks before a fix lands, and memory faithfully stores what people said. When the model summarised, the loudest sentence won. It also cited #71281, which isn't in our dataset at all.

Memory was doing its job. I was asking it the wrong question. Hindsight is very good at deciding what is relevant to an upgrade. It was never meant to be a database of the current state of 98 GitHub issues, and I already had that database, in a committed snapshot.

So I split the question in two.

Ask memory only what it is good at

The briefing now asks reflect() about relevance and nothing else. Hindsight's response_schema parameter makes that enforceable: the schema has no status field at all, so the model has nowhere to put an opinion about whether something is fixed.

const data = await call("/reflect", {
  query,              // "...list what is likely to break for this stack,
                      //  with the GitHub issue numbers it comes from.
                      //  Do not say whether something is fixed."
  budget: "high",
  response_schema: RISK_SCHEMA, // title, why, confidence, issue_numbers
  ...(context ? { context } : {}),
});
Enter fullscreen mode Exit fullscreen mode

What comes back is structured: a title, one sentence on why it affects this stack, a confidence level, and the issue numbers it came from.

Let the snapshot decide what is true

Every issue number then goes through a lookup against the committed snapshot. The status a user sees comes from GitHub's own data, never from the model, and anything that can't be found is dropped instead of shown.

for (const n of numbers) {
  const hit = lookup(n);       // the committed 98-issue snapshot
  if (hit) found.push(hit);
  else dropped.push(n);        // can't verify it, so it isn't shown
}
const anyOpen = found.some((i) => i.state === "open");
Enter fullscreen mode Exit fullscreen mode

I made a deliberate call here: a warning you can't click through to is worse than one fewer warning. So a risk with no verifiable source disappears entirely.

The same rule, applied to the playbook

Hindsight also supports mental models: a standing summary it writes by running a question over the whole bank, and can rewrite on its own after new memories are consolidated. We keep one called "What breaks upgrading Next.js 14.1 to 14.2 with the App Router and Prisma".

Its first draft repeated the old failure. It labelled closed issues as open, and it cited "#84901675" for the sitemap problem, a number that can't exist. The real issue is #66248. So the playbook gets the same treatment as the briefing: any status label the model wrote is stripped, every number is checked, the real status is stamped next to it, and anything unverifiable is removed.

t = t.replace(/#(\d{3,9})\b/g, (_m, num: string) => {
  const hit = lookup(Number(num));
  if (!hit) return GONE;               // removed from the text
  return `#${num} · ${hit.state === "open" ? "still open" : "fixed"}`;
});
Enter fullscreen mode Exit fullscreen mode

What it looks like now

Ask about Next.js 14.1 to 14.2 with the App Router and Prisma, and you get back a handful of warnings, usually around eight. One is CSS resolving in a different order in production than in development, #64921, still open. Another is sitemap.ts generating a broken XML namespace, #66248, fixed.

A Memory Inspector shows what memory did for that answer. A typical question recalls about 80 memories: 48 world facts, and 32 observations that Hindsight consolidated from the raw reports on its own. The same question with memory switched off, sent to a capable general model, cited between zero and five issue numbers across our runs, and none of them could ever be verified against our data. With memory on, it cites seven to twelve, and every one shown is verified.

Then the learning loop. Mark the sitemap warning "this hit us" and another one "didn't affect us", then ask again. The sitemap warning moves to the top with "Confirmed by 1 developer who did this upgrade", and the other drops to the bottom with low confidence. Within 22 seconds, the playbook had rewritten itself to mention the report.

Lessons

Memory is an index of relevance, not a source of truth for state you already hold. If you have ground truth, stamp it in code. Don't ask a model to rediscover it from comments.

Use response_schema to remove questions, not just to format answers. The most reliable way to stop a model guessing a field is to not give it that field.

Verify every reference and drop what fails. Checking citations against a snapshot is cheap, and it catches exactly the failures that read best.

When a model won't follow instructions reliably, move that logic into code. Developer feedback was first passed to the model as context. It ignored the instruction to mention confirmations, and one report was counted six times because Hindsight splits a report into several memories. Now feedback is stored as tagged memories and applied deterministically.

Count reports, not facts. Anything you tally from memory needs a stable document ID.

If you're building on agent memory, Hindsight's documentation is the place to start, and Vectorize's explainer on what agent memory is is a good framing of why it's different from retrieval. The most useful thing I learned is also the simplest: let memory decide what's relevant, and let your data decide what's true.

Built with Shreya, Chandana and Meghana.

Top comments (0)