DEV Community

Ahmed Raza
Ahmed Raza

Posted on Fully Autonomous

Teaching a Code Reviewer Team Decisions with Hindsight

Exact account mapping: 0/3 without memory, 3/3 with Hindsight. That was the clearest result from our eight-case synthetic comparison of ReviewMemory. We ran each case once per condition, with the same reviewer and settings. All 16/16 paced requests completed.

Fixed check Baseline With memory
Exact account mapping 0/3 3/3
All required literal fixes in a defect case 1/4 3/4
Exact await-call syntax 3/3 2/3
Boundary cases with zero facts and citations 2/2 2/2

These checks overlap; adding their numerators would invent an aggregate score. The account result concerns three specific suggestions, and the await result regressed. The two boundary cases test evidence isolation. None establishes general review accuracy.

Hindsight retains explicit team decisions and retrieves facts for later reviews. Our illustrative billing service requires a particular account mapping that the diff does not reveal. The public repository contains the fixtures, raw responses and fixed scorer. No customer code or measured time savings are involved.

Start with a patch and an empty workspace

The sample refund handler starts a database operation and returns a success response without awaiting the write. This excerpt comes from the actual example diff:

+  void recordRefund({
+    invoiceId,
+    amountCents,
+    currency,
+    eventId: event.id,
+  });
Enter fullscreen mode Exit fullscreen mode

You can load this patch in the workbench and select the baseline review. Baseline mode skips memory retrieval. It uses the same review model, settings, and system instructions as memory mode; the difference is the supplied evidence.

In the recorded public-app walkthrough, the baseline returned one finding, zero supplied facts and zero memory citations. It identified the unawaited write. After retaining the decisions, the same diff received three facts and two findings, both citing evidence. This walkthrough is separate from the fixed eight-case evaluation.

The patch omits the helper implementation and route wrapper. A maintainer can supply their contract as feedback, then inspect whether the next review uses it.

Baseline review showing one finding and zero supplied memory facts

Baseline review of the illustrative refund patch: one finding and zero supplied facts. Baseline mode bypasses saved decisions.

Save the reasoning behind a decision

The example workflow offers three decisions to preview and save. One says that a route wrapper checks webhook signatures before invoking handlers under src/webhooks/**. Another describes durable deduplication: the wrapper commits its processed marker after the handler resolves, so the handler must await database writes. The third requires these handlers to pass connectedAccount: event.account into the billing helpers.

That third decision adds information unavailable from the patch alone. In the illustrative contract, omitting the optional account field falls back to a legacy platform account. An invoice identifier does not select the correct connected account.

A useful saved decision includes its rationale, file scope, review outcome, and source. “Don't report this again” would leave too much unresolved. The signature decision instead explains where verification happens and limits the convention to the wrapper's handlers. It does not give unrelated endpoints permission to accept unverified requests.

You can also use “Teach a decision” from a finding and write your own feedback. Saving requires an explicit action. Generated suggestions enter memory through user-saved decisions, and the submitted diff is not stored as a decision.

After sending a synchronous retain request, the backend checks that Hindsight returns extracted facts for the saved document. This is the final step in the actual retain implementation:

const memories = await this.list(bank, decision.repository, documentId);
return { documentId, retained: memories.length > 0, memories };
Enter fullscreen mode Exit fullscreen mode

The interface can therefore distinguish a confirmed memory from a button click. A failed provider request produces an error rather than a fabricated saved result.

Retrieve evidence for the next review

Hindsight handles retention and retrieval; Groq serves the review model, openai/gpt-oss-120b, in the recorded run. The application passes retrieved facts into the model request as teamEvidence. Each fact carries an identifier, text, document reference when available, and file scope.

The recall request uses repository and decision tags together. This excerpt is from the same backend:

types: ['world', 'experience'],
tags: [repositoryTag(input.repository), 'review-decision'],
tags_match: 'all_strict',
budget: 'mid',
max_tokens: 4096,
include: { entities: null },
Enter fullscreen mode Exit fullscreen mode

The server then filters recalled facts against the changed file paths before calling the reviewer. A rule scoped to src/webhooks/** cannot enter the prompt for a patch that changes only src/routes/refund-callback.ts. For a diff containing several files, the server also checks that each finding's cited evidence applies to that finding's file.

The Hindsight source repository is the reference for the memory system itself. ReviewMemory adds application-specific boundaries around it: what users may save, which repository a request concerns, which paths a decision covers, and how evidence references are validated.

Test the feedback on fixed cases

A separate evaluation compared eight prewritten synthetic scenarios: four defective handlers, two handlers satisfying the stated contracts, and two scope boundaries. Each ran once without memory and once after retaining the three decisions. All 16/16 review requests completed in this paced run on September 28, 2026.

The model, prompt, fixtures, and scoring rules stayed fixed. The scorer checked literal suggestions for the expected file, including connectedAccount: event.account and await followed by the helper call. It used no LLM judge and did not execute proposed patches.

The opening table uses four defect cases for the all-fixes check and three cases each for account mapping and await syntax. A separate zero-findings check on two contract-compliant handlers passed 1/2 baseline and 2/2 with memory.

All three matching account suggestions cited the corresponding retained decision. That is the clearest measured change: the reviewer supplied a team-specific mapping absent from the diff. The await count moved the other way. One memory response recommended awaiting the helper in prose but omitted the literal call syntax, so it remained a failure under the unchanged scorer.

All six applicable memory requests received three facts each. Both boundary cases received zero facts and zero memory citations. Baseline also supplied zero facts to both boundaries because it skipped recall. The memory condition tests exclusion after retention.

The earlier unpaced pilot had 26/32 requests fail with PROVIDER_BUSY. Its raw results are preserved and excluded from the paired comparison. Before the next run, an amendment reduced repeats from two to one and added a 65-second gap after each review. Sanitized errors did not establish the exact upstream quota responsible.

Median round-trip time was 1,659.5 ms baseline and 3,067 ms with memory, eight completed requests per condition. Median recall took 715 ms. These exclude the 65-second scheduling gaps and do not characterize cold starts. One domain and one repeat cannot establish reliability or superiority over a reviewer given the same context directly.

Comparison of baseline and memory reviews for the same refund patch

The captured memory review has three facts and two findings; baseline has one finding. This interactive example is separate from the fixed evaluation above.

A citation still needs inspection

The qualitative inspection found flaws even in responses passing the literal checks. Some memory responses predicted duplicate processing after a premature deduplication commit, although the stated wrapper would skip a retry. Others invented downstream validation failures. Correct mappings and real citations did not make those explanations correct.

Read the source fact beside a finding. Check the condition, the affected file, and the proposed fix. A remembered convention can explain a team's intent without proving that its wrapper exists, that a security check works, or that the model's causal explanation is sound.

Try the workflow on a small change

The landing page offers sign-up, sign-in and an explicit guest demo. Account onboarding saves a workspace name, repository and role; PostgreSQL stores account workspaces and review history, while Hindsight stores saved decisions. Returning users can reopen history after signing in. Guests can try the samples without an account. Use non-confidential code: reviews send code to Groq, and recall sends a diff excerpt to Hindsight.

Inspect each response beside the patch and its evidence, then try the later-change scenario. ReviewMemory does not fetch repository files, execute tests or post review comments; check a proposed change before applying it.

For local setup, configure PostgreSQL with DATABASE_URL and the public origin with APP_URL, alongside Node.js, provider credentials and a stable SESSION_SECRET. The live verification script exercises actual retention, later recall, evidence references, and isolation. It creates a workspace and writes the disclosed sample decisions, so running it makes real provider requests.

Vectorize's agent memory overview provides background. The measured result here is specific: account mapping 0/3 to 3/3, all fixes 1/4 to 3/4, await syntax 3/3 to 2/3. Keep the gains, misses and original decisions together when assessing a suggestion.

Related developer tooling: Code.in.

Top comments (0)