DEV Community

Cover image for Stopping an Agent from Inventing Stats Using Hindsight Recall
Mohammed Fahad Ali
Mohammed Fahad Ali

Posted on

Stopping an Agent from Inventing Stats Using Hindsight Recall

By Mohd Fahad Ali

An early version of our agent told a test brand that a topic "increases engagement by 43%." No such number existed anywhere in our data. The model made it up because a confident percentage sounds like good strategy.

What I own

I built the reasoning layer for a content strategist. Hindsight supplies memory. My code takes a question, pulls what the brand has learned, and returns a recommendation that has to justify itself. The public interface is small:

askAgent(brandId, question) // -> { recommendation, reasoning[], format, memoriesUsed }
Enter fullscreen mode Exit fullscreen mode

The pipeline: recall, reflect, build prompt, call the LLM (Groq, with OpenAI as a fallback), validate JSON, return.

Through-line: grounding is a contract, not a hope

Telling a model "be accurate" doesn't work. I made grounding structural in four ways.

1. Number the evidence

Recalled memories go into the prompt as a labeled list, so the model can cite them and I can check the citations.

const evidence = memories.map((m, i) => `[${i + 1}] ${m}`).join('\n');

const system = `You are a content strategist.
Use ONLY facts in the numbered memory below. Never invent statistics,
dates, or outcomes. Each reasoning item must cite a memory like [3].
If memory is insufficient, say so instead of guessing.`;
Enter fullscreen mode Exit fullscreen mode

2. Force structured output

I ask for JSON and validate it with zod. If it fails, I retry once with a stricter reminder, then throw a typed AgentError rather than return junk.

const Strategy = z.object({
  recommendation: z.string().min(3),
  reasoning: z.array(z.string()).min(1),
  format: z.string(),
});

const raw = await llm.json({ system, user: `${evidence}\n\nQuestion: ${question}`, temperature: 0.3 });
const parsed = Strategy.parse(JSON.parse(stripFences(raw)));
return { ...parsed, memoriesUsed: memories.length };
Enter fullscreen mode Exit fullscreen mode

Note that memoriesUsed is set by my code from the recall result, not by the model. The model never gets to report its own evidence count.

3. Check the output against the input

After generation, I scan every reasoning line for numbers and verify each one appears in the recalled memory:

const nums = (s: string) => s.match(/\d+(\.\d+)?%?/g) ?? [];
const memoryText = memories.join(' ');
const ungrounded = parsed.reasoning.filter(line =>
  nums(line).some(n => !memoryText.includes(n)));
if (ungrounded.length) throw new AgentError('UNGROUNDED_CLAIM', ungrounded);
Enter fullscreen mode Exit fullscreen mode

Combined with the retry, this removed the invented percentages. It's a crude filter, and it can't catch a wrong claim that contains no number, but it catches the most tempting failure.

4. Lower the temperature

Reasoning runs at 0.3. Creativity helps with copy, so content generation is a separate function at a higher temperature, given the brand voice and audience models from Hindsight's reflect rather than raw memories.

Before and after

No memory (count is 0): the agent lists generic ideas and flags that memory is empty, so the interface can say "generic answer."

With brand history: recommendation "Why Most AI Agents Forget Everything," with reasons such as "technical posts earned stronger saves [2]" and "audience asked about agent memory [5]." Every line points to something that was recalled.

After new feedback ("we need more practical examples"): the same question yields a carousel built on a real implementation example. I changed no prompt text. The difference is what agent memory returned.

(Add screenshots here: the prompt with numbered memories, a validated JSON response, and the rejected ungrounded line in logs.)

"What did we learn?"

A small summary function reads the five reflected models and produces four to six concrete statements, such as "audience prefers practical technical content." It has the same rule: only what the models contain.

Lessons

  1. Cite by index. Numbered evidence makes claims checkable.
  2. Set trust boundaries in code. The model doesn't get to decide memoriesUsed.
  3. Validate outputs, not just inputs. Schema plus a grounding check beat prompt wording.
  4. Separate reasoning from writing. Different tasks want different temperatures.

Where it still fails

The grounding check is string matching. A paraphrased wrong claim ("engagement doubled") passes if it has no digits. A second pass that asks a model to verify each claim against the memory list would be stronger, at the cost of another call per request. I also rely on Groq's speed for latency, so a slow provider day is visible to users immediately.

Top comments (0)