There's a difference between knowing what an AI assistant says about a brand and knowing what the brand should do next. That distinction shaped our GEO visibility agent.
The GEO problem
Generative Engine Optimization is about how brands appear in AI-generated answers. Imagine a customer asking an assistant for the best project management tool for a small team. Several brands get recommended. If yours consistently doesn't, that's a visibility problem, and it's different from a traditional search ranking problem.
Our Scan Agent approaches it by repeatedly generating realistic customer queries, sending them to AI engines, and analyzing responses for the target brand and competitors. The result is structured data you can compare over time.
But comparison only becomes useful when the system also remembers what happened between scans.
Why a history of actions matters
A founder gets a recommendation and implements it. On the next scan, visibility changes. That outcome is valuable information. If the system forgets it, the next recommendation starts almost from zero.
Hindsight holds two kinds of history:
- Scan results: what the AI engines said.
- Actions and outcomes: what the founder tried and what happened afterward.
Together they're more useful than either alone.
The recommendation loop
The Recommendation Agent receives the current scan plus the Hindsight record for the brand. With no action history, it treats the output as a baseline. Once actions exist, it can refer to previous attempts.
| Stage | What's available |
|---|---|
| Scan 1 | No previous experience |
| Scan 5 | Some actions and outcomes |
| Scan 10 | Rich history for specific recommendations |
Data contracts
The least visible but most important decision was defining fixed JSON structures between components. The Scan Agent emits a known scan shape, Hindsight stores it, the Recommendation Agent consumes it, and the frontend displays it.
This matters when several people build different parts of one system. If one component renames a field unexpectedly, another breaks even though its own code is correct. Fixed structures act as a shared contract.
Building in parallel
- The AI pipeline is tested against fake memory data.
- The memory layer is tested against fake scan records.
- The frontend is built against hardcoded JSON.
Integration becomes a controlled step instead of a bottleneck.
Showing the difference
The benefit of memory isn't obvious from a single interaction, so the demo uses a synthetic multi-scan history for one brand. The audience sees a baseline recommendation, then one with several past actions behind it, then one backed by a rich history. The point is to make one idea visible: the current recommendation is influenced by what happened before.
What we learned
It's tempting to think of an agent as a prompt wrapped around a model. Ours is different. The Scan Agent provides observations, Hindsight provides experience, the Recommendation Agent turns both into advice, the founder supplies real-world action and outcome, and the frontend makes the process understandable. Each part has a clear responsibility, and the more history the system has, the more it has to work with for the next decision.
- Sufiya Maheen markdown --- title: "Engineering the Scan and Recommendation Agents for a Memory-Driven GEO System" published: false description: "Why we split the AI pipeline into two modules, and how we test whether memory actually changes the output." tags: ai, llm, agents, testing cover_image: ---
The AI pipeline turns a brand name into two things: a measurable GEO visibility result and a recommendation for what to do next. I worked on both the Scan Agent and the Recommendation Agent, and we kept them as separate modules because they solve different problems.
- Scan Agent: gathers and structures evidence.
- Recommendation Agent: reasons over that evidence and the history in Hindsight.
Module A: Scan Agent
It starts with a brand name and category. Instead of asking a model "do you know this brand?", it generates questions around a customer's actual need. Those go to models representing engines like ChatGPT and Perplexity, and each response is analyzed for:
- whether the target brand was mentioned
- which competitors were mentioned
Raw responses are kept too. The output has a fixed shape:
{
"brand": "Acme",
"timestamp": "2026-01-15T10:00:00Z",
"queries_tested": ["..."],
"mentions": 3,
"total_queries": 20,
"competitors_mentioned": ["..."],
"raw_snippets": ["..."]
}
Why structured data?
It would be easy to pipe raw model responses into the next prompt. We didn't, because a scan should be reusable evidence. A structured scan can be stored, displayed, compared with another scan, and passed to the recommendation layer without that layer knowing how the responses were collected. It also tells the frontend exactly which fields to expect.
Module B: Recommendation Agent
Inputs:
- the current scan
- the Hindsight record (scan history plus actions log)
The actions log matters most because it records not just what was recommended, but what was actually tried and what happened. One rule is built in:
If the actions log is empty, we're at scan 1. Give a baseline recommendation and acknowledge there's no history yet.
If actions exist, the recommendation must consider what was tried and whether it worked.
Why they stay separate
The Scan Agent asks, "What is the current visibility situation?" The Recommendation Agent asks, "Given the situation and what happened before, what should we recommend?" One big script would make testing and integration harder. Separate modules let me test the Scan Agent with sample inputs without Hindsight running, and test the Recommendation Agent with a fake scan and progressively richer histories.
Testing scan 1, 5, and 10
The key question isn't "does it produce a recommendation?" It's "does the recommendation change when the memory changes?"
- Scan 1: empty history. The output should be general and honest about it.
- Scan 5: earlier actions and outcomes. It should notice patterns.
- Scan 10: richer history. It should be specific and cite previous actions and measured outcomes.
Similar precedents
Hindsight also exposes a similar-precedent function. Given a new scan, an AI call picks a relevant past action or outcome from the same brand or a similar one, with a short explanation of why it's useful. The agent doesn't have to rely only on fixed rules, since a past situation can be useful even when it isn't identical.
Fake data first
We built against hardcoded data before connecting real components, so problems in the pipeline showed up before integration.
The interface is simple: current scan + remembered history β recommendation.The pipeline doesn't need to know how Hindsight stores things, and Hindsight doesn't generate scans.
Top comments (0)