How I Used Hindsight to Stop Repeating Failed SEO Experiments
Draft article — add verified repository snippets in the marked placeholders
The Problem
I wanted to build an SEO agent that could do more than analyze the current state of a page and generate another recommendation from general SEO patterns. The missing piece was memory.
An SEO experiment has a useful lifecycle: observe a signal, make a decision, run an experiment, measure the outcome, and use that outcome when making the next decision. Without persistent memory, the agent can repeatedly arrive at recommendations that look reasonable in isolation but conflict with what the same system has already learned from measured results.
YourWebMind is built around that loop. It compares current SEO signals with historical experiments, uses Hindsight to retrieve relevant evidence, reasons over that evidence, proposes a recommendation, waits for human approval, measures the experiment, and retains the outcome for future decisions.
What I Wanted the Agent to Remember
The important part was not simply storing previous prompts or conversations. I wanted the system to remember decisions and their measured consequences.
For SEO, that means retaining information such as the page and keyword involved, the type of change made, the baseline metric, the measured result, and whether the outcome was positive, negative, neutral, or inconclusive. A failed experiment is therefore not discarded. It becomes evidence that can affect a later recommendation.
The demo uses a synthetic SEO dataset covering June 1, 2026 through September 28, 2026. It contains pages, keywords, rankings, observations, decisions, experiments, and outcomes.
The Memory Loop
The core loop is:
current SEO data → Hindsight Recall → historical evidence → Hindsight Reflect → recommendation → human approval → experiment → measured outcome → Hindsight Retain → future recommendation.
This separates three memory operations with different jobs. Retain records what happened. Recall retrieves relevant historical information. Reflect synthesizes the retrieved evidence into reasoning that can support the next decision.
The human approval step is intentional. YourWebMind recommends; a human decides whether an experiment should proceed.
How Retain, Recall, and Reflect Fit Together
The application uses Hindsight as the memory layer rather than treating the language model as the only source of historical context.
File:
https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/services/hindsight_service.py
Function/Class: self._embedded = HindsightEmbedded(
profile=self.settings.HINDSIGHT_PROFILE,
llm_provider=self.settings.HINDSIGHT_API_LLM_PROVIDER,
llm_api_key=self.settings.HINDSIGHT_API_LLM_API_KEY,
llm_model=self.settings.HINDSIGHT_API_LLM_MODEL,
llm_base_url=self.settings.HINDSIGHT_API_LLM_BASE_URL,
database_url=self.settings.HINDSIGHT_API_DATABASE_URL,
)
# Ensure the daemon starts and verify reachability
daemon_url = self._embedded.url
logger.info(f"Hindsight Embedded daemon successfully running at {daemon_url}")
self._initialized = True
Purpose: Show the actual Hindsight initialization used by the project.
RETAIN
File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/memory/retain_adapter.py
Function/Class: def retain_experiment_outcome(
self,
payload: Dict[str, Any],
bank_id: str = "searchmind-demo",
) -> RetainResult:
"""Execute real Hindsight retain operation using Prompt 03 verified schema."""
service = self._get_service()
doc_id = payload.get("document_id")
exp_id = payload.get("metadata", {}).get("experiment_id") or doc_id or "unknown"
t0 = time.time()
resp = service.retain(
bank_id=bank_id,
content=payload["content"],
context=payload.get("context"),
metadata=payload.get("metadata"),
tags=payload.get("tags"),
document_id=doc_id,
timestamp=payload.get("timestamp"),
update_mode=payload.get("update_mode", "replace"),
)
latency = round(time.time() - t0, 3)
Purpose: Show the smallest real code path that retains an experiment outcome.
RECALL
File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/memory/recall_adapter.py
Function/Class: def recall(
self,
query: str,
bank_id: str = "searchmind-demo",
budget: str = "mid",
max_tokens: int = 2000,
) -> List[RecalledMemoryItem]:
"""Execute live Hindsight recall and map results to RecalledMemoryItem with provenance."""
service = self._get_service()
resp = service.recall(
bank_id=bank_id,
query=query,
budget=budget,
max_tokens=max_tokens,
)
recalled_items: List[RecalledMemoryItem] = []
for item in getattr(resp, "results", None) or []:
text = getattr(item, "text", "") or ""
doc_id = getattr(item, "document_id", None)
context = getattr(item, "context", None)
meta = getattr(item, "metadata", None) or {}
recalled_items.append(
RecalledMemoryItem(
source="hindsight",
memory_bank=bank_id,
query=query,
memory_text=text,
experiment_id=extract_experiment_id(doc_id, text, context),
outcome=extract_outcome(text, meta),
)
)
return recalled_items
Purpose: Show how the current SEO context is turned into a memory query and how relevant memories are retrieved.
REFLECT
File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/memory/reflect_adapter.py
Function/Class: def recall(
self,
query: str,
bank_id: str = "searchmind-demo",
budget: str = "mid",
max_tokens: int = 2000,
) -> List[RecalledMemoryItem]:
"""Execute live Hindsight recall and map results to RecalledMemoryItem with provenance."""
service = self._get_service()
resp = service.recall(
bank_id=bank_id,
query=query,
budget=budget,
max_tokens=max_tokens,
)
recalled_items: List[RecalledMemoryItem] = []
for item in getattr(resp, "results", None) or []:
text = getattr(item, "text", "") or ""
doc_id = getattr(item, "document_id", None)
context = getattr(item, "context", None)
meta = getattr(item, "metadata", None) or {}
recalled_items.append(
RecalledMemoryItem(
source="hindsight",
memory_bank=bank_id,
query=query,
memory_text=text,
experiment_id=extract_experiment_id(doc_id, text, context),
outcome=extract_outcome(text, meta),
)
)
return recalled_items
Purpose: Show the real Reflect call or adapter used to synthesize reasoning from recalled evidence.
A Concrete SEO Example
The demo focuses on a Waterproof Alpine Hiking Boots page targeting the keyword “waterproof hiking boots.”
The current CTR is 1.80%, compared with a previous CTR of 3.20%. The current ranking is 3.1, with a search volume of 15,400 and 277 clicks.
When the system looks at historical evidence, it finds nine relevant experiments: two positive, three negative, two neutral, and two inconclusive. One useful precedent is EXP_001, a title optimization experiment with a +50.0% CTR change. Another is EXP_014, a promotional modifier experiment with a -20.1% CTR change. EXP_018 is inconclusive.
Based on that evidence, YourWebMind recommends:
“Move the primary keyword closer to the beginning of the title, retain the brand, and avoid aggressive promotional discount modifiers.”
The recommendation is shown with 57% confidence and cites nine historical experiments. It is still gated behind human approval rather than being applied automatically.
The Code That Makes This Work
The article should use small, real snippets from the repository rather than large implementation dumps. The snippets should demonstrate the actual boundary between the SEO agent and Hindsight.
[CODE SNIPPET — HINDSIGHT RETAIN]
• Primary File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/memory/retain_adapter.py
• Supporting File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/experiments/lesson.py
• Class / Function: HindsightRetainAdapter.retain_experiment_outcome
def retain_experiment_outcome(
• self,
• payload: Dict[str, Any],
• bank_id: str = "searchmind-demo",
• ) -> RetainResult:
• """Execute real Hindsight retain operation using Prompt 03 verified schema."""
• service = self._get_service()
• doc_id = payload.get("document_id")
• exp_id = payload.get("metadata", {}).get("experiment_id") or doc_id or "unknown"
•
• t0 = time.time()
• resp = service.retain(
• bank_id=bank_id,
• content=payload["content"],
• context=payload.get("context"),
• metadata=payload.get("metadata"),
• tags=payload.get("tags"),
• document_id=doc_id,
• timestamp=payload.get("timestamp"),
• update_mode=payload.get("update_mode", "replace"),
• )
• latency = round(time.time() - t0, 3)
[CODE SNIPPET — HINDSIGHT RECALL]
• Primary File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/memory/recall_adapter.py
• Supporting File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/agents/workflow.py
• Class / Function: HindsightRecallAdapter.recall
• def recall(
• self,
• query: str,
• bank_id: str = "searchmind-demo",
• budget: str = "mid",
• max_tokens: int = 2000,
• ) -> List[RecalledMemoryItem]:
• """Execute live Hindsight recall and map results to RecalledMemoryItem with provenance."""
• service = self._get_service()
• resp = service.recall(
• bank_id=bank_id,
• query=query,
• budget=budget,
• max_tokens=max_tokens,
• )
•
• recalled_items: List[RecalledMemoryItem] = []
• for item in getattr(resp, "results", None) or []:
• text = getattr(item, "text", "") or ""
• doc_id = getattr(item, "document_id", None)
• context = getattr(item, "context", None)
• meta = getattr(item, "metadata", None) or {}
•
• recalled_items.append(
• RecalledMemoryItem(
• source="hindsight",
• memory_bank=bank_id,
• query=query,
• memory_text=text,
• experiment_id=extract_experiment_id(doc_id, text, context),
• outcome=extract_outcome(text, meta),
• )
• )
• return recalled_items
[CODE SNIPPET — HINDSIGHT REFLECT]
• Primary File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/memory/reflect_adapter.py
• Supporting File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/agents/workflow.py
• Class / Function: HindsightReflectAdapter.reflect
python
resp = service.reflect(
bank_id=bank_id,
query=query,
context=context,
budget="low",
include_facts=False,
)
latency = round(time.time() - t0, 3)
raw_text = (
getattr(resp, "text", None)
or getattr(resp, "response", None)
or str(resp)
)
sanitized_text = sanitize_non_dogmatic_reflection(raw_text)
sections = parse_reflection_sections(sanitized_text)
[CODE SNIPPET — HUMAN APPROVAL GATE]
• Primary File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/experiments/service.py
• Supporting File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/api/agent.py
• Class / Method: ExperimentExecutionService.create_experiment_from_recommendation
python
# Enforce human approval gateway
is_approved = (
approval_status == ApprovalStatus.APPROVED
or (hasattr(approval_status, "value") and approval_status.value == "approved")
or str(approval_status).lower() == "approved"
)
if not is_approved:
raise PermissionError(
f"Cannot create experiment from recommendation '{recommendation.recommendation_id}' "
f"with status '{approval_status}'. Operator approval is required."
)
[CODE SNIPPET — EXPERIMENT OUTCOME RETENTION]
• Primary File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/experiments/service.py
• Supporting File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/experiments/lesson.py
• Class / Method: ExperimentExecutionService.execute_learning_loop and build_memory_retain_payload
python
# 1. Generate canonical document ID
doc_id = f"seo_{experiment.experiment_id}" # Resolves to 'seo_exp_021_demo'
# 2. Package experiment, measured delta (+50.0%), and lesson into retain payload
payload = build_memory_retain_payload(
experiment=experiment,
outcome=outcome,
lesson=lesson,
project_id=experiment.project_id,
)
# 3. Retain into memory bank
return retain_adapter.retain_experiment_outcome(payload, bank_id=target_bank)
Why Negative Precedents Matter
A useful property of this design is that negative outcomes remain available as evidence.
For example, the historical record does not only contain successful title changes. It also contains an experiment where a promotional modifier was associated with a 20.1% CTR decline. That does not prove that promotional language always causes a decline. It does mean the result is available as a relevant precedent in the same decision context.
This changes the recommendation process from “What is a common SEO best practice?” to “What happened in comparable experiments, and what evidence should influence the next test?”
That distinction is important because an agent can otherwise produce a plausible recommendation without knowing that a similar change has already produced an undesirable result.
Deterministic Evaluation and Live Hindsight
I also separated deterministic evaluation from live Hindsight verification.
The normal evaluator/demo mode runs without making live LLM calls. This makes the end-to-end workflow reproducible and keeps the demo independent of external model quota. The repository includes a dedicated demo-mode configuration and startup path.
[ADD VERIFIED DETERMINISTIC DEMO MODE SNIPPET HERE]
File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/app/main.py
Function/Class: Lifespan
@asynccontextmanager
async def lifespan(app: FastAPI):
"""Application lifespan manager handling startup and shutdown of Hindsight."""
logger.info("=== Starting YourWebMind Backend ===")
if settings.SEARCHMIND_DEMO_MODE:
logger.info(
"SEARCHMIND_DEMO_MODE is active. Skipping live Hindsight Embedded daemon initialization "
"and external LLM provider calls. Operating in deterministic offline evaluator mode."
)
yield
logger.info("=== Shutting down YourWebMind Backend (Demo Mode) ===")
return
hindsight = get_hindsight_service()
Purpose: When SEARCHMIND_DEMO_MODE=true, the backend bypasses embedded daemon startup and external LLM provider calls entirely. As shown in backend/app/api/agent.py, the agent runtime instead injects InMemoryRecallAdapter, InMemoryReflectAdapter, and InMemoryRetainAdapter populated from verified offline fixtures in data/demo/, enabling repeatable, zero-token evaluation..
Separately, I verified the Hindsight lifecycle using an isolated live-test memory bank named searchmind-live-test. The test exercised Retain, Recall, and Reflect. The persistent demo bank, searchmind-demo, remained unchanged at 133 memory units during that verification.
[ADD VERIFIED LIVE HINDSIGHT TEST SNIPPET HERE]
File: https://github.com/Gowtham3424/yourwebmind/blob/master/backend/tests/test_hindsight_integration.py
Test: test_hindsight_smoke_test_lifecycle
Purpose: Show the real integration test covering the Hindsight lifecycle.
What Happened After the Experiment
After human approval, the demo experiment changes the measured CTR from 1.80% to 2.70%.
That is a +0.90 percentage-point change and a +50.0% relative increase. The outcome is classified as POSITIVE and retained as memory under seo_exp_021_demo.
The important part is not treating this single result as universal proof. The retained lesson is scoped to the demonstrated context: the result was observed within the commercial hiking footwear category and should be treated as evidence for future recommendations, not as a guarantee that the same change will work everywhere.
[ADD VERIFIED OUTCOME-RETENTION SNIPPET HERE]
File: ______________________________
Function/Class: Part 1: Metric Calculation & Empirical Classification
• File: backend/app/experiments/outcome.py (Lines 81–101)
• Function: compute_experiment_outcome
baseline_value = experiment.baseline_metrics.get(experiment.primary_metric, 0.0)
abs_change = round(measured_value - baseline_value, 4)
# Relative percentage change calculation
pct_change = round((measured_value - baseline_value) / abs(baseline_value) * 100.0, 2)
outcome_cls = classify_outcome_deterministically(
baseline=baseline_value,
measured=measured_value,
percentage_change=pct_change,
primary_metric=experiment.primary_metric,
caveats=caveats,
)
What happens here:
When evaluated with baseline CTR 1.80% and measured CTR 2.70%:
• abs_change = 2.70 - 1.80 = +0.90 percentage points
• pct_change = (0.90 / 1.80) * 100 = +50.0% relative lift
• outcome_cls = OutcomeClass.POSITIVE
Part 2: Dynamic Document ID & Payload Construction
• File: backend/app/experiments/lesson.py (Lines 88–101)
• Function: build_memory_retain_payload
python
def build_memory_retain_payload(
experiment: SEOExperiment,
outcome: ExperimentOutcome,
lesson: str,
project_id: str = "searchmind-demo-store",
tags: Optional[List[str]] = None,
) -> Dict[str, Any]:
"""Construct the complete Hindsight Retain payload for persistent storage."""
doc_id = f"seo_{experiment.experiment_id}"
exp_type = experiment.experiment_type
metric = experiment.primary_metric
before = outcome.primary_metric_before
after = outcome.primary_metric_after
pct = outcome.percentage_change
pct_str = f" ({pct:+.2f}%)" if pct is not None else ""
ocls_str = (
outcome.outcome_class.value
if hasattr(outcome.outcome_class, "value")
else str(outcome.outcome_class)
).upper()
How seo_exp_021_demo is generated:
The document identifier is generated dynamically:
python
doc_id = f"seo_{experiment.experiment_id}"
Because the experiment is created with experiment_id="exp_021_demo" in
backend/app/api/agent.py, doc_id automatically evaluates to seo_exp_021_demo.
Part 3: Committing to Hindsight Persistent Memory
• File: backend/app/experiments/service.py (Lines 150–170)
• Class / Method: ExperimentExecutionService.retain_in_memory
python
def retain_in_memory(
self,
experiment: SEOExperiment,
outcome: ExperimentOutcome,
lesson: str,
retain_adapter: BaseRetainAdapter,
bank_id: Optional[str] = None,
) -> RetainResult:
"""Format memory document and retain into Hindsight via retain_adapter."""
target_bank = bank_id or self.default_bank_id
payload = build_memory_retain_payload(
experiment=experiment,
outcome=outcome,
lesson=lesson,
project_id=experiment.project_id,
)
return retain_adapter.retain_experiment_outcome(payload, bank_id=target_bank)
How the Measured Experiment Becomes Retained Memory
- Measurement & Classification: compute_experiment_outcome calculates the empirical lift (1.80% → 2.70%, +0.90 pp, +50.0% relative change) and deterministically classifies it as POSITIVE.
- Contextual Lesson Synthesis: generate_experiment_lesson frames the finding as contextual evidence rather than a universal claim (stating that CTR improved on this specific hiking-boot page under purchase-intent title alignment).
- Persistent Ingestion: build_memory_retain_payload attaches identifier seo_exp_021_demo and structures the observation into natural-language memory, which retain_adapter.retain_experiment_outcome commits into Hindsight so future SEO recall queries retrieve it as an active precedent.
Purpose: Show the exact code that records the measured experiment outcome into memory.
What I Learned
The main lesson was that agent memory is useful when it changes the decision process, not simply when it stores more information.
Retain gives the system a durable record of what happened. Recall makes relevant historical evidence available at decision time. Reflect provides a way to synthesize that evidence into contextual reasoning.
The other important lesson was evaluation discipline. Keeping a deterministic evaluator separate from the live Hindsight integration made it possible to test the complete product workflow without depending on model availability or API quota. The live integration test could then verify the actual Hindsight lifecycle independently.
The current test suite includes 118 deterministic tests, with the live Hindsight lifecycle verified separately.
Limitations
This system does not prove that historical correlation is causation. A positive or negative CTR result can have multiple contributing factors, so a retained experiment should be treated as evidence rather than universal truth.
The demo dataset is also synthetic. The recommendation and measured result demonstrate the memory workflow and decision loop; they are not a claim about real-world search performance across all websites or industries.
The confidence value is part of the recommendation workflow, not a statistical probability that the recommendation will succeed.
Conclusion
The goal of YourWebMind is straightforward: make an SEO agent remember what happened before it makes the next recommendation.
The useful difference is the closed loop. Current data provides the signal. Hindsight Recall brings back relevant precedent. Reflect helps reason over that evidence. A human approves the proposed experiment. The measured outcome is then retained so that the next decision has more context than the previous one.
For me, the practical value of Hindsight was not simply persistent storage. It was the ability to turn previous experiments—including failed ones—into part of the agent’s future decision context.
Required Hindsight References
Hindsight GitHub:
Hindsight Documentation: https://hindsight.vectorize.io/
Vectorize — What is Agent Memory?: https://vectorize.io/what-is-agent-memory



Top comments (0)