Our audit agent flagged the same harmless sale banner every single time. It also had no idea that a hidden checkout fee we'd "fixed" had quietly returned. Same agent, same model, and it started every audit with amnesia. Adding Hindsight agent memory fixed both problems.
What the system does
Dark patterns are design tricks that push shoppers into paying or staying: countdown timers that reset, fees that appear only at the last step, pre-ticked paid add-ons, buttons that read "No, I don't want to save money," and cancel links buried in grey text. Regulators are paying attention to them, so companies need to audit their own pages.
We built an audit agent for a fictional Indian shop, UrbanKart, with three versions of its pages. The stack is small:
A FastAPI backend that sends page HTML to an LLM on Groq and asks for structured findings.
A React frontend where a human reviewer marks each finding as confirmed, false alarm, or accepted design.
Hindsight as the memory layer between audits.
A stateless version of this agent is easy to build and mediocre to use. It detects tricks reasonably well, but it can't learn.
The problem with a stateless auditor
store_v1 contains five planted tricks plus one legitimate element that looks suspicious to a model reading raw HTML. Without memory, this happens on every single run:
The agent flags the legitimate element as a problem.
The reviewer says it's fine.
The next audit flags it again, because nothing was remembered.
The second failure is worse. store_v2 fixes two issues, and store_v3 reintroduces one of them. A stateless agent audits v3 as if v2 never happened, so it can't say "this came back."
The design: recall before, retain after
The whole integration comes down to two habits. Before each audit, the agent recalls what it knows about this shop. After each review, it saves what the human decided. We use one memory bank per site, so all three versions of UrbanKart share the same history.
Saving a reviewer decision is one plain sentence, built from the decision type and the exact evidence the agent found:
python
if dec == "confirmed":
sentence = f"Reviewer confirmed {finding_type} in {version} as a real dark pattern{ev_part}.{note_part}"
elif dec == "false_alarm":
sentence = f"Reviewer marked {finding_type} in {version} as a false alarm{ev_part}.{note_part}"
elif dec == "accepted":
sentence = f"Reviewer accepted {finding_type} in {version} as a legitimate design pattern{ev_part}.{note_part}"
await client.aretain(
bank_id=target_bank,
content=sentence,
metadata={"version": version, "finding_type": finding_type, "decision": dec},
)
Before an audit, we pull relevant memories with two targeted queries and feed the results into the prompt:
python
queries = [
"reviewer decisions and false alarms",
"issues fixed in earlier audits",
]
for query in queries:
resp = await client.arecall(bank_id=target_bank, query=query, max_tokens=2048)
for item in resp.results:
recalled_texts.append(item.text)
Audit summaries are retained the same way, but with the version order spelled out explicitly, so recall never has to guess which version came first:
python
summary = (
f"Audit of {version} (later version than {prev_version}) for {target_bank}: "
f"found issues: {current_str}. "
f"Issues fixed compared to earlier version {prev_version}: {fixed_str}."
)
if regression_types:
summary += (
f" Regression alert: {', '.join(regression_types)} was fixed in an "
f"earlier version but reappeared in later version {version}."
)
The prompt then gives the model three rules: don't report anything a reviewer already marked as a false alarm or accepted; label a finding a regression if it was fixed in an earlier version and is present again; otherwise call it new.
I chose free-text sentences over a database with columns because reviewer notes are messy ("real sale, ends next month, not a fake countdown"). This is documented on the Hindsight docs site, where retain and recall are described as separate operations that don't require hand-written matching rules — a note about one banner can still match a similar banner elsewhere because recall works on meaning, not exact text. Each stored sentence also carries structured metadata (version, finding type, decision) alongside the free text, so the memory is both human-readable and filterable later if we need it.
What it does now
Audit 1, store_v1, empty memory: 5 findings, 0 memory records to start. Each finding carries evidence pulled straight from the page, not a guess — including a timer that showed the same value on two separate page loads:
I mark the legitimate-looking element a false alarm and confirm the rest.
Audit 2, store_v2: the memory panel shows the reviewer's earlier decision, and the false alarm is no longer flagged. The agent also notes which issues are now fixed.
The version history view makes the improvement easy to see at a glance:
Audit 3, store_v3: a previously fixed issue returns, and the agent labels it a regression, because memory says it was fixed in the previous version.
I made the memory panel visible in the interface on purpose. When you can see which memories were used for a result, the agent's behavior stops looking like luck.
Lessons learned
Test data will contaminate production memory if you let it. Early on, calls made while building and testing the /review endpoint were writing straight into our real memory bank. The agent later recalled contradictory decisions about the same issue — accepted in one memory, confirmed in another — which could easily have suppressed a genuine regression. The fix was a small guard function that inspects every incoming review call for test markers (a pytest environment, a TESTING flag, or keywords like "unit test" in the note or evidence) and redirects anything suspicious to a separate urbankart-test bank, regardless of what bank id was requested. If your code can be called by both a human reviewer and an automated test, assume it eventually will be, and defend the real bank accordingly.
"Change the bank name" doesn't always mean "start clean." We assumed switching bank ids gave us a fresh slate for each demo run. It usually did — but every accidental re-run of an already-reviewed version, and every partially-applied fix, still landed in the same bank and compounded. The practical lesson: treat a bank as disposable and easy to respin, but never re-run an audit you've already reviewed in that bank. If you need to look at a result again, read it back — don't regenerate it.
Compare by version order, not by run order. Our first version compared each audit against whichever audit had run immediately before it. After a few retries during testing, a later version ended up being compared with itself, and regressions silently disappeared. The fix was retaining explicit language — "later version than X", "earlier version" — directly in the summary text, so recall doesn't have to infer ordering from timestamps.
The agent's language model doesn't always make the same mistake twice. A page element we expected to be flagged as a false positive on every run sometimes wasn't. That's a property of language models, not a bug: if you want a repeatable "false alarm corrected by memory" moment for a demo, the ambiguous element needs to be genuinely ambiguous, not just occasionally so.
Reading a saved HTML file only gets you so far. Real dark patterns are often dynamic — a timer that resets on reload, or a fee that only shows up at the final checkout step, doesn't show in a static snapshot. We added a small browser-driven check that loads a page twice and compares the countdown value between loads, which turns "this looks like fake urgency" into "the timer showed 09:59 on load 1 and 09:59 on load 2" — an observation, not a guess. Our test pages are still synthetic, so we treat all of this as a controlled test of the memory mechanism, not a claim about any real shop.
Why I'd use agent memory again
The tricks the agent finds are the easy part. The value is in what it stops doing (repeating false alarms) and what it starts doing (noticing regressions). If you're building an agent that gets reviewed by humans, agent memory turns each review into something the next run benefits from. The Hindsight quickstart gets you a working retain/recall loop in a few lines.









Top comments (0)