A dark-pattern detector can find a suspicious button today. The harder
problem is making sure it remembers why a human reviewer said that
button was harmless yesterday---and still notices when a previously
fixed problem comes back.
I built an audit agent around that problem. It analyzes online-shop
pages for five classes of dark patterns, lets a human reviewer confirm
or reject findings, and uses Hindsight as persistent agent memory so
those decisions can influence later audits.
The interesting part was not getting an LLM to recognize a hidden fee.
It was designing the boundary between memory and current evidence.
The Problem With Stateless Audits
A stateless auditor has a simple failure mode: every run starts from
zero.
Suppose an online store contains a legitimate banner saying that a
Diwali sale ends on a real calendar date. An LLM may find the language
suspicious because it contains phrases such as "ENDS SOON." A human
reviewer can decide that it is a legitimate fixed-date promotion.
If the next audit forgets that decision, the same false alarm comes
back.
The opposite problem is more dangerous. Imagine a hidden checkout fee is
removed in one version of a site and returns in the next. A system that
only asks, "What looks suspicious on this page?" sees another hidden
fee. It does not know that this is a regression.
That led me to a different design:
Recall before auditing. Retain after reviewing.
The agent gets historical context before it analyzes a page, and every
human decision becomes persistent memory that can affect future audits.
What I Built
The system has a FastAPI backend, a React/Vite frontend, Groq for
structured LLM analysis, Playwright for runtime browser observations,
and Hindsight for persistent memory.
The backend exposes three main operations:
- POST /audit recalls memory, observes the page, runs the LLM audit, classifies findings by version, retains an audit summary, and stores local history.
- POST /review receives confirmed, false_alarm, or accepted and retains that decision in Hindsight.
- GET /history returns finding counts by version. For the demonstration, I use three versions of a fictional Indian shopping site, UrbanKart. Version 1 contains fake urgency, a hidden convenience fee, a pre-checked paid add-on, confirm-shaming language, and a hard-to-cancel subscription. It also contains a legitimate Diwali sale banner that deliberately looks suspicious. Version 2 removes the hidden fee and pre-checked add-on. Version 3 brings the hidden fee back. That gives the system something important to reason about: not just what exists now, but what changed. Recall Before the LLM The /audit endpoint starts by recalling historical memory before calling the auditor. recalled_memories, recall_warning = await recall_audit_memories(bank_id=bank_id) if recall_warning: warnings.append(recall_warning) The memory service performs two recall queries: queries = [ "reviewer decisions and false alarms", "issues fixed in earlier audits" ]
for query in queries:
resp = await client.arecall(
bank_id=target_bank,
query=query,
max_tokens=2048
)
I separate reviewer decisions from earlier audit information because
they are different kinds of context. A reviewer decision is direct human
judgment; an audit summary is historical context.
The recalled results are then converted into explicit instructions for
the LLM. The agent is told that the most recent reviewer decision is
authoritative and that an earlier decision must not suppress a different
element unless its evidence matches.
Memory Is Not Permission to Ignore the Page
This became one of the most important rules in the implementation.
The audit prompt says:
Audit strictly and ONLY what is currently present in the provided
HTML and dynamic observations. Never report an issue that does not
exist in the current page just because it was mentioned in past memories.
It also adds a cross-version evidence rule:
Never let a decision about one version's evidence suppress a finding
in another version UNLESS the evidence text matches.
I enforce the rule in Python after the LLM responds:
if should_suppress_finding(
finding_type=f_type,
evidence=f_ev,
current_version=current_version,
authoritative_decisions=authoritative_decisions
):
continue
filtered_findings.append(f)
This second check is important because I don't want the safety rule to
exist only inside a prompt.
A false_alarm or accepted decision can suppress a matching finding,
but if the current version contains different evidence, the old decision
does not automatically suppress it.
So memory provides context; current evidence remains the authority for
what is actually on the page.
Retaining Human Judgment
After a reviewer clicks a decision in the frontend, /review writes a
plain-English memory.
For a false_alarm:
sentence = (
f"Reviewer marked {finding_type} in {version} "
f"as a false alarm{ev_part}.{note_part}"
)
For a confirmed finding:
sentence = (
f"Reviewer confirmed {finding_type} in {version} "
f"as a real dark pattern{ev_part}.{note_part}"
)
The retained information includes an evidence snippet when available.
The actual Hindsight write is small:
await client.aretain(
bank_id=target_bank,
content=sentence,
metadata=metadata
)
I did not want to store an opaque application-state blob and call that
"memory." The retained information should be understandable when
recalled: what the reviewer decided, on which version, about which
finding, and with what evidence.
The system also handles corrections. If a reviewer changes a decision,
it retains an explicit correction sentence such as "Reviewer changed
decision on ... from accepted to confirmed." The memory parser
recognizes these corrections so the newer decision can become
authoritative.
Detecting a Regression
Hindsight provides memory, but it does not by itself decide whether
something is a regression.
I keep version ordering explicitly:
VERSION_ORDER = ["store_v1", "store_v2", "store_v3"]
The backend compares versions according to that order rather than
relying on the order in which someone happened to click the UI.
The classification is:
if f_type in earlier_fixed_set:
status = "regression"
elif f_type in prev_version_types:
status = "still present"
else:
status = "new"
That creates three useful states:
- NEW: the pattern was not present in the immediately previous version.
- STILL PRESENT: it existed in the previous version and remains.
- REGRESSION: it disappeared in an earlier version and has returned. The system also computes issues fixed since the immediately previous version, which are shown separately in the UI. This changes the meaning of an audit. Instead of producing only a list of warnings, the system produces a history of change. Adding Runtime Evidence HTML alone is not always enough. A countdown timer can look suspicious in static HTML, but stronger evidence comes from observing what happens when the page actually runs. The backend uses Playwright to observe the local sample page: target_url = f"http://localhost:9000/{base_name}"
try:
observations = await asyncio.to_thread(observe, target_url)
except Exception as e:
warnings.append(
f"Dynamic observation via Playwright failed ({type(e).name}: {e}); "
"fell back to HTML-only audit."
)
Those observations are passed to the LLM alongside the HTML and recalled
memory.
The prompt treats browser observations as runtime facts---for example,
whether a checkbox is actually checked on initial load or whether a
timer behaves consistently across page loads.
The agent therefore combines three kinds of information:
- What the current HTML contains.
- What the browser actually does.
- What previous human reviewers decided. The third one is where Hindsight changes the behavior of the system over time. The UrbanKart Sequence Version 1: Start With No Memory The first audit starts with an empty or nearly empty memory bank. The agent can identify the planted patterns, including the suspicious-looking Diwali banner. A human reviewer then marks the banner as a false_alarm and confirms genuine problems. That decision is retained in Hindsight with its evidence. Version 2: The Agent Remembers Version 2 fixes the hidden fee and pre-checked add-on. When the audit runs again, the agent recalls previous reviewer decisions and historical audit information. The fixed issues can be reported separately, while the previously dismissed Diwali banner is not repeatedly reported when its evidence matches the remembered false-alarm decision. This is the point where persistent memory becomes visible. The agent is no longer producing another independent LLM response. Version 3: The Problem Comes Back Then the hidden fee returns. The system has historical evidence that the issue was absent from the previous version and was fixed earlier. When the finding appears again, its status becomes REGRESSION. That is more useful than a generic "hidden fee detected." The question changes from: "What is wrong with this page?"
to:
"What changed, what was previously fixed, and what has come back?"
Keeping Tests Away From Real Memory
Persistent memory introduces another engineering concern: tests should
not pollute production memory.
The project detects test execution and redirects automated test data to
a dedicated urbankart-test bank:
if is_test_environment() or is_test_data or target_bank == TEST_BANK_ID:
return TEST_BANK_ID
The tests cover bank isolation, conflicting reviewer decisions,
cross-version evidence matching, and mock Hindsight retention/recall.
One test protects a particularly important rule: a reviewer decision
about the Diwali banner in store_v1 must not suppress a different
countdown timer in store_v2 merely because both are classified as fake
urgency.
The evidence has to match.
If Hindsight Goes Down
I also did not want a memory-service failure to make the auditor
unusable.
Recall and retain operations return warnings rather than forcing the
entire audit to fail. If Hindsight cannot be reached, the audit can
continue and the frontend can expose a warning.
Memory improves the audit; it should not become a single point of
failure for the audit itself.
What I Learned
- Memory only matters when it changes behavior Adding a memory database to an agent is easy to describe and easy to make useless. Here, memory causes observable behavior: suppressing a reviewed false alarm, retaining corrections, and providing historical context for regression classification.
- Evidence needs to travel with the decision "Fake urgency is a false alarm" is too broad. A decision tied to evidence is safer: this particular banner, with this particular text, was reviewed as legitimate. That is why reviewer memories include evidence snippets.
- Version order should be explicit Audit execution order is not version order. Someone can run store_v3 before store_v2, rerun an old version, or audit the same version multiple times. Regression logic should still follow the site's actual version sequence.
- Prompts should not be the only enforcement layer The LLM receives detailed memory rules, but the application also verifies suppression decisions in Python. That gives the system another layer of protection against an incorrect memory interpretation.
- Human review is part of the system The reviewer is not simply correcting the model once. Their decisions become persistent context for future audits. The loop is: audit → human judgment → retain → recall → next audit Conclusion The most useful part of this project was not making an LLM recognize dark patterns. Models can already recognize suspicious interfaces reasonably well. The harder engineering problem was giving the agent memory without giving that memory unlimited authority. Hindsight became the persistent layer for reviewer decisions and audit history. The application then adds the rules around that memory: evidence must match before an old decision suppresses a new finding, version order determines regressions, tests stay isolated from the real bank, and an unavailable memory service does not stop an audit. The result is an auditor that does more than ask what is suspicious on a page right now. It can remember what a reviewer decided, distinguish a legitimate design from a previously dismissed false alarm, and recognize when a problem that was fixed has returned. For this kind of agent, memory is not an optional feature. It is part of the reasoning loop. Hindsight on GitHub Hindsight documentation What is agent memory? --- Vectorize
Top comments (0)