<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Poojitha Boinapalli</title>
    <description>The latest articles on DEV Community by Poojitha Boinapalli (@poojitha_boinapalli_4c094).</description>
    <link>https://dev.to/poojitha_boinapalli_4c094</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148923%2F7ee5e647-2775-44cf-9b7e-f46f1cc0e153.png</url>
      <title>DEV Community: Poojitha Boinapalli</title>
      <link>https://dev.to/poojitha_boinapalli_4c094</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/poojitha_boinapalli_4c094"/>
    <language>en</language>
    <item>
      <title>Why My Audit Agent Needed Hindsight, Not More Prompts</title>
      <dc:creator>Poojitha Boinapalli</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:22:57 +0000</pubDate>
      <link>https://dev.to/poojitha_boinapalli_4c094/why-my-audit-agent-needed-hindsight-not-more-prompts-47dd</link>
      <guid>https://dev.to/poojitha_boinapalli_4c094/why-my-audit-agent-needed-hindsight-not-more-prompts-47dd</guid>
      <description>&lt;p&gt;A dark-pattern detector can find a suspicious button today. The harder&lt;br&gt;
problem is making sure it remembers why a human reviewer said that&lt;br&gt;
button was harmless yesterday---and still notices when a previously&lt;br&gt;
fixed problem comes back.&lt;br&gt;
I built an audit agent around that problem. It analyzes online-shop&lt;br&gt;
pages for five classes of dark patterns, lets a human reviewer confirm&lt;br&gt;
or reject findings, and uses Hindsight as persistent agent memory so&lt;br&gt;
those decisions can influence later audits.&lt;br&gt;
The interesting part was not getting an LLM to recognize a hidden fee.&lt;br&gt;
It was designing the boundary between memory and current evidence.&lt;br&gt;
The Problem With Stateless Audits&lt;br&gt;
A stateless auditor has a simple failure mode: every run starts from&lt;br&gt;
zero.&lt;br&gt;
Suppose an online store contains a legitimate banner saying that a&lt;br&gt;
Diwali sale ends on a real calendar date. An LLM may find the language&lt;br&gt;
suspicious because it contains phrases such as "ENDS SOON." A human&lt;br&gt;
reviewer can decide that it is a legitimate fixed-date promotion.&lt;br&gt;
If the next audit forgets that decision, the same false alarm comes&lt;br&gt;
back.&lt;br&gt;
The opposite problem is more dangerous. Imagine a hidden checkout fee is&lt;br&gt;
removed in one version of a site and returns in the next. A system that&lt;br&gt;
only asks, "What looks suspicious on this page?" sees another hidden&lt;br&gt;
fee. It does not know that this is a regression.&lt;br&gt;
That led me to a different design:&lt;br&gt;
Recall before auditing. Retain after reviewing.&lt;br&gt;
The agent gets historical context before it analyzes a page, and every&lt;br&gt;
human decision becomes persistent memory that can affect future audits.&lt;br&gt;
What I Built&lt;br&gt;
The system has a FastAPI backend, a React/Vite frontend, Groq for&lt;br&gt;
structured LLM analysis, Playwright for runtime browser observations,&lt;br&gt;
and Hindsight for persistent memory.&lt;br&gt;
The backend exposes three main operations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;POST /audit recalls memory, observes the page, runs the LLM audit,
classifies findings by version, retains an audit summary, and stores
local history.&lt;/li&gt;
&lt;li&gt;POST /review receives confirmed, false_alarm, or accepted
and retains that decision in Hindsight.&lt;/li&gt;
&lt;li&gt;GET /history returns finding counts by version.
For the demonstration, I use three versions of a fictional Indian
shopping site, UrbanKart.
Version 1 contains fake urgency, a hidden convenience fee, a pre-checked
paid add-on, confirm-shaming language, and a hard-to-cancel
subscription. It also contains a legitimate Diwali sale banner that
deliberately looks suspicious.
Version 2 removes the hidden fee and pre-checked add-on.
Version 3 brings the hidden fee back.
That gives the system something important to reason about: not just what
exists now, but what changed.
Recall Before the LLM
The /audit endpoint starts by recalling historical memory before
calling the auditor.
recalled_memories, recall_warning = await recall_audit_memories(bank_id=bank_id)
if recall_warning:
warnings.append(recall_warning)
The memory service performs two recall queries:
queries = [
"reviewer decisions and false alarms",
"issues fixed in earlier audits"
]&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;for query in queries:&lt;br&gt;
    resp = await client.arecall(&lt;br&gt;
        bank_id=target_bank,&lt;br&gt;
        query=query,&lt;br&gt;
        max_tokens=2048&lt;br&gt;
    )&lt;br&gt;
I separate reviewer decisions from earlier audit information because&lt;br&gt;
they are different kinds of context. A reviewer decision is direct human&lt;br&gt;
judgment; an audit summary is historical context.&lt;br&gt;
The recalled results are then converted into explicit instructions for&lt;br&gt;
the LLM. The agent is told that the most recent reviewer decision is&lt;br&gt;
authoritative and that an earlier decision must not suppress a different&lt;br&gt;
element unless its evidence matches.&lt;br&gt;
Memory Is Not Permission to Ignore the Page&lt;br&gt;
This became one of the most important rules in the implementation.&lt;br&gt;
The audit prompt says:&lt;br&gt;
Audit strictly and ONLY what is currently present in the provided&lt;br&gt;
HTML and dynamic observations. Never report an issue that does not&lt;br&gt;
exist in the current page just because it was mentioned in past memories.&lt;br&gt;
It also adds a cross-version evidence rule:&lt;br&gt;
Never let a decision about one version's evidence suppress a finding&lt;br&gt;
in another version UNLESS the evidence text matches.&lt;br&gt;
I enforce the rule in Python after the LLM responds:&lt;br&gt;
if should_suppress_finding(&lt;br&gt;
    finding_type=f_type,&lt;br&gt;
    evidence=f_ev,&lt;br&gt;
    current_version=current_version,&lt;br&gt;
    authoritative_decisions=authoritative_decisions&lt;br&gt;
):&lt;br&gt;
    continue&lt;/p&gt;

&lt;p&gt;filtered_findings.append(f)&lt;br&gt;
This second check is important because I don't want the safety rule to&lt;br&gt;
exist only inside a prompt.&lt;br&gt;
A false_alarm or accepted decision can suppress a matching finding,&lt;br&gt;
but if the current version contains different evidence, the old decision&lt;br&gt;
does not automatically suppress it.&lt;br&gt;
So memory provides context; current evidence remains the authority for&lt;br&gt;
what is actually on the page.&lt;br&gt;
Retaining Human Judgment&lt;br&gt;
After a reviewer clicks a decision in the frontend, /review writes a&lt;br&gt;
plain-English memory.&lt;br&gt;
For a false_alarm:&lt;br&gt;
sentence = (&lt;br&gt;
    f"Reviewer marked {finding_type} in {version} "&lt;br&gt;
    f"as a false alarm{ev_part}.{note_part}"&lt;br&gt;
)&lt;br&gt;
For a confirmed finding:&lt;br&gt;
sentence = (&lt;br&gt;
    f"Reviewer confirmed {finding_type} in {version} "&lt;br&gt;
    f"as a real dark pattern{ev_part}.{note_part}"&lt;br&gt;
)&lt;br&gt;
The retained information includes an evidence snippet when available.&lt;br&gt;
The actual Hindsight write is small:&lt;br&gt;
await client.aretain(&lt;br&gt;
    bank_id=target_bank,&lt;br&gt;
    content=sentence,&lt;br&gt;
    metadata=metadata&lt;br&gt;
)&lt;br&gt;
I did not want to store an opaque application-state blob and call that&lt;br&gt;
"memory." The retained information should be understandable when&lt;br&gt;
recalled: what the reviewer decided, on which version, about which&lt;br&gt;
finding, and with what evidence.&lt;br&gt;
The system also handles corrections. If a reviewer changes a decision,&lt;br&gt;
it retains an explicit correction sentence such as "Reviewer changed&lt;br&gt;
decision on ... from accepted to confirmed." The memory parser&lt;br&gt;
recognizes these corrections so the newer decision can become&lt;br&gt;
authoritative.&lt;br&gt;
Detecting a Regression&lt;br&gt;
Hindsight provides memory, but it does not by itself decide whether&lt;br&gt;
something is a regression.&lt;br&gt;
I keep version ordering explicitly:&lt;br&gt;
VERSION_ORDER = ["store_v1", "store_v2", "store_v3"]&lt;br&gt;
The backend compares versions according to that order rather than&lt;br&gt;
relying on the order in which someone happened to click the UI.&lt;br&gt;
The classification is:&lt;br&gt;
if f_type in earlier_fixed_set:&lt;br&gt;
    status = "regression"&lt;br&gt;
elif f_type in prev_version_types:&lt;br&gt;
    status = "still present"&lt;br&gt;
else:&lt;br&gt;
    status = "new"&lt;br&gt;
That creates three useful states:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;NEW: the pattern was not present in the immediately previous
version.&lt;/li&gt;
&lt;li&gt;STILL PRESENT: it existed in the previous version and remains.&lt;/li&gt;
&lt;li&gt;REGRESSION: it disappeared in an earlier version and has returned.
The system also computes issues fixed since the immediately previous
version, which are shown separately in the UI.
This changes the meaning of an audit. Instead of producing only a list
of warnings, the system produces a history of change.
Adding Runtime Evidence
HTML alone is not always enough.
A countdown timer can look suspicious in static HTML, but stronger
evidence comes from observing what happens when the page actually runs.
The backend uses Playwright to observe the local sample page:
target_url = f"&lt;a href="http://localhost:9000/%7Bbase_name%7D" rel="noopener noreferrer"&gt;http://localhost:9000/{base_name}&lt;/a&gt;"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;try:&lt;br&gt;
    observations = await asyncio.to_thread(observe, target_url)&lt;br&gt;
except Exception as e:&lt;br&gt;
    warnings.append(&lt;br&gt;
        f"Dynamic observation via Playwright failed ({type(e).&lt;strong&gt;name&lt;/strong&gt;}: {e}); "&lt;br&gt;
        "fell back to HTML-only audit."&lt;br&gt;
    )&lt;br&gt;
Those observations are passed to the LLM alongside the HTML and recalled&lt;br&gt;
memory.&lt;br&gt;
The prompt treats browser observations as runtime facts---for example,&lt;br&gt;
whether a checkbox is actually checked on initial load or whether a&lt;br&gt;
timer behaves consistently across page loads.&lt;br&gt;
The agent therefore combines three kinds of information:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What the current HTML contains.&lt;/li&gt;
&lt;li&gt;What the browser actually does.&lt;/li&gt;
&lt;li&gt;What previous human reviewers decided.
The third one is where Hindsight changes the behavior of the system over
time.
The UrbanKart Sequence
Version 1: Start With No Memory
The first audit starts with an empty or nearly empty memory bank.
The agent can identify the planted patterns, including the
suspicious-looking Diwali banner.
A human reviewer then marks the banner as a false_alarm and confirms
genuine problems.
That decision is retained in Hindsight with its evidence.
Version 2: The Agent Remembers
Version 2 fixes the hidden fee and pre-checked add-on.
When the audit runs again, the agent recalls previous reviewer decisions
and historical audit information.
The fixed issues can be reported separately, while the previously
dismissed Diwali banner is not repeatedly reported when its evidence
matches the remembered false-alarm decision.
This is the point where persistent memory becomes visible. The agent is
no longer producing another independent LLM response.
Version 3: The Problem Comes Back
Then the hidden fee returns.
The system has historical evidence that the issue was absent from the
previous version and was fixed earlier. When the finding appears again,
its status becomes REGRESSION.
That is more useful than a generic "hidden fee detected."
The question changes from:
"What is wrong with this page?"&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;to:&lt;br&gt;
"What changed, what was previously fixed, and what has come back?"&lt;/p&gt;

&lt;p&gt;Keeping Tests Away From Real Memory&lt;br&gt;
Persistent memory introduces another engineering concern: tests should&lt;br&gt;
not pollute production memory.&lt;br&gt;
The project detects test execution and redirects automated test data to&lt;br&gt;
a dedicated urbankart-test bank:&lt;br&gt;
if is_test_environment() or is_test_data or target_bank == TEST_BANK_ID:&lt;br&gt;
    return TEST_BANK_ID&lt;br&gt;
The tests cover bank isolation, conflicting reviewer decisions,&lt;br&gt;
cross-version evidence matching, and mock Hindsight retention/recall.&lt;br&gt;
One test protects a particularly important rule: a reviewer decision&lt;br&gt;
about the Diwali banner in store_v1 must not suppress a different&lt;br&gt;
countdown timer in store_v2 merely because both are classified as fake&lt;br&gt;
urgency.&lt;br&gt;
The evidence has to match.&lt;br&gt;
If Hindsight Goes Down&lt;br&gt;
I also did not want a memory-service failure to make the auditor&lt;br&gt;
unusable.&lt;br&gt;
Recall and retain operations return warnings rather than forcing the&lt;br&gt;
entire audit to fail. If Hindsight cannot be reached, the audit can&lt;br&gt;
continue and the frontend can expose a warning.&lt;br&gt;
Memory improves the audit; it should not become a single point of&lt;br&gt;
failure for the audit itself.&lt;br&gt;
What I Learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Memory only matters when it changes behavior
Adding a memory database to an agent is easy to describe and easy to
make useless.
Here, memory causes observable behavior: suppressing a reviewed false
alarm, retaining corrections, and providing historical context for
regression classification.&lt;/li&gt;
&lt;li&gt;Evidence needs to travel with the decision
"Fake urgency is a false alarm" is too broad.
A decision tied to evidence is safer: this particular banner, with this
particular text, was reviewed as legitimate.
That is why reviewer memories include evidence snippets.&lt;/li&gt;
&lt;li&gt;Version order should be explicit
Audit execution order is not version order.
Someone can run store_v3 before store_v2, rerun an old version, or
audit the same version multiple times. Regression logic should still
follow the site's actual version sequence.&lt;/li&gt;
&lt;li&gt;Prompts should not be the only enforcement layer
The LLM receives detailed memory rules, but the application also
verifies suppression decisions in Python.
That gives the system another layer of protection against an incorrect
memory interpretation.&lt;/li&gt;
&lt;li&gt;Human review is part of the system
The reviewer is not simply correcting the model once. Their decisions
become persistent context for future audits.
The loop is:
audit → human judgment → retain → recall → next audit
Conclusion
The most useful part of this project was not making an LLM recognize
dark patterns. Models can already recognize suspicious interfaces
reasonably well.
The harder engineering problem was giving the agent memory without
giving that memory unlimited authority.
Hindsight became the persistent layer for reviewer decisions and audit
history. The application then adds the rules around that memory:
evidence must match before an old decision suppresses a new finding,
version order determines regressions, tests stay isolated from the real
bank, and an unavailable memory service does not stop an audit.
The result is an auditor that does more than ask what is suspicious on a
page right now. It can remember what a reviewer decided, distinguish a
legitimate design from a previously dismissed false alarm, and recognize
when a problem that was fixed has returned.
For this kind of agent, memory is not an optional feature. It is part of
the reasoning loop.
Hindsight on GitHub
Hindsight documentation
What is agent memory? ---
Vectorize&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>software</category>
    </item>
  </channel>
</rss>
