<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: shiva mani</title>
    <description>The latest articles on DEV Community by shiva mani (@mani_hp_87012ecdf3e415034).</description>
    <link>https://dev.to/mani_hp_87012ecdf3e415034</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150195%2F43f9a3d6-ce2c-4ccb-a370-4636d57554fc.png</url>
      <title>DEV Community: shiva mani</title>
      <link>https://dev.to/mani_hp_87012ecdf3e415034</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mani_hp_87012ecdf3e415034"/>
    <language>en</language>
    <item>
      <title>I Gave a CAPTCHA Memory Instead of Making the Model Bigger</title>
      <dc:creator>shiva mani</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:10:33 +0000</pubDate>
      <link>https://dev.to/mani_hp_87012ecdf3e415034/i-gave-a-captcha-memory-instead-of-making-the-model-bigger-434j</link>
      <guid>https://dev.to/mani_hp_87012ecdf3e415034/i-gave-a-captcha-memory-instead-of-making-the-model-bigger-434j</guid>
      <description>&lt;p&gt;A CAPTCHA normally gets one chance to decide whether an interaction is human.&lt;/p&gt;

&lt;p&gt;That bothered me.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsbulj3d1b3vynqaj0yfr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsbulj3d1b3vynqaj0yfr.png" alt=" " width="800" height="858"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If someone interacts with a CAPTCHA once, the system has some behavioral evidence. But when the same kind of interaction happens again, most CAPTCHA systems don't really remember what happened before. Every attempt starts from zero.&lt;/p&gt;

&lt;p&gt;So while working on SwipeCHA, I tried a different approach: keep the existing behavioral classifier, but give the security system memory.&lt;/p&gt;

&lt;p&gt;That is where Hindsight came in.&lt;/p&gt;

&lt;p&gt;The part that was already working&lt;/p&gt;

&lt;p&gt;SwipeCHA is a behavioral CAPTCHA. Instead of asking users to identify objects or type distorted characters, it asks them to slide a handle across a track.&lt;/p&gt;

&lt;p&gt;While the slider is being moved, the browser records pointer movement and timing information. Those events are converted into ten features:&lt;/p&gt;

&lt;p&gt;avg_mouse_speed&lt;br&gt;
mouse_path_entropy&lt;br&gt;
click_delay&lt;br&gt;
task_completion_time&lt;br&gt;
idle_time&lt;br&gt;
micro_jitter_variance&lt;br&gt;
acceleration_curve&lt;br&gt;
curvature_variance&lt;br&gt;
overshoot_correction_ratio&lt;br&gt;
timing_entropy&lt;/p&gt;

&lt;p&gt;Those features go to the existing Random Forest classifier.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;I didn't start by assuming the model needed to become an LLM agent. The classifier already has a clear job: evaluate the current behavioral signal.&lt;/p&gt;

&lt;p&gt;So I kept it.&lt;/p&gt;

&lt;p&gt;The problem I wanted to solve was everything around that single prediction.&lt;/p&gt;

&lt;p&gt;The problem with starting from zero&lt;/p&gt;

&lt;p&gt;The original security path was essentially:&lt;/p&gt;

&lt;p&gt;Swipe&lt;br&gt;
  ↓&lt;br&gt;
Feature extraction&lt;br&gt;
  ↓&lt;br&gt;
Hard-rule checks&lt;br&gt;
  ↓&lt;br&gt;
Random Forest&lt;br&gt;
  ↓&lt;br&gt;
Human / Bot&lt;/p&gt;

&lt;p&gt;Every verification was mostly an isolated event.&lt;/p&gt;

&lt;p&gt;Suppose the first interaction looks human. The next interaction also looks human. A third one is slightly unusual, but not unusual enough to trigger the hard rules.&lt;/p&gt;

&lt;p&gt;Looking at only the third interaction gives the system one piece of evidence.&lt;/p&gt;

&lt;p&gt;Looking at all three gives it a pattern.&lt;/p&gt;

&lt;p&gt;I wanted the second kind of system.&lt;/p&gt;

&lt;p&gt;Why I didn't replace the Random Forest&lt;/p&gt;

&lt;p&gt;The obvious architecture would have been:&lt;/p&gt;

&lt;p&gt;CAPTCHA → LLM → Human / Bot&lt;/p&gt;

&lt;p&gt;I rejected that approach.&lt;/p&gt;

&lt;p&gt;The behavioral features are numerical, the existing model is trained specifically for them, and throwing that away would make the system harder to reason about rather than better.&lt;/p&gt;

&lt;p&gt;Instead, I split the responsibilities:&lt;/p&gt;

&lt;p&gt;Random Forest — evaluates the current behavioral signal.&lt;/p&gt;

&lt;p&gt;Hindsight — provides relevant historical security experiences.&lt;/p&gt;

&lt;p&gt;Security Agent — interprets the current evidence in that historical context.&lt;/p&gt;

&lt;p&gt;Decision Policy — turns those signals into a controlled action.&lt;/p&gt;

&lt;p&gt;That separation became the central design of the project.&lt;/p&gt;

&lt;p&gt;Memory should store experiences, not mouse movements&lt;/p&gt;

&lt;p&gt;There was another decision I had to make before integrating memory.&lt;/p&gt;

&lt;p&gt;I could have stored every pointer event.&lt;/p&gt;

&lt;p&gt;I didn't want to.&lt;/p&gt;

&lt;p&gt;A raw swipe can contain a large number of coordinates and timestamps. Putting all of that into long-term memory would give the memory layer a lot of detail without necessarily giving it useful security context.&lt;/p&gt;

&lt;p&gt;Instead, the system retains a distilled security experience.&lt;/p&gt;

&lt;p&gt;A simplified example looks like this:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "prediction": "Human",&lt;br&gt;
  "confidence": 0.98,&lt;br&gt;
  "risk_level": "low",&lt;br&gt;
  "reason_codes": ["ZERO_HISTORY_BASELINE"],&lt;br&gt;
  "recommended_action": "allow"&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;The actual retained information depends on the verification path. The important idea is that the memory represents what happened from a security perspective, rather than becoming a replayable dump of the browser session.&lt;/p&gt;

&lt;p&gt;The first interaction is supposed to be boring&lt;/p&gt;

&lt;p&gt;On the first interaction there is no history.&lt;/p&gt;

&lt;p&gt;So the system should not pretend there is.&lt;/p&gt;

&lt;p&gt;The first turn is evaluated using the current behavioral evidence:&lt;/p&gt;

&lt;p&gt;Turn 1&lt;br&gt;
Hindsight memories: 0&lt;/p&gt;

&lt;p&gt;Random Forest:&lt;br&gt;
Human — 0.98&lt;/p&gt;

&lt;p&gt;Reason:&lt;br&gt;
ZERO_HISTORY_BASELINE&lt;/p&gt;

&lt;p&gt;Decision:&lt;br&gt;
Allow&lt;/p&gt;

&lt;p&gt;Then the experience is retained.&lt;/p&gt;

&lt;p&gt;That gives the next interaction something to retrieve.&lt;/p&gt;

&lt;p&gt;The interesting part happens on the next turn&lt;/p&gt;

&lt;p&gt;On a later verification, Hindsight can recall relevant previous experiences before the Security Agent evaluates the current result.&lt;/p&gt;

&lt;p&gt;Now the agent has two kinds of information:&lt;/p&gt;

&lt;p&gt;Current interaction&lt;br&gt;
        +&lt;br&gt;
Previous security experiences&lt;br&gt;
        ↓&lt;br&gt;
Contextual assessment&lt;/p&gt;

&lt;p&gt;For a consistent human interaction pattern, historical evidence can reinforce the current assessment. The implementation records this kind of context with reason codes such as HISTORICAL_HUMAN_CONSISTENCY.&lt;/p&gt;

&lt;p&gt;For suspicious or borderline behavior, previous experiences can provide context that a single swipe cannot.&lt;/p&gt;

&lt;p&gt;The important point is subtle:&lt;/p&gt;

&lt;p&gt;The Random Forest didn't suddenly become a better classifier. The decision became contextual.&lt;/p&gt;

&lt;p&gt;I tested whether the memory was actually persistent&lt;/p&gt;

&lt;p&gt;This was one of the first things I wanted to verify.&lt;/p&gt;

&lt;p&gt;If the application restarts and all the previous experiences disappear, then the system isn't demonstrating the kind of persistent memory I was trying to build.&lt;/p&gt;

&lt;p&gt;So I ran the interaction sequence, restarted the application, and ran it again.&lt;/p&gt;

&lt;p&gt;The observed sequence was:&lt;/p&gt;

&lt;p&gt;Turn 1 → 0 memories → RETAIN&lt;br&gt;
Restart&lt;br&gt;
Turn 2 → historical memory recalled → RETAIN&lt;br&gt;
Turn 3 → previous experiences recalled&lt;/p&gt;

&lt;p&gt;That was a useful checkpoint because it separated two ideas that are easy to confuse: state in a running process and persistent experience memory.&lt;/p&gt;

&lt;p&gt;The current implementation uses the official Hindsight client with a local Hindsight-compatible deployment for the development/staging environment. I am deliberately not presenting a live Hindsight Cloud deployment as verified.&lt;/p&gt;

&lt;p&gt;The agent doesn't get unlimited authority&lt;/p&gt;

&lt;p&gt;I also didn't want the Security Agent to become a black box sitting in front of the CAPTCHA.&lt;/p&gt;

&lt;p&gt;It produces structured context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;risk level&lt;/li&gt;
&lt;li&gt;historical context&lt;/li&gt;
&lt;li&gt;reason codes&lt;/li&gt;
&lt;li&gt;recommended action&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The final result still passes through a decision policy.&lt;/p&gt;

&lt;p&gt;That gives the system controlled actions such as:&lt;/p&gt;

&lt;p&gt;Allow&lt;br&gt;
Block&lt;br&gt;
Challenge Again&lt;/p&gt;

&lt;p&gt;There is also a deterministic hard-rule path for obvious automation. If the interaction clearly matches that pattern, the system can reject it without making memory retrieval part of every decision.&lt;/p&gt;

&lt;p&gt;That keeps the memory layer useful without making it a single point of failure.&lt;/p&gt;

&lt;p&gt;The part I underestimated: deciding what to remember&lt;/p&gt;

&lt;p&gt;My first instinct was to think that adding memory would be the difficult part.&lt;/p&gt;

&lt;p&gt;It wasn't.&lt;/p&gt;

&lt;p&gt;The harder question was: what is actually worth remembering?&lt;/p&gt;

&lt;p&gt;If I retained everything, the system would have more data but not necessarily more useful context.&lt;/p&gt;

&lt;p&gt;That changed the integration.&lt;/p&gt;

&lt;p&gt;I started thinking about memory in terms of security experiences rather than raw telemetry:&lt;/p&gt;

&lt;p&gt;What happened?&lt;br&gt;
What did the current classifier say?&lt;br&gt;
Was there historical context?&lt;br&gt;
What risk did the agent identify?&lt;br&gt;
What action followed?&lt;/p&gt;

&lt;p&gt;That made Hindsight part of the reasoning loop instead of just another storage service.&lt;/p&gt;

&lt;p&gt;Failure handling is part of the design&lt;/p&gt;

&lt;p&gt;A security check shouldn't stop working because its memory dependency has a bad day.&lt;/p&gt;

&lt;p&gt;So the integration has a fallback path.&lt;/p&gt;

&lt;p&gt;If Hindsight or the agent layer is unavailable or times out, the system can fall back to the existing Random Forest path.&lt;/p&gt;

&lt;p&gt;The integration also includes a circuit breaker so repeated memory-service failures don't keep adding latency to verification.&lt;/p&gt;

&lt;p&gt;This wasn't the most visually impressive part of the project, but it was one of the parts I cared about most once the architecture started looking like a real security system rather than a demo.&lt;/p&gt;

&lt;p&gt;What the final flow looks like&lt;/p&gt;

&lt;p&gt;The resulting pipeline is:&lt;/p&gt;

&lt;p&gt;User&lt;br&gt;
  ↓&lt;br&gt;
SwipeCHA&lt;br&gt;
  ↓&lt;br&gt;
Behavioral features&lt;br&gt;
  ├── Hard-rule gate&lt;br&gt;
  ↓&lt;br&gt;
Random Forest&lt;br&gt;
  ↓&lt;br&gt;
Hindsight Recall&lt;br&gt;
  ↓&lt;br&gt;
Security Agent&lt;br&gt;
  ↓&lt;br&gt;
Decision Policy&lt;br&gt;
  ├── Allow&lt;br&gt;
  ├── Block&lt;br&gt;
  └── Challenge Again&lt;br&gt;
  ↓&lt;br&gt;
Hindsight Retain&lt;/p&gt;

&lt;p&gt;The Random Forest still evaluates the live interaction.&lt;/p&gt;

&lt;p&gt;Hindsight adds historical context.&lt;/p&gt;

&lt;p&gt;The Security Agent interprets that context.&lt;/p&gt;

&lt;p&gt;The decision policy keeps the final behavior controlled.&lt;/p&gt;

&lt;p&gt;That division is more important to me than simply saying “the CAPTCHA now has AI memory.”&lt;/p&gt;

&lt;p&gt;What I would not claim&lt;/p&gt;

&lt;p&gt;There are a few things I don't want to overstate.&lt;/p&gt;

&lt;p&gt;Behavioral signals are not identity.&lt;/p&gt;

&lt;p&gt;Historical consistency is not proof that an interaction is legitimate.&lt;/p&gt;

&lt;p&gt;Memory doesn't magically make the underlying classifier perfect.&lt;/p&gt;

&lt;p&gt;And the current environment demonstrates Hindsight integration and persistent retain/recall using a local Hindsight-compatible deployment; I would not describe that as a verified production Hindsight Cloud deployment.&lt;/p&gt;

&lt;p&gt;Being precise about those boundaries makes the system easier to trust.&lt;/p&gt;

&lt;p&gt;What actually changed&lt;/p&gt;

&lt;p&gt;The most interesting result wasn't a new model accuracy number.&lt;/p&gt;

&lt;p&gt;It was a change in the question the system can ask.&lt;/p&gt;

&lt;p&gt;Without memory:&lt;/p&gt;

&lt;p&gt;“What does this swipe look like?”&lt;/p&gt;

&lt;p&gt;With memory:&lt;/p&gt;

&lt;p&gt;“What does this swipe look like, and does it fit the security experiences I've already seen?”&lt;/p&gt;

&lt;p&gt;That is a small architectural change, but it changes the type of system you can build.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fah0eu12t2kdc9lxv9z1c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fah0eu12t2kdc9lxv9z1c.png" alt=" " width="800" height="858"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I started with a CAPTCHA that evaluated interactions.&lt;/p&gt;

&lt;p&gt;I ended with a security pipeline that can retain experiences, recall relevant history, and use that history as context for later decisions.&lt;/p&gt;

&lt;p&gt;The lesson I took away is pretty simple:&lt;/p&gt;

&lt;p&gt;Sometimes the next useful improvement isn't a bigger model. It's giving the existing system something worth remembering.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>security</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
