DEV Community

Cover image for HINDY: Hunting Wolves in Sheep's Clothing
Siva Balaji B K
Siva Balaji B K

Posted on

HINDY: Hunting Wolves in Sheep's Clothing

A database server sends roughly 500 GB out of the building at exactly the hour its nightly backup runs. Same host, similar size, same time of night. A memory system that searches past cases finds last night's backup, which analysts correctly closed as benign, and reports a strong match.

But this transfer goes to an external IP address, with no scheduled job and an unfamiliar initiating process. The past decision does not apply, and a system that only measures similarity will never notice.

That is the problem I built HINDY to solve. HINDY is a SOC investigation platform that remembers past cases and still makes every new alert prove itself.

HINDY's memory layer is built on Hindsight, an open-source agent memory system. See the Hindsight documentation for how it works, and Vectorize's overview of agent memory for the bigger picture.

The problem

A mid-size security operations center investigates thousands of alerts a day, and most turn out to be noise. That is why the dangerous one is easy to miss: it can look just like last night's backup or a sign-in mistake an analyst has closed three times.

I wanted to avoid two bad options: forgetting past investigations, or trusting history so much that a familiar-looking attack gets waved through. HINDY is a React workspace on a FastAPI backend, with a Hindsight memory bank, structured alert data and a reasoning model. Each screen answers one question: the Dashboard shows what needs attention, Investigation what HINDY knows about one alert, Memory what the team has learned, Replay how that was used, and Evaluation whether it helped.

New alert -> recall experience -> compare context -> reason -> safety check
          -> analyst decision -> retain learning -> future alert
Enter fullscreen mode Exit fullscreen mode

HINDY never closes an alert on its own. It helps an analyst decide whether the past really applies to the present.

Signing in

The analyst is a named person, and every decision made later is recorded under that identity.

The HINDY sign-in portal. The labels beside the mascot, recall, analyze and protect, are the three verbs the product is built around.

The Dashboard starts with work

After login the analyst sees the live queue, ordered by severity, with user, host, timestamp and status on every row. Here there are 285 pending alerts, 2 of them critical. Memory is kept off this screen on purpose: old closed cases stay available as knowledge but should not crowd the queue. From here the analyst clicks Investigate and the alert's context carries over.

The Dashboard: the active alert queue ordered by severity, with one-click investigation

Familiar is not the same as safe

A SOC sees the same alert families again and again, and a rule that trusts anything familiar lets a lookalike attack borrow the shape of a harmless event. The dataset has exactly this case. The host db-prod-01 sends a large backup to an internal server every night inside a scheduled window, and analysts correctly closed those alerts. Then ALRT-00663 arrives: roughly 500 GB leaves the same host at the expected time, but to an external IP address, with no matching scheduled job and an unfamiliar initiating process. That is the wolf in sheep's clothing, and it produced my central rule:

Memory found does not mean memory applies.

Where Hindsight fits

Hindsight is the memory bank. HINDY uses it in two places: recall at the start of every investigation, and retain at the end, once a human has confirmed a decision.

Two design choices matter here. The recall query is built from category, user, host and destination, and volatile fields such as execution IDs are excluded so equivalent events still match. And only human-confirmed decisions are retained, so HINDY's early mistakes cannot reinforce themselves.

Investigation turns a match into evidence

The analyst opens the alert and presses Analyze with HINDY. A SOCMemoryAgent runs it in stages rather than one prompt: Hindsight recall, lookup of the full past records, a deterministic context comparison, model reasoning, and finally safety rules.

recall_results = recall_history(self.client, self.bank_id, query)
history_records = lookup_history(extract_alert_ids(recall_results))
compared_cases = [compare_context(alert, hist, history_records)
                  for hist in history_records]
llm_output = reason_with_llm(alert, compared_cases, best_match, ...)
final_assessment = apply_safety_rules(alert, best_match,
                                      compared_cases, llm_output)
Enter fullscreen mode Exit fullscreen mode

The workspace shows the recalled past cases with current-versus-previous signals side by side.

Relevant past investigations, with matches and differences shown next to each other

The original detector severity stays separate from HINDY's assessment, so it is obvious the assessment is advice and the decision belongs to the analyst.

An investigation opens with the detector's severity and, separately, HINDY's assessment

An assistant scoped to the selected alert answers questions such as why the assessment came out this way and what differs from past cases.

The alert-scoped assistant, opened from the investigation page

With HINDY and without it: the 500 GB transfer

This is the core technical story. Recall alone is a solved problem, since any vector database returns the old backup when you search with the new alert. The hard part is deciding whether the old decision still applies.

Signal Past case (closed as benign) New alert ALRT-00663
Source host db-prod-01 db-prod-01
Volume 457 GB roughly 500 GB
Destination backup-srv-02, internal External IP address
Scheduled job BK-NIGHTLY-01 None
Initiated by backup-orchestrator-01 Unfamiliar process
Sensitive-data flag, account usage Backup service account, normal Differ from the past case

Without a context check, a similarity-only system sees the same host, a similar volume and the same hour, and reports a strong match to a benign precedent. That is an illustration of the failure I designed against, not a measured result.

With HINDY, the match is only a starting point, and compare_context compares signal by signal. This is the real destination check; the same pattern covers host, user, job, process, data sensitivity and account behavior.

p_class = classify_destination(past_dst)
c_class = classify_destination(curr_dst)
if p_class == c_class and p_class != "missing":
    matches.append("destination")
else:
    differences.append({"signal": "destination",
        "past": f"{past_dst} ({p_class})", "current": f"{curr_dst} ({c_class})",
        "is_key_signal": True})
Enter fullscreen mode Exit fullscreen mode

Here destination lands in differences as a key signal, while the host stays in matches, which is honest because the source really is the same. The safety pass then stops the model's wording from overriding it:

if key_differences and final_state == "green":
    final_state = "yellow"
Enter fullscreen mode Exit fullscreen mode

Once a key signal differs, the assessment can no longer be green. Critical alerts stay critical, an alert with no precedent is treated as novel, and even a perfect match only earns "Analyst quick-confirm."

The analyst decides, and it is kept

The analyst records a final disposition, either expected or needing a security response, with a written rationale, and can disagree with HINDY. That confirmed decision is what gets retained as new memory.

The disposition panel: HINDY suggests, the analyst decides and explains why

The whole loop, as the product draws it, feeds the last box back into the memory bank at the start.

The learning loop: past case, memory bank, new alert, signal diff, reasoning, human decision, memory ingest

Memory: what the team has learned

Each retained experience is laid out as what happened, what was observed, what the analyst found, how it was resolved and the lesson to carry forward. The bank held 515 memories: 420 verified benign baselines, 14 confirmed attack precedents and 1 created live from an analyst override. The reasoning matters more than the label, because "benign" alone tells a future analyst little, while "approved under a named job, from a known host, inside the organization" tells them whether a later event deserves the same conclusion.

The Memory page, separating benign baselines, attack precedents and analyst-evolved learning

A retained experience, with its telemetry, context signals and the analyst's reasoning

Replay: how memory was used

Replay rebuilds an investigation as a sequence: the original alert, the recalled precedent, the context comparison, HINDY's reasoning, the analyst's decision and the lesson kept. It is the fastest way to explain the system to someone new, and each recorded trace begins with the original telemetry.

Replay catalogue: each recorded trace begins with the original alert telemetry

The recalled precedent and the context comparison. Every signal matches, so HINDY recommends a quick analyst confirmation and keeps the original decision visible

Evaluation: does the memory help?

The Evaluation page runs memory-assisted and no-memory analysis on the same 94 alerts: 10 attacks, 44 lookalikes and 40 recurring benign alerts, with seed 42. The agent never reads the ground truth; only the evaluation code does.

Measure Without memory With HINDY memory
Attacks caught 10 / 10 10 / 10
False greens (threats marked benign) 23 0
Unnecessary escalations of benign alerts 1 / 40 (2.5%) 17 / 40 (42.5%)
Human-review load 32 alerts (34.0%) 71 alerts (75.5%)
Average time per alert 3.61 s 10.11 s

Three caveats change what this table means.

  1. It is not a held-out test. I tuned the V2 comparison rules after seeing failures on this same sample, so the numbers are for diagnosis, not a claim about unseen data.
  2. The baseline is not a clean same-model comparison. Some no-memory alerts ran on smaller fallback models after a provider rate limit, so the 23-to-0 false-green gap should not be credited to memory alone. A single-model rerun is still to do.
  3. Memory made things worse in one respect. Unnecessary escalations rose from 1 in 40 to 17 in 40, review load from 34% to 75.5%, and each alert took longer. That is the price of refusing green when a key signal differs. I think the trade is right for a security tool, but analysts will feel it.

The Evaluation summary comparison, with the review-load tradeoff stated beside the wins

The per-alert comparison: what changed and where memory altered the context

Settings

Settings ties the system to a person: identity, role and session, with decisions recorded under the authenticated analyst.

Settings: account identity, role and session

What I learned

  1. Similarity is not safety. The old backup and the 500 GB transfer look almost identical to retrieval; applicability must be tested separately.
  2. Differences should be data. A changed destination is an item in a list, not a sentence buried in model output.
  3. Put rules around the model. It explains well, but whether a threat can be called safe is decided deterministically.
  4. A degraded baseline taught me more about evaluation hygiene than about the agent. Next time both sides run on the same model.
  5. Only human-confirmed decisions become memory. Otherwise HINDY's early mistakes would reinforce themselves.
  6. Report the cost beside the win. Zero false greens looks great until the review load sits next to it.

HINDY does not make a SOC trust the past more easily. It makes the past available, explainable and open to challenge at the moment an analyst needs it, so it can recognize the sheep, remember the wolf, and still ask a person to decide whether they are the same animal.

Links

Top comments (0)