The DIVD breach, disclosed September 29, 2026, was carried out by an autonomous agent that narrated its own decisions into the logs. Your telemetry already records that signature. Here is how to hunt for it.
The Dutch Institute for Vulnerability Disclosure says it ran seven years without a security incident. Then an attacker broke in and narrated the whole thing in the logs, out loud, in real time. The attacker was an AI agent. DIVD's summary of the attack is the part that should change your detection rules: "The attack itself was loud and very very messy."
What DIVD reported
On September 29, 2026, BleepingComputer published DIVD's statement: a breach carried out by an automated AI agent operating with decisional autonomy during post-exploitation. After every action, the agent decided the next step by itself, at high speed, with significant operational errors. DIVD called the incident unprecedented in its experience.
The specifics read like a parody of tradecraft. The agent interfered with its own adversary-in-the-middle attack by triggering password-spraying operations that worked against the primary objective. DIVD's verdict on that was "some pretty dumb things." Then it got worse for the attacker: the agent extensively commented on its own decisions inside the system logs, producing an anomalous volume of forensic recordings. DIVD judged the agent "poorly trained and configured for such operations."
The honest version of this story includes what the record does not say. DIVD denied the exploited vulnerability involves Citrix NetScaler and named no alternative. There is no CVE, no confirmed data exfiltration, no confirmed intrusion duration, and no word on whether the agent arrived with the initial payload or was deployed after compromise. A detailed technical update is promised for October 1. Treat DIVD's account as working facts and the rest as open questions, because a single-source incident is a starting point for detection hypotheses, not a settled playbook.
The gap every agent-security post skips
Most agent-security writing, including mine, answers one question: how do I secure the agents I deploy? Confused deputies, zero standing privilege, policy gates at the tool boundary. All necessary. All one-sided.
The DIVD breach asks the mirror question: how do I detect the agent someone else deploys against me? Almost nobody is writing about that. Four decades of detection engineering assume human operators and scripted malware, and an autonomous agent behaves like neither. The gap is a detection vocabulary for agent-shaped attacks.
Here is the thesis in one line. The properties that make agents dangerous when you deploy them are the properties that make them detectable when attackers deploy them.
I want to steelman the skepticism before going further, because it is reasonable. One incident, one source, one sloppy configuration. A well-built attack agent will not narrate its plans into your logs, and defenders should not bet the program on attackers staying sloppy. Fair. But sloppy is the default state of new tooling, and the first generation of any attack technique is always the loud one. The question is not whether attackers will get quieter. It is whether you have a detector running while they are still loud.
Study the loop, not the model
Forget the cinematic version of an AI attacker. An attack agent is a loop, and the loop is where the evidence comes from.
┌─────────────────────────────────────────────────────┐
│ THE AGENT LOOP │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │
│ │ OBSERVE │───▶│ DECIDE │───▶│ ACT │ │
│ │ read the │ │ pick the │ │ run the tool │ │
│ │ last │ │ next │ │ call │ │
│ │ output │ │ action │ │ │ │
│ └──────────┘ └──────────┘ └──────────────┘ │
│ ▲ │ │
│ └────── at machine speed ────────────┘ │
│ │
│ A human runs this loop in minutes. An agent runs │
│ it in seconds, and writes down what it was │
│ thinking. │
└─────────────────────────────────────────────────────┘
Three detection hooks fall out of that loop.
First, cadence. A human operator thinks between commands. Minutes pass. Coffee happens. An agent fires the next tool call the moment the last one returns, so actions arrive in bursts with near-uniform spacing. There is no thinking pause, only inference latency, and inference latency has a very regular shape.
Second, narration. Reasoning is the agent's working memory, and careless configurations persist that working memory wherever output goes, including logs: "I will now attempt...", "The previous step failed because...", "Next, I will try...". DIVD's attacker did exactly this, extensively. Traditional malware minimizes traces. A chatty agent maximizes them.
Third, plan churn. A human follows a playbook. An agent re-derives its plan after every observation, which means it can pursue two incompatible subgoals at once. Password spraying in the middle of an adversary-in-the-middle attack is precisely that failure: one subgoal sabotaging the other. The errors are not noise around the signal. They are the signal.
The analogy I keep coming back to: a human intruder is a burglar who works in the dark. An attack agent is a burglar live-narrating a podcast while picking your lock, pausing every thirty seconds to reconsider whether burglary is the right strategy. You do not need a new appliance to catch that. You need to listen for the narration.
Hunt these four signatures
The pipeline is deliberately boring. Collect, normalize, score, triage.
raw logs ──▶ normalize ──▶ score sessions ──▶ triage
│
┌───────────┴────────────┐
│ 1. burst cadence │
│ 2. narration density │
│ 3. subgoal conflict │
│ 4. playbook entropy │
└────────────────────────┘
normalize to JSONL:
ts, session, kind, msg
Signature one: burst cadence. Many actions per minute from a single session, sustained, with low variance in inter-arrival times. Humans cluster their bursts with long gaps between them. Agents burst continuously.
Signature two: narration density. The fraction of a session's log lines matching reasoning markers. Humans typing commands do not write "let me try a different approach" into auth.log. Agents and their wrappers do.
Signature three: subgoal conflict. One session touching technique pairs that undermine each other, like credential spraying (MITRE ATT&CK T1110.003) alongside session hijacking or adversary-in-the-middle (T1557). Conflict is expensive for a human to produce by accident and cheap for an agent to produce by re-planning.
Signature four: playbook entropy. A human following a runbook produces a low-entropy action sequence: recon, access, persist, in that order. An agent that reorders its plan after every step produces high-entropy sequencing. Messy is measurable.
The detail that makes this cheap enough to run: none of it needs a model. Inter-arrival histograms and a regular expression get you to version one. You are not classifying attackers. You are ranking sessions by weirdness and reading the top ten every morning.
Score sessions with stdlib Python
This script is standard library only. It reads JSONL traces, one event per line, scores each session on the four signatures, and prints the weirdest sessions first. The --demo flag generates two synthetic sessions, one human-shaped and one agent-shaped, so you can watch the heuristics work before pointing them at real logs.
#!/usr/bin/env python3
"""agent_shape.py: rank sessions by agent-shaped attack behavior.
Standard library only.
python3 agent_shape.py --demo # score two synthetic sessions
python3 agent_shape.py trace.jsonl # score your own JSONL trace
JSONL schema, one object per line:
{"ts": 1729999999.0, "session": "s1", "kind": "auth|cmd|http", "msg": "..."}
"""
import json
import random
import re
import statistics
import sys
from collections import defaultdict
NARRATION = re.compile(
r"\b(i will|i'll|next,? i will|attempting|the previous (step|attempt)|"
r"because the|let me (try|check|look)|decided to|plan is to|"
r"trying a different|step \d+ of \d+)",
re.IGNORECASE,
)
SPRAY = re.compile(r"spray|credential stuff|password list|common password", re.IGNORECASE)
HIJACK = re.compile(
r"arp spoof|mitm|adversary.in.the.middle|session token|sslstrip|cookie",
re.IGNORECASE,
)
def score(events):
by_session = defaultdict(list)
for e in events:
by_session[e["session"]].append(e)
rows = []
for sid, evs in by_session.items():
evs.sort(key=lambda e: e["ts"])
n = len(evs)
span_min = max((evs[-1]["ts"] - evs[0]["ts"]) / 60.0, 1.0 / 60.0)
cadence = n / span_min
gaps = [b["ts"] - a["ts"] for a, b in zip(evs, evs[1:])]
steadiness = 1.0 / (1.0 + (statistics.pstdev(gaps) if len(gaps) > 1 else 0.0))
narration = sum(1 for e in evs if NARRATION.search(e["msg"])) / n
kinds = [e["kind"] for e in evs]
transitions = len(set(zip(kinds, kinds[1:])))
entropy = transitions / max(n - 1, 1)
text = " ".join(e["msg"] for e in evs)
conflict = bool(SPRAY.search(text) and HIJACK.search(text))
weird = (
0.35 * min(cadence / 30.0, 1.0)
+ 0.30 * narration
+ 0.20 * entropy
+ 0.10 * steadiness
+ (0.25 if conflict else 0.0)
)
rows.append(
{
"session": sid,
"weird": weird,
"events": n,
"per_min": cadence,
"narration": narration,
"entropy": entropy,
"conflict": conflict,
}
)
return sorted(rows, key=lambda r: r["weird"], reverse=True)
def report(rows):
print(
f"{'session':<10}{'weird':>7}{'events':>8}{'/min':>8}"
f"{'narr%':>8}{'entropy':>9}{'conflict':>10}"
)
for r in rows:
print(
f"{r['session']:<10}{r['weird']:>7.2f}{r['events']:>8}"
f"{r['per_min']:>8.1f}{r['narration'] * 100:>7.0f}%"
f"{r['entropy']:>9.2f}{str(r['conflict']):>10}"
)
def demo():
random.seed(7)
t0 = 1_729_000_000.0
events = []
# A human operator: 14 terse commands over 38 minutes.
human_msgs = [
("auth", "ssh login ok from 198.51.100.14"),
("cmd", "sudo systemctl restart web"),
("cmd", "tail -n 200 /var/log/app.log"),
("cmd", "grep -i error /var/log/app.log | head"),
("cmd", "df -h"),
("cmd", "sudo journalctl -u web --since '1 hour ago' | tail"),
("cmd", "cp /etc/web/config.yaml /etc/web/config.yaml.bak"),
("cmd", "sudo vim /etc/web/config.yaml"),
("cmd", "sudo systemctl restart web"),
("cmd", "curl -s -o /dev/null -w '%{http_code}' localhost:8080/health"),
("cmd", "rm /etc/web/config.yaml.bak"),
("auth", "scp backup-2026-09-30.tgz backup-host:/srv/"),
("cmd", "uptime"),
("auth", "ssh logout"),
]
t = t0
for kind, msg in human_msgs:
events.append({"ts": t, "session": "op-114", "kind": kind, "msg": msg})
t += random.uniform(60, 300) # a human thinks between commands
# An agent: 120 narrated actions in 5 minutes, with conflicting subgoals.
agent_templates = [
("auth", "I will now attempt SSH login with the next credential on the list"),
("auth", "The previous attempt failed because the password was rejected, trying the next one"),
("auth", "Decided to spray common passwords across accounts to avoid lockouts"),
("cmd", "Next, I will try arp spoofing to intercept the session in the middle"),
("http", "Step 3 of 12: attempting to steal the session token from the response"),
("cmd", "The previous step failed because the interface was busy, let me check again"),
("http", "I will now replay the captured cookie against the admin panel"),
("cmd", "Plan is to pivot: let me try a different approach on port 445"),
("auth", "Attempting login as administrator with the common password list"),
("cmd", "Because the spray is slow, I will also keep the spoof running"),
]
t = t0
for i in range(120):
kind, msg = agent_templates[i % len(agent_templates)]
events.append({"ts": t, "session": "op-207", "kind": kind, "msg": msg})
t += random.uniform(1.5, 4.0) # inference latency, no thinking pauses
return events
def main(argv):
if len(argv) == 2 and argv[1] == "--demo":
events = demo()
elif len(argv) == 2:
events = []
with open(argv[1]) as fh:
for line in fh:
line = line.strip()
if line:
events.append(json.loads(line))
else:
print("usage: python3 agent_shape.py --demo | trace.jsonl")
return 2
report(score(events))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv))
Run it and the separation is immediate:
$ python3 agent_shape.py --demo
session weird events /min narr% entropy conflict
op-207 0.88 120 22.7 100% 0.04 True
op-114 0.05 14 0.5 0% 0.23 False
A few things are worth noting about this example. First, the weights are illustrative, not tuned; the right thresholds come from your own logs, and you find them by running the scorer for a week and reading the top ten. Second, the entropy term is a ratio, so tiny sessions inflate it, which is why the human session scores higher on entropy than the agent. In production, gate scoring on a minimum session size or weight the term down. Third, the conflict pairs are the part you will customize most: encode the technique pairs your environment actually sees, like noisy scanning paired with quiet persistence. Fourth, yes, the demo trace is synthetic. Real traces are messier, which is exactly why the thresholds are yours to tune and not mine to dictate.
Now the part that touches your real infrastructure. Two commands that cost nothing and sometimes pay immediately.
# Failed SSH logins per source IP over the last 24 hours.
# Spray-shaped traffic shows up as many accounts, one source.
sudo journalctl -u sshd --since "24 hours ago" -o short \
| grep "Failed password" \
| awk '{for(i=1;i<=NF;i++) if($i=="from") print $(i+1)}' \
| sort | uniq -c | sort -rn | head -20
# Narration-shaped lines in your own application logs.
# If your agents log here, attackers' agents log in the same style.
grep -rEi "i will now|attempting to|the previous step|next,? i will|let me (try|check)" \
/var/log/myapp/ 2>/dev/null | wc -l
"The attack itself was loud and very very messy. We could see the agent working automated, because after every action it decided the next step itself, at the speed of light and sloppy logic or pattern."
(DIVD, via BleepingComputer)
Where this breaks
- One incident, one source, details pending. If the October 1 technical update contradicts DIVD's initial account, the signatures above need revisiting.
- A well-configured attack agent will not narrate. Narration is a configuration choice, not a law of nature, and the quiet agents are already being built.
- Speed is normalizable. Attackers can and will add jitter and pacing once defenders start keying on cadence. Today's signature is tomorrow's evasion homework.
- Your own automation looks bursty. CI pipelines, Terraform runs, and config management produce high-cadence, low-narration sessions. Tune on your baseline or drown in false positives.
- Endpoint logs are attacker-writable. An agent with shell access can muzzle its own logging. Server-side logs, network telemetry, and identity-provider logs are the surfaces the attacker does not control.
- Agent-shaped is not AI-attributed. A sloppy script with verbose echo statements scores the same as a reasoning model. The detector finds weird sessions. Attribution is a human job.
Build it if / Skip it if
Build it if you run internet-facing authentication, if your logs already land in one place, or if you operate a honeypot and want to know what the next generation of visitors looks like.
Skip it if you have no centralized logging yet, because this detector is a query layer and you need the layer underneath first. Skip it if nobody will read the output; this produces a ranked list, not an auto-remediation, and an unread list is just disk usage.
The minimal viable version fits in an afternoon. Normalize one week of SSH logs to the JSONL schema, run the scorer, tune one threshold until the top ten are all worth a glance, and put that glance on a daily rotation. Measure how many flagged sessions turn out to be your own automation, then add those patterns to an allowlist. That is the whole program.
Point it at your logs this week
The DIVD breach is the first documented case, not the last attempt. Attack tooling always starts loud and gets quiet, which means the window where simple detectors work is the window that is open right now. Run the scorer against your own telemetry, find out what your baseline looks like, and decide from data whether agent-shaped traffic is already visiting.
What is the noisiest intruder your logs have ever caught, automated or otherwise?
Resources
- BleepingComputer: "Automated AI agent used to breach cybersecurity nonprofit DIVD": the original reporting, September 29, 2026.
- DeafNews: "Autonomous AI Agent Hits DIVD: Operational Errors Leave Forensic Trail": careful analysis, including the single-source caveats.
- MITRE ATT&CK T1110.003: Password Spraying
- MITRE ATT&CK T1557: Adversary-in-the-Middle
- NIST AI Risk Management Framework: the Map function is where detecting novel AI-driven threats belongs.
Top comments (0)