DEV Community

Charles Givre
Charles Givre

Posted on Originally published at gtkcyber.com

Where to Learn Applied ML for Incident Response: Start at Scoping

The incident has one confirmed host. The CISO's first question is not "how did they get in." It is "how many more are there," and every hour you spend answering it is an hour the attacker keeps their access. Scoping is where machine learning earns its place in incident response, and it is the skill most training skips.

Most "ML for IR" material teaches classifiers: label malware, label phishing. Responders rarely need a classifier mid-incident. They need two answers fast: which hosts could the attacker have reached, and which hosts are behaving like the one we know is bad. Both are a few dozen lines of Python on logs you already collect.

Question one: who could they have reached?

Lateral movement over SMB, RDP, or WMI (T1021.002, T1021.001, T1047) with valid accounts (T1078) leaves Windows Event ID 4624 on the destination host. Logon types 3 (network) and 10 (RemoteInteractive) are the ones to keep. Treat each logon as a directed, timestamped edge from source host to destination host. A host is in scope only if a path reaches it from patient zero in time order: a logon to host B at 01:00 cannot carry an attacker who reached host A at 02:10.

import pandas as pd

logons = pd.read_json("4624.jsonl", lines=True)
logons["ts"] = pd.to_datetime(logons["TimeCreated"], utc=True)
logons = logons[logons["LogonType"].isin([3, 10])
                & ~logons["TargetUserName"].str.endswith("$", na=False)]

# IpAddress is the source; Computer is the host that logged the event.
# ip_to_host comes from DHCP leases or DNS for the incident window.
logons["src"] = logons["IpAddress"].map(ip_to_host)
edges = (logons.dropna(subset=["src"])
               .query("src != Computer")
               .sort_values("ts"))

def reachable(edges, patient_zero, t0):
    reached = {patient_zero: t0}
    for e in edges[edges["ts"] >= t0].itertuples():
        if e.src in reached and e.Computer not in reached:
            reached[e.Computer] = e.ts
    return pd.Series(reached, name="earliest_possible").sort_values()

t0 = pd.Timestamp("2026-09-14 02:10", tz="UTC")
in_reach = reachable(edges, "WS-0412", t0)
Enter fullscreen mode Exit fullscreen mode

Because the edges are processed in time order, a single pass gives every host's earliest possible compromise time. Field names depend on your export (Winlogbeat nests them under winlog.event_data), so adjust the column names.

Run it unfiltered and the answer is usually "everything within an hour." A backup server or a vulnerability scanner authenticates to the whole fleet, and an attacker who reaches one of those inherits all its paths. That result is honest, but it is useless for triage. Restrict edges to the accounts you have evidence were compromised, rerun, and the set shrinks to something you can work. The output is an upper bound on scope, not a list of findings.

Question two: who looks like patient zero?

Reachability says who could be compromised. Behavior says who probably is. Sysmon Event ID 1 records every process with its parent. Reduce each to a parent>child token, keep only tokens that are new to each host since the intrusion started, and you have a document per host. scikit-learn's TfidfVectorizer weights those documents so a pair present on every machine (explorer.exe>chrome.exe) counts for almost nothing and a rare one (wmiprvse.exe>powershell.exe, services.exe> a random eight-character binary from a PsExec-style install, T1569.002) dominates. NearestNeighbors with cosine distance then ranks hosts by similarity to patient zero.

from pathlib import PureWindowsPath
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.neighbors import NearestNeighbors

procs = pd.read_json("sysmon_eid1.jsonl", lines=True)
procs["ts"] = pd.to_datetime(procs["UtcTime"], utc=True)
exe = lambda p: PureWindowsPath(p).name.lower() if isinstance(p, str) else "?"
procs["pair"] = procs["ParentImage"].map(exe) + ">" + procs["Image"].map(exe)

base = procs[procs["ts"].between(t0 - pd.Timedelta(days=30), t0)]
new = (procs[procs["ts"] >= t0]
       .merge(base[["Computer", "pair"]].drop_duplicates(),
              how="left", indicator=True)
       .query("_merge == 'left_only'"))

docs = new.groupby("Computer")["pair"].apply(list)
vec = TfidfVectorizer(analyzer=lambda pairs: pairs, sublinear_tf=True)
X = vec.fit_transform(docs)

nn = NearestNeighbors(metric="cosine").fit(X)
i = docs.index.get_loc("WS-0412")
dist, idx = nn.kneighbors(X[i], n_neighbors=min(20, len(docs)))
lookalikes = pd.Series(1 - dist[0], index=docs.index[idx[0]], name="similarity")
Enter fullscreen mode Exit fullscreen mode

Always print the terms behind the score before you act on it: sorted(zip(X[i].toarray()[0], vec.get_feature_names_out()), reverse=True)[:10]. If the top terms for patient zero are attacker tradecraft, the neighbors sharing them are your next images. If the top terms are ccmexec.exe> something, the similarity is a software deployment, not an intrusion.

Hosts that appear in both in_reach and the top of lookalikes go to the front of the queue. That intersection is the deliverable: a ranked list a lead responder can hand to the people pulling disk images.

Where this breaks

  • VDI and freshly imaged hosts have no 30-day baseline, so every pair is "new" and they float to the top of every similarity list. Baseline them against a golden-image host instead.
  • Change windows. An SCCM or Intune push during the incident window produces a cluster of nearly identical new pairs on hundreds of hosts. Check the change calendar before you trust a large cluster.
  • Hands-on-keyboard attackers using RDP and built-in tools may leave few distinctive process pairs. The logon graph carries more weight then, and the similarity list carries less.
  • Identity outside Windows. Entra ID sign-ins, VPN sessions, and SSH to Linux hosts are not in 4624. If the attacker moved through the cloud tenant, this graph misses that hop entirely.

None of it replaces forensic confirmation. It decides the order you collect evidence in, which matters because Mandiant's M-Trends 2025 still puts global median dwell time at 11 days: an attacker who has been inside for a week and a half has had time to spread. NIST's SP 800-61 Rev. 3 frames scoping as part of continuous detection and response rather than a one-time step, and this is the kind of analysis worth rerunning as new logs arrive.

Where to learn it

You can teach yourself most of this with public data. Three sources are worth your time:

  • OTRF Security-Datasets: recorded attack simulations with Sysmon and Security logs, including lateral movement over WMI, PsExec, and remote services.
  • Splunk BOTS v3: a multi-host environment with a full scenario, which is what you need to practice scoping rather than single-host triage.
  • EVTX-ATTACK-SAMPLES: raw .evtx files organized by ATT&CK tactic. Good for checking that your parser handles the real fields.

Run Chainsaw or Hayabusa over the same files first. Sigma-rule hits make good labels for checking whether your similarity ranking surfaces the hosts the rules flagged.

If you want instruction, judge a course by its lab data. A course that teaches IR analytics on one host's logs, or on a Kaggle intrusion dataset with no hostnames and no timestamps, cannot teach scoping. Ask whether the labs include many hosts, real Windows event fields, and a question with a time dimension. For the timeline-building and command-line clustering half of the job, see data science for incident responders.

We teach the building blocks (pandas on Windows and network logs, TF-IDF, nearest neighbors, anomaly detection) in Threat Hunting with Data Science and on day four of Applied Data Science & AI, in Jupyter in the AI Training Dojo. The same notebooks you write for a hunt are the ones you open at 3 a.m. when there is a patient zero and a CISO waiting on a number.

Top comments (0)