TL;DR
In part 1 I showed how an eBPF monitor catches ransomware behavior: hook execve/openat/unlink/mkdir, stream events to userspace, score a 1-second sliding window per PID.
This is the other half: what happens after the verdict fires. Not another alert. Not a dashboard row. The process dies, in the same second, by design.
Three parts: why a plain SIGKILL from userspace is enough (no LSM, no kernel module), how I handle the race conditions, and how the agent sandboxes itself so a bug can't turn your EDR into the attack.
All code below is from Talus, MIT, written in Rust.
The decision point: a sliding window, not a signature
First, the verdict itself. Ransomware has one loud fingerprint: it opens files in bulk. So the monitor keeps a per-PID deque of timestamps and scores it every event:
let now = Instant::now();
let window = self.windows.entry(ev.pid).or_default();
let cutoff = now - Duration::from_secs(WINDOW_SECS);
while window.front().is_some_and(|t| *t < cutoff) {
window.pop_front(); // evict events older than the window
}
window.push_back(now);
if self.threshold > 0 && stats.window_opens == self.threshold {
stats.alerts += 1;
// verdict: this PID is mass-opening files RIGHT NOW
...
}
WINDOW_SECS is 1. If a process crosses the open-rate threshold inside that second, it gets a verdict. The same pattern the big EDR vendors describe in marketing blogs - except here you can read the whole engine in one file.
Why SIGKILL from userspace is enough
The obvious objection: "if your detection runs in the kernel, why not kill in the kernel?" You can do that with an LSM or a bpf_send_signal() helper. I deliberately don't.
Reasons:
-
kill(2)is one syscall. The event already traveled kernel → userspace through a per-CPU perf buffer with zero-copy handoff. The round-trip adds microseconds, not seconds. The ransomware's own encryption loop is orders of magnitude slower than that. -
Policy lives in userspace. Thresholds, allowlists, the neural sidecar (MeMLP), audit logging - all of it is ordinary, debuggable Rust. Kernel code is where correctness goes to die and
panicis not an option. - The kernel side stays dumb. A dumb kernel program is a safe kernel program. All the eBPF side does is watch and report.
So the response path is exactly what you'd write by hand:
/// Send SIGKILL to a process. Returns true on success.
fn kill_process(pid: u32) -> bool {
// SIGKILL cannot be caught, so the target process will terminate.
let rc = unsafe { libc::kill(pid as i32, libc::SIGKILL) };
rc == 0
}
And it's wired straight into the scoring loop:
if self.auto_kill {
let result = kill_process(ev.pid);
outputs.push(Output::Action(ResponseAction {
ts: ev.ts.clone(),
pid: ev.pid,
comm: ev.comm.clone(),
action: format!("SIGKILL sent to PID {} ({})", ev.pid, ev.comm),
success: result,
}));
}
(Listings abridged - error paths and logging trimmed.)
The race conditions nobody talks about
Auto-kill means your monitor is now allowed to end processes. Two races matter.
Race 1: kill on verdict, not on first event. A naive design kills the process the moment it sees one suspicious open. That's how you eat a false positive and murder your database. The verdict only fires when the rate crosses the threshold inside the window - so a process that legitimately touches many files (a build, updatedb, a backup job) either stays under the threshold or gets killed because it genuinely looks like ransomware, which is a tuning conversation, not a coin flip.
Race 2: PID reuse. The PID comes from the kernel as u32 and goes back into libc::kill as i32. Between the event and the kill, that PID could theoretically die and be reused. The window is tiny (the verdict is computed on the same event that triggered it, microseconds later), and SIGKILL to a recycled PID is bad but bounded. A full fix would verify /proc/<pid>/comm matches the reported process name before pulling the trigger - it's on the roadmap, and I'd rather ship the honest version than claim a perfect one.
What you do not get: catching file #1. The first opens of a real ransomware run will land. The design goal is narrower and honest: the attacker gets seconds, not your whole disk, and definitely not your backup job that started at 2 AM.
Testing auto-kill without wrecking your machine
You cannot test a response engine by running real ransomware. You don't need to. Ransomware's fingerprint is "opens files in bulk", so the test is just that:
# fake ransomware: 500 file opens as fast as the shell can go
for i in $(seq 1 500); do
touch /tmp/v$i.enc && cat /tmp/v$i.enc > /dev/null
done
Run the monitor in observe mode with a low threshold first, confirm the alert fires, then flip on --auto-kill and run the same loop. The loop gets SIGKILLed mid-run - usually inside the first second, well before 500.
To keep even that controlled, run the victim inside a throwaway scope:
systemd-run --user --scope \
bash -c 'for i in $(seq 1 500); do touch /tmp/v$i.enc; cat /tmp/v$i.enc >/dev/null; done'
Same syscall pattern, same verdict, and if anything goes wrong the blast radius is one scope you can systemctl away. I run this against my own desktop daily - a week of normal usage in observe mode (browsers, builds, editors) produced zero false positives, which is the number that actually matters for an auto-killing agent.
The agent must be unable to become the attack
Here's the part I care about most. Your response engine just got permission to kill arbitrary processes, loaded an eBPF program, and holds capabilities. If someone compromises the agent, they don't need ransomware anymore - they have your EDR.
So Talus strips itself down before it touches a single event. Step 1: drop every capability it doesn't strictly need:
let mut keep = HashSet::new();
keep.insert(CAP_BPF as u32); // load/attach eBPF programs
keep.insert(CAP_PERFMON as u32); // perf buffers
keep.insert(CAP_NET_ADMIN as u32); // some eBPF program types
let to_drop: Vec<u32> = caps.difference(&keep).copied().collect();
for cap in &to_drop {
// PR_CAPBSET_DROP = 24
let ret = unsafe { libc::prctl(24, *cap as i32, 0, 0, 0) };
...
}
CAP_SYS_ADMIN, CAP_DAC_OVERRIDE, everything else - gone, via the capability bounding set, so no child process can get them back either.
Step 2: seccomp. The agent builds a syscall allowlist (it needs a surprisingly short one: perf setup, kill, logging, that's mostly it) and installs a BPF filter so anything outside the list returns EPERM.
Step 3: Landlock (kernel ≥ 5.13). The filesystem gets a read-only ruleset over exactly the paths eBPF needs:
let allowed_paths: Vec<(&str, u64)> = vec![
("/sys/kernel/debug", LANDLOCK_ACCESS_FS_READ_DIR),
("/sys/fs/bpf", LANDLOCK_ACCESS_FS_READ_DIR),
("/proc", LANDLOCK_ACCESS_FS_READ_DIR),
("/dev/null", LANDLOCK_ACCESS_FS_READ_FILE),
("/dev/urandom", LANDLOCK_ACCESS_FS_READ_FILE),
("/tmp", LANDLOCK_ACCESS_FS_READ_DIR),
];
A compromised agent on a Landlocked kernel can read BPF maps and process info. It cannot write your disks. That asymmetry - "can kill, can't touch data" - is the whole design.
Numbers
From my desktop, sustained: ~280k events/s at ~7.6% CPU, per-CPU perf buffers, zero-copy handoff. Single static binary, ~2 MB, no runtime dependencies. Install path: pip install talus-process-monitor && talus-monitor install.
Try it / next steps
- Repo (MIT): https://github.com/BartoszOsiej/talus-process-monitor
- Part 1 (detection): Detecting ransomware with eBPF in Rust
If you run Linux servers and you try it in observe mode, I genuinely want to hear what your open-rate distribution looks like - threshold defaults are the hardest part of this whole thing, and real fleets beat my desktop every time.
(For teams that want the response layer managed - web dashboard, license management, support - that's the Enterprise edition. The detector and everything you just read is MIT.)
Top comments (0)