At the end of February 2024, CrowdStrike shipped a new template for its Falcon sensor that declared 21 input fields, while the code feeding it supplied only 20. Nobody noticed for almost five months, because every rule written against the template used a wildcard for the last slot. Then, at 04:09 UTC on July 19, a rule arrived that actually checked field 21, and a kernel driver on millions of Windows machines tried to read a 21st value from an array that held 20. The financial aftermath of that morning, and of the Okta, F5 and SonicWall breaches on either side of it, is traced in When the Security Vendor Becomes the Incident, a case study of the bills that investors, customers and courts sent afterward. The engineering story underneath those headlines is just as instructive, because nearly every root cause in it has a twin in ordinary application code.
Security products are the code we trust most and inspect least. They run with kernel or administrator rights, update on someone else's schedule and sit in the path of every boot, login and packet. Read their post-mortems back to back and three patterns keep returning. None of them is exotic, and all three probably exist in the repository you will open tomorrow.
Configuration Is Code That Skipped the Pipeline
CrowdStrike customers could stage and pin new sensor versions. Rapid Response Content, the behavioral rules the sensor pulls down between releases, was treated as data, and data went to every online host at once. Two safety nets failed on the way. The Content Validator checked the new rules against the template's definition of 21 fields, so it faithfully confirmed the wrong contract, and the interpreter had no runtime bounds check to catch what the validator missed. For many machines, recovery meant booting into Safe Mode or the Windows Recovery Environment, often typing in a BitLocker recovery key, and deleting one channel file by hand.
Microsoft's own estimate put the affected devices at less than one percent of all Windows machines. Airlines still grounded flights and hospitals still postponed non-urgent care, because, as Microsoft noted, CrowdStrike is used by enterprises that run many critical services. Fleet percentage is a comforting number and an almost useless one: blast radius depends on which machines fail, not how many. Microsoft's longer answer has been architectural. It has been working with CrowdStrike and other vendors on a Windows endpoint security platform that lets antivirus and EDR products run in user mode, outside the kernel, and announced a private preview for partners in mid-2025.
If that sounds like a kernel-driver problem, Cloudflare's outage on November 18, 2025 says otherwise. At 11:05 UTC, a deliberate ClickHouse permissions change made a metadata query, one that never filtered by database name, start returning duplicate rows. That query builds the feature file for Cloudflare's Bot Management model, which is regenerated every five minutes and pushed to the whole network. The file doubled in size and blew through a hard cap of 200 features, a limit set well above normal use because the proxy preallocates memory for them. The new Rust proxy, FL2, checked the count, got an error back and called unwrap() on it, and the resulting panic surfaced as HTTP 5xx errors across Cloudflare's network. Customers still on the older proxy got no 5xx errors; instead every request received a bot score of zero, so sites with bot-blocking rules began turning away real visitors. Because only part of the database cluster had the change at first, good and bad files alternated, the network kept failing and recovering, and the team's first suspicion was a massive DDoS attack. Cloudflare called it its worst outage since 2019.
Put the two incidents side by side and the lesson gets sharper. CrowdStrike's interpreter had no bounds check and read past the end of an array. Cloudflare's code had the check, in a memory-safe language, and the network went down anyway, because the only response it had to a failed check was to crash. Validation is half the job. The other half is deciding what happens to a bad file once you catch it, and "keep serving with the last good version" beats "panic" almost every time. Sixteen months apart, both failures had the same shape: a file produced by the vendor's own tooling and trusted because it came from inside, a consumer that assumed the file's shape, and a distribution system built for speed.
The first item on Cloudflare's remediation list is to ingest its own generated configuration files with the same suspicion it applies to user input. Feature flags, WAF rules, ML feature lists and policy bundles all belong in that bucket, and plenty of teams still ship them through a faster, less guarded lane than code.
The Front Door Was Somebody's Browser Profile
Okta's October 2023 breach did not begin with an exploit. A service account for its customer support system had its username and password saved in an employee's personal Google profile, signed into Chrome on a company laptop, and Okta concluded that the credential most likely leaked through that personal account or device. Inside the support system sat HAR files, the recordings of browser sessions that customers upload so engineers can reproduce a bug. Some still held live session tokens, and the attacker used them to hijack the sessions of five customers. Okta's fixes read like a checklist for any SaaS team: personal profiles blocked in managed Chrome, and administrator session tokens bound to network location so a stolen one stops working from somewhere else. Within a week of Okta's disclosure, Cloudflare, one of the targeted customers, open-sourced a HAR sanitizer that strips session cookies and JSON Web Tokens in the browser before a file is shared.
LastPass learned the same lesson at a higher price. After a 2022 intrusion into its development environment, attackers went after the home computer of a senior DevOps engineer, one of only four engineers who could reach the decryption keys for the company's production backups. A vulnerable third-party media package on that machine gave them remote code execution. A keylogger captured the master password as it was typed, after multi-factor authentication had already succeeded, and the corporate vault behind it held the keys to backups stored in Amazon S3. Those backups held copies of customer vault data.
Neither attack broke any cryptography. Both walked through convenience: browser sync, a home media app, a debugging file with a live cookie inside. If an engineer can reach production from the same machine that runs weekend side projects, that machine is part of your threat model whether or not your architecture diagram shows it.
Your Vendor's Build Room Is Part of Your Attack Surface
F5 found intruders in its network on August 9, 2025. When it disclosed the breach in October, it said a nation-state actor had held long-term, persistent access to its BIG-IP product development environment and engineering knowledge management platform, and had taken portions of BIG-IP source code along with details of vulnerabilities its engineers were still working on. F5 reported no evidence of tampering with its build and release pipelines. The risk was quieter than a poisoned update: an adversary holding the source and the bug backlog can find and weaponize flaws faster than customers can patch them.
CISA's Emergency Directive 26-01 turned that risk into concrete work for federal agencies: inventory every BIG-IP product, check whether management interfaces are reachable from the public internet, install F5's updates within about a week, and disconnect devices that have reached end of support. Nothing on that list requires access to F5's systems. It is the half of a vendor breach that customers control.
SonicWall's 2025 incident makes the same point from the storage side. In September, the company found that attackers had accessed firewall configuration files stored in its MySonicWall cloud backup service, and in November it attributed the intrusion to a state-sponsored actor. The credentials inside the files were encrypted, yet the files still described how each firewall was set up, which SonicWall itself warned could make targeted attacks easier, and its remediation guidance had customers reset admin passwords, VPN pre-shared keys and TOTP bindings anyway. A configuration backup is a secrets file with a friendlier extension.
What to Change Before Your Next Deploy
None of these vendors were careless amateurs, which is exactly why their post-mortems are worth stealing from. Here is what they add up to for a team shipping its own software:
- Map your privileged dependencies. List every agent, driver, browser extension and appliance that runs with kernel, root or admin rights and updates itself, and write down who controls its update schedule.
- Turn on update rings. If a vendor lets you stage content updates as well as agent versions, put a few non-critical machines in the first ring and let them soak before the rest follow.
- Distrust your own config. Validate schema, field counts and size where a file is consumed, keep the last known-good version, fall back to it instead of crashing, and give every fast-moving feature a kill switch.
- Price out the bad-boot day. Know where recovery keys live, test out-of-band console access, and time how long it takes to fix ten machines by hand.
- Scrub debug artifacts. Strip cookies, tokens and authorization headers from HAR files and logs before they reach any support portal, and keep session tokens short-lived and bound to a device or network where your stack allows it.
- Shrink the management plane. Keep admin interfaces off the public internet and patch edge appliances on a clock measured in days, not quarters.
- Rotate after every vendor breach. Include tokens, API keys and config secrets you believe are idle, since those are the ones nobody remembers to change.
Find Your 21st Field First
Every company in these stories sells protection, and each one became, for a while, the thing its customers needed protecting from. That is not an argument for ripping out endpoint agents or identity providers, since a fleet without them is worse off. It is an argument for designing as if the most trusted code in your stack will someday misbehave: bounded inputs, a safe answer for when a bound is crossed, staged rollouts, scrubbed debug files, short-lived credentials and a recovery path someone has actually walked. Somewhere in your own system a template expects 21 fields and a caller sends 20. It is far cheaper to find it on a quiet Tuesday in a canary ring than at 04:09 on a Friday, on every machine you own.
Top comments (0)