The CrowdStrike outage of July 19, 2024 is the biggest IT failure most of us have lived through, and its root cause fits in one sentence: a detection template declared 21 input fields, and the code that fed it supplied 20. CrowdStrike pushed a content update that used the 21st field to every online Windows machine running its Falcon sensor, and about 8.5 million devices blue-screened and boot-looped. If you ship anything that downloads configuration and parses it in a privileged place, an agent, a driver, a sidecar, this is your incident too.
TL;DR
- The bad file, Channel File 291, went out at 04:09 UTC and was reverted at 05:27, a 78-minute window. Machines that had already loaded it kept crashing on every boot.
- CrowdStrike's root cause analysis: the 21st value was read past the end of a 20-element array, "an out-of-bounds memory read", inside a kernel driver, where Windows cannot catch it.
- The mismatch had shipped in February. It stayed hidden for months because every rule until July 19 used a wildcard in field 21, so nothing ever read it.
- The fix was manual: Safe Mode, delete one file, reboot; on BitLocker machines, type the recovery key first. About 99 % of sensors were back by July 29.
- Delta says it cancelled about 7,000 flights and is claiming at least $500 million. CrowdStrike's own finding number six: "Template Instances should have staged deployment".
What caused the CrowdStrike outage? Template types and channel files
The Falcon sensor has two kinds of detection logic, and the difference is the whole story.
Template Types are code. They are compiled into the sensor, released as sensor versions, and customers can stage those (run the newest version, or one or two behind). In February 2024, sensor 7.11 shipped a new IPC Template Type for detecting attacks that abuse Windows named pipes.
Template Instances are data. CrowdStrike calls them Rapid Response Content: configuration pushed from its cloud to every sensor through "channel files", so new detections can go out in minutes. In the RCA's words, Rapid Response Content "is configuration data; it is not code or a kernel driver." That sentence is why it did not go through the same staged rollout as code.
The bug sat between the two. From the RCA, page 3: "The new IPC Template Type defined 21 input parameter fields, but the integration code that invoked the Content Interpreter with Channel File 291's Template Instances supplied only 20 input values."
A simplified sketch of the shape of the bug (illustrative, not CrowdStrike's code):
/* simplified sketch */
const void *inputs[20]; /* integration code supplies 20 values */
/* the template type declares 21 fields; an instance matches on field 21 */
const void *v = inputs[20]; /* index 0x14 = the 21st slot: past the end */
/* with no bounds check, v is whatever memory follows the array, and it gets dereferenced */
Why it stayed hidden for five months
The template type went live on February 28. A stress test on March 5 passed, and between March and April four Template Instances shipped through Channel File 291 without trouble. The RCA explains why: the mismatch survived "in part due to the use of wildcard matching criteria for the 21st input during testing and in the initial IPC Template Instances". A wildcard matches anything, so the interpreter never needed to read field 21, so it never went out of bounds.
On July 19 two new instances shipped, and one "introduced a non-wildcard matching criterion for the 21st input parameter". The cloud-side Content Validator checked it and passed it. Finding 4 of the RCA says why: the validator "based its assessment on the expectation that the IPC Template Type would be provided with 21 inputs". It checked the rule against the declaration, and the declaration was the thing that was wrong. Per the preliminary review, it shipped on "trust in the checks performed in the Content Validator".
Then: "The Content Interpreter expected only 20 values. Therefore, the attempt to access the 21st value produced an out-of-bounds memory read beyond the end of the input data array and resulted in a system crash."
CrowdStrike outage timeline
| When (UTC) | What happened |
|---|---|
| Feb 28, 2024 | Sensor 7.11 ships the IPC Template Type, declared with 21 fields |
| Mar 5 | Stress test passes; first instance via Channel File 291 |
| Apr 8–24 | Three more instances, all with a wildcard in field 21 |
| Jul 19, 04:09 | Two new instances pushed to all online Windows sensors; one uses field 21 |
| Jul 19, 05:27 | Content reverted, 78 minutes later |
| Jul 19, 09:45 | CEO George Kurtz's first post |
| Jul 25 / Jul 27 | Runtime bounds check added / field-count check at compile time in production |
| Jul 29 | About 99 % of Windows sensors back online |
| Aug 6 | Full root cause analysis published |
July 19 was a Friday, and the preliminary review opens with exactly that: "On Friday, July 19, 2024 at 04:09 UTC".
Why a CrowdStrike update caused a blue screen of death
In a normal program, reading a bad pointer throws an exception or kills one process. The Falcon sensor runs as csagent.sys, a kernel driver, and the kernel has no one above it to clean up. The crash dump in the RCA shows the bugcheck PAGE_FAULT_IN_NONPAGED_AREA (50), which Windows describes as invalid system memory that "cannot be protected by try-except". The preliminary review makes the same point from the other side: the Content Interpreter "is designed to gracefully handle exceptions", and this one "could not be gracefully handled, resulting in a Windows operating system crash (BSOD)."
The RCA walks the dump: the faulting instruction is mov r9d,dword ptr [r8], and "register r11 indicates that the input to be retrieved is at index 0x14, i.e., the 21st element". Dumping the array shows 20 valid pointers and then ffffd603'0000006a, "which does not point to valid memory". Security researcher Patrick Wardle had read the same thing off a crash dump on the afternoon of the outage:
Why the boot loop? A security driver loads early in boot on purpose, to catch early-starting malware. A machine that had downloaded the bad file crashed, rebooted, loaded the driver, read the same file and crashed again. The revert at 05:27 could only reach machines that stayed up long enough to download it.
How the CrowdStrike outage was fixed, one machine at a time
CrowdStrike's workaround from July 19:
- Reboot, ideally on a wired network, in case the good file arrives first.
- If it crashes again, boot into Safe Mode or the Windows Recovery Environment.
- Go to
%WINDIR%\System32\drivers\CrowdStrike, find the file matchingC-00000291*.syswith the 04:09 UTC timestamp, and delete it. - Boot normally. "Note: Bitlocker-encrypted hosts may require a recovery key."
For cloud VMs the advice was to detach the disk, delete the file from another machine and reattach it. That note about BitLocker is where the day went: each encrypted laptop needed its 48-digit recovery key typed by hand, and the key could sit on a server that was itself blue-screening. Per the remediation hub, more than 97 % of Windows sensors were back online by July 24 and about 99 % by July 29.
The blast radius
Microsoft's David Weston estimated that the update "affected 8.5 million Windows devices, or less than one percent of all Windows machines". Under one percent is the point: that one percent runs airlines, hospitals and banks.
- Delta reported "approximately 7,000 flight cancellations over five days" in an 8-K filing, and CEO Ed Bastian said Delta was "pursuing legal claims against CrowdStrike and Microsoft to recover damages caused by the outage, which total at least $500 million". Those are claims.
- Fortune 500 losses were estimated at $5.4 billion, excluding Microsoft, by the insurer Parametrix, per the Guardian, with only a fraction of that insured.
- The Hacker News thread on the day reached 4,489 points and 3,859 comments.
The CEO's first post, five and a half hours after the push, was accurate and not what anyone wanted to hear:
The apology came later: "We apologize unreservedly" in the RCA summary, and "we are deeply sorry" in testimony to Congress in September. The irony the episode ended on: an update meant to detect novel attack techniques delivered one. The third-party review did find the bug "not exploitable by a threat actor".
Who is to blame? The six findings and the kernel question
CrowdStrike's RCA lists six findings, to its credit in plain language:
| # | Finding | Fix |
|---|---|---|
| 1 | Field count not validated at sensor compile time | in production Jul 27 |
| 2 | No runtime array bounds check in the Content Interpreter | added Jul 25 |
| 3 | Template type tests should cover more kinds of matching criteria | test changes |
| 4 | The Content Validator had a logic error | fixed by Aug 19 |
| 5 | Validation should include running content in the interpreter | test changes |
| 6 | "Template Instances should have staged deployment" | canary, rings, customer control |
My blame split in the episode: CrowdStrike 65 %, for all six. The kernel-mode design 25 %: CrowdStrike's RCA argues that "Significant work remains for the Windows ecosystem to support a robust security product that doesn't rely on a kernel driver for at least some of its functionality", while Microsoft's blog notes "this was not a Microsoft incident". Both are half right, and the page fault didn't care. The Friday 10 %, which is the joke slice and still true.
What developers should learn from the CrowdStrike outage
- Configuration is code once something parses it. The content went out as "configuration data; it is not code". The parser that read it ran in the kernel. Anything a privileged process parses deserves code-level review, tests and rollout.
- Stage everything that ships to everyone. With a canary ring, the crash would have stopped at a small first ring of hosts; without one, it hit every Windows machine online in those 78 minutes. This is finding 6, and it applies to feature flags, rules files and model weights as much as to binaries.
- Validate against the code that consumes it. The validator and the template agreed with each other and both disagreed with the code. Run test content through the real parser (finding 5), and fuzz it; in the kernel, the missing bounds check takes down the whole machine, where user code would only get an exception.
- Plan recovery for machines that can't boot. Remote rollback only helps hosts that stay up. Know where your recovery keys live, and keep a copy somewhere other than the servers that just went down.
Verdict: SHIP IT, narrowly
The Postmortem verdict is on the response, and CrowdStrike's was fast and honest: a runtime bounds check in six days, a compile-time check in eight, a 12-page root cause analysis with the crash dump in 18 days, two outside reviews, staged rings with customer control over content updates, and the sentence "Template Instances should have staged deployment" written by the company itself. Narrowly, because everything on that list was standard practice before July 19. The Monday line: whatever your agent pulls from the cloud gets a canary ring.
FAQ
What caused the CrowdStrike outage?
A Falcon content update used the 21st input field of a template whose calling code supplied only 20 values. The sensor read past the end of the array in kernel mode and crashed Windows.
How many computers did CrowdStrike crash?
Microsoft estimated 8.5 million Windows devices, less than one percent of all Windows machines. Mac and Linux were not affected.
What is Channel File 291?
The Falcon channel file that delivers Template Instances for the named-pipe (IPC) template type; the bad version was C-00000291*.sys timestamped 04:09 UTC on July 19, 2024.
How long did the CrowdStrike outage last?
The bad file was live for 78 minutes, but affected machines needed manual repair; about 99 % of sensors were back online by July 29.
Sources
- CrowdStrike, Root Cause Analysis, Channel File 291 (Aug 6, 2024): https://www.crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf
- CrowdStrike, Executive Summary of the RCA: https://www.crowdstrike.com/wp-content/uploads/2024/08/Executive-Summary_Root-Cause-Analysis_Channel-File-291.pdf
- CrowdStrike, Preliminary Post Incident Review: https://www.crowdstrike.com/en-us/blog/falcon-content-update-preliminary-post-incident-report/
- CrowdStrike, Remediation and Guidance Hub: https://www.crowdstrike.com/falcon-content-update-remediation-and-guidance-hub/
- CrowdStrike, July 19 statement and workaround (archive): https://web.archive.org/web/20240719145915/https://www.crowdstrike.com/blog/statement-on-falcon-content-update-for-windows-hosts/
- Microsoft, Helping our customers through the CrowdStrike outage: https://blogs.microsoft.com/blog/2024/07/20/helping-our-customers-through-the-crowdstrike-outage/
- Delta Air Lines, Form 8-K (Aug 8, 2024): https://www.sec.gov/Archives/edgar/data/27904/000168316824005369/delta_8k.htm
- Adam Meyers, House Homeland Security testimony (Sep 24, 2024): https://homeland.house.gov/wp-content/uploads/2024/09/2024-09-24-HRG-CIP-Testimony-Meyers.pdf
- George Kurtz on X: https://x.com/George_Kurtz/status/1814235001745027317
- Patrick Wardle on X: https://x.com/patrickwardle/status/1814343502886477857
- The Guardian on Fortune 500 losses: https://www.theguardian.com/technology/article/2024/jul/24/crowdstrike-outage-companies-cost
- Hacker News, July 19, 2024: https://news.ycombinator.com/item?id=41002195
- Hacker News, Wardle's analysis: https://news.ycombinator.com/item?id=41021366
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.


Top comments (0)