DEV Community

Cover image for What is BGP? How Facebook's 2021 outage took it off the internet

What is BGP? How Facebook's 2021 outage took it off the internet

On October 4, 2021, Facebook, Instagram, WhatsApp and Messenger disappeared from the internet for almost six hours. The cause was not a hack or a DNS provider going down. A routine maintenance command cut every link between Facebook's data centres, and Facebook's own DNS servers then did what they were designed to do: stop announcing their BGP routes. If you have ever asked what BGP is and why it matters to you, this is the outage that answers it, and the lessons apply to any system with a health check and a single network.

TL;DR

  • A command meant "to assess the availability of global backbone capacity" took down all of Facebook's backbone connections. The audit tool that should have blocked it had a bug, per Meta's postmortem.
  • Facebook's DNS servers withdraw their BGP announcements when they cannot reach the data centres. With the whole backbone gone, every one of them did it at once. The servers were running; nobody on the internet could find them.
  • From outside, Cloudflare saw facebook.com stop resolving at about 15:50 UTC and come back at 21:20 UTC. Resolvers worldwide handled 30 times their usual queries as apps retried.
  • Internal tools, out-of-band access and, per the New York Times, office badges ran on the same network, so engineers had to reach the data centres in person.
  • My verdict: NEEDS REVIEW. A clear, fast postmortem; a thin list of fixes.

What is BGP?

The internet is tens of thousands of independent networks, called autonomous systems (AS), each with a number. Facebook's is AS32934. BGP, the Border Gateway Protocol, is how those networks tell each other which blocks of IP addresses (prefixes) they can deliver traffic to. An announcement says "send traffic for this prefix to me". A withdrawal says "I can't reach this any more, forget the route".

Every router on the internet builds its map from those messages. If nobody announces a prefix, it drops out of every routing table and packets for it have nowhere to go. That is the whole mechanism of this outage: Facebook's name servers were still running, but the prefixes they lived in were no longer announced, so to the rest of the world they did not exist.

DNS sits on top. To open facebook.com, your resolver (your ISP, 1.1.1.1, 8.8.8.8) asks Facebook's authoritative name servers for the address. If the resolver cannot reach any of them, it returns SERVFAIL.

The Facebook outage 2021 timeline

All times UTC, October 4, 2021, from Cloudflare's post, Meta's two notes and the posts linked below.

Time What happened
≈15:39 The backbone command runs; every link between data centres goes down
≈15:40 Cloudflare sees "a peak of routing changes from Facebook. That's when the trouble began."
15:45 Hacker News: "Facebook-owned sites were down" (2,589 points)
≈15:50 facebook.com stops resolving on 1.1.1.1
15:51 Cloudflare opens an internal incident: "Facebook DNS lookup returning SERVFAIL", fearing its own resolver is broken
15:58 Cloudflare confirms Facebook "had stopped announcing the routes to their DNS prefixes"
16:07 Facebook spokesman Andy Stone posts on Twitter
18:51 NYT reporter Sheera Frenkel: employees can't enter buildings
19:52 CTO Mike Schroepfer apologises
≈21:00 Renewed BGP activity from Facebook, peaking at 21:17
21:20 facebook.com resolves on 1.1.1.1 again; by 21:28 "Facebook appears to be reconnected"

Cloudflare checked the public route collectors and found the prefixes of Facebook's DNS servers simply gone. From its post:

route-views>show ip bgp 185.89.218.0/23
% Network not in table
route-views>

route-views>show ip bgp 129.134.30.0/23
% Network not in table
route-views>
Enter fullscreen mode Exit fullscreen mode

Cloudflare's post at 15:58 UTC: Facebook had stopped announcing the routes to their DNS prefixes, with the route-views output showing

Other Facebook prefixes stayed routed, Cloudflare notes, but they "weren't particularly useful" without DNS: nobody could look up the addresses to reach them.

The first public word came 27 minutes in, on a competitor's platform:

Andy Stone on Twitter, 16:07 UTC:

"Some people" was, per the New York Times, the more than 3.5 billion who use Facebook, Instagram, Messenger and WhatsApp.

Why did Facebook's DNS servers withdraw their own routes?

This is the part of the Facebook BGP story that is a deliberate design decision. Facebook's authoritative DNS servers sit in smaller edge facilities and check their own health. Meta's postmortem, verbatim:

"To ensure reliable operation, our DNS servers disable those BGP advertisements if they themselves can not speak to our data centers, since this is an indication of an unhealthy network connection. In the recent outage the entire backbone was removed from operation, making these locations declare themselves unhealthy and withdraw those BGP advertisements. The end result was that our DNS servers became unreachable even though they were still operational."

In code, the logic is roughly this. A simplified sketch, not Facebook's:

# illustrative example of a per-site health check that controls a BGP announcement
def tick(site):
    if site.can_reach_data_centers():
        site.announce(DNS_PREFIX)      # healthy: attract DNS traffic here
    else:
        site.withdraw(DNS_PREFIX)      # unhealthy: let another site take it
Enter fullscreen mode Exit fullscreen mode

For one bad site this is correct: withdraw, and routing sends users to another, healthy site. The rule has no idea what to do when every site fails the check at once, because the data centres themselves are unreachable. Each site made a locally sensible decision, and together they removed Facebook's DNS from the internet. A per-site safety feature with no global brake turned a network fault into an existence fault.

Why a backbone command could reach everything

The trigger, again from Janardhan's post: "During one of these routine maintenance jobs, a command was issued with the intention to assess the availability of global backbone capacity, which unintentionally took down all the connections in our backbone network, effectively disconnecting Facebook data centers globally."

And the part that should have stopped it: "Our systems are designed to audit commands like these to prevent mistakes like this, but a bug in that audit tool prevented it from properly stopping the command."

Meta did not publish the command, the bug or why a capacity check could take links down. The first note, the evening of the outage, said the cause was "configuration changes on the backbone routers" and that there was "no malicious activity behind this outage" and "no evidence that user data was compromised".

Locked out: out-of-band access on the same network

Once the backbone was down, fixing it from a desk was impossible. The postmortem: "it was not possible to access our data centers through our normal means because their networks were down, and second, the total loss of DNS broke many of the internal tools we'd normally use to investigate and resolve outages." And: "Our primary and out-of-band network access was down, so we sent engineers onsite to the data centers."

The physical layer didn't help either:

Sheera Frenkel on Twitter, 18:51 UTC: employees were unable to enter buildings because their badges weren't working to access doors

Once inside, the hardware fought back as designed: data centres are "hard to get into, and once you're inside, the hardware and routers are designed to be difficult to modify even when you have physical access to them." Security hardening built to slow an attacker slowed the owner by the same amount. Meta, to its credit, says so: "it was interesting to see how that hardening slowed us down … I believe a tradeoff like this is worth it".

Why recovery took hours: the power problem

Restoring the backbone did not mean flipping everything back on. "Individual data centers were reporting dips in power usage in the range of tens of megawatts, and suddenly reversing such a dip in power consumption could put everything from electrical systems to caches at risk." Facebook ramped traffic up deliberately, using procedures from its "storm" drills. But, the post admits, "we've never previously run a storm that simulated our global backbone being taken offline".

Meanwhile the rest of the internet was hammering on the door. Cloudflare saw "DNS resolvers worldwide handling 30x more queries than usual" because every app retried and every user reloaded. It also saw more DNS queries for Twitter, Signal and other messaging apps as people went elsewhere. Telegram founder Pavel Durov claimed more than 70 million new users that day.

git blame: who is at fault

My split, each slice pinned to Meta's own postmortem:

Share Who Why
55 % Facebook's change process "a bug in that audit tool prevented it from properly stopping the command"; one command reached the whole backbone
25 % the DNS design self-withdrawal is right for one site, catastrophic for all of them at once
15 % one network for everything tools, out-of-band access and, per the NYT, badges shared the network that failed
5 % the hardened racks "designed to be difficult to modify even when you have physical access"

Blast radius: 3.5 billion users, about five and a half hours on Cloudflare's resolver and "nearly six hours" in the press, per The Verge. The stock closed down 4.9 %, and Forbes estimated Mark Zuckerberg's paper loss at $5.9 billion. The timing was bad too: the day after whistleblower Frances Haugen's 60 Minutes interview and the day before her Senate testimony. Zuckerberg's note to employees called it "the worst outage we've had in years".

Lessons from the Facebook BGP outage for your own systems

  • Give health checks a sense of proportion. If a check can withdraw a service, ask what happens when every instance fails it at the same moment. Often the right answer is "keep serving, stale"; a fleet that disappears all at once is the outage.
  • Out-of-band means a different network. Consoles, runbooks, chat, the VPN and the door should work when production is gone. If they depend on your own DNS or backbone, they are in-band with extra steps.
  • Audit the auditor. A safety tool that blocks dangerous commands is itself production code. Test that it actually refuses the worst command you can think of.
  • Drill the total failure. Facebook had storm drills and had never simulated the backbone disappearing. That is the one worth rehearsing.
  • Watch BGP from outside. Your own monitoring may be behind the failure. Public route collectors and third-party resolvers saw this within minutes.

Verdict: NEEDS REVIEW

I stamped the response NEEDS REVIEW. The communication was good: a statement the same evening and, the next day, a named-author postmortem that admits the audit-tool bug and the design trade-off in plain English. The fix list was thin: "strengthen our testing, drills, and overall resilience". Nothing about moving out-of-band access and the doors off the production backbone, and no global brake on DNS servers withdrawing themselves. Monday line: put your out-of-band access, and your door, on a network you do not operate.

FAQ

What caused the Facebook outage on October 4, 2021?
A maintenance command took down all of Facebook's backbone links; an audit tool bug failed to stop it. Facebook's DNS servers then withdrew their BGP routes, making facebook.com unresolvable.

Was the Facebook outage a hack?
No. Meta said there was "no malicious activity" and "no evidence that user data was compromised".

How long was Facebook down in 2021?
About five and a half hours on Cloudflare's resolver (15:50 to 21:20 UTC), reported as nearly six hours.

What does SERVFAIL mean?
A DNS resolver's answer when it cannot get a valid response from the domain's authoritative name servers. On October 4 it was the answer for facebook.com everywhere.

Sources


This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Top comments (0)