DEV Community

Cover image for The Meta outage had no bad decision in it
Trust Boundary
Trust Boundary

Posted on Originally published at trustboundarystudio.com

The Meta outage had no bad decision in it

Originally published at trustboundarystudio.com. The video version, with diagrams, is on YouTube.

Key facts

  • When: 4 October 2021, from about 15:39 UTC. Roughly six hours to full restoration.
  • Scale: Facebook, Instagram and WhatsApp unreachable worldwide. The servers ran the whole time.
  • Trigger: A routine command to assess backbone capacity took down every backbone connection. A bug in the audit tool that should have stopped it did not.
  • Why it became global: Edge DNS servers withdraw their BGP routes when they cannot reach the data centres, by design. With the backbone gone, every one withdrew at once.
  • Why it took six hours: Remote access needed the network. Internal tools needed DNS. The data centres are hard to enter and modify by design. Every recovery path ran through the thing that was broken.
  • Primary sources: Meta Engineering, 4 and 5 October 2021.

On 4 October 2021, Facebook, Instagram and WhatsApp did not go down. They
disappeared. For six hours the rest of the internet could not find out where
they were, and the servers were running the whole time.

Nobody attacked them. And more uncomfortably, nobody made a mistake that looks
like a mistake.

Two pieces of infrastructure

Meta's backbone is the private network connecting their data centres to
each other and out to the smaller facilities at the edge. Every internal system
rides on it.

BGP is how networks tell the rest of the internet which addresses they can
reach. It is not a lookup service. It is continuous advertisement: send traffic
for these addresses to me.
If you stop advertising, you stop existing, as far
as everyone else is concerned.

Meta's edge locations advertise the routes to their DNS servers, and DNS is what
turns facebook.com into an address.

Hold that shape. The backbone carries everything internal, and BGP tells the
world where to find the front door.

The command, and the guardrail that had a bug

Routine maintenance. In Meta's words, a command issued "with the intention to
assess the availability of global backbone capacity." A capacity check.
Read-only in intent. The kind of command that runs constantly at that scale.

It "unintentionally took down all the connections in our backbone network."

Meta had anticipated exactly this, and it is the part most retellings skip. They
had a system whose entire job was to catch commands like this before they
executed: "Our systems are designed to audit commands like these to prevent
mistakes like this."

And then: "a bug in that audit tool prevented it from properly stopping the
command."

The guardrail existed. It was built for precisely this scenario. It had a bug.

So the command ran, and every data centre disconnected from every other data
centre, globally, at once. That is already a serious outage. It is not yet a six
hour disappearance from the internet.

The safety mechanism that worked perfectly

Meta's edge locations run DNS servers, and those servers have a safety
mechanism.

If a DNS server cannot reach the data centres, it might answer with stale or
wrong information. Answering wrongly is worse than not answering. So the design
is: if you cannot talk to the data centres, declare yourself unhealthy and
withdraw your BGP advertisement. Stop attracting traffic you cannot serve
properly.

That is good engineering. If one edge location loses connectivity, it removes
itself and traffic goes elsewhere. It is the correct behaviour, and if you were
reviewing the design you would approve it.

Except the backbone was gone. So every edge location asked the same question at
the same moment, and every one of them got the same answer.

Meta's words: "the entire backbone was removed from operation, making these
locations declare themselves unhealthy and withdraw those BGP advertisements."

Every DNS server withdrew. Simultaneously. Worldwide.

With that, Meta's name servers stopped being reachable from the internet. Not
down. Unreachable. "Our DNS servers became unreachable even though they were
still operational."

The machines were fine. The data was fine. There was simply no longer any route
that led to them.

Then the retries started: billions of devices asking again, and asking harder,
piling load onto DNS infrastructure across the whole internet for a name that
could no longer be answered. If your own service felt slow that afternoon and
you never worked out why, that is why.

Nothing here malfunctioned. The health check did precisely what it was designed
to do. It was designed for one edge location failing, and it was handed all of
them.

Every way back in was already gone

This is where it becomes genuinely uncomfortable.

They could not reach the data centres remotely: "it was not possible to access
our data centers through our normal means because their networks were down."

They could not use their tooling: "the total loss of DNS broke many of the
internal tools we'd normally use to investigate and resolve outages."

Read that twice. The tools for diagnosing an outage were themselves resolved by
DNS. When DNS went, so did the ability to see what was wrong.

So, physical access. And Meta's data centres are "hard to get into, and once
you're inside, the hardware and routers are designed to be difficult to modify
even when you have physical access to them."

A correction, since this is the most repeated detail of the whole incident.
You will read everywhere that engineers could not badge into the building.
Meta's post-mortem does not say that. It says the facilities are hard to enter
and the hardware hard to modify, by design. That is enough to make the same
point, and it has the advantage of being what they actually wrote.

Every one of those properties is a security control working correctly. Hardened
facilities, hardened hardware, no easy console access. On any other day, that is
exactly what you want.

Engineers travelled to data centres, debugged on site, and restarted systems by
hand. Then brought things back gradually, because flipping everything on at once
risked "a new round of crashes due to a surge in traffic." There was a second
reason too: individual data centres were "reporting dips in power usage in the
range of tens of megawatts," and suddenly reversing a dip that size could put
"everything from electrical systems to caches at risk."

Six hours.

What actually failed

Not the command. Not the audit tool bug. That is the trigger, and triggers are
interchangeable. Something will always eventually get through.

The failure was circular dependency. Every path to fixing the problem ran
through the thing that was broken. Remote access needed the network. The tools
needed DNS. The monitoring that would tell you what was wrong was inside the
system that was wrong.

A second failure sits alongside it: an automated safety mechanism with no floor.
Withdrawing routes when unhealthy is right. Withdrawing every route globally,
with nothing that says if every location is failing at the same moment then
this is not a local fault, hold the last known good state and escalate to a
human
, is what turned a bad internal day into a global disappearance.

Three controls, and the order matters.

Out-of-band access. A management path that does not depend on the production
network. Separate connectivity, separate credentials, separate DNS or none at
all. Expensive, boring, unused for years at a time. Also the only thing that
works when the primary path is the thing that failed.

A floor on automated withdrawal. If every health check in the world fails at
once, the problem is probably not the world.

Recovery tooling with no dependency on production. Including the runbook. If
your incident documentation lives behind a single sign-on that depends on the
data centre that is down, you do not have documentation.

Why this one is harder than the others

Meta published a technical post-mortem the next day naming their own guardrail
failure. Almost everything above comes from their document, and that is unusual
enough to be worth saying.

But the reason this incident matters more than its scale suggests is that there
is no villain in it.

The command was routine. The audit tool existed and was the right idea. The DNS
health check was correct design. The hardened data centres were correct
security. Every individual decision was defensible, and several were best
practice.

They composed into six hours of a company not existing.

That is the version of a trust boundary failure that is hardest to find before
it happens, because there is no bad decision to go looking for. There is only a
set of good ones that share a dependency nobody drew on the same diagram.

So the question for your environment is not what could go wrong. It is this:

What would I need in order to fix it, and does that thing depend on what
broke?

Your VPN. Your single sign-on. Your runbook. Your monitoring. Your password
manager.

Go and check. Most people find at least one loop.

Sources

Outage timings are corroborated by third-party network telemetry. No revenue
figure appears here because none of the widely quoted estimates are primary.

Corrections are welcome, and any made are listed, dated, at the end of this article.

The video version

Questions this answers

Was Facebook hacked on 4 October 2021?

No. A routine maintenance command intended to assess backbone capacity took down all the connections in Meta's backbone network, and a bug in the audit tool that should have blocked the command did not. No attack was involved, and Meta's servers were operational throughout.

Why did Facebook disappear from the internet rather than just go down?

Meta's edge DNS servers are designed to withdraw their BGP advertisements when they cannot reach the data centres, so they never answer with stale information. With the backbone gone, every edge location withdrew at the same moment, and the rest of the internet no longer had a route to Meta's name servers. Unreachable, not down.

Why did the Facebook outage take six hours to fix?

Every recovery path depended on the system that had failed. Remote access needed the network that was down. The internal tools needed DNS, which was unreachable. The data centres are hard to get into and the hardware hard to modify by design. Engineers had to travel on site, restart systems by hand, and restore gradually to avoid a fresh round of crashes.

Could engineers not badge into the buildings?

That claim is widely repeated but is not in Meta's post-mortem. What Meta wrote is that the facilities are hard to get into and, once inside, the hardware and routers are designed to be difficult to modify even with physical access. That is enough to make the same point, and it is what they actually said.

Top comments (1)

Collapse
 
eternaclarity profile image
Jesse Gamble •

The 'no bad decision' framing is exactly right, and the badge-in correction matters because the myth hides the real lesson. The closing question belongs on every architecture review: what would I need to fix this, and does it depend on what broke?