On June 12, 2025, a policy row with blank fields reached a code path with no null check, and Google Cloud's API layer crashed in every region at once. The Google outage lasted three hours, touched about 70 Cloud products and 10 Workspace apps, and took Cloudflare's Workers KV down with it, which in turn broke Cloudflare Access, WARP, Workers AI and more. Both companies published detailed postmortems within 30 hours, and together they are one of the cleanest lessons in how a first-week programming bug becomes a global incident.
TL;DR
- Trigger: a policy change with "unintended blank fields" was written to the Spanner tables behind Service Control, the binary that checks every Google Cloud API request. It replicated globally "within seconds".
- Bug: a quota-check code path added on May 29 had no error handling and no feature flag. The blank field hit a null pointer and every Service Control binary went into a crash loop. External APIs returned 503.
- Recovery: root cause in 10 minutes, a "red-button" kill switch fully rolled out in 40. us-central1 took about 2 h 40 min because restarting tasks overloaded the same Spanner table with no randomized backoff.
- Cloudflare: Workers KV's central store sat on "a third-party cloud provider". For 2 h 28 min, 90.22 % of KV requests failed and Access failed 100 % of identity logins.
- Both reports blame themselves. Google: "This incident should not have happened." Cloudflare: "this was a failure on our part".
What caused the Google Cloud outage?
Every call to a Google Cloud API passes a policy check: is this project allowed, is it within quota. In Google's words, "The core binary that is part of this policy check system is known as Service Control. Service Control is a regional service that has a regional datastore that it reads quota and policy information from. This datastore metadata gets replicated almost instantly globally". That sentence contains the whole incident.
On May 29, 2025, Google added a feature to Service Control for additional quota policy checks. It went out through the normal region-by-region rollout, "but the code path that failed was never exercised during this rollout due to needing a policy change that would trigger the code." It shipped with a red-button to turn that path off. From Google's incident report:
The issue with this change was that it did not have appropriate error handling nor was it feature flag protected. Without the appropriate error handling, the null pointer caused the binary to crash.
On June 12 at about 17:45 UTC, "a policy change was inserted into the regional Spanner tables that Service Control uses for policies. Given the global nature of quota management, this metadata was replicated globally within seconds. This policy data contained unintended blank fields." Every regional Service Control read the row, reached the new path, hit the null pointer and crashed. Then it restarted, read the same row and crashed again. Google's summary of the symptom: "increased 503 errors in external API requests".
Google also wrote the counterfactual down: "If this had been flag protected, the issue would have been caught in staging."
Why did one blank field go global in seconds?
Three design choices lined up.
The check is in the request path. Service Control sits in front of every API call. When it crashes, the API is not degraded, it is gone. A policy system that fails closed is safe until the policy system itself is the thing that fails.
Quota data is global and fast. Quota has to be consistent across regions, so the data replicates almost instantly. Code went out region by region; data did not. The rollout protected Google from a bad binary, and nothing protected it from a bad row. That is why one of the fixes on Google's list is about data: "data replication needs to be propagated incrementally".
The dangerous path was dormant. Staged rollouts only test the code paths that real traffic hits. This one needed a specific kind of policy change to run, so for two weeks it passed every stage while never executing. A feature flag that ships off, then turns on gradually, is the tool for exactly that case.
A simplified sketch of the pattern, not Google's code:
# illustrative example: dormant path, missing guard
def check_quota(request, policy):
if new_quota_checks_enabled(): # feature flag, default off
limit = policy.extra_limits.get(request.resource)
if limit is None: # the null check
log.warning("policy row missing field; skipping new check")
return ALLOW # fail open for this check only
if request.usage > limit:
return DENY
return legacy_check(request, policy)
Either guard alone would have been enough.
Timeline of the June 12 Google outage (UTC)
| Time (UTC) | Event |
|---|---|
| May 29 | New quota-policy check ships in Service Control; failing path never exercised; red-button yes, null check no, feature flag no |
| ~17:45 | Policy change with blank fields lands in Spanner, replicates globally within seconds |
| 17:49 | Google incident start: Service Control crash-loops in every region, APIs return 503 |
| 17:51 | SRE triaging "within 2 minutes" |
| 17:52 | Cloudflare incident start: new WARP device registrations fail |
| ~17:59 | Root cause identified; red-button being put in place |
| 18:11 | HN "GCP Outage" thread starts (1,468 points by the end) |
| ~18:29 | Red-button rollout complete; smaller regions recover first |
| 18:46 | Google's first status post, about an hour in |
| 19:32 | @Google support tells a user there are no known disruptions |
| 19:48 | All regions except us-central1 mitigated |
| 20:28 | Cloudflare impact ends (2 h 28 min) |
| 20:49 | Google incident end (3 h) |
| 22:00 | Cloudflare postmortem published |
| Jun 13, 23:45 | Google's full incident report |
The response was fast. "Within 2 minutes, our Site Reliability Engineering team was triaging the incident. Within 10 minutes, the root cause was identified". The red-button was ready at about 25 minutes and fully out at 40.
The status page was slower, for a reason that belongs in every SRE onboarding deck: "We posted our first incident report to Cloud Service Health about ~1h after the start of the crashes, due to the Cloud Service Health infrastructure being down due to this outage." The top comment on Hacker News: "The status page is green, but there are outages reported". And at 19:32, the official @Google account replied to a user:
us-central1 and the thundering herd
Most regions were back by 19:48. us-central1 took up to about 2 h 40 min. From the report: "as Service Control tasks restarted, it created a herd effect on the underlying infrastructure it depends on (i.e. that Spanner table), overloading the infrastructure. Service Control did not have the appropriate randomized exponential backoff implemented to avoid this."
Tasks that restart together hit the same table together, fail together and retry together. Exponential backoff spreads retries out; randomization ("jitter") keeps them from staying synchronised. Without it, recovery itself becomes the load.
The fix is a few lines. An illustrative version:
# illustrative example: exponential backoff with full jitter
import random, time
def backoff_retry(fn, base=0.5, cap=60, attempts=10):
for n in range(attempts):
try:
return fn()
except TransientError:
time.sleep(random.uniform(0, min(cap, base * 2 ** n)))
raise RuntimeError("gave up")
Why the Cloudflare outage happened the same day
At 17:52 UTC, three minutes after Google's crash started, Cloudflare's WARP team saw new devices fail to register. Gergely Orosz posted the question everyone had:
They were not independent. From Cloudflare's postmortem: the cause "was due to a failure in the underlying storage infrastructure used by our Workers KV service, which is a critical dependency for many Cloudflare products … Part of this infrastructure is backed by a third-party cloud provider, which experienced an outage today". The post never names Google.
Workers KV is Cloudflare's key-value store, and many Cloudflare products store their own config and identity data in it: "One of our principles is to build Cloudflare services on our own platform as much as possible". KV is built as a "coreless" service, "However, Workers KV today relies on a central data store to provide a source of truth for data." And the part Cloudflare admitted outright: "Workers KV is in the process of being transitioned to significantly more resilient infrastructure for its central store: regrettably, we had a gap in coverage which was exposed during this incident. Workers KV removed a storage provider as we worked to re-architect KV's backend, including migrating it to Cloudflare R2". Mid-migration, KV was down to one provider, and that provider was the one having the outage.
What failed, per Cloudflare:
| Product | Impact |
|---|---|
| Workers KV | 90.22 % of requests failing |
| Access | 100 % of identity-based logins failed ("designed to fail closed") |
| WARP | no new clients could connect or sign up |
| Workers AI | "All inference requests to Workers AI failed" |
| Pages | errors peaked at ~100 %, all builds failed |
| Stream | error rate above 90 %, Stream Live 100 % |
DNS, cache, the proxy and WAF were not affected, and "No data was lost".
Access fails closed on purpose, the right default for a security product, and the reason a storage outage locked people out of their own apps.
Who is to blame? The git blame split
As in every postmortem episode, my split of the production failure. Each slice is a finding from the primary reports; the percentages are my read.
- Google, 55 %. No null check, no feature flag, no randomized backoff. Their own report says a flag would have caught it in staging.
- Global replication, 20 %. One row, every region, seconds. Google's own fix: changes "propagated incrementally with sufficient time to validate".
- Cloudflare, 20 %. Half a product line on one vendor's store, a gap it knew about. Its own words: "we are ultimately responsible for our chosen dependencies and how we choose to architect around them."
- The status page, 5 %. Hosted on the thing it reports on.
What Google and Cloudflare changed after the outage
Google's list, from the full report: modularise Service Control so it "fails open"; audit "all systems that consume globally replicated data"; "enforce all changes to critical binaries to be feature flag protected and disabled by default"; improve static analysis and testing; "audit and ensure our systems employ randomized exponential backoff"; improve external communication; and keep monitoring and comms "operational … even when Google Cloud … [is] down".
Cloudflare's: remove "the dependency on any single provider" for KV, blast-radius work per product, and tooling to "progressively re-enable namespaces" so Access and WARP can come back "without risking a denial-of-service against our own infrastructure". That last one is the thundering-herd lesson again, learned independently the same afternoon: at 20:23 Cloudflare's recovery hit "infrastructure rate limits due to the influx of services repopulating caches".
For your own systems, the checklist is short:
- New code paths ship behind a flag that defaults to off.
- Validate data where it comes in, and treat config and policy data like code and roll it out in stages; a global push in seconds is how this one spread.
- Decide per dependency whether it fails open or closed, and write the decision down.
- Every retry loop gets exponential backoff with jitter.
- Host your status page and your monitoring somewhere your outage cannot reach.
- Know which of your "own" services sit on someone else's cloud.
Verdict: SHIP IT
I stamped the response SHIP IT. Both postmortems landed within about 30 hours, both blame themselves in plain words, and Google's fixes are concrete: fail open, flags off by default, randomized backoff, incremental replication. CEO Thomas Kurian posted that night: "We regret the disruption this caused our customers." The one thing I'd hold against Cloudflare is a matter of style: a postmortem that says "third-party cloud provider" about the outage everyone was watching.
The Monday line from the episode: new code behind a flag that ships off, and a null check where the data comes in.
FAQ
What caused the Google Cloud outage on June 12, 2025?
A policy change with blank fields replicated globally and hit a new Service Control code path with no null check and no feature flag. Service Control crashed in every region, and Google Cloud APIs returned 503 errors for up to three hours.
Why was Cloudflare down at the same time as Google Cloud?
Workers KV's central data store relied on a single third-party cloud provider during a migration. When that provider went down, KV failed, and Cloudflare products that depend on KV, including Access and WARP, failed with it.
What is randomized exponential backoff?
A retry strategy where each retry waits exponentially longer, with random jitter so many clients do not retry at the same moment. Google's report names its absence as the reason us-central1 recovered last.
Sources
- Google Cloud incident report, June 12–13, 2025: https://status.cloud.google.com/incidents/ow5i3PPK96RduMcb1SsW
- Cloudflare, "Cloudflare service outage June 12, 2025": https://blog.cloudflare.com/cloudflare-service-outage-june-12-2025/
- Cloudflare status, "Broad Cloudflare service outages": https://www.cloudflarestatus.com/incidents/25r9t0vz99rp
- Hacker News, "GCP Outage": https://news.ycombinator.com/item?id=44260810
- Hacker News, "Cloudflare was down": https://news.ycombinator.com/item?id=44261064
- Hacker News, Google Cloud incident report thread: https://news.ycombinator.com/item?id=44274563
- Gergely Orosz on X: https://x.com/GergelyOrosz/status/1933240698716729511
- @Google support reply on X: https://x.com/Google/status/1933246051512644069
- Thomas Kurian on X: https://x.com/ThomasOrTK/status/1933337436970709493
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.



Top comments (0)