Geographic redundancy is the architectural claim most mission-critical infrastructure makes and the one most rarely tested — and on August 1, 2026, the Old Trails Fire came within a third of a mile of Mann-Grandstaff VA Medical Center in Spokane, Washington and tested it.
The Old Trails Fire — one of three blazes collectively known as the Spokane Complex Fire, which would go on to force roughly 67,000 people from their homes and destroy 833 structures — never reached the building. Staff evacuated 27 inpatients; patients were transferred to other hospitals and area shelters while the site remained unavailable for 10 days, according to the Spokesman-Review's reporting on the reopening. The medical center — a 46-bed facility that also administers records, pharmacy management, and specialty care including cardiology, oncology, ophthalmology, and urology for a regional VA clinic network — closed completely on August 1. It began reopening August 11. Full operations resumed August 12.
Mann-Grandstaff wasn't destroyed. It didn't need to be.
The facility became unavailable because the location became inaccessible. That's the architectural distinction: physical survival is not operational availability.
That immediately raises the infrastructure question: where does the function run when the primary location cannot?
To be clear about what the public record does and doesn't tell us: we don't know what recovery architecture existed behind Mann-Grandstaff, and this isn't a critique of the VA's continuity planning. What the record does tell us is what happened when the primary location could no longer be used — and that's exactly the scenario every infrastructure architect designs against and rarely tests for.
Multiple Locations Are Not the Same as Geographic Redundancy
Here's the assumption this incident quietly tests: that having more than one location is geographic redundancy.
It isn't. It's the raw material geographic redundancy is built from — nothing more.
Spokane's VA network had multiple sites. The outpatient clinics on East Front Avenue, East 2nd Avenue, and in Spokane Valley remained open through the event. Mann-Grandstaff, however, was the network's 46-bed inpatient facility and housed a broad range of specialty services including cardiology, oncology, ophthalmology and urology. When that location became inaccessible, those functions could not simply continue at the outpatient sites as though nothing had happened.
This is the gap between "we have multiple locations" and "we have geographic redundancy," and it's a gap enterprise infrastructure teams fall into constantly — usually without noticing, because it only becomes visible during the failure. It's the same gap this site has already named at the multi-cloud layer: most claimed failover capability doesn't survive contact with a real event, which is why multi-cloud failover is mostly theater more often than architects want to admit.
The Redundancy Has to Follow the Function
You don't achieve geographic redundancy because a second facility exists 100 miles away. You achieve it when that facility can assume the specific function that failed — with the capacity, the dependencies, and the operating conditions required to actually keep that function running. This is the same assumption underlying most data protection architecture: that continuity, containment, and functional survivability all inherit from having a second site, when they actually have to be engineered into it. The same distinction shows up as a named recovery domain in the Disaster Recovery & Failover Architecture stage — where a technically clean failover can still fail if the dependencies underneath it were never tested independently.
That's a chain, not a checkbox: location → function → capacity → dependencies → recovery. Break any link and the redundancy is theoretical. This is close to the exact shape of what the dependencies recovery plans forget — the chain fails quietly at the link nobody load-tested, not at the link everyone diagrams.
Most infrastructure claims about resilience stop at the first link and call it done. Here's what each claim actually proves, and doesn't:
| Claim | Does it prove geographic redundancy? |
|---|---|
| We have another facility | No |
| We replicate our data elsewhere | No |
| We have a DR site | Not necessarily |
| The DR site is outside the primary hazard zone | Better — still insufficient |
| The alternate site has the required capacity | Now we're getting somewhere |
| Required dependencies can operate there | Stronger |
| The workload or function can actually fail over there | Yes |
| It can sustain the required service level for the required duration | That's the actual test |
Many resilience claims stop at row one or two. The questions that separate real geographic redundancy from an inventory of buildings are further down the table: how much of the failed workload can the secondary location actually absorb, how fast, for which specific services, and under which dependencies — staff, network paths, identity systems, data, physical resources — sustained for how long. A system can pass every failover test in the lab and still fail this way in production, which is the same structural gap the system recovered but the recovery boundary didn't describes from the technical side: recovery succeeding at the component level tells you almost nothing about whether the dependencies crossed the boundary with it.
The Secondary Site Has to Survive the Same Failure
There's a sharper version of this problem, and it's the one that should worry architects who think they've already solved it: two facilities in the same hazard domain aren't geographic redundancy, even when they're in different buildings, different neighborhoods, or different parts of a metro area.
If both locations depend on the same power infrastructure, network carriers, regional workforce, transportation routes, utility infrastructure, cloud region, or supplier base — a regional event can take out the "redundant" pair together. Distance between two sites means nothing if the failure domain is bigger than the distance. This is the same principle that makes blast radius the thing worth measuring instead of distance at the cloud layer — the failure domain is defined by shared dependency, not by how many miles separate two dots on a map.
And the hazard itself is almost beside the point. A wildfire, a hurricane, a flood, an earthquake, a regional power failure, a civil emergency, or an evacuation order can all produce the identical architectural condition: the primary location becomes unavailable while the systems inside it remain perfectly healthy. A recovery architecture tested only against component and system failure has a blind spot when the systems themselves remain healthy but the location containing them becomes unavailable.
Recovery Boundary
Here's the sharper version of what just happened: the systems at Mann-Grandstaff didn't fail. Access to the location containing them did. That's a different boundary than the one most recovery architecture is built around.
Component and system recovery assumes the boundary of failure is the system itself — a server, a cluster, a database. This incident sits outside that boundary entirely: the systems were healthy, but the recovery boundary had already been crossed at the geographic layer. Geographic recovery capacity has to exist outside that boundary, independent of it, or it isn't recovery capacity at all.
This is what Framework #163, the Continuity Execution Boundary names directly: the point at which recovery has been validated as executable and survivable, but the dependencies and ownership decisions required to actually resume operating haven't been tested separately from it. Below that boundary, a successful recovery is treated as evidence that continuity succeeded. Above it, continuity is its own architectural claim — dependent on the recovery that precedes it, but not identical to it.
Mann-Grandstaff is the geographic version of the same gap. The framework's usual failure mode is a clean technical failover undone by an untested dependency — DNS, identity, a certificate, an ownership decision nobody made. Here the untested dependency was physical: the people, equipment, and clinical capability were fine. Access to the location containing them wasn't. Whether that gap had been closed in advance isn't something the public record can tell us — but the shape of the gap is exactly what #163 describes: recovery and continuity are different claims, and only one of them gets tested by default.
Architect's Verdict
Geographic redundancy isn't another building. It's continuity of the function.
Mann-Grandstaff's ten-day interruption illustrates the architectural question that matters: when a critical location becomes unavailable, where does the function run? If the answer is unclear, the redundancy may exist on paper without existing operationally.
The fire was Spokane's. The architectural question isn't.
Originally published at rack2cloud.com



Top comments (0)