On March 14, 2024, four subsea cables off Côte d'Ivoire failed almost simultaneously, and Microsoft started warning customers in its South Africa North and South Africa West regions about rising latency and dropped packets. Nobody had pushed a bad config, and no certificate had expired. Some operators failed over automatically while others spent hours buying capacity on cables that still worked, but the broken glass itself had to wait for the small, aging fleet of repair ships that quietly keeps the internet running, and that wait is the one part of our stack almost nobody writes into a design doc. If your product crosses an ocean, there is a ship somewhere in your dependency graph, and it is worth knowing how that ship schedules its work.
Four Cables, One Failure Domain
WACS, MainOne, SAT-3 and ACE look like four-way redundancy on a network diagram. On the seabed they were neighbors. MainOne's preliminary analysis ruled out human activity, given the depth of the fault, and pointed to movement of the seabed itself, which means a single geological event took out every "independent" path at once. Engineers meet this pattern constantly at smaller scale: two availability zones fed by one substation, two DNS providers managed through one registrar account, two "diverse" fiber routes squeezed through the same bridge. Redundancy is not a count of paths. It is a claim that their failures are uncorrelated, and geography is extremely good at falsifying that claim.
The seabed correlates failures in time, too. Research published in Nature Communications followed underwater avalanches of sand and mud, known as turbidity currents, down the Congo Canyon off West Africa. One flow in January 2020 ran for more than 1,130 kilometers, accelerating from about 5 to 8 meters per second, and snapped two seabed telecom cables, slowing the internet from Nigeria to South Africa. Those cables had gone roughly 18 years without such a break, and then they broke again in March 2020, April 2021 and January 2022. The trigger was not an earthquake but the biggest Congo River floods in decades, with the destructive flows arriving weeks to months after the flood peak, often at spring tides. Some cables even survived flows of similar speed that snapped their neighbors, most likely because erosion was patchy and concentrated around steep steps in the canyon floor. Read that with on-call eyes: failures arrive in bursts after an observable precursor, and small differences in placement decide which replica lives.
Mean Time to Repair Is a Queue
Logical recovery happens in routers. Physical recovery happens on a ship, and ships are a scarce shared resource with a scheduler. In July 2026, the final report of an international advisory body on submarine cable resilience, set up by the ITU (the UN agency whose X.509 standard defines the certificates in your TLS stack) and the International Cable Protection Committee, put hard numbers on that scheduler. Between 2012 and 2024, the average time to begin a repair more than doubled, from under 20 days to over 50, even though the yearly number of repairs and of available ships stayed flat; permitting delays take most of the blame. About 70 vessels can technically fix a cable, yet only 20 to 25 are dedicated to repair work or on long-term charter, and a significant share were built in the 1980s or 1990s. Asia-Pacific and the North Atlantic each host roughly 40% of the active repair fleet, while Africa, the South Atlantic and the Pacific islands share about 13%. A fault in the North Atlantic is often handled within 3 to 5 days, while a comparable one off West Africa can take 12 to 18.
Africa shows the sharpest version of this. The whole continent has one permanently stationed, dedicated repair ship, Orange Marine's Léon Thévenin, built in 1983 and based in Cape Town. In 2024 it left port on March 19 for the Côte d'Ivoire breaks, where it took SAT-3 while another vessel, CS Sovereign, was assigned the other three cables. It was back in Cape Town by April 25. On May 12, Seacom and EASSy broke off the KwaZulu-Natal coast, cutting all subsea capacity between South Africa and East Africa, and on May 14 the same ship sailed again. One server, two coasts, strictly serialized. Orange Marine has since ordered two new ships to replace it and a 1987-built vessel in Italy, with delivery planned for 2028 and 2029.
If you have ever watched a worker pool saturate, you know the curve. In the textbook single-server queue, the average wait equals utilization divided by one minus utilization, measured in service times: at 50% busy a new job waits about one repair, and at 90% busy it waits nine. Cable repairs are long jobs, since every splice keeps the ship holding position for a day or more and needs a weather window to do it. That is why correlated faults hurt twice. They arrive together, and then they queue together.
What a Cut Looks Like From Your Side
A cable cut rarely shows up as a red outage banner. It shows up as gray failure: health checks pass, error rates look tolerable, and users in one geography quietly have a terrible day. The physics explains why. Light in fiber covers roughly 200,000 km per second, so every extra 1,000 km of detour adds about 10 ms of round-trip time. A single TCP connection can move at most one window of data per round trip, so doubling the RTT roughly halves what that connection can push, before you even count congestion on the detour. NetBlocks' research director made exactly that point during the March 2024 outage: as networks route around damage, they can drain capacity that other countries depend on, so the first problem may be physical while the ones that follow are technical. Meanwhile your client timeouts, tuned to last month's p99, start firing, retries pile onto an already congested path, and a distant event on the ocean floor becomes a load spike you generated yourself.
A Pre-Mortem for Anything That Crosses an Ocean
You cannot buy a cable ship, but you can stop assuming repair capacity is infinite. A few habits make the difference:
-
Map physical paths, not logical ones. Run
mtr -rwc 100 <host>from where your users actually are, compare the hops with a public submarine cable map, and ask your providers which systems carry your traffic. - Budget timeouts for the detour. Derive deadlines from a latency budget that survives an extra 100 to 150 ms of RTT instead of copying last month's p99.
- Make retries polite. Use exponential backoff with jitter and a retry budget, so a reroute does not turn into a denial-of-service attack you launched on yourself.
-
Rehearse the slow path. On a staging host,
sudo tc qdisc add dev eth0 root netem delay 150ms 20ms loss 1%simulates a long detour; watch what breaks first, then clean up withsudo tc qdisc del dev eth0 root. - Alert on latency steps by region. A sudden, flat jump in RTT from one geography is the signature of a reroute, and error-rate alerts alone will sleep through it. Public trackers such as Cloudflare Radar usually show within hours whether the problem is yours or the whole region's.
- Write the runbook in weeks, not hours. Decide in advance what you will degrade for users behind a cut, such as upload sizes, media quality or sync frequency, because the fix might be a ship on the other side of a continent.
The Slowest Component Wins
Our services recover in milliseconds, routing converges in minutes, and the seabed heals as fast as an old ship can clear its queue. That last number has been growing for more than a decade, and the incidents above show it does not care how many lines your architecture diagram has. Treat repair capacity like any other shared dependency you do not control: know where you lean on it, make the slow path survivable, and design as if one day your ticket will be third in line.
Top comments (1)
tr.ee/dev-to