DEV Community

Cover image for Telstra outage explained: how one GPS card set the network to 2006

Telstra outage explained: how one GPS card set the network to 2006

The Telstra outage of July 8, 2026 started with a routine repair. At 02:50 an engineer in Melbourne powered a timing chassis back on, and by breakfast Australia's biggest mobile network agreed it was November 2006. Nearly nine million customers lost calls, texts or data. The external investigation Telstra published on September 2 traces it to one GPS card and an NTP server design that let every other clock agree with it. If you run anything that trusts a clock, which is everything, this is the postmortem to read.

TL;DR

  • A GPS card in a Melbourne NTP chassis rebooted during a power-supply swap. Its firmware lacked the GPS week-rollover fix, so it came up 1,024 weeks in the past: November 2006.
  • Nine months earlier, that card had been switched on to stop a flapping time source. It made Melbourne stratum 1, the same rank as Australia's national time source, and nobody reviewed the change.
  • NTP did what it is designed to do: the lowest stratum wins, outliers get outvoted. Every source that could have disagreed sat downstream of the bad one.
  • Result: 45 % of all calls and data sessions affected that day, 604 Triple Zero calls with errors, trains and card terminals down. A possible fine of up to $30 million, over a firmware update the vendor had published.
  • Telstra hired outside auditors, published the report and moved all three sites off the old servers. I stamped the response SHIP IT.

What happened in the Telstra outage: the timeline

The report is by Technology Audit Partners (TAP), 10 pages. Telstra CEO Vicki Brady's summary post opens with: "It was an outage that should not have happened." The best technical read is by Sven-Christian Ebenhag at Netnod, which runs Sweden's national time distribution. The story spans sixteen years:

When What happened
2010 Design: three national (NMI) stratum-1 sources feed two stratum-2 servers (Sydney, Melbourne), which feed three stratum-3 servers (Sydney, Melbourne, Perth). Time only flows down. TAP: "fit for purpose".
2020 New timing chassis. It cannot feed stratum 3 from stratum 2 in the same box, so Sydney and Melbourne get cross-wired and each is left with one stratum-2 source. A GPS card is installed but never configured. The architecture moves from client/server to peering. TAP found "no evidence" anyone investigated the looping risk.
Oct 2025 Melbourne stratum 3 loses its Sydney source "multiple times per day" and shows up as stratum 5, taking time from a client of its own tree. No problem ticket. The fix: activate the idle GPS card. Melbourne becomes stratum 1.
Jul 8, 02:50 Planned change to replace a faulty power supply. The chassis power-cycles, the GPS card resets, and its old firmware picks the previous GPS epoch: November 2006. The Melbourne stratum-2 server was already switched off, so no peer could contradict it.
~04:00 The Mobility Management Entities (MMEs), which handle phones attaching to the network, start resetting. The remaining capacity cannot cope.
04:20 "NTP alarm storm". Those alarms were not in standard monitoring; they were "checked during business hours … by a limited number of individuals".
04:29 Telstra detects the outage from customer reports, "based on customer web searches for Telstra outages".
04:47 Major incident bridge opens. Finding the root cause "took several hours".
12:12 Most postpaid services restored; residual issues until July 11.

One detail from §1.9 of the report explains the several hours:

TAP findings §1.9: the two primary NTP engineers who were involved in the planned change were on mandatory standdown by the time the impact of that change became known

The people who knew what had changed were on mandatory rest. The rule is sensible. It just means your documentation has to know what the engineers know, and TAP found "no golden configuration, no documented configuration record".

How does an NTP server decide which clock is right?

NTP, the network time protocol, arranges clocks in strata. Stratum 0 is a reference clock (an atomic clock, a GPS receiver). A server attached to one is stratum 1. Anything syncing from a stratum-1 server is stratum 2, and so on. When a server has several sources, two defences matter here:

  • Lower stratum carries more weight. A stratum-1 source beats a stratum-3 one.
  • Outliers get voted out. If one source disagrees with the rest, NTP treats it as the false ticker.

Netnod's summary of why neither helped: "NTP does not ask whether a date is plausible; it asks whether a source disagrees with the others." On July 8 the GPS card made the broken source the highest-ranked clock in the tree, and every clock that might have disagreed took its time from Melbourne. There was nobody left to outvote it. Netnod's line: "The protocol worked. The architecture did not."

The 2020 switch to peering is what made loops possible. In ntpd syntax the difference is one word per line. A simplified illustrative config, not Telstra's:

# client/server: take time from above, never give it back
server ntp1.upstream.example iburst
server ntp2.upstream.example iburst

# peering: two servers at the same level may sync from each other
peer   ntp-mel.internal.example
Enter fullscreen mode Exit fullscreen mode

Peering is meant for servers of equal rank to back each other up. When one peer has quietly become the only real source for the other, which is what the 2020 cross-wiring did, the network can end up voting with its own echo. Netnod again: "Two servers fed by the same GNSS receiver is still one source, but counted twice."

GPS week rollover: why the card thought it was 2006

GPS broadcasts the week number in a 10-bit field. Ten bits count to 1,023, so every 1,024 weeks, just under twenty years, the counter wraps to zero. GPS time started in January 1980. The wraps so far were August 1999 and April 2019.

A receiver that is running when the wrap happens just keeps counting. A receiver that cold-starts has to guess which 1,024-week era it is in, and that guess lives in firmware. Older firmware assumes an older era. Per TAP §2.2, the Melbourne card came up 1,024 weeks in the past, which from July 2026 lands in November 2006.

The irony: the card survived the real rollover in 2019, because nobody turned it off. What killed it was a repair. The fix was in vendor bulletins from 2022 and January 2026, per the Guardian's Senate coverage, and never installed. The CEO told the Senate the box was a 2011 timing chassis costing about $30,000 to replace. (The press disagrees on the model; ACS Information Age names a different one.)

Why did the whole network fail along with the clocks?

A mobile core runs on agreed time. When it jumped nineteen years, the MMEs reset, and the rest of the capacity could not absorb the load; TAP says "the network was overloaded". Worse, the Diameter Signalling Controller "incrementally corrupted" about 30,000 IP address prefixes by writing them with the wrong date. Those stayed blocked regardless of capacity and had to be repaired with the vendor.

The blast radius, from the Guardian and The Register:

  • 45 % of all calls and data sessions on the day, per Brady's Senate evidence. Nearly nine million customers.
  • Triple Zero: 58,835 emergency calls connected, 604 "experienced an error".
  • Payments: about 80,000 businesses' Tyro card terminals on 4G could not take payments.
  • Transport: V/Line suspended all regional rail in Victoria into Thursday.
  • Money: about 8,000 compensation claims in the first week, and a possible fine of up to $30 million under the law passed after the Optus outage, for a free update missing from a $30,000 box.

It was not a surprise risk, either. The ABC reported that government alerts in 2024 and October 2025 warned about dependence on satellite timing, and that a Swinburne professor had pitched this scenario to Telstra earlier in the year.

git blame: who is at fault

My split, from TAP's findings:

Share Who Why
55 % Telstra Nobody owned the clock, so nobody reviewed the October change; two vendor bulletins unread; alarms read in business hours
25 % Peering mode The 2020 switch let servers vote with their own echo
15 % The firmware The rollover fix was published and never installed
5 % The power supply For failing at ten to three on a Wednesday

TAP's first finding is the one that matters: Telstra did not treat network timing as a critical, owned capability, a "sovereign function" in the report's words. Everything else followed. Netnod: "Each decision solved the problem in front of it. Nobody was asked to look at the sum of all actions." On Hacker News, rcaught quoted that line and added: "This is the perfect description of Telstra as a company."

How to check your own NTP sources

The HN comment I would pin is ipython's: "has no one heard of ntptrace?" You can do this Monday:

ntpq -p          # ntpd: your peers, their stratum and offset
chronyc sources  # chrony: the same view
ntptrace         # ntpd tools: follow the chain up to stratum 1
Enter fullscreen mode Exit fullscreen mode

Then ask the questions TAP asked:

  • Count independent sources, and ignore the server count. Two servers behind one GPS receiver are one source wearing two hats.
  • Draw the tree. If any server can end up taking time from its own clients, you have a loop waiting for a bad day.
  • Put time alarms in 24/7 monitoring. Telstra's alarm storm started nine minutes before customer searches told it something was wrong.
  • Treat a stratum change as a change. The October 2025 fix made one card the most trusted clock in the country, with no review.
  • Read the firmware bulletins for anything with a GPS receiver. The rollover is known, dated and patched.

Verdict: SHIP IT

I stamped Telstra's response SHIP IT; the outage itself gets no stamp. Telstra hired outside auditors, published their report with the line that it never treated time as critical, and had moved all three sites off the old NTP servers "to our strategic system" before it landed, with new alarms and lab testing of changes. Netnod wrote that "the way Telstra handled the aftermath deserves praise", and I agree. The regulator's investigation is still open and the fine is still possible.

FAQ

What caused the Telstra outage in July 2026?
A GPS card in a Melbourne NTP server rebooted with old firmware and set the time to November 2006. The network's design let that clock outrank and outvote all others.

What is the GPS week rollover?
GPS counts weeks in a 10-bit field that wraps every 1,024 weeks. A receiver with outdated firmware can cold-start in the wrong 1,024-week era, about twenty years off.

What is an NTP timing loop?
When servers can sync from each other, one can end up taking time from its own downstream clients. Then the network agrees with itself instead of an external reference.

How many people did the Telstra outage affect?
Nearly nine million customers, and 45 % of all calls and data sessions that day, per Telstra's CEO.

Sources


This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Top comments (0)