DEV Community

VLAD
VLAD

Posted on

Ariane 5, Flight 501: The Exception That Destroyed a Rocket

On June 4th, 1996, Europe launched Ariane 5 for the first time. About forty seconds after the flight sequence began, the rocket veered off course, broke up and destroyed itself. The engines were fine. The weather was fine. What failed was a data conversion in code that had no reason to be running.

This story gets retold a lot, usually as "an integer overflow blew up a rocket". The inquiry board's report (chaired by Jacques-Louis Lions, published July 19th, 1996) tells a more useful story. The check actually worked. What nobody had decided was what should happen when it fired.

The box that knows which way is up

To fly straight, a rocket has to know which way it's pointing and how it's moving. On Ariane, that's the job of the inertial reference system, called SRI in the report after its French name. It has laser gyros, accelerometers and its own computer. It sends attitude and velocity data to the main on-board computer, and that computer steers by swiveling the engine nozzles.

Ariane 5 carried two SRIs with identical hardware and identical software. One was active, the other was a hot standby. If the active one failed, the main computer would switch to the backup.

The report also notes that the SRI design, and especially its software, was practically the same as on Ariane 4. It had flown for years. Nobody saw a reason to change software that worked.

Code that had no job

Before launch, the SRI aligns itself: it finds the direction of gravity and finds north from the rotation of the Earth. That only makes sense while the rocket stands still on the pad.

On Ariane 4, the alignment function kept running for a while after the SRI went into flight mode. That was on purpose. If a countdown was held in its last seconds, a full realignment took 45 minutes or more, and the launch window could be lost. Keeping alignment alive a little longer let the team restart the countdown quickly. According to the report, the feature was used once, in 1989.

Ariane 5 had a different preparation sequence and didn't need this. The requirement was kept anyway, to stay in common with Ariane 4. So on Ariane 5 the alignment code ran for about 40 seconds after lift-off, computing values that no longer meant anything.

The number that couldn't get too big

Inside the alignment code there's a value the report calls BH, for horizontal bias. It's related to the horizontal velocity the platform senses. It was held as a 64-bit floating point number, and at one point the Ada code converted it to a 16-bit signed integer. A 16-bit signed integer tops out at 32,767.

The developers knew conversions like this were a risk. The report says they analysed every operation that could raise an exception and found seven variables at risk. But there was a constraint: a maximum workload target of 80% for the SRI computer. So they protected four of the seven conversions. The other three, including BH, were left unprotected, because the reasoning was that those values were either physically limited or had a large safety margin.

For Ariane 4 that reasoning was correct. But:

  • The justification wasn't in the source code. The report says the assumption, although agreed, was in practice hidden from any external review.
  • The Ariane 5 trajectory data was not part of the SRI's requirements and specification.
  • Ariane 5 builds up horizontal velocity about five times faster than Ariane 4 in the early part of the flight.

What happened in flight

The rocket flew normally until about 36 seconds after the start of the main engine ignition sequence. Then BH exceeded what 16 bits can hold. The conversion raised an Operand Error, and nothing in the code handled it.

For an unhandled exception, the SRI specification said: indicate the failure on the databus, store the failure context in memory, and shut the processor down. The report explains where that rule came from. The programme was designed around random hardware failures. If a part breaks at random, you switch the unit off and the backup takes over.

This wasn't random. The backup ran the same software with the same inputs. It had already stopped, one 72-millisecond data cycle earlier, for the same reason. As the report puts it, two critical units that were still healthy got switched off.

Then it got worse. The failed active SRI put a diagnostic bit pattern on the bus, and the main computer interpreted it as flight data. Based on that, it commanded full nozzle deflections on both solid boosters and then on the main engine. The rocket turned at an angle of attack of more than 20 degrees, the boosters started to separate from the core stage, and that correctly triggered the self-destruct system at about 4 km altitude.

The debris fell on about 12 square kilometres of swamp and savanna near the launch site. Both SRIs were recovered, and the failure context was read out of their memory. The four Cluster science satellites on board were lost. The loss is widely reported at about $370 million; that figure is not from the inquiry report.

The same line in C

Here's the part that matters for those of us who will never write flight software. In C#, the same conversion is one cast:

double bh = 40000.0;

short a = (short)bh;            // compiles without a warning
short b = checked((short)bh);   // throws OverflowException
Enter fullscreen mode Exit fullscreen mode

By default, C# evaluates non-constant expressions in an unchecked context. For a double to integer conversion that's out of range, the documentation says the result is an unspecified value of the destination type. No exception, no warning. Your program keeps going with a number that means nothing. In a checked context, the same cast throws OverflowException.

So which one is safer? It's tempting to say "checked, obviously". But Ariane's code was the checked one. The conversion raised the error, and the error is what shut both computers down. The value itself wasn't needed any more. The board points out that the SRIs could have kept providing their best estimates of the attitude data instead.

That doesn't make silent garbage the right answer. It means a check is only half a decision. The other half is what the system does when the check fails: crash, retry, fall back to a safe value, or keep going in a degraded mode. On flight 501, that half had been decided years earlier, for a different kind of failure.

What to take from this

The board made 14 recommendations. Three ideas from the report apply to ordinary code:

  1. Code that isn't needed shouldn't be running. Recommendation R1 is exactly this: switch off the alignment function immediately after lift-off, and more generally, run no software function during flight unless it's needed.
  2. Reused code brings its old assumptions with it. "This value can't get that big" was true for Ariane 4. It wasn't visible in the code, and nobody tested it against Ariane 5's trajectory. The report notes that a ground test with simulated Ariane 5 flight data would have exposed the failure. That test was never run on the SRI.
  3. A backup that runs the same code is no backup against a bug. Redundancy protects you from random hardware faults. Two copies of the same software fail the same way at the same moment.

There's also one position in the report that still sounds bold. The project had treated software as correct until shown to be at fault. The board argued for the opposite view: assume software is faulty until accepted best-practice methods can demonstrate that it's correct.

Ariane 5 recovered. According to ESA, it flew 117 times between 1996 and 2023, and it launched the James Webb Space Telescope. The Cluster satellites were rebuilt and launched in 2000 on Soyuz rockets.

What's the oldest assumption you've found hiding in code you reused? Tell me in the comments.


I make Vlad's Stack, stories and explainers about how the tools you use every day actually work, for people who write code: https://www.youtube.com/channel/UCUO8Uo5LsEy1b9eRkrH6JNg

Top comments (0)