On August 1st, 2012, Knight Capital lost more than $400 million in about 45 minutes. It wasn't a hack or a bad bet. It was old code that nobody remembered was still running on one of eight servers.
Everything below comes from the official record: the SEC order against Knight Capital (Release No. 34-70694, October 2013). It reads like a postmortem, and most of its lessons apply far outside finance.
The machine: parent orders, child orders, and a counter
Knight handled about one in ten trades in U.S. listed stocks in 2011–2012. One of its core systems was SMARS, an automated router. It took a big "parent" order, split it into smaller "child" orders, and sent those to exchanges.
The key part was a counter. SMARS tracked how many shares of the parent order had already been filled, and stopped sending children when the order was complete. Remember that counter.
The ghost in the code
Inside SMARS lived an old feature called Power Peg. Knight stopped using it in 2003, but the code was never deleted. It stayed on the servers, still callable.
In 2005, the counter that told Power Peg when to stop was moved to a different place in the code. Power Peg was never retested after that. Nobody checked whether it still knew when to stop. It didn't.
In 2012, the NYSE launched a new retail program starting August 1st, and Knight wrote new code for it. The new code needed a switch, a flag on each order that meant "use the new feature". Instead of adding a new flag, it reused the old flag that used to activate Power Peg. The plan was to replace the old code everywhere, so the flag would simply mean something new.
The deploy: seven out of eight
SMARS ran on eight servers. Starting July 27th, a technician copied the new code to them by hand over several days. Seven servers got it. The eighth didn't, so it still had Power Peg.
There was no second person reviewing the deployment, and no written procedure requiring one. Other teams at Knight had written procedures; the SMARS team didn't.
97 warnings nobody read
At 8:01 a.m., before the market opened, a Knight system started emailing a group of employees about SMARS orders with the error "Power Peg disabled". By 9:30 it had sent 97 of them. They weren't designed as alerts, and nobody generally read them. They named the exact problem.
9:30: 212 orders become millions
When the market opened, orders arrived carrying the reused flag. Seven servers ran the new code. The eighth ran Power Peg, which kept sending child orders because its counter wasn't where it expected. Another part of Knight's system knew the orders were filled, but that information never reached SMARS.
212 customer orders turned into millions of child orders: about 4 million executions in 154 stocks, over 397 million shares.
People inside saw positions piling up in one account with a $2 million limit. But the risk tool only displayed numbers to humans. It wasn't connected to the order system, it didn't raise automated alerts, and nothing cut the machine off.
The fix that made it worse
The engineers reasoned that the trouble started with the new code, so they removed it from the seven servers where it was working. The orders still carried the reused flag. Now all eight servers woke up Power Peg.
The bill
When it stopped, Knight held about $3.5 billion of stock it never wanted to buy and had sold about $3.15 billion it didn't own. Knight first estimated the loss at about $440 million; the SEC later put it at over $460 million. In 37 stocks, prices moved more than 10% with Knight doing most of the trading.
Five days later, investors put in $400 million to keep the firm alive. Less than a year later it merged with a rival. The SEC fined it $12 million for not having the controls the market access rule requires.
What a developer should take from this
None of the individual decisions looked dangerous on the day they were made. That's the point.
- Delete dead code. Code that's switched off is still code that can be switched on.
- One flag, one meaning. Give every new feature its own flag, and plan the day you'll remove it. Reusing a flag makes old code reachable in ways nobody is thinking about.
- Verify every deploy. Automate it, and make every instance report which version it runs. The eighth server can't hide from a script.
- Build the kill switch before you need it. If your code can spend money or send messages at machine speed, stopping it must not depend on first understanding the bug.
- An alert nobody reads is not an alert. If a message can say "Power Peg disabled" 97 times without anyone noticing, it's noise, and noise hides the one message that matters.
What's the oldest dead code you've ever found in production?
I make Vlad's Stack — how the tools you use every day actually work, for people who write code: https://www.youtube.com/@VladsStack
Top comments (0)