DEV Community

Cover image for The Trade Looked Perfect Until the System Disagreed
Matthew Carlino
Matthew Carlino

Posted on

The Trade Looked Perfect Until the System Disagreed

For a long time, I thought the hardest part of building trading software was the strategy itself. Find a useful signal, test it against historical data, manage the risk, and execute quickly enough. On paper, that sounds like most of the job.

Then I started looking more closely at what actually happens when real money moves through a system.

I remember watching a trading system during a fairly volatile market session. The strategy looked fine. It detected the signal, generated the order, and the dashboard showed what looked like a normal entry. Nothing unusual.

A few seconds later, the numbers stopped lining up.

The exchange showed one position, our internal database showed another, and the risk service was calculating exposure from state that was already outdated. The trading idea itself wasn't wrong. The system around it had lost track of reality.

That was one of the moments that changed how I think about trading technology. A trading platform is not just a strategy engine with a chart and a few APIs attached to it. In practice, it is a distributed system where several services are constantly trying to agree on what happened.

Market data is arriving from one place. Orders are going somewhere else. WebSocket messages are updating account state. Another service is calculating risk. The UI is trying to show everything in real time. Each component has its own delays, failures, retries, and assumptions.

And networks are never as clean as they look during development.

Imagine sending an order to an exchange and losing the connection before the response arrives. Now you have a surprisingly difficult question: did the trade happen?

If you simply send the order again, you might create a duplicate position. If you do nothing, your system may remain uncertain about what it owns. This is why things like unique order IDs, idempotent requests, event logs, and reconciliation processes become extremely important in financial software.

They sound like backend implementation details until one of them fails.

Market data creates another problem. Receiving a price does not automatically mean you received the correct market state. Messages can arrive late or in the wrong order. Connections can drop for a few seconds and reconnect without you immediately realizing that some updates were missed.

Suppose your application sees prices move from 100.21 to 100.28 and then back to 100.24. That looks simple enough, but what if the second message was delayed? What if part of the order book changed while the connection was recovering?

Your trading algorithm may still run perfectly. It is just running on incorrect information.

That is probably one of the more dangerous types of failure because nothing necessarily crashes. The application remains online, logs continue appearing, and the dashboard still looks alive. It is simply becoming wrong quietly.

Latency has a similar problem. People usually talk about trading latency as if the only goal is getting the smallest possible number. Five milliseconds sounds better than twenty milliseconds, so naturally everyone wants five.

But consistency matters just as much.

I would often rather have a system that responds in twenty milliseconds consistently than one that normally responds in five but occasionally takes half a second. Those unexpected delays can be much more damaging than a slightly slower average response.

A signal may be generated using fresh market data, but by the time an order reaches the venue, the market may already have moved. The user interface may still show the original price while the execution system is working with something completely different.

This is why I think the useful measurement is not simply API latency. You need to understand the entire path from receiving market data, calculating a signal, checking risk, sending the order, getting confirmation, and updating the final portfolio state.

Risk management is another area where trading software becomes very different from ordinary applications.

You cannot design the system assuming the strategy will always behave correctly. Eventually there will be a bad deployment, a strange exchange response, an API failure, or some piece of logic that behaves differently under real market conditions.

The risk layer has to assume those things will happen.

Position limits, daily loss limits, exposure checks, circuit breakers, rate limits, and emergency kill switches should not depend entirely on the strategy behaving properly. They need to exist as separate protections around it.

The goal isn't to build software that can never fail. That is unrealistic. The goal is to make sure failures happen inside boundaries you already understand.

AI is making this area even more interesting.

There are already useful applications for AI in trading research, anomaly detection, market summaries, portfolio analysis, and assisting traders with large amounts of information. I expect that to grow quickly.

But I still think there should be a very clear line between AI interpretation and actual financial execution.

An AI model can say that current market behavior looks unusual. It can summarize news, identify patterns, or propose a possible trade. But when real capital is involved, the final execution and risk checks should come from deterministic systems using verified state.

Models are probabilistic. Your account balance should not be.

That separation is going to matter more as trading platforms become increasingly automated.

What I find interesting is that most of this complexity is invisible to the person using the product. A trader may only see a chart, a position, a P&L number, and a buy or sell button.

Behind that simple screen, the system is constantly asking questions.

Did this order already execute? Is this price still valid? Did we miss an event? Is the exchange state different from our state? Are we still inside our risk limits? If this service crashes right now, can we reconstruct everything correctly?

That is where a lot of the real engineering lives.

Writing a strategy that says “buy when this signal crosses a threshold” is relatively easy. Making sure that the word “buy” still behaves correctly during network failures, duplicated events, partial fills, delayed market data, exchange outages, and sudden volatility is much harder.

And I think that is where trading technology is becoming more interesting.

As systems become more automated, connected to more venues, and increasingly supported by AI, the biggest advantage may not come only from having a smarter prediction model.

It may come from building infrastructure that still knows exactly what happened when everything around it starts behaving unpredictably.

Top comments (0)