DEV Community

Sam Hartley
Sam Hartley

Posted on

My Drawdown Breaker Replayed to Zero Fires — Because It Was Watching a Number That Never Moved

My Drawdown Breaker Replayed to Zero Fires — Because It Was Watching a Number That Never Moves

I built a rolling drawdown breaker for my trading bot a few weeks ago. The idea is boring and correct: if equity drops more than X percent from its recent peak, stop opening positions. I wired it in, ran it in paper, and then did the responsible thing — I replayed the last month of real trades through it to see how it would have behaved.

Zero fires.

I sat with that for a while, and the conclusion I almost wrote down was: the breaker is useless, the strategy never drew down enough to need it. That conclusion would have been wrong, and it would have been wrong for a reason I've now added to a checklist.

"Zero" is an ambiguous result

There are two reasons a guard can have never fired:

  1. Nothing ever crossed the threshold — the guard is fine, the market was kind.
  2. The input the guard watches can't move — the guard is dead and you've been testing a corpse.

They look identical in the output. Both print 0. And the second one is much worse, because it doesn't just mean "no protection" — it means every replay, every backtest, and every paper run you did while believing in that guard was measuring nothing.

So the question I should have asked first isn't "did it fire?" It's "could this number have triggered it at all?"

Auditing the input

The breaker reads an equity series. Mine came out of the trader state and was synced from the exchange's account endpoint. On my exchange's futures API there are two different totals, and I had grabbed the wrong one:

  • total — the account's margin balance. It includes realized funding and PnL.
  • cross_margin_balance — the balance marked to market, i.e. total plus the open position's unrealized PnL.

I was feeding the breaker total. That number moves when a trade closes, when funding settles, when you deposit. It does not move when your open position bleeds in real time. From the breaker's point of view, a position could halve in price and the equity curve would sit perfectly flat until the moment of exit — and a rolling drawdown that only observes realized points will, by construction, almost never cross a threshold.

I checked the correlation directly, because that's a claim I could verify rather than reason about:

import numpy as np

# delta of state equity (what the breaker watched)
# delta of the same-moment unrealized PnL (what actually happened)
corr = np.corrcoef(d_eq, d_unrealised)[0, 1]
print(corr)   # -> 0.015
Enter fullscreen mode Exit fullscreen mode

0.015. The two series are statistically unrelated. My equity curve had a 6-figure open loss sitting in it and recorded a gentle flat line. The breaker was reading a number that, at the scale and frequency of the drawdowns I cared about, essentially cannot move.

Re-running against a live input

Then I re-ran the exact same replay, same trades, same thresholds, but fed it mark-to-market equity (equity + unrealized):

input: state.equity (as shipped)     -> 0 fires
input: equity + unrealised (MTM)     -> 17 fires
Enter fullscreen mode Exit fullscreen mode

Seventeen. And not scattered noise — clustered on one position, on one afternoon:

With a 5% threshold, the breaker fires repeatedly on 30.09 between 15:00 and 17:00. That is exactly the window where a single open position went from a normal drawdown to a multi-hundred-dollar loss.

So the breaker wasn't useless. It was correct. It just had no eyes.

The trap underneath the trap

And here's the part that still annoys me. Once the input was fixed, my first instinct was to ask "so would it have helped?" I ran the counterfactual — what would the account look like if the breaker had force-closed at the moment it fired?

  • Exit at 30.09 17:00: −801.82
  • Actually held to now: −576.87

Holding won by about +225. Which made me want to write "glad I didn't have a working breaker!" — and that would have been a second mistake stacked on the first. Two reasons:

One: survivorship in a sample of one. I have exactly one live position that went deep red. A sweep I ran over stop widths 5–30% showed no stop beating no-stop on the window I measured — but the moment you include that one still-open loss in the PnL, the sign flips:

closed trades only:     +144.01
closed trades + open:   -498.93
Enter fullscreen mode Exit fullscreen mode

Same trades. Opposite verdict. A backtest that excludes an open drawdown isn't measuring a strategy, it's measuring the part of the strategy that already ended.

Two — the real point — the live account has no brake at all. The rolling breaker only exists in the paper path. The live path has a separate check_drawdown() that exits on the first trigger and is even stricter. So "live didn't stop the loss" is not evidence that holding was the right call. It's evidence that the mechanism was never wired. I was about to draw a lesson from an absence.

I also, while auditing, found my stop-width sweep had an inverted label bug — it was reporting "stop helped" as "stop hurt." One line. It had been quietly feeding me the wrong conclusion for a week.

How I now test a guard's input

The fix took a few lines. The finding is the checklist. For anything that enforces a threshold, before I trust a "0 events" result I now do this:

  1. Prove the input can move. Feed the guard a synthetic worst-case through the real code path and assert it fires. If I can't make it trigger in a test, I have no evidence it can trigger in production.
  2. Name what the number means, not where it came from. total and cross_margin_balance are both "the account balance." Only one is marked to market. equity is not a definition.
  3. Correlate input against the thing it's supposed to track. corr(Δequity, Δunrealised) = 0.015 took ten seconds and would have saved a week. A guard's input should be strongly correlated with the phenomenon it guards against; if it isn't, the guard is decorative.
  4. Never conclude from a sample that excludes the live case. Open positions are exactly the data most likely to carry the signal you're looking for. A sweep that drops them is a sweep that drops the finding.
  5. Re-check the input after every "fix." My first fix swapped a missing value for a remote field that reports zero in cross-margin mode. The number was now structurally zero instead of accidentally zero. Same dead sensor, better provenance.

The assertion I keep coming back to, across all of this:

# a guard that has never fired is not "safe" — it's unproven.
# make it fire on purpose, or you've tested the test.
triggered = run_breaker(replay_with_synthetic_drawdown())
assert triggered, "breaker did not fire on a guaranteed drawdown -> input is dead"
Enter fullscreen mode Exit fullscreen mode

The boring takeaway

I don't trust a guard because it's running. I trust it because I've watched it fire on an input I constructed to break it. "It returned 0" and "it returned 0 because it couldn't return anything else" are the same string in a log line, and only one of them means you're protected.

The breaker code got a two-line change. My process got a rule: audit the sensor before you celebrate the silence.


Curious if this is common: has anyone else shipped a guard that was technically correct and completely blind? I'm especially interested in cases where the input looked right — a balance, a rate, a percentage — but meant something different than you assumed. Drop it in the comments.

Part of my Building in Public series — previously: the margin guard that read zero, a single int() that disabled my risk system, the assertion that catches runs that do nothing, and the circuit breaker that caught three outages.

Top comments (0)