DEV Community

Sam Hartley
Sam Hartley

Posted on

My Margin Guard Was Reading Zero. Four Separate Bugs Kept It That Way.

My Margin Guard Was Reading Zero. Four Separate Bugs Kept It That Way.

There's a number I keep in my head from this build. Five positions open, $933.89 of my account tied up in margin, against a rule that says never more than 50%. The rule existed. The rule was running every hour. The rule read zero.

I want to write this one down carefully, because the fix is about four lines and the finding took a week — and because "the guard read zero" turned out to be four independent bugs that only worked as a team.

The guard

A scheduled scan wakes up every hour, looks at funding rates across a few hundred contracts, and can open up to three positions. Each position takes a fixed slice of equity as margin, and there's a cap: total margin in use should never cross 50% of the account.

The check is the shape of every risk guard I've ever written:

margin_cap = equity * 0.50
if used >= margin_cap - 0.01:
    return  # account is full; don't open anything else
Enter fullscreen mode Exit fullscreen mode

A number, a threshold, a comparison. Simple enough that I stopped reading it. Which is exactly the profile of the bugs that hurt.

Bug one: None became 0.0, and 0.0 means "you're fine"

used came out of the saved state:

used = state.get("margin_used") or 0.0
Enter fullscreen mode Exit fullscreen mode

On a state where that key was missing — and it was missing from most of them — or 0.0 turned it into zero. That line reads like defensive programming. It's actually a safety failure with good manners. The guard's whole job is to compare a number; when the number is absent, the honest answer is I don't know, don't open. The code said zero, go ahead. A fail-open default in the one place where fail-closed is the entire point.

Bug two: the number, when it existed, was in the wrong unit

Behind the None bug sat a worse one. When margin_used was populated, it wasn't always margin — some code paths wrote the position's notional value (size × leverage) into it and then compared that against a cap expressed in account percentage. Two units, one comparison.

This is why the bug survived so long: a unit error doesn't crash. It just makes the guard either fire early (annoying, you notice) or fire never (quiet, you don't).

Bug three: the exchange told me the truth, and the truth was zero

At some point I "fixed" bug one by syncing margin_used from the exchange instead of trusting local state. That felt rigorous. It wasn't.

My account runs in cross-margin mode. In cross mode the collateral sits at the account level, not on each position — and the API returns position_margin = 0 and order_margin = 0 for every open position. I checked it read-only: three positions open, both fields zero, real cross initial margin $460.97. My new source of truth was structurally zero.

I had outsourced a safety decision to a field whose semantics I never verified in my own account mode. The field wasn't lying. It was answering a different question than the one my guard was asking.

Bug four: the guard ran once, the loop opened five

This is the one I'm least proud of, because it's a bug I've written about before in a different costume.

The guard was evaluated once per scan — before the entry loop, from a single sync — and then the loop went to work opening positions. Nothing re-checked. Whether it's a rate limiter, a budget, or a margin cap: if you read your limit before the loop and mutate your usage inside the loop, your limit is a snapshot, not a limit.

The demonstration

So I stopped guessing and ran it end-to-end against the real code: five identical entry signals in one scan, everything stubbed except the guard.

Old: the guard saw $0.00 on every iteration. Five positions opened. Margin in use: $933.89 — 83.5% of the account, against a 50% cap.

New: the guard saw 0 → 186.78 → 373.56 → 560.34. It blocked the fourth entry. Three positions, $560.34 — 50.1%.

The fix

One idea: derive the number from the positions I actually hold, in the unit the cap is actually in, and read it inside the loop where it matters.

def margin_in_use(positions):
    """Sum of the margin each open position actually ties up."""
    return sum(
        float(p.get("capital_at_entry", 0)) / max(float(p.get("leverage", 1)), 1.0)
        for p in positions
    )

for candidate in candidates:
    # re-read every iteration, from the live position list
    used = margin_in_use(get_active_positions(state))
    if used >= margin_cap - 0.01:
        break
    execute_entry(candidate)
Enter fullscreen mode Exit fullscreen mode

The exchange numbers stay in the code, but only as diagnostics — logged, printed, never used to make the decision. The rule I took away: don't outsource a safety decision to a remote field whose semantics you haven't verified for your own account configuration.

The test that actually proved it

This is the part I'd hand to anyone building a guard like this.

My first test rebuilt the guard formula inside the test file and compared its output against itself. It passed. It proved nothing — a tautology with a green checkmark. A test that re-implements the logic under test is only measuring its own author.

So I rewrote it to pull the guard block verbatim out of the source file, run the same scenarios through both the unpatched and the patched version, and — this is the important bit — assert that the two disagree. A control probe. If patched and unpatched produce identical output, the test isn't sensitive to the change, and any pass is meaningless.

scenario: margin_used missing
  unpatched: stops after 2 positions   (buggy — should be blocked elsewhere)
  patched:   opens 3, then blocks      (correct)
  control:   patched != unpatched ✅
Enter fullscreen mode Exit fullscreen mode

That control line has caught more bad tests for me than any assertion since. It's the same instinct as checking that a scheduled job actually did something: make the test fail when nothing changes, if you want it to mean something when it passes.

What's still imperfect

I'd rather write the remaining gap down than pretend it's closed. The guard checks used, not used + next_position. So the worst-case overshoot is one full slot: with 16.7% slots, three of them is 50.1% against a 50% cap — 0.1pp over, arithmetic, and deliberate. With 30% slots it would be +10pp, and that's a decision I'd revisit before I ever changed the slot size.

What I'd copy

If you have anything that enforces a limit — margin, spend, rate, quota — the pattern is boring, and it took me far too long to land on it:

  1. Absence is not zero. A missing value in a safety check should stop the action, never permit it.
  2. Check units, not just numbers. A notional value and a margin value can both be "the right number" and still be incomparable.
  3. Don't trust a remote field you haven't verified in your own mode. Cross margin reports zero at the position level. That's not a bug in their API; it's a question I wasn't asking.
  4. Re-read the limit inside the loop. A limit evaluated against a snapshot is a snapshot.
  5. Make your test prove it can fail. Extract the logic under test, and assert that the old and new versions disagree.

The cap was correct on paper the entire time. What it needed was a number that meant what I thought it meant — and the humility to stop treating "no value" as good news.


Same question I keep asking on these: has anyone else been bitten by a guard whose default was the unsafe path? I keep going back and forth on whether a missing state value should halt the job entirely or degrade gracefully — drop a comment with what's worked for you.

Part of my Building in Public series — previously: budget caps that held at 2am, the assertion that catches runs that do nothing, timeout means no on approval gates, the guardrail stack before going live, and a circuit breaker that caught three outages.

Top comments (0)