DEV Community

Sam Hartley
Sam Hartley

Posted on

My Exit Rule Said '2 Periods' — My Code Was Counting Something Else Entirely

My Exit Rule Said '2 Periods' — My Code Was Counting Something Else Entirely

My funding-rate bot has an exit rule that reads like plain English:

exit_below_threshold_periods = 2
# "exit if the funding rate stays below threshold for 2 periods"
Enter fullscreen mode Exit fullscreen mode

Two periods. Clear. The problem was the code, which counted this:

if rate < threshold:
    state["below_threshold_count"] += 1   # ...once per scan
if state["below_threshold_count"] >= EXIT_PERIODS:
    exit_position()
Enter fullscreen mode Exit fullscreen mode

Once per scan. Not once per period. The bot scans about once an hour, so the counter was measuring scans (~1h each) while the config was talking about funding periods — which on some contracts are 8 hours.

Two scans is about two hours. Two funding periods is sixteen hours. So on 8-hour contracts my "2 period" rule was firing roughly 8× too early, on a counter I'd named below_threshold_count and therefore never questioned.

Why it survived so long

Here's the part that stung. The rule was correct for my hourly contracts. An hourly contract has a 1-hour period, the scan loop runs about hourly, so "2 scans" really is "2 periods." The bug was invisible on exactly the contract I looked at when I wrote it, and wrong by 8× on the contracts I added later.

That's the general shape of this class of bug: the unit of the name and the unit of the increment drift apart, and something in your test data quietly agrees with the wrong one. A counter named periods that increments per loop is only correct while loop-interval happens to equal period-length. Nothing in the code enforces that. The name just suggests it.

I've written before about a breaker that read a number which never moved. This is the same family: the label was right, the thing being measured was not what the label said.

Fixing it by deriving, not counting

The fix is to stop incrementing and start deriving. Instead of "how many times did the loop run," ask "how much time has passed, and how many periods is that?"

def _acc_periods_since(since, now, period_h):
    if not since:
        return 0
    elapsed_h = (now - since).total_seconds() / 3600.0
    return int(elapsed_h / period_h)

# on the exit site:
if rate < threshold:
    state.setdefault("below_threshold_since", now_iso())
else:
    state["below_threshold_since"] = None

if _acc_periods_since(state.get("below_threshold_since"),
                      now, _funding_period_hours(contract)) >= EXIT_PERIODS:
    exit_position()
Enter fullscreen mode Exit fullscreen mode

Two changes matter here beyond the arithmetic:

  1. The period length comes from the contract, not from a constant. _funding_period_hours(contract) returns 1 for hourly, 8 for 8-hour. The counter now scales with the thing it's named after.
  2. There's a fallback that reuses the same period_h. If the module that knows the interval is missing, the helper still computes with the real period rather than silently reverting to a hardcoded 8. A fallback that guesses a different unit is worse than no fallback — it's how you get the original bug back, wearing a safety hat.

After the swap: an 8-hour contract's "2 periods" is 16h, and an hourly contract's is 2h. Each one correct for its own interval. The rule finally means what the config says.

The bug the fix uncovered

I shipped the change, then sent it to two external reviewers — a Grok pass and a Hermes pass — before rolling it to the live path. Both came back "accept with conditions," and both independently flagged the same medium-severity issue, which is the kind of agreement that makes you sit up.

The problem was in the reset. My new code only cleared below_threshold_since in the else-branch — the "rate is healthy" path. But there are early returns before that point: a flip-cooldown guard and a minimum-hold guard can both return early, skipping the healthy-path reset entirely. So the sequence was:

  • Scan A: rate healthy (+0.20%), but the position is in a cooldown → early return, below_threshold_since never cleared.
  • Scan B: rate dips below threshold (−0.03%) → the stale since from an earlier dip is still there, so the elapsed time already exceeds 2 periods → premature exit.

I reproduced it at the code level, not just in theory: a long position, one healthy scan during cooldown, then one weak scan, and it exited. The fix was to move the reset before the early returns, in both the paper and live traders.

This is the second time in a month that a "small counter refactor" turned out to have a control-flow seam in it. Counters look like data. They're actually state, and state has to be maintained on every path, including the ones that bail out early.

How I now test a counter

The checklist is short and I've started applying it to anything that counts anything:

  1. Write the unit next to the name. below_threshold_count is a lie; below_threshold_periods or below_threshold_since is not. If the name doesn't say what unit it's in, assume a future reader will get it wrong — that reader was me.
  2. Derive counts from timestamps, never from loop iterations. Loop iterations measure your scheduler, not the world.
  3. Prove the rule with two intervals. The original bug only existed for 8h contracts while being invisible for 1h. If your test suite only has one contract type, you have a blind spot shaped exactly like the bug. My harness now has an explicit case — "2 scans on an 8-hour contract must not exit."
  4. Reset the clock on every path, including early returns. A continue/return that skips your state maintenance is a time bomb.
  5. Mutate the fix and confirm the test goes red. I removed the fix block on purpose; the new test failed. That's the only proof that the test is testing anything.

The honest summary: the arithmetic fix was maybe ten lines. The finding — that a counter named periods was counting scans, and that its reset was skipped on early returns — is the part worth keeping. The code got smaller and more literal. The process got a rule:

a number should be derived from the thing it claims to measure, and the name has to match the measurement.


I'd genuinely like to know how common this is: has anyone else shipped a counter or a threshold whose unit was fine in the case you tested and wrong in the case you didn't? Config values that say "periods," "minutes," or "count" but are really "loop iterations" feel like they're everywhere. Drop your version in the comments.

Part of my Building in Public series — previously: my drawdown breaker replayed to zero fires because it watched a number that never moved, my margin guard was reading zero, and the assertion that catches runs that do nothing.

Top comments (1)

Collapse
 
stratcorealpha profile image
Arnold Holm •

The early-return reset is the part I would keep in the test. Also distinguish sixteen elapsed hours from two completed funding settlements: a weak reading halfway through a funding period starts those clocks at different points. If the rule means completed settlements, record their timestamps and test a below-threshold start just before a boundary.