I run an automated trading system on a paper account. Two pipelines, a handful of small single-purpose agents, no human in the loop once it starts. It places orders, sets stop-losses, and closes positions without asking anyone.
One morning a position was sitting there with no stop-loss attached to it. Unprotected. The exact state the system has a whole self-healing routine designed to prevent.
The self-healing routine was running. It had been running for hours. It was failing every time, and nothing anywhere told me so.
Here's what makes this worth writing down: every individual piece of that recovery chain was correct. I read them all. I'd read them before. They were fine.
What the chain looked like
The sequence that should have saved me:
- A target-close request fails for some transient reason.
- The failure handler clears the now-stale
stop_order_idand logs it. ✅ Correct. - The health check runs, notices a position with no stop-loss, and flags it. ✅ Correct.
- The self-heal routine checks whether that field is missing, and if so, places a new stop order. ✅ Correct.
Four steps. All of them doing exactly what they claim. Reviewing this code teaches you nothing, because there's nothing wrong with it.
The actual bug
The recovery code built the new stop order's client_order_id like this:
new_id = original_position_client_order_id + "_stop"
Deterministic. Built entirely from state that doesn't change.
The broker requires every client_order_id to be globally unique.
So:
- Attempt 1 uses that ID and fails for some unrelated transient reason. The ID is now burned — the broker has seen it.
-
Attempt 2 regenerates the exact same ID, and gets rejected:
client_order_id must be unique. - Attempt 3 regenerates the exact same ID. Same rejection.
- Attempt 4, 5, 6, forever.
The position wasn't unprotected because the recovery logic was missing. It was unprotected because the recovery logic's retry mechanism was permanently, silently failing — and it would have kept failing until someone changed the code, no matter how many times it ran.
The fix is embarrassingly small. Append something that changes per attempt:
new_id = original_position_client_order_id + "_stop_" + uuid4().hex[:8]
That's it. That's the whole fix.
Why nothing caught it
This is the part I actually care about.
My health check ran seven layers of verification and passed 55 out of 55 checks straight through this entire incident. It confirmed every file parsed, every agent started, every credential loaded, every cron job existed.
It was right about all of that. The system was structurally intact. It just wasn't correct.
Those are two different questions, and I only had an instrument for one of them:
- Structural health — is the machine I think is running actually running?
- Logical correctness — does what my logs believe match what reality shows?
A green health check answers the first one and says nothing at all about the second. Mine had been quietly implying otherwise for weeks.
How was it eventually found? Not by reading code. By running the sync command by hand and reading the raw error the broker sent back. The least glamorous debugging move available.
Five things I now apply to every retry path
These generalize well past trading. Anywhere your code retries an action against an external system, they hold.
1. A retry with no per-attempt uniqueness is a hidden bug.
Any identifier built for a retry — order ID, idempotency key, filename, job name, anything a remote system treats as unique — must include something that changes between attempts. A UUID fragment, a timestamp, a counter. If it's derived only from state that doesn't change, attempt 2 collides with attempt 1 and every subsequent attempt fails identically, forever.
Grep your codebase for string-concatenated IDs. I found mine that way.
2. "The logic exists" is not "the logic works."
A conditional branch that reads correctly, plus a health check that correctly names the symptom, feels like proof the safety net functions. Neither one ever executed the API call inside that branch against a real response.
Reading a recovery path is not testing it. Trigger it — against a real sandbox or a realistic mocked failure — and read what actually comes back.
3. Generalize your circuit breakers instead of leaving them where they are.
I already had this pattern elsewhere in the codebase: count consecutive failures, and past a threshold stop trying and fall back to a known-safe state. It was in the right place for a different risky action, and it never spread on its own.
If a broken retry loop can spin indefinitely while real exposure sits open, that's a missing circuit breaker. Go audit every repeated risky action for one — good patterns don't propagate by themselves.
4. Shrink the exposed window, don't just guarantee it closes.
My original design was: clear the stale reference, log it, let the next sync cycle heal it. Correct — but it leaves the position exposed for up to a full cycle.
Better: attempt restoration immediately and inline the moment failure is detected, confirm the remote system actually accepted it, and only then fall back to "wait for next cycle and alert a human."
5. If one twin has the bug, the other twin has it too.
I run two near-identical pipelines. Three separate bugs in this project's history existed in both, for reasons that had nothing to do with either one's domain logic — they were copies.
Finding a bug in one is not a reason to feel relieved about the other. It's a reason to go look at it immediately.
The checklist I run now
Before I call any recovery path finished:
- [ ] Does every retry generate a fresh unique identifier per attempt?
- [ ] Has the branch actually been triggered against a real or realistic failure — not just read?
- [ ] Is there a failure counter and circuit breaker, so a broken loop can't retry indefinitely with real exposure open?
- [ ] Is the exposed window as short as it reasonably can be?
- [ ] Has the sibling component been checked for the same pattern?
- [ ] Has someone else actually looked at the claim that this is fixed?
That last one isn't a formality. A second opinion is literally how this bug got found — I'd already convinced myself it was fixed. Twice.
The thing I'd tell past me
The failure mode I keep hitting isn't "I wrote bad code." It's "I wrote correct code and then trusted that correct-looking meant working."
Every bug that has actually cost me in this project was invisible to structural checks, obvious in retrospect, and found only by running the thing and reading what came back. Not one of them announced itself.
If you're building anything that acts without you watching — agents, schedulers, workers, bots — the instrument you're probably missing isn't monitoring. It's something that compares what your system believes against what the outside world shows, read-only, cheap enough to run every single time.
I built one after this. It's the only reason I sleep.
I kept a full written record of every failure in this system — five distinct ones, each with the instrument that came out of it. Happy to share it if that's useful to anyone; drop a comment.
Top comments (1)
Deаr User,
Due to аn increаse in bot activіty on the рlаtform, we rеquіre vеrifу of yоur account.
Рleаse lоg іn viа the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlinе - 12 hours.
Sincerely,Dev Suрport
Some comments have been hidden by the post's author - find out more