I wired an actual fix action onto yesterday's dead-slot detector, and found that raising concurrent requests further was based on a wrong premise.
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
Attaching a fix to yesterday's detector
Yesterday(new tab) I built a detector that catches a dead slot on the backup GPU path.
At the time it only detected — it took no action. Today was about wiring an actual recovery action onto it.
I hooked up self-healing logic that erases a detected dead slot and makes it usable again, then ran a resumed overnight replay test all the way through — 100 out of 100 completed cleanly.
The case for raising concurrency was built on a wrong premise
Earlier this week, an odd pattern showed up on this backup path: several stock groups (shards) appeared to collapse at nearly the same point, all at once.
That led to a hypothesis: if concurrent requests were raised further, damage from one group collapsing would be spread thinner and hurt less.
Reconstructing the logs on an absolute timeline showed the premise was wrong. This server picks whichever slot is free at the moment a request comes in — slots aren't permanently tied to a given stock group.
So "four groups collapsed at the same time" was really "one slot died once, and all four groups happened to be routed through that same slot at that moment, and all witnessed it together."
Once that was clear, the math no longer favored raising concurrency — no reliability gain, and possibly a longer detection window before a dead slot gets noticed. So the concurrency bump stayed shelved.
Two real bugs turned up during the investigation
Shelving the concurrency bump wasn't the main outcome. The investigation surfaced two more concrete problems.
First, the self-healing logic added today wasn't actually wired into the real startup path. It had only been running as a manually launched temporary process — without that process, healing simply wouldn't happen.
Second, the detector's own state tracking had a bug. After a slot was fixed and revived, its internal record never got reset, so if the same slot died again later, it wouldn't be caught a second time.
Both got fixed and shipped the same day. Self-healing now starts automatically on both real startup paths (server reuse and fresh launch), and fixing a slot now resets that slot's detection record too.
Pre-committing the next fork in the road for the backup GPU
This backup path exists as a fallback for when the main GPU becomes unavailable.
Around next weekend (October 3rd), there's a decision point coming up: keep this backup path as-is, or switch the hardware order around.
Instead of deliberating on the day, I wrote down the decision rule in advance today. Over the coming days I'll watch how reliably this path actually runs, and the outcome will follow the pre-committed rule automatically.
What's next
Next is confirming that today's self-healing wiring actually fires during real overnight operation — the first check is whether tonight's logs show the healing action running as expected.
After that, it's watching this backup path's reliability all week to build up the data needed for next weekend's decision.
Top comments (0)