The specifics change from team to team, but this exact failure mode plays out the same way often enough to be worth walking through in full.
Hot-patching — swapping a running function's implementation in-place without restarting the process — feels like a superpower the first few times it works. Fix a bug, apply the patch, nobody notices a blip. We'd used it exactly twice before, both times for genuine emergencies, both times successfully. The third time, applying what looked like an identical emergency patch to an identical kind of bug, the process didn't crash and didn't recover. It just hung. Forever. No errors, no stack trace, no CPU usage — a process that was technically alive and functionally dead.
The wrong turns
First theory: the patch itself had a bug. We reviewed the patched function line by line. It was correct — we'd tested the same logic in isolation, it behaved exactly as intended. Not the patch's content.
Second theory: a deadlock from the patch being applied mid-execution. Hot-patching a function while it's actively running, mid-call, is a known hazard — if you swap the implementation while execution is paused inside the old version, you can get undefined behavior. We checked our patching tool's logs: it had correctly waited for the function to be idle before swapping. Not this either, at least not in the way we expected.
Third theory: memory corruption from the patch allocation. We attached a debugger to a non-production instance and reproduced the hang. The instruction pointer was sitting inside the function we'd just patched — except it was executing bytes that belonged to neither the original function nor our new patch. It was executing a jump instruction into a location that made sense for neither version.
What had actually happened
This function had been hot-patched once before, months earlier, by a different engineer, for a different incident, and nobody had ever unpatched it — "temporary" patches have a way of becoming permanent when they work. Hot-patching typically works by redirecting the function's entry point with a jump instruction to the new code, keeping the old code in memory. Our tooling assumed, reasonably for a first patch, that the function's entry point contained the original, untouched prologue. It did not. It contained a jump instruction left over from the first patch.
Our new patch wrote its own redirect jump at the function's entry point — correctly, as far as it knew — but the target it jumped to still contained a reference back to the function's original address for context it needed, and that original address now pointed into the middle of the first patch's jump instruction rather than the real original code. Two patches, each individually reasonable, had been stacked without either one knowing about the other, and the result was a jump chain that led into the middle of an unrelated instruction.
The fix
1. Hot-patch tooling now records every patch applied to a given function, in a registry checked before any new patch is allowed.
def apply_hotpatch(function_addr: int, new_impl: bytes, registry: dict) -> None:
existing = registry.get(function_addr)
if existing:
raise RuntimeError(
f"Function at {hex(function_addr)} already has an active hotpatch "
f"(applied {existing['timestamp']}, reason: {existing['reason']}). "
f"Unpatch it first or deploy a real fix instead of stacking patches."
)
original_bytes = read_original_prologue(function_addr)
write_jump(function_addr, new_impl)
registry[function_addr] = {
"timestamp": now(),
"original_bytes": original_bytes,
"reason": current_patch_reason(),
}
2. Every hot-patch now has a mandatory expiry and a tracked owner. A patch that's still active thirty days later triggers an alert, not silence. The real fix was never meant to be "live forever as a patch" — it's meant to buy time until the real fix ships in a real deploy.
3. We test every hot-patch, including emergency ones, against a disposable clone of the actual running production process state first — not a clean restart, the actual live memory layout, including any prior patches that might already be sitting there. I'm the founder of Krova Cloud, and this is one of the uses I'm most fond of for disposable VMs: snapshot the exact running state you're about to touch, including its full history of prior modifications, clone it, apply the risky change there first, and only touch the real thing once you've watched it survive on the clone. Takes a couple of minutes, costs almost nothing, and would have caught this stacked-patch problem instantly, because the clone would have hung exactly the same way, with zero customer impact.
Lessons
- A hot-patch that works today becomes permanent infrastructure the moment nobody schedules its removal. Treat "temporary" patches as a liability with a clock on it, not a free fix.
- Tooling that assumes a clean, unpatched starting state will behave unpredictably the second that assumption is false. If your system can be modified at runtime more than once, your tooling needs to know about every modification, not just the next one it's about to make.
- "It's the same kind of fix we did last time" is not the same as "it's safe for the same reason it was safe last time." The context around the fix — what else has touched this code since — matters as much as the fix itself.
- Testing a risky live change against an actual clone of the current running state, not a fresh restart, catches an entire category of bug that only exists because of accumulated history you've forgotten about.
If your systems support any form of runtime patching, it's worth asking right now whether you actually know everything that's currently patched, or whether you're trusting memory.
I'm Rohit, founder of Krova Cloud — disposable VMs you can snapshot and clone from actual running state, so risky changes get tested against reality, not a clean assumption. If you want more deep debugging stories like this one, I write regularly over at debugly.dev too.
Top comments (0)