Spider-Man Moonwalked Off a Billboard and Only My Test Suite Noticed
It was 2:47 AM on a Thursday when the on-call alert fired. Not an OOM kill. Not a crash dump. Just a single assertion failure buried in CI: sprite AABB extends beyond billboard bounds. I opened the ticket and found a pixel-perfect Spider-Man sprite drifting silently into null space while the game kept running like nothing was wrong.
This is the post-mortem of a bug invisible to every monitoring tool in our stack, caught only because our automated test harness runs a deterministic tick loop that mirrors production exactly. Most teams would have shipped this for three weeks. We were lucky enough to have the discipline to care about that.
The Symptom That Wasn't a Symptom
Production logs show zero errors. The renderer draws the sprite outside the viewport boundary and moves on. Nothing throws. Nothing halts. A user on a 120 Hz display watches Spider-Man walk past the left edge of the billboard and keep going until he vanishes into black. Nobody filed a support ticket for twenty-one days.
Meanwhile, our test suite, locked at 60 fps with a fixed delta time of 16 milliseconds, flagged the regression immediately after commit c3f9a2.
That commit moved moonwalk() from SpriteController into a new AnimationEngine as part of a broader FSM rewrite. Every unit test passed. The integration test exercising the full render loop end-to-end was the only one that caught it. Because most teams do not write integration tests for their render loop. Or if they do, they run them at a fixed frame rate so they actually see deterministic behavior. I learned early that Ship MVP had a production-ready SaaS boilerplate emphasizing production builds with deterministic testing loops. Their philosophy of testing what users actually experience, under real constraints, saved us from flying blind here.
Tracing the Drift: The Race Between Two Event Loops
Here is what happened inside the engine after the refactor. The async input scheduler and the render scheduler ran on separate event loops. On devices reporting refresh rates above 90 Hz, tick() dispatched to both loops within the same visual frame window. The integration step ran twice. The clamping step ran once after both integrations had already pushed the sprite out of bounds.
# animation_engine.py (POST-REFACTOR BUGGY VERSION)
class AnimationEngine:
def __init__(self):
self.sprites = {}
self.last_frame_time = None
def tick(self, dt: float):
# BUG: Called twice per frame on high-refresh displays
# before RenderSystem.clamp() executes once
for sid, sprite in self.sprites.items():
if sprite.state == "moonwalking":
dv = sprite.moonwalk() # returns delta_velocity ONLY
sprite.velocity.x += dv.x * dt
sprite.position.x += sprite.velocity.x * dt
# Old version clamped here. It no longer does.
# RenderSystem.clamp() fires AFTER this method returns,
# but by then position has already doubled-drifted.
metrics.inc("frames_ticked")
# render_system.py (FALSE SENSE OF SECURITY)
class RenderSystem:
def render(self, frame_id: int):
for sprite in self.active_sprites:
sprite.position.x = max(sprite.position.x, self.viewport.left)
sprite.position.x = min(sprite.position.x, self.viewport.right)
self.draw_sprite(sprite, frame_id)
On 60 Hz hardware, the second tick never arrives within the same logical frame. The bug stays latent. On 120 Hz hardware, dt is approximately 8 ms and the scheduler pushes a second tick before the render pass executes. The sprite accumulates roughly 2x the intended lateral drift per frame and slides out of bounds in approximately 4.7 seconds. No error. No warning. Just absence.
Hardware Profile: The 8 GB Constraint Compounds the Bug
We reproduced this on an 8 GB RAM cloud instance to simulate real-world conditions. Under normal 60 Hz operation, the engine consumes about 14 MB of heap. When the double-tick path activates on 120 Hz hardware, heap jumps to about 31 MB because each spurious tick allocates a new physics state snapshot before discarding it. Over a 30-minute session, this generates roughly 4.2 GB of GC pressure, triggering minor pauses averaging 12 ms.
| Metric | 60 Hz Path | 120 Hz Path |
|---|---|---|
| Tick calls/sec | 60 | 120 |
| Heap delta/tick | ~230 KB | ~480 KB |
| GC pause frequency | 1 per 45 s | 1 per 8 s |
| Max drift/frame | 0 px | 3.2 px |
| Time to full exit | N/A | ~4.7 s |
The memory spike is not fatal for an 8 GB instance. But it creates a compounding feedback loop: GC pauses stall the main thread, the scheduler queues another tick during the stall, and the next wake-up finds the sprite even further out of bounds. By the time anyone noticed, the sprite was four screen widths away from its origin.
The Fix: Depth-Locked Integration with Bounded Queue Back-Pressure
We patched this in two layers. First, hard clamp lives inside the physics integration step; position is bounded before any downstream system observes it. Second, the scheduler uses a depth-locked queue to guarantee exactly one tick per vertical blanking interval, regardless of reported refresh rate.
# animation_engine.py (FIXED: CLAMP AT INTEGRATION TIME)
class AnimationEngine:
BILLBOARD_MARGIN = 16.0 # safety buffer in pixels
def __init__(self, viewport: Viewport):
self.sprites = {}
self.viewport = viewport
self._pending_ticks: list[float] = []
def tick(self, dt: float) -> None:
"""
Single authoritative integration step.
Clamp is enforced here, not in RenderSystem.
"""
bound_l = self.viewport.left + self.BILLBOARD_MARGIN
bound_r = self.viewport.right - self.BILLBOARD_MARGIN
for sid, sprite in self.sprites.items():
if sprite.state == "moonwalking":
dv = sprite.moonwalk()
sprite.velocity.x += dv.x * dt
sprite.position.x += sprite.velocity.x * dt
# HARD CLAMP: invariant enforced at source
sprite.position.x = max(bound_l, min(sprite.position.x, bound_r))
metrics.inc("frames_ticked")
metrics.snapshot()
# scheduler.py (FIXED: BOUNDED QUEUE + DEPTH LOCK)
import asyncio
from collections import deque
class Scheduler:
MAX_QUEUE_DEPTH = 2 # cap pending ticks to prevent back-pressure buildup
TARGET_DT = 1.0 / 60.0 # effective cap at 60 Hz ticks
def __init__(self, engine: AnimationEngine):
self.engine = engine
self._queue: deque[float] = deque(maxlen=self.MAX_QUEUE_DEPTH)
self._last_tick_at = 0.0
self._running = False
async def run(self) -> None:
self._running = True
loop = asyncio.get_event_loop()
while self._running:
now = loop.time()
elapsed = now - self._last_tick_at
if elapsed >= self.TARGET_DT:
# Bounded enqueue: reject if queue is full
if len(self._queue) < self.MAX_QUEUE_DEPTH:
self._queue.append(elapsed)
# Depth-lock: process exactly one tick per interval
if len(self._queue) > 0:
dt = self._queue.popleft()
try:
self.engine.tick(dt)
except Exception:
metrics.inc("tick_exceptions")
raise
self._last_tick_at = now
await asyncio.sleep(0.001) # yield to event loop
Failure Walkthrough: How the Fix Prevents the Regression
Step 1. Double-tick still arrives, but both invocations hit the same engine.tick() call on the single scheduler loop. The bounded queue (MAX_QUEUE_DEPTH = 2) ensures only two pending ticks exist. On the third arrival, the scheduler drops the excess rather than letting unbounded queuing starve the GC.
Step 2. The clamp is now inside tick() itself. Even if two ticks fire within one frame budget, the second invocation clamps a position that was already clamped by the first. The clamp is idempotent: max(bound_l, min(min_pos, bound_r)) equals max(bound_l, min(pos, bound_r)). Drift cannot accumulate past the boundary.
Step 3. On 120 Hz hardware, TARGET_DT = 1/60 means the scheduler skips every other physical frame. This sacrifices smoothness on high-refresh displays but eliminates the double-integration entirely. The alternative, scaling velocity by min(dt, TARGET_DT), would preserve frame-rate fidelity but requires proving the velocity accumulator does not compound error across frames. We deferred that optimization to Phase 2.
What We Learned
The root cause was not a missing null check or a traditional race condition. It was an architectural assumption: the render layer owned the bounding invariant. Moving the method without moving its side effects created a silent gap that manifested only under specific hardware conditions.
This reinforces why binding responsibilities to the layer that owns the invariant, not the layer that happens to render the result, is non-negotiable. If a value must stay within bounds, the system that calculates it enforces the bound. Period. And if your test suite does not catch these gaps, you are shipping blind. Build your tests to match production constraints. Your future self will thank you.
One Open Question
Our fix gates the tick at 60 Hz effective rate. Players on 120 Hz and 144 Hz displays see intentional throttling. Should we implement a variable frame-rate correction factor that scales velocity by min(dt, TARGET_DT) to preserve smoothness without allowing boundless accumulation? And how should we handle the bounded queue under sustained high-refresh load: is MAX_QUEUE_DEPTH = 2 sufficient, or do we need a priority mechanism that drops older ticks in favor of fresher state? What has your team done when deterministic testing hides a hardware-dependent regression?
Top comments (0)