DEV Community

Cover image for Ghost in the Wire: Our Two-Day Hunt for a Network Bug That Didn't Exist
Flude team
Flude team

Posted on Originally published at blog.flude.guide

Ghost in the Wire: Our Two-Day Hunt for a Network Bug That Didn't Exist

Ghost in the wire

Our self-hosted runner serves CI for five repositories on one physical VM. It started acting completely crazy. The engine job sat queued in GitHub Actions for over a day. It just hung there doing absolutely nothing. The supervisor meant to pick it up was restarting itself every two and a half to three minutes around the clock. We called this endless loop "the carousel" and wasted two days chasing ghosts.

The Symptom, and the First False Lead

We already had a documented network bug in this exact infrastructure. A classic VirtualBox NAT "black hole" used to kill TCP connections inside the guest VM silently. The guest OS would consider the connection alive forever. We tracked down and closed that issue a week and a half earlier.

When the supervisor started hanging on the host side, our reflexes kicked in. We assumed it was the exact same disease in a fresh disguise. That feeling was convincing enough to steer the whole investigation sideways for a full day. The watchdog detected the hang, killed the process, and brought it back up. Every new instance hung reliably after 150-190 seconds. That unnatural precision should have made us suspicious way earlier.

Escalation: Packets, Memory Dumps, Power Settings

A thorough and completely useless sweep through network theories followed. We disabled a third-party VirtualBox NDIS filter attached to the Wi-Fi adapter. We reset the MIMO power-saving settings. Nothing made a dent in the problem.

Running pktmon during an actual hang showed a completely healthy network. New TCP connections to GitHub established in 60 milliseconds. The "hung" process supposedly couldn't send a single byte. We wrongly concluded the hang was happening before the packet got sent.

A memory dump of the process via WinDbg showed twenty-two threads. None of them were sitting inside an HTTP call. We spotted suspicious modules like antivirus drivers and crypt32 certificate validators. Every single one of these plausible candidates turned out to be a dead end.

We did notice one genuinely interesting detail. The host machine dropped into Modern Standby right when the hangs occurred. We disabled sleep and screen timeouts entirely. The restarts kept happening on the exact same schedule. The correlation was totally real—but completely unrelated.

A Second Pair of Eyes Solves Half the Problem

Another session looked at the logs and caught what we missed. The watchdog was treating plain silence as a fatal hang. The GitHub API poll writes nothing when the queue is empty. It simply sleeps for twenty seconds and tries again. A few quiet cycles looked exactly like a dead process to the watchdog.

We verified this by pulling the I/O counters directly. The bytes were genuinely moving. The process was perfectly healthy with nothing to log. We updated the watchdog to check those exact counters before killing anything. The restarts stopped completely. The engine job still refused to dispatch.

The Real Cause

API filter bug

We ran a direct test using the exact token the supervisor uses. Querying runs?status=queued for engine returned zero results. Running the same query with a personal token showed the job instantly. Access rights were perfectly fine.

The answer was buried in the API response itself. The run had a "pending" status. The job inside that run was correctly marked as "queued". Our code filtered runs at the top level. GitHub had split the run's status from the job's status. One bad line of filtering made the code completely blind to a waiting job.

We dropped the top-level filter completely. The nested check handled everything perfectly on its own. The very first poll picked up the stuck jobs and ran the full suite beautifully.

A Small Postscript: Fixing Our Own Fix

Ghost In The Wire Part 3

Removing the filter spawned a new headache. The code started pulling dozens of recent runs and making separate job requests for each one. The call volume spiked massively. The carousel of restarts came back because I/O genuinely dropped to zero. We added client-side filtering to skip runs marked as completed. Everything stabilized immediately.

What We're Taking Away From This

Solid diagnostic tools will gladly prove the wrong theory. WinDbg and packet sniffers correctly showed a healthy network. We just kept interpreting those results to fit our initial bias. A simple API call with the right token proved infinitely more useful than all the heavy debugging gear.

Next time, we get into legal boundaries and clean room design. We ended up moving an entire renderer family into a private plugin. It was the only way to firmly separate our core engine from logic shaped by an external client format.


Originally published on our blog: https://blog.flude.guide/blog/ghost-in-the-wire

Also read us:

Top comments (0)