DEV Community

Cover image for Our Deploys Were Dead for Six Hours and It Wasn't Our Bug
Rohit Bhadani
Rohit Bhadani Subscriber

Posted on

Our Deploys Were Dead for Six Hours and It Wasn't Our Bug

Wednesday, 2 PM. A hotfix for a billing bug sat in a PR, approved, green checks pending for merge. Except the checks never finished. Not failed — just stuck, "queued," for forty minutes, then an hour. Someone said "GitHub's probably just slow today," which is the engineering equivalent of "it's probably nothing."

It was not nothing. By 3 PM our status page had three separate people asking if deploys were broken. They weren't broken. They were queued behind a worldwide GitHub Actions degradation, and we had no way to tell the difference from the inside.

The wrong turns

First instinct: it's our workflow config. Someone had touched a .github/workflows/deploy.yml file two days earlier to add a caching step. We reverted it. Jobs still sat queued. Not it.

Second instinct: our self-hosted runners died. We run a handful of self-hosted runners for the slower integration suite. We checked them — alive, idle, polling GitHub fine, just never getting assigned work. That's backwards from "our runners are broken." It looked more like GitHub's scheduler itself wasn't handing out jobs.

Third instinct: a permissions or billing issue on our org. We checked our Actions usage dashboard. Nothing unusual, no quota warnings, no billing alerts. At this point someone finally checked GitHub's own status page instead of assuming the problem was local, and there it was: Actions and several other services in degraded or major outage mode, worldwide, several hours running.

What made it actually bad

The outage itself being someone else's problem didn't make it our problem go away. Here's the part that stung: almost every pipeline we had — tests, builds, container pushes, the deploy trigger itself — ran exclusively on GitHub-hosted infrastructure. We had no fallback path. When GitHub's scheduler stopped handing out runner capacity, our ability to ship anything, including a billing hotfix customers were actively affected by, went to zero. Not because our code was broken. Because our entire delivery pipeline had a single point of failure we'd never actually stress-tested, because it had never failed before.

We ended up manually building the container image on a laptop, pushing it to the registry by hand, and running the deploy script locally against production credentials — the kind of thing that works in an emergency and that you never, ever want to be your Tuesday-afternoon normal.

The fix

We didn't try to replace GitHub Actions. We built a narrow, boring escape hatch that only exists for exactly this scenario: a self-hosted runner path that can build and deploy the two or three services that actually matter during an incident, independent of GitHub's hosted runner capacity.

# .github/workflows/deploy.yml
jobs:
  build-and-push:
    runs-on: [self-hosted, emergency-capable]
    timeout-minutes: 20
    steps:
      - uses: actions/checkout@v4
      - name: Build image
        run: docker build -t registry.internal/api:${{ github.sha }} .
      - name: Push image
        run: docker push registry.internal/api:${{ github.sha }}
Enter fullscreen mode Exit fullscreen mode

The emergency-capable label points at runners that live outside GitHub's hosted fleet entirely — not borrowed capacity, actual machines we control, that poll for jobs over an outbound-only connection and never need an inbound port opened to reach them. We keep two of these running permanently, small and cheap, specifically so "GitHub Actions is degraded" stops being synonymous with "we cannot ship." I'm the founder of Krova Cloud, and these runners live on it for a simple reason: they're outbound-only by default, so there's nothing for an attacker to find even if the image somehow leaked, they're billed per minute so running two idle standby runners costs close to nothing most months, and spinning up a third one during an actual incident takes under a minute from a snapshot instead of provisioning a new box from scratch while customers wait.

A script now checks GitHub's status API before every deploy, and routes to the self-hosted path automatically if Actions availability drops below a threshold — no manual decision required at 3 PM while people are also trying to fix the actual bug.

Lessons

  • "It's probably GitHub, not us" is a hypothesis to check, not an assumption to make — check the status page early, not as a last resort after three wrong internal theories.
  • A CI/CD pipeline that depends entirely on one vendor's hosted infrastructure, with zero fallback, isn't actually resilient just because that vendor is reliable most of the time. Most of the time isn't all of the time.
  • The cheapest insurance against a vendor outage isn't migrating off the vendor — it's keeping a small, boring, independent path alive for the handful of services that actually need to ship during an incident.
  • Decide the failover trigger and threshold before the incident. At 3 PM, with customers messaging support, is the worst possible time to debate whether "degraded" counts as "bad enough to switch."

If your entire deploy pipeline lives on one platform's hosted runners, it might be worth the hour it takes to build the boring version of what we built — before the day you actually need it.


I'm Rohit, founder of Krova Cloud — cheap, disposable cloud VMs with outbound-only networking by default, perfect for self-hosted CI runners that need to exist without becoming an attack surface. If you want more deep debugging stories like this one, I write regularly over at debugly.dev too.

Top comments (0)