DEV Community

Cover image for The real cost of self-hosting your own CI/CD runner
Amit Shukla
Amit Shukla

Posted on Originally published at amitshuklabag.hashnode.dev

The real cost of self-hosting your own CI/CD runner

It starts with a reasonable decision. Your hosted CI minutes bill keeps climbing, or your builds need more memory than the default runners give you, or your tests need to reach a database inside your private network. So someone spins up a VM, installs a runner agent, registers it with your Git platform, and the first build on it is twice as fast and costs almost nothing.

Three months later that same runner is the reason nobody can merge on a Monday morning. The disk is full of old Docker images. The agent is two versions behind and has stopped picking up jobs. A build that passed on Friday fails today because a previous job left a global package installed that it should not have. The compute was cheap. Everything around it was not.

Self-hosting a CI runner is one of the most common ways teams take on an ops burden without noticing. Here is what that burden actually looks like.

The machine needs everything a production server needs

A CI runner is a server that executes code from your repository with access to your secrets, your registry, and often your internal network. It deserves the same care as anything else you run in production, and it rarely gets it.

That means OS security patches applied on a schedule, a firewall that only allows what the runner needs, monitoring on disk, memory, and CPU, and alerts that reach a human before the next build fails. It means someone owns the machine. On most teams, the runner belongs to whoever set it up, and it quietly stops being maintained the week that person gets busy.

Disk fills up, every time

Every build pulls base images, writes layers, downloads dependencies, and leaves artifacts behind. Hosted runners throw the whole machine away after each job. Your runner does not, so all of that piles up until a build dies halfway through with no space left on device.

The usual fix is a cleanup job:

# /etc/cron.daily/docker-cleanup
docker system prune -af --filter "until=72h"
docker volume prune -f
Enter fullscreen mode Exit fullscreen mode

Prune too aggressively and you throw away the layer cache that made your builds fast in the first place. Prune too little and the disk fills again in a week. There is no setting that is right forever, because your image sizes and build volume keep changing. Someone has to keep watching it.

Persistent runners leak state between jobs

On a long lived runner, every job runs on the same machine as the last one. Anything a job writes outside its workspace, a globally installed tool, a modified config file, a leftover process holding a port, is still there for the next job.

That produces the worst kind of CI failure: builds that pass or fail depending on what ran before them. A test suite that only passes because a previous job happened to install a CLI globally. A job that fails because another one left a container running on port 5432.

The clean answer is ephemeral runners, where each runner takes exactly one job and then deregisters, so every build starts from a fresh environment:

./config.sh --url https://github.com/your-org/your-repo \
  --token "$RUNNER_TOKEN" \
  --ephemeral
Enter fullscreen mode Exit fullscreen mode

Ephemeral runners fix state leakage, but now you need something that creates a fresh runner for every job, cleans up the old ones, and scales up when ten pull requests land at once. That is an autoscaling system you now own, with its own failure modes, on top of the runner itself.

The agent has to stay current

Runner agents are not install once software. Git platforms ship new runner versions regularly, and older versions eventually stop being accepted. GitHub, for example, requires self hosted runners to be updated within 30 days of a new release when automatic updates are turned off, after which the runner stops receiving jobs.

So either you let the agent update itself on a machine you are trying to keep stable, or you build a process to roll updates out yourself. Both are work. Neither shows up in the original plan to save money on CI minutes.

Security is the cost nobody prices in

This is the one that matters most. A runner executes whatever is in the pipeline definition, with whatever access the machine has.

On a self hosted runner, that usually means access to deploy credentials, a container registry, cloud provider keys, and the private network the runner sits in. If anyone who can open a pull request can get a job onto that runner, they can run code with all of that access. GitHub's own documentation recommends against using self hosted runners with public repositories for exactly this reason, since a pull request from a fork can run arbitrary code on your machine.

Even on private repositories, a persistent runner means secrets from one job can end up on disk where a later job can read them. Locking this down properly means isolating jobs, scoping credentials per pipeline, restricting which repositories and branches can use which runners, and auditing all of it regularly. That is real security engineering, and it is easy to skip when the runner was supposed to be a quick cost saving.

Adding it up

Here is a rough, honest tally for a single team running their own runners:

OS patching and reboots          ~1 hr / month
Disk cleanup and tuning          ~1 hr / month
Runner agent updates             ~1 hr / month
Debugging flaky, stateful builds ~3 hrs / month
Security review and credentials  ~2 hrs / month
Incident when it breaks          ~4 hrs, a few times a year
Enter fullscreen mode Exit fullscreen mode

Call it eight to ten hours of senior engineering time every month, before counting the time everyone else spends blocked when the runner is down. Multiply that by your engineering hourly cost and compare it to the CI minutes bill you were trying to cut. For a lot of teams, the "free" runner turns out to be the most expensive line item in the pipeline.

And none of that time goes into your product. It goes into keeping a build machine alive.

Run the CI platform without running the platform work

The reason teams self host CI is still a good one. You want control over your build environment, your data staying where you choose, enough machine for real builds, and no per seat or per minute bill growing with your team. The goal is to keep all of that and hand off the operations underneath it.

That is what Elestio does. Pick the CI and Git tools you want from a catalog of 400+ open source templates, choose where they run, a major cloud provider and region of your choice, your own virtual machines, or your own hardware on premise, and Elestio deploys them on a dedicated instance, not a shared box.

Then the work from this article is handled for you:

  • OS and application updates applied for you, so the platform stays current without you scheduling it
  • Automated encrypted backups, off host, with retention, so a bad day does not cost you your pipelines or history
  • Monitoring and alerts on the instance, so disk and memory problems surface before a build fails
  • TLS certificates issued and renewed automatically
  • Firewall and DDoS protection in front of your instance
  • Built in CI/CD for deploying your own code alongside
  • Support from people who run these tools every day

You keep full root access. It is your instance and your data, exportable any time. Billing is a flat fee on top of compute, not per user, so adding developers does not grow your CI bill.

Getting started takes minutes

Pick a template, choose a region, and Elestio provisions the instance with backups, monitoring, updates, TLS, and a firewall already in place. Your team gets the control and cost profile of self hosted CI from the first build, without a runbook to write before you can go live.

Try Elestio free and put your engineering hours back into the product.

Takeaway

A self hosted runner is never just a VM and an agent. It is patching, disk management, version updates, state isolation, and security, every month, for as long as you run it. Keep the control and the savings, and let Elestio carry the operations, so your pipeline is fast, safe, and nobody's weekend project.

Top comments (0)