kiln v0.2.5: self-hosted GitHub Actions on hardware you already own, with one fresh rootless QEMU/KVM VM per job.
Why build this
Two of our repositories ran their PR smoke gates on Google Cloud Build. Over a 27-day window that came to 396 builds and 2,192 build-minutes: about 81 minutes a day, or roughly $39 a month on E2_HIGHCPU_8. That is not a big number. But a box with an i9-14900KF (32 threads) and 61 GB of RAM was already sitting on the tailnet, mostly idle, and the hosted bill grows with every repo you add.
Money was only half of it. The other half was isolation. A long-lived self-hosted runner keeps state between jobs: stray processes, a dirty Docker daemon, files in $HOME. Container-based runners are better, but every job still shares the host kernel, and things like services:, sudo apt install, real Docker and nested KVM turn into special cases. I wanted what GitHub-hosted runners give you, a clean machine for every job, without paying per minute.
kiln is the result. It is a single Rust binary (kiln serve) with an embedded dashboard. It polls GitHub for queued jobs, boots a throwaway Ubuntu 24.04 VM for each one, lets a single-use runner take the job, and deletes the VM afterwards. It needs no root, no tap devices and no bridges. It runs as a systemd user service and needs only /dev/kvm and QEMU.
Architecture
GitHub REST API <-- poll queued jobs (ETag-cached), generate-jitconfig,
^ delete runner, run/job details
|
+------+-------------------------------------------------------+
| kiln serve |
| scheduler (tick every poll_secs) --> VM supervisor per VM |
| Docker Hub mirror on 127.0.0.1:5000 qemu-img overlay, |
| dashboard + API on :7878 (tailnet) QEMU, console, cache |
+------------------------------------------+-------------------+
| ttyS0: console + control
| ttyS1: step logs
v
+-------------------------------------------+
| job VM (Ubuntu 24.04, q35, KVM) |
| vda = overlay of base.qcow2 |
| vdb = overlay of the repo's cache disk |
| kiln-job -> run.sh --jitconfig |
+-------------------------------------------+
Direct kernel boot on qcow2 overlays
kiln bake builds images/base.qcow2 from the Ubuntu 24.04 cloud image and a cloud-init recipe (guest/user-data.yaml). It installs packages, the actions/runner release and the Node versions you ask for, then freezes the image. The kernel is extracted next to it as base.vmlinuz.
For each job, qemu-img create makes a disk.qcow2 backed by the frozen base. QEMU boots the kernel directly, with no initrd and no bootloader. virtio and ext4 are built into the Ubuntu kernel, so the runner is listening about 4 seconds after launch. panic=1 with -no-reboot turns a kernel panic into a dead VM instead of a hung one. The overlays use cache=unsafe because they get thrown away anyway. Deleting the overlay is the cleanup.
On the reference box, kiln's own self-test and CI jobs run in about 45 to 75 seconds end to end. An optional warm pool (warm: {"owner/name": N}) keeps N idle VMs ready per repo, which cuts job start-up to about a second.
JIT runners
kiln never stores a registration token. For each VM it calls generate-jitconfig and gets back a single-use runner that deregisters itself. The config is written with mode 0600 and passed to the guest as an SMBIOS OEM string (-smbios type=11,path=...). That works with the stock cloud kernel, and path= keeps the secret out of ps. kiln deletes the file as soon as the guest prints its first console line, which proves QEMU has read it.
Count-based scheduling, and what it means for fork PRs
There is a GitHub detail that shapes the whole design: a queued job goes to any idle runner whose labels match, not to the runner you started for it. So kiln never ties a VM to a job id. Per repo and per VM size it computes:
launch = queued - (booting + idle)
That number is then capped by max_vms, a memory and vCPU budget (QEMU allocates guest RAM lazily, so MemAvailable alone would let a burst overcommit), and gates for low memory and low disk. A reaper drops surplus idle VMs, and deregistration decides the race: if GitHub refuses because the runner just picked up a job, the VM lives. Sizes come from labels (kiln, kiln-2cpu up to kiln-16cpu), and a big VM never registers the plain kiln label, so it cannot take a default-size job.
The security consequence matters. In v0.2.0 the scheduler refuses fork pull requests: their jobs never count as demand, so they never boot a VM, and the dashboard lists them as refused. That is not enough on its own. Any JIT runner, booted for some other job, can be handed a fork PR's queued job if the labels match. So the guarantee has to live where the job actually lands. Every VM runs a runner job-started hook (kiln-prejob.sh) before the job's first step. If the event is a pull_request or pull_request_target from another repository, or a workflow_run triggered by one, the hook fails the job on the spot, and kiln kills the VM when it sees the assignment. It fails closed: a deleted fork, a missing head repository or an unreadable event payload are all refused.
The hook has a known blind spot. It sees where the event came from, not what the workflow checks out afterwards. A /test comment bot triggered by issue_comment that checks out a PR head can still run fork code, so on public repos those workflows belong on GitHub-hosted runners.
The cache: one trusted writer, throwaway readers
Fresh VMs are clean, but they are also cold. kiln gives each repo one persistent cache disk, cache/<owner>__<name>.qcow2. Inside the guest it is bind-mounted over /var/lib/docker, ~/.cache, ~/.npm, ~/.cargo, ~/.rustup, Go modules, Gradle and Maven caches, and (new in v0.2.0) /var/cache/apt/archives. Docker layers and toolchains are warm with no actions/cache round trip. For kiln's own repo the cache disk is about 924 MB.
Every job, PR or branch, gets a private qcow2 overlay of that disk, so any number of jobs can read it at once. An overlay is merged back with qemu-img commit only when all of these hold:
- the job succeeded according to GitHub's API, not the console. A job is root in its VM and can print anything, so the console alone never earns a commit;
- the event is
push; - the branch is the default branch or one of the repo's
cache_branches(v0.2.0; for teams that integrate ondev); - the pushed commit really is on that branch. kiln resolves the tip through
git/ref/heads/<branch>and checks it with the compare API, because a pushed tag nameddevalso arrives asevent=pushwith that name.
Everything else is discarded with the overlay, and the job page says why ("cache not saved: pull_request event"). Committing under a live reader would corrupt that reader's overlay, so a trusted overlay is first parked and committed once no other VM of that repo has the cache attached. qemu-img commit takes the image write lock, so a commit that races a QEMU process fails cleanly and the next tick retries it. Per-repo sizes (repo_cache_gb) arrived in v0.2.0.
Network egress
By default (egress: "open") QEMU uses user-mode networking, and a job reaches whatever the host reaches, including the LAN and the tailnet. That is fine for your own private repos.
egress: "filtered" is for anything less trusted. Each VM's QEMU starts inside rootlesskit --net=slirp4netns, a new unprivileged network namespace. A sh -e wrapper loads an nftables ruleset first, then execs QEMU through setpriv with every capability dropped and no-new-privs set, so QEMU cannot change the rules. The output chain drops by default. It allows DNS to slirp's resolver and the Docker mirror at 10.0.2.2:5000, drops RFC 1918, CGNAT (the tailnet), loopback, link-local (cloud metadata) and SMTP, and then accepts the public internet. IPv6 is off. Each of these steps fails closed, and a readiness probe gates launches: while it fails, kiln launches nothing rather than falling back to open.
The SECURITY.md is explicit about the residual risks: DNS is not filtered (so it can be used to exfiltrate data), IPv6 is disabled rather than filtered, and a router that hairpins its WAN address can expose services you forwarded.
GitHub App authentication
Until v0.2.0, kiln needed a PAT from a repo admin, which meant a durable, highly privileged credential parked on the CI box. v0.2.0 adds GitHub App auth, and it is now the recommended setup:
- One-click creation. Settings › GitHub › Create GitHub App uses GitHub's manifest flow. GitHub shows the App with its permissions preset (Administration: write, Actions: write, Contents: read, Metadata: read) and redirects back to the dashboard, which trades the code for the App id and private key.
-
Short-lived tokens. Only
app.pemis stored (mode 0600, written atomically). Installation tokens are minted per installation, expire after an hour and are refreshed before they do. - Repos come from installations. The repos kiln serves are the repos the App is installed on, refreshed every 5 minutes (every 30 s while it serves none).
-
Only the owner's installations. A public App can be installed by anyone, so kiln serves only installations on the App owner's account (plus an explicit
app_accountslist). Other installations are ignored and shown as "Skipped".
One pitfall is worth calling out. The obvious way to check "is this installation the owner's?" is to compare logins. But GitHub logins can be renamed, so a login comparison silently stops serving after the owner renames their account. kiln reads the owner from GET /app at every refresh and matches installations by the numeric account id, which survives renames. A unit test pins this down: same id, new login, still served. A call for a repo the App does not serve never borrows another installation's token.
Signed over-the-air updates
A CI box you never SSH into should not need SSH to upgrade. Every v0.2.0 release ships two builds, kiln-0.2.0-x86_64-linux.tar.gz (glibc 2.39+) and kiln-0.2.0-x86_64-linux-musl.tar.gz (fully static, any x86_64 Linux), each with a .sha256 and a 64-byte Ed25519 .sig. Settings › Updates shows a new release and its notes; one click does the rest:
flowchart LR
A[Newer release seen] --> B[Download the tarball for this build]
B --> C{Ed25519 signature valid<br/>against the key compiled in?}
C -- no --> X[Refuse, nothing written]
C -- yes --> D[Extract only kiln, check it reports the new version]
D --> E[Drain: no new VMs, wait for running jobs and bakes]
E --> F[Atomic swap, keep .prev]
F --> G[Re-exec in place, same PID under systemd]
G --> H{Starts cleanly?}
H -- yes --> I[Done, page reloads itself]
H -- "no, twice" --> R[Restore .prev and skip that version]
Three details carry the weight. The signature covers the exact bytes that get unpacked, and they never touch the disk before extraction. The version is bound three ways (the tag, the tarball's top directory, and what the new binary prints), so a genuine old build relabelled as a new release is refused. And signing happens in its own job on a fresh VM, with the key in a GitHub Environment only v* tags can deploy to; the build job, where build scripts, proc macros and the toolchain installer run, never sees it.
The drain step is not optional. That lesson came from operating the thing (see below): restarting kiln kills running jobs.
The dashboard, now a PWA
The dashboard is one HTML file embedded with include_str! and served by axum on :7878. It answers only tailnet peers whose Tailscale identity is the box owner (or listed in allowed_users). Requests from the box itself need a key, because job VMs reach the host through QEMU's NAT and look like local traffic.
It has an Overview with one "chamber" per VM slot, a Jobs page with a queue, boot, wait and job timeline plus live console, steps (mirrored over a second serial port) and GitHub logs, Repos with workflow dispatch, rerun and cancel, and Settings with Diagnostics (the same checks as kiln doctor). v0.2.0 redesigns it with elevated surfaces instead of outlines, an icon set, tooltips on states and actions, zero layout shift between pages, and Settings › Appearance (theme, density, accent, reduced motion). It is also installable as a PWA, so the firing log sits on your phone's home screen.
Operational lessons
- Don't restart the box mid-job. On SIGTERM kiln stops launching, kills its VMs, deregisters runners that never got a job, and waits up to 20 s for cleanup. A job in flight is lost. Hence the drain step in OTA.
-
Gate launches on the recipe version, not just image age. The fork-refusal hook lives in the image. An image baked by an older recipe lacks it, so kiln records
recipeinbase.jsonand launches nothing, warm VMs included, until the image is rebaked.auto_rebake(on by default) handles it within minutes. The same mechanism rebakes when the runner version goes stale, or whenbake_node_versionsorbake_apt_packageschange. -
Bake what you always need. Node 20 was a hard requirement for prod parity, and
setup-nodedownloading it on every job never warms, because the tool cache is deliberately off the cache disk.bake_node_versions: ["20", "24"]givesFound in cache @ /opt/hostedtoolcache/node/20.20.2/x64while barenode -vstays on 24. Tarballs are checked against nodejs.org'sSHASUMS256.txt. - Recycle idle VMs when security settings change. Each VM records a fingerprint of its boot-time security settings. Switch egress to filtered and idle open-egress VMs are deregistered before they can take a job.
Getting started
The repository is currently private, so the release download uses an authenticated gh. On the CI box:
v=0.2.0 flavor=x86_64-linux # or x86_64-linux-musl for any Linux
cd "$(mktemp -d)"
gh release download "v$v" -R Bunty9/kiln -p "kiln-$v-$flavor.tar.gz*"
sha256sum -c "kiln-$v-$flavor.tar.gz.sha256"
tar -xzf "kiln-$v-$flavor.tar.gz"
install -Dm755 "kiln-$v-$flavor/kiln" ~/.local/bin/kiln
install -Dm644 "kiln-$v-$flavor/deploy/kiln.service" ~/.config/systemd/user/kiln.service
kiln doctor # KVM, tools, disk, memory, auth, mirror, Tailscale
kiln bake # one time, about 5 minutes
systemctl --user daemon-reload && systemctl --user enable --now kiln
loginctl enable-linger $USER
Open http://<box>:7878 from the tailnet and the first-run stepper walks you through creating the GitHub App, the image, a repo and a first job. Then opt workflows in:
jobs:
test:
if: github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository
runs-on: [self-hosted, kiln] # or kiln-2cpu ... kiln-16cpu
The if: guard is a second layer on top of kiln's own fork refusal.
Limits and roadmap
- Runners are per repo. With an App, kiln serves every repo it is installed on, but org-level runner groups are not implemented.
-
Network throughput. slirp tops out well below line rate and burns CPU on large
docker pulls. Candidates:passt(still rootless), or a one-time root setup of tap devices. - Single host, x86_64, Ubuntu 24.04 guests. No macOS, Windows or GPU. Scheduling is a per-repo count with no priorities or fair-share.
- Filtered egress is opt-in and does not filter DNS.
Closing
A kiln does one thing well: it takes something raw, fires it in a controlled chamber, and gives you back a finished piece, or a cracked one you throw away. That is the model here. Every job gets a fresh chamber, the cache is glazed only by trusted pushes, and nothing from a fork goes near the heat. If you have a spare Linux box and a hosted CI bill, I would like to hear how your workloads behave on it. The repo is public for now, reach out to @bunty9 for contributing.
Top comments (0)