DEV Community

Cover image for I didn't want my homelab to deploy on every push, so I built a checkbox
Robbe Verhelst
Robbe Verhelst

Posted on

I didn't want my homelab to deploy on every push, so I built a checkbox

The Sluiceway dashboard issue on my homelab repo: 18 pending stacks, 35 in sync, and a warning that 3 pending stacks delete or replace resources

My homelab runs on Pulumi. Proxmox on the metal, a Kubernetes cluster on top, and every app deployed from one repository. Today that repository has 54 stacks.

For a long time I deployed them by hand. Merge a PR, open a terminal, run a preview, run the deploy. That works when you have five stacks. With fifty, I stopped knowing which ones I had actually deployed. Some stacks went weeks without a deploy, and when I finally ran a preview on one, it showed a pile of changes I didn't expect. Some were mine from three PRs ago. Some I couldn't place at all.

The obvious fix is CI: deploy on every push to main. I sat with that idea for a while and didn't like it. The whole problem was changes I didn't expect. Deploying on every push doesn't remove those surprises. It just ships them to production before I see them.

So I wanted something in between. After a merge, show me what is waiting for every stack. Then let me pick which ones go out.

A dashboard in a GitHub issue

That's what Sluiceway is. After a merge, it runs a preview for every stack in the repo and writes the result into one GitHub issue. One row per stack, with the number of creates, updates and deletes, and the list of what changes.

A pending row expanded: the garden stack with 21 creates and 2 updates, each change listed

Every pending row has a checkbox. Tick it, and GitHub Actions deploys exactly that stack. Nothing else.

That's the whole interaction. I merge, I look at the issue, I tick what I'm happy with. Stacks I'm not sure about stay pending until I am.

Some things it does on top of that, because I needed them:

  • It warns about deletes and replaces. As I write this, my dashboard has a caution block at the top: three pending stacks would delete or replace resources, one of them the cluster upgrade itself. That's the row I don't tick on a Friday evening.
  • It shows where a change came from. A row says which PR caused it, and whether the stack also picked up changes from outside that PR.
  • It notices stacks that won't settle. If a stack is pending again right after deploying the same change, a value in the program probably differs on every run. The row says so, instead of asking me to deploy it forever.

The caution block and the three stacks that would replace resources, including the cluster upgrade and the Proxmox VM templates

The tick that deployed nothing

What I approve should be what runs. So when I tick a box, Sluiceway doesn't just replay the preview from earlier. It runs a fresh preview, compares it to the row I ticked, and only deploys if they match.

On 24 September I ticked workspaces/apps/arc:prod. Between the preview on the dashboard and my tick, a value in that stack had changed, one that the row didn't show. So nothing was deployed. The bot left a comment, updated the row to what the change looked like now, and asked me to look again.

The bot comment: I ticked workspaces/apps/arc:prod, a value the row does not show changed since the tick, so nothing was deployed

A deploy-on-push pipeline would have shipped that. That's exactly the kind of change I built this to catch.

How it runs

There are two ways to set it up, and they use the same engine.

The GitHub App is the quickest. Install it, pick the repos, and it opens an onboarding PR with the workflow already written for what it found in your repo. Merge that, and the first scan writes the dashboard. With the app installed, a tick lands in about a second instead of the 17 to 23 seconds it takes GitHub to start a workflow for an issue edit. It's free for up to 5 repos.

The action on its own is open source (Apache-2.0) and complete for a single repo. You add the workflow yourself. The core of it is one job with one step: a push or the daily schedule scans, a tick deploys.

on:
  push:
    branches: [main]
  schedule:
    - cron: "0 6 * * *"
  issues:
    types: [edited]

jobs:
  sluiceway:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: pulumi/actions@v7 # or OpenTofu, Terraform, Helm, kubectl
      # load your credentials here
      - uses: sluiceway/sluiceway@v0
Enter fullscreen mode Exit fullscreen mode

The full version with permissions and comments is in the README, and npx sluiceway init writes a first one from what it finds in your repo. It finds Pulumi and OpenTofu/Terraform stacks on its own, and also handles Terragrunt, CDKTF, Helm and plain Kubernetes manifests.

Either way, the deploys run in your own GitHub Actions runners, with your own credentials. The app never holds cloud credentials and can't start a deploy by itself. It only keeps facts about stacks (state, counts, who ticked, which commit, when), never resource names, diffs or values. Nothing deploys unless a person ticks, or you set a stack to deploy: on-merge yourself. Destroys always wait for a person.

My homelab is the first user, through the app. The app catches each tick and starts the workflow, which runs on my own runners in the cluster. Renovate opens PRs on the repo every day. Before, every merged Renovate PR meant a manual preview and deploy. Now they show up as pending rows, and I tick the ones that look boring.

And at work

I built this for my homelab, but the same problem exists on a platform team, just with more people and more repos. Who is allowed to deploy, did they see what they deployed, and what is waiting where?

The action already covers who may tick: a tick rule in sluiceway.yaml (for example tickers: maintain) narrows who can deploy, or you put the deploy behind a GitHub Environment with required reviewers. Every deploy is written as a GitHub deployment: who ticked it, which commit, when.

The app adds what you need once there are many repos:

  • A console across repos. One view of what's pending, failed or drifting in every repo, instead of opening dashboards one by one.
  • An audit log of every tick and deploy, with CSV export.
  • Insights: lead time, deploy frequency, failure rate and drift per repo.
  • Edit mode: change sluiceway.yaml from the console, and it opens a PR for you.
  • A required second approver, on the paid Team plan. It's built on GitHub's deployment protection rules, so it only really binds on GitHub Enterprise or public repos.

How it compares

When I explain this, people ask how it's different from what's already out there. Roughly:

  • A deploy-on-merge pipeline ships everything that lands on main. Fast, but no moment to look.
  • Atlantis is self-hosted and Terraform-focused. You comment atlantis apply on the PR, so you deploy before you merge, one PR at a time.
  • Spacelift, env0, HCP Terraform and Pulumi Cloud are hosted platforms that generally run the deploys for you, so they need access to your cloud.
  • Sluiceway deploys after the merge, per stack, from one place, in your own runners.

None of these is wrong. I wanted the one where main is the truth, nothing ships without me looking, and my credentials never leave my runners.

Try it

The quickest way is the GitHub App: sluiceway.dev. It's free for up to 5 repos, and the onboarding PR does the setup. Paid plans start when you have more repos than that.

If you'd rather run it yourself, the action is free and open source: github.com/sluiceway/sluiceway. Docs are at docs.sluiceway.dev.

It's still 0.x. If you try it and something doesn't fit how you deploy, I'd like to hear about it.

Top comments (0)