<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: david</title>
    <description>The latest articles on DEV Community by david (@dwoitzik).</description>
    <link>https://dev.to/dwoitzik</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3933869%2F1fb8aa5b-2239-46a7-bf78-b5352809883c.png</url>
      <title>DEV Community: david</title>
      <link>https://dev.to/dwoitzik</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dwoitzik"/>
    <language>en</language>
    <item>
      <title>Postgres 16 18: Why You Can't Just Swap the Image</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 05 Oct 2026 12:44:42 +0000</pubDate>
      <link>https://dev.to/dwoitzik/postgres-16-18-why-you-cant-just-swap-the-image-1j4j</link>
      <guid>https://dev.to/dwoitzik/postgres-16-18-why-you-cant-just-swap-the-image-1j4j</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/postgres-16-18-migration-fresh-pvc/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;PostgreSQL major version upgrades are not swap-the-image operations. You can't change &lt;code&gt;postgres:16&lt;/code&gt; to &lt;code&gt;postgres:18&lt;/code&gt; in a Deployment and expect it to work. The data directory format changes, system catalogs are incompatible, and extensions need to be rebuilt.&lt;/p&gt;

&lt;p&gt;When CNPG (CloudNativePG) upgraded from PG16 to PG18, the procedure required a fresh PVC, a full dump and restore, and a role discovery bug that nearly locked me out of the &lt;a href="https://dev.to/blog/k3s-authelia-proxmox-homelab/"&gt;Authelia&lt;/a&gt; database.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why You Can't Swap the Image
&lt;/h2&gt;

&lt;p&gt;PostgreSQL stores data in a format specific to its major version. The &lt;code&gt;PG_VERSION&lt;/code&gt; file in the data directory tells PostgreSQL which version created it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /var/lib/postgresql/data/PG_VERSION
&lt;span class="c"&gt;# → 16&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When PostgreSQL 18 starts and finds &lt;code&gt;PG_VERSION = 16&lt;/code&gt;, it refuses to start:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FATAL: data directory has wrong ownership
HINT: The data directory was initialized by PostgreSQL 16. 
      Upgrade by running pg_upgrade.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pg_upgrade&lt;/code&gt; is the official tool for in-place major version upgrades. It copies data files from the old format to the new format, rewriting system catalogs and tuple headers. It works well on bare-metal PostgreSQL where you have direct filesystem access.&lt;/p&gt;

&lt;p&gt;On Kubernetes with CNPG, &lt;code&gt;pg_upgrade&lt;/code&gt; is not the right approach because:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;CNPG manages the data directory through its operator — manual modifications are reverted&lt;/li&gt;
&lt;li&gt;The PVC is bound to the cluster definition — changing the PostgreSQL version in the CRD doesn't automatically upgrade the data&lt;/li&gt;
&lt;li&gt;CNPG's recommended migration path is dump-and-restore, not in-place upgrade&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Migration Procedure
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Dump from PG16
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Exec into the CNPG pod&lt;/span&gt;
kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; database &lt;span class="nt"&gt;-it&lt;/span&gt; postgres-authelia-0 &lt;span class="nt"&gt;--&lt;/span&gt; bash

&lt;span class="c"&gt;# Dump the database&lt;/span&gt;
pg_dump &lt;span class="nt"&gt;-U&lt;/span&gt; postgres &lt;span class="nt"&gt;-Fc&lt;/span&gt; authelia &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /tmp/authelia.dump

&lt;span class="c"&gt;# Copy the dump to a temporary location&lt;/span&gt;
kubectl &lt;span class="nb"&gt;cp &lt;/span&gt;database/postgres-authelia-0:/tmp/authelia.dump ./authelia.dump
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;-Fc&lt;/code&gt; flag produces a custom-format dump that's compressed and can be restored with &lt;code&gt;pg_restore&lt;/code&gt;. The dump includes all data, schemas, roles, and extensions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Create Fresh PG18 Cluster
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kubernetes/system/postgres/cluster.yml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgresql.cnpg.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Cluster&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres-authelia&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;database&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;imageName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io/cloudnative-pg/postgresql:16.4&lt;/span&gt;  &lt;span class="c1"&gt;# temporary — will be updated&lt;/span&gt;
  &lt;span class="na"&gt;instances&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;
  &lt;span class="na"&gt;storage&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;size&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;2Gi&lt;/span&gt;
    &lt;span class="na"&gt;storageClass&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;local-path&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wait — the image is still PG16. That's intentional. CNPG creates the cluster with PG16 first, then upgrades the image to PG18 after the data is restored. This ensures the PVC and the operator agree on the initial state.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Restore into PG18
&lt;/h3&gt;

&lt;p&gt;After the cluster is running with PG16, update the image to PG18:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;imageName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io/cloudnative-pg/postgresql:18.4&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CNPG detects the image change, creates a new pod with PG18, and the old PG16 pod is terminated. The data directory is still PG16 format, so the new pod fails to start — which is expected.&lt;/p&gt;

&lt;p&gt;Now restore the dump into the fresh PG18 data directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Port-forward to the CNPG service&lt;/span&gt;
kubectl port-forward &lt;span class="nt"&gt;-n&lt;/span&gt; database svc/postgres-authelia 5432:5432 &amp;amp;

&lt;span class="c"&gt;# Restore the dump&lt;/span&gt;
pg_restore &lt;span class="nt"&gt;-U&lt;/span&gt; postgres &lt;span class="nt"&gt;-d&lt;/span&gt; authelia &lt;span class="nt"&gt;--clean&lt;/span&gt; &lt;span class="nt"&gt;--if-exists&lt;/span&gt; ./authelia.dump
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pg_restore&lt;/code&gt; handles the format conversion — it reads PG16-format data and writes it in PG18 format. The &lt;code&gt;--clean --if-exists&lt;/code&gt; flags drop existing objects before restoring, ensuring a clean state.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Update Authelia
&lt;/h3&gt;

&lt;p&gt;Update the Authelia deployment to use the new Postgres 18 connection string (same host, same port, same database — the connection string doesn't change):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Verify Authelia connects to the new database&lt;/span&gt;
&lt;span class="s"&gt;kubectl logs -n apps -l app=authelia --tail=20&lt;/span&gt;
&lt;span class="c1"&gt;# → "Successfully connected to PostgreSQL 18.4"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Role Discovery Bug
&lt;/h2&gt;

&lt;p&gt;During the restore, &lt;code&gt;pg_restore&lt;/code&gt; reported:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pg_restore: error: could not open input file "/tmp/authelia.dump": No such file or directory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dump file was at a different path than expected. The real problem: the restore was running from a pod that had a different filesystem layout than the dump pod.&lt;/p&gt;

&lt;p&gt;After fixing the path, the restore completed but Authelia failed to start:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FATAL: role "oc_dw" does not exist
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;oc_dw&lt;/code&gt; role was a leftover from the original CNPG cluster initialization — CNPG creates a default operator role that isn't visible in &lt;code&gt;pg_dump&lt;/code&gt; output because it's a replication role, not a regular database role.&lt;/p&gt;

&lt;p&gt;The fix: create the missing role before restoring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;oc_dw&lt;/span&gt; &lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="n"&gt;REPLICATION&lt;/span&gt; &lt;span class="n"&gt;LOGIN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lesson: CNPG creates internal roles (&lt;code&gt;oc_dw&lt;/code&gt;, &lt;code&gt;streaming_replica&lt;/code&gt;) that aren't included in &lt;code&gt;pg_dump&lt;/code&gt; output. After a dump-and-restore migration, these roles must be recreated manually.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fresh PVC Requirement
&lt;/h2&gt;

&lt;p&gt;The critical step that most tutorials skip: &lt;strong&gt;you need a fresh PVC.&lt;/strong&gt; You can't restore a PG16 dump into a PG16 data directory and then expect PG18 to read it. The data directory must be empty — PG18 creates its own data directory format on first start, and &lt;code&gt;pg_restore&lt;/code&gt; populates it.&lt;/p&gt;

&lt;p&gt;In CNPG, this means:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Delete the existing PVC (data loss — you need the dump)&lt;/li&gt;
&lt;li&gt;Let CNPG create a new PVC with the PG18 image&lt;/li&gt;
&lt;li&gt;Restore the dump into the fresh cluster&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The PVC deletion is the scary part. If the dump is corrupted or incomplete, the data is gone. The verification before deletion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Verify dump is complete&lt;/span&gt;
pg_restore &lt;span class="nt"&gt;--list&lt;/span&gt; ./authelia.dump | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
&lt;span class="c"&gt;# → Should show hundreds of objects (tables, sequences, functions)&lt;/span&gt;

&lt;span class="c"&gt;# Verify dump integrity&lt;/span&gt;
pg_restore &lt;span class="nt"&gt;--verbose&lt;/span&gt; &lt;span class="nt"&gt;--no-owner&lt;/span&gt; &lt;span class="nt"&gt;--no-privileges&lt;/span&gt; &lt;span class="nt"&gt;--dry-run&lt;/span&gt; ./authelia.dump
&lt;span class="c"&gt;# → Should complete without errors&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What I'd Change
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Use CNPG's Backup/Restore instead of manual dump.&lt;/strong&gt; CNPG supports &lt;code&gt;pg_basebackup&lt;/code&gt; and WAL archiving to S3. A CNPG backup includes the operator roles and can be restored directly without the &lt;code&gt;oc_dw&lt;/code&gt; gotcha. The manual dump approach was chosen because the existing cluster wasn't configured for CNPG backups at the time of migration.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Test the migration on a non-production cluster first.&lt;/strong&gt; The &lt;code&gt;oc_dw&lt;/code&gt; role discovery happened during the Authelia migration — the only SSO for every service. If the restore had failed, every service would be unreachable.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;PostgreSQL major version upgrades on Kubernetes are the same challenge as Azure Database for PostgreSQL Flexible Server upgrades: Azure handles the in-place upgrade automatically, but the same data directory format incompatibility exists. The difference is that Azure abstracts the dump-and-restore behind a API call, while Kubernetes requires you to do it manually. The underlying PostgreSQL constraint is identical: major versions are not backward-compatible at the storage layer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/go/4gEAhv9"&gt;The Linux Command Line*&lt;/a&gt; is worth having on the shelf for exactly this kind of migration - &lt;code&gt;pg_dump&lt;/code&gt;, &lt;code&gt;pg_restore&lt;/code&gt;, &lt;code&gt;kubectl cp&lt;/code&gt;, and a dozen other shell tools chained together under time pressure go a lot smoother when the shell itself isn't also something you're learning in the moment.&lt;/p&gt;



</description>
      <category>kubernetes</category>
      <category>postgres</category>
      <category>homelab</category>
      <category>migration</category>
    </item>
    <item>
      <title>Renovate in Kubernetes: OOM, Hashicorp Downloads, and the 6GB Container</title>
      <dc:creator>david</dc:creator>
      <pubDate>Sun, 04 Oct 2026 11:30:39 +0000</pubDate>
      <link>https://dev.to/dwoitzik/renovate-in-kubernetes-oom-hashicorp-downloads-and-the-6gb-container-56n8</link>
      <guid>https://dev.to/dwoitzik/renovate-in-kubernetes-oom-hashicorp-downloads-and-the-6gb-container-56n8</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/renovate-kubernetes-oom-6gb-container/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Renovate is the dependency update bot that keeps my k3s cluster current, and it's the "auto-update" half of &lt;a href="https://dev.to/blog/operating-model-auto-update-human/"&gt;the operating model&lt;/a&gt; that decides what gets merged automatically vs. what waits for a human. It runs as a Kubernetes CronJob every 2 hours, checks for new container images, Helm chart versions, and Terraform provider releases, and opens PRs on GitHub.&lt;/p&gt;

&lt;p&gt;For two weeks, it was failing silently on every run. The CronJob reported "Completed." The GitHub commits showed no new PRs. The logs showed nothing — because Renovate exits cleanly on config validation errors, and the CronJob's &lt;code&gt;successfulJobsHistoryLimit&lt;/code&gt; kept only the last successful run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure 1: Invalid Preset
&lt;/h2&gt;

&lt;p&gt;The first failure was a config error. The &lt;code&gt;renovate.json&lt;/code&gt; file referenced an invalid preset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"extends"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"config:base"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;":enableHelpfulPre-commit"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;:enableHelpfulPre-commit&lt;/code&gt; is not a real Renovate preset. Renovate's config validation caught the error and exited with a zero exit code — because config validation errors are treated as "nothing to do," not "something broke."&lt;/p&gt;

&lt;p&gt;The CronJob's exit code was 0. Kubernetes considered the job successful. No alert fired. Renovate did nothing on every run.&lt;/p&gt;

&lt;p&gt;The fix: remove the invalid preset and replace the deprecated &lt;code&gt;config:base&lt;/code&gt; with &lt;code&gt;config:recommended&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"extends"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"config:recommended"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Failure 2: OOM at 512Mi
&lt;/h2&gt;

&lt;p&gt;After fixing the config, Renovate started running — and immediately OOM'd.&lt;/p&gt;

&lt;p&gt;The container limit was 512Mi. Renovate's Node.js runtime, plus the GitHub API client, plus the Terraform provider registry client, plus the container image metadata parser, consumed more than 512Mi during a full scan.&lt;/p&gt;

&lt;p&gt;The symptom: the Pod restarted with &lt;code&gt;OOMKilled&lt;/code&gt; status. Kubernetes restarted it. The next run OOM'd again. After 3 failures, the CronJob marked the job as failed — but &lt;code&gt;failedJobsHistoryLimit: 3&lt;/code&gt; meant only the last 3 failures were visible, and they were being garbage-collected faster than I checked.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get &lt;span class="nb"&gt;jobs&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; apps &lt;span class="nt"&gt;-l&lt;/span&gt; &lt;span class="nv"&gt;app&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;renovate
&lt;span class="c"&gt;# NAME              COMPLETIONS   DURATION   AGE&lt;/span&gt;
&lt;span class="c"&gt;# renovate-28197    0/1           ...        2m&lt;/span&gt;
&lt;span class="c"&gt;# renovate-28196    0/1           ...        2h&lt;/span&gt;
&lt;span class="c"&gt;# renovate-28195    0/1           ...        4h&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fix: raise the memory limit to 1Gi and set &lt;code&gt;NODE_OPTIONS&lt;/code&gt; to limit the V8 heap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;NODE_OPTIONS&lt;/span&gt;
    &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--max-old-space-size=896"&lt;/span&gt;
&lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;6Gi&lt;/span&gt;
  &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1Gi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;NODE_OPTIONS&lt;/code&gt; limit of 896MB (leaving ~100MB for native code and runtime overhead) prevents V8 from consuming the entire container memory during garbage collection cycles. Without it, V8 allocates memory up to the container limit and gets OOMKilled before the next GC cycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure 3: Terraform Provider Downloads
&lt;/h2&gt;

&lt;p&gt;The third failure was the most subtle. Renovate was consuming 6GB of memory during Terraform provider update checks. The reason: Renovate downloads Terraform provider binaries to verify version compatibility.&lt;/p&gt;

&lt;p&gt;For each Terraform provider in the repo (&lt;code&gt;routeros&lt;/code&gt;, &lt;code&gt;proxmox&lt;/code&gt;, &lt;code&gt;cloudflare&lt;/code&gt;, &lt;code&gt;garage&lt;/code&gt;), Renovate:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Queries the Terraform Registry API for available versions&lt;/li&gt;
&lt;li&gt;Downloads the provider binary for the current and latest versions&lt;/li&gt;
&lt;li&gt;Compares binary compatibility (provider schema, API version)&lt;/li&gt;
&lt;li&gt;Opens a PR if a newer version is available&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 2 is the memory problem. Terraform provider binaries are 50-200MB compressed. With 4 providers, plus their dependencies, Renovate downloads ~500MB of provider binaries during a full scan. The decompressed binaries and their metadata consume 2-3x that in memory.&lt;/p&gt;

&lt;p&gt;The fix was two-part:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Raise the memory limit to 6Gi&lt;/strong&gt; — enough for the full scan including provider downloads&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Limit concurrent requests&lt;/strong&gt; to the Terraform Registry:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"terraform"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"concurrentRequestLimit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;concurrentRequestLimit: 2&lt;/code&gt;, Renovate downloads at most 2 provider binaries simultaneously, reducing peak memory usage from ~6GB to ~3GB. The scan takes longer, but it completes within the 6Gi limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Silent Failure Pattern
&lt;/h2&gt;

&lt;p&gt;The common thread across all three failures: Renovate exited cleanly. The CronJob reported success. No alert fired. The only evidence of failure was the absence of new PRs on GitHub.&lt;/p&gt;

&lt;p&gt;This is the silent failure pattern:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The application fails internally but exits with code 0&lt;/li&gt;
&lt;li&gt;Kubernetes considers the job successful&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;successfulJobsHistoryLimit&lt;/code&gt; preserves the "successful" job&lt;/li&gt;
&lt;li&gt;Nobody checks because the job "succeeded"&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The fix: add a post-run check that verifies Renovate actually did something:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Post-run check: verify PRs were opened&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Verify Renovate activity&lt;/span&gt;
  &lt;span class="na"&gt;script&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;PR_COUNT=$(gh pr list --repo dwoitzik/homelab-infrastructure \&lt;/span&gt;
      &lt;span class="s"&gt;--author "app/renovate" --state open --json number --jq 'length')&lt;/span&gt;
    &lt;span class="s"&gt;if [ "$PR_COUNT" -eq 0 ] &amp;amp;&amp;amp; [ "$(date +%u)" -le 5 ]; then&lt;/span&gt;
      &lt;span class="s"&gt;echo "WARNING: No open Renovate PRs on a weekday"&lt;/span&gt;
      &lt;span class="s"&gt;# Alert via Discord webhook&lt;/span&gt;
    &lt;span class="s"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Current Configuration
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kubernetes/apps/renovate/renovate.yml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;batch/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;CronJob&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;renovate&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*/2&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;  &lt;span class="c1"&gt;# every 2 hours&lt;/span&gt;
  &lt;span class="na"&gt;jobTemplate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;renovate&lt;/span&gt;
              &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io/renovatebot/renovate:43.245.0&lt;/span&gt;
              &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RENOVATE_TOKEN&lt;/span&gt;
                  &lt;span class="na"&gt;valueFrom&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                    &lt;span class="na"&gt;secretKeyRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;renovate-token&lt;/span&gt;
                      &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;github-pat&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;NODE_OPTIONS&lt;/span&gt;
                  &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--max-old-space-size=896"&lt;/span&gt;
              &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;6Gi&lt;/span&gt;
                  &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;2000m&lt;/span&gt;
                &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1Gi&lt;/span&gt;
                  &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;100m&lt;/span&gt;
          &lt;span class="na"&gt;restartPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Never&lt;/span&gt;
      &lt;span class="na"&gt;backoffLimit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;restartPolicy: Never&lt;/code&gt; with &lt;code&gt;backoffLimit: 2&lt;/code&gt; means the CronJob retries twice on failure, then gives up. Combined with the memory limit and heap size cap, Renovate completes its scan within resource bounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Lesson
&lt;/h2&gt;

&lt;p&gt;Silent failures in CronJobs are the most dangerous failure mode in a Kubernetes cluster — the same "reports success but did nothing real" pattern as ArgoCD showing &lt;code&gt;Synced/Healthy&lt;/code&gt; on an Application that never actually applied. The job "succeeds," the operator doesn't check, and the dependency update pipeline silently stops working. Weeks pass without updates, security patches don't apply, and nobody notices until a vulnerability is disclosed for a package that Renovate would have updated.&lt;/p&gt;

&lt;p&gt;The fix isn't just more memory — it's observability on the pipeline itself. Every CronJob that performs a critical function should have a post-run verification that confirms it actually did its job, not just that it exited cleanly.&lt;/p&gt;




&lt;p&gt;CronJob observability is the same problem in Azure DevOps: a pipeline that succeeds but produces no artifact is indistinguishable from a pipeline that didn't run. Azure Monitor's Pipeline Analytics tracks success rate, not "did the pipeline produce meaningful output." The fix in both cases is the same: add a verification step that checks for the expected outcome (PRs opened, artifacts published, deployments completed) and alerts on absence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/go/4f4RHQf"&gt;The Phoenix Project*&lt;/a&gt; tells the same story from the human side of this exact failure mode - the automated system that everyone assumes is working, until someone finally checks.&lt;/p&gt;



</description>
      <category>kubernetes</category>
      <category>gitops</category>
      <category>homelab</category>
      <category>debugging</category>
    </item>
    <item>
      <title>Longhorn to NFS: Why Distributed Storage Didn't Make Sense Here</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:44:58 +0000</pubDate>
      <link>https://dev.to/dwoitzik/longhorn-to-nfs-why-distributed-storage-didnt-make-sense-here-5530</link>
      <guid>https://dev.to/dwoitzik/longhorn-to-nfs-why-distributed-storage-didnt-make-sense-here-5530</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/longhorn-to-nfs-distributed-storage/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Longhorn is designed for multi-node Kubernetes clusters. It replicates PersistentVolume data across nodes, so if one node dies, the data survives on the other two. It's a great distributed storage system.&lt;/p&gt;

&lt;p&gt;I run three k3s nodes on a single physical host. All three nodes share the same NVMe. When Longhorn replicates a PVC across three nodes, all three replicas live on the same physical disk. There is no distribution. There is no redundancy. And when a node failed, Longhorn's own high-availability logic created a problem that wouldn't exist without it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Multi-Attach Error
&lt;/h2&gt;

&lt;p&gt;When a k3s node becomes unreachable (OomKill, kubelet crash, network blip), Longhorn marks its replicas as "rebuilding" and tries to attach the volume to a healthy node. But Longhorn's RWO (ReadWriteOnce) volumes can only be attached to one node at a time.&lt;/p&gt;

&lt;p&gt;The sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;k3s-12 becomes unreachable (kubelet lost lease)&lt;/li&gt;
&lt;li&gt;Longhorn detaches the PVC from k3s-12&lt;/li&gt;
&lt;li&gt;Longhorn tries to attach the PVC to k3s-13&lt;/li&gt;
&lt;li&gt;k3s-12 recovers and comes back online&lt;/li&gt;
&lt;li&gt;Longhorn sees both nodes claiming the volume&lt;/li&gt;
&lt;li&gt;Multi-Attach error: the volume is "attached" to two nodes simultaneously&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The pod that needs the volume (Postgres, for example) can't start on either node because Longhorn refuses to mount a volume that's in a Multi-Attach state. The fix requires manually detaching the volume from both nodes and letting Longhorn re-attach it cleanly — a manual step during an outage, exactly when automation should be working.&lt;/p&gt;

&lt;p&gt;On a multi-host cluster, this is a real HA scenario — the volume legitimately needs to failover to a different physical disk. On a single-host cluster, the "failover" is to the same NVMe, and the Multi-Attach error is pure overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  The NFS Alternative
&lt;/h2&gt;

&lt;p&gt;NFS (Network File System) doesn't have a Multi-Attach problem because it's not a block storage system. NFS exports a directory over the network. Any number of clients can mount it simultaneously. There's no "attachment" state to conflict.&lt;/p&gt;

&lt;p&gt;The NFS server runs on a dedicated LXC (&lt;code&gt;ct-srv-nfs-01&lt;/code&gt;) with a ZFS-backed dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# terraform/stacks/proxmox/lxc.tf&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"ct_srv_nfs_01"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;vm_id&lt;/span&gt;   &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;220&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"ct-srv-nfs-01"&lt;/span&gt;
  &lt;span class="nx"&gt;memory&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;dedicated&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2048&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="c1"&gt;# ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The k3s NFS provisioner creates PVCs on this NFS server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kubernetes/system/nfs-provisioner/application.yml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argoproj.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Application&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;helm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;nfs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10.0.20.100&lt;/span&gt;
          &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/archive&lt;/span&gt;
        &lt;span class="na"&gt;storageClass&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nfs-client&lt;/span&gt;
          &lt;span class="na"&gt;defaultClass&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When k3s-12 goes down, the pods reschedule to k3s-11 or k3s-13, mount the same NFS export, and continue where they left off. No Multi-Attach error, no manual detachment, no volume attachment state machine. NFS handles concurrent mounts natively.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-offs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What NFS Gives Up
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Longhorn&lt;/th&gt;
&lt;th&gt;NFS&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Replication across nodes&lt;/td&gt;
&lt;td&gt;Yes (3x by default)&lt;/td&gt;
&lt;td&gt;No (single server)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Snapshots&lt;/td&gt;
&lt;td&gt;Yes (per-volume)&lt;/td&gt;
&lt;td&gt;Yes (ZFS snapshots)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ReadWriteMany&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CSI driver&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes (nfs-subdir-external-provisioner)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Performance&lt;/td&gt;
&lt;td&gt;Block-level (faster)&lt;/td&gt;
&lt;td&gt;Network-mounted (slower)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Node failure tolerance&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No (single server)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;NFS gives up replication and node-failure tolerance. If the NFS server dies, every PVC it serves is unavailable. On a single physical host, this is the same risk as Longhorn — both storage systems live on the same NVMe. Longhorn's replication doesn't help when the underlying disk is the single point of failure.&lt;/p&gt;

&lt;h3&gt;
  
  
  What NFS Gives Back
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No Multi-Attach errors.&lt;/strong&gt; Pods reschedule freely without volume attachment conflicts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Simpler failure modes.&lt;/strong&gt; NFS is either up or down. No "rebuilding," "degraded," or "reverted" states.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ZFS snapshots for backup.&lt;/strong&gt; PBS backs up the NFS server's ZFS dataset, giving point-in-time recovery without Longhorn's snapshot overhead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lower resource usage.&lt;/strong&gt; Longhorn runs a per-node process (manager + driver) consuming ~200MB RAM per node. NFS is a single process on one LXC.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Storage Decision Matrix
&lt;/h2&gt;

&lt;p&gt;After migrating from Longhorn to NFS, I defined the rule:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Application&lt;/th&gt;
&lt;th&gt;Storage Class&lt;/th&gt;
&lt;th&gt;Reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PostgreSQL (CNPG)&lt;/td&gt;
&lt;td&gt;local-path&lt;/td&gt;
&lt;td&gt;CNPG manages its own replication&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Garage S3 metadata&lt;/td&gt;
&lt;td&gt;local-path&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dev.to/blog/nfs-vs-local-path-sqlite-trap/"&gt;SQLite file-locking (the NFS trap)&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mealie, Home Assistant&lt;/td&gt;
&lt;td&gt;local-path&lt;/td&gt;
&lt;td&gt;SQLite databases&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Everything else&lt;/td&gt;
&lt;td&gt;nfs-client&lt;/td&gt;
&lt;td&gt;Default, survives pod rescheduling&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The principle: &lt;strong&gt;embedded databases go on local-path (file-locking), everything else goes on NFS (survivability).&lt;/strong&gt; PostgreSQL doesn't fit either category because CNPG manages its own PV independently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Migration
&lt;/h2&gt;

&lt;p&gt;Moving PVCs from Longhorn to NFS required:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Deploy the NFS server LXC with the correct ZFS dataset&lt;/li&gt;
&lt;li&gt;Install the &lt;code&gt;nfs-subdir-external-provisioner&lt;/code&gt; Helm chart&lt;/li&gt;
&lt;li&gt;Copy data from Longhorn volumes to NFS (using &lt;code&gt;kubectl cp&lt;/code&gt; or a temporary pod)&lt;/li&gt;
&lt;li&gt;Update each application's &lt;code&gt;storageClassName&lt;/code&gt; from &lt;code&gt;longhorn&lt;/code&gt; to &lt;code&gt;nfs-client&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Delete the old Longhorn PVC (data is already on NFS)&lt;/li&gt;
&lt;li&gt;Remove Longhorn from the cluster&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 5 is the nerve-wracking part — deleting a PVC that contains production data. The verification before deletion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Verify data exists on NFS&lt;/span&gt;
kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; temp-pod &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-la&lt;/span&gt; /mnt/nfs/authelia-data/
&lt;span class="c"&gt;# → Confirm database files, config, etc.&lt;/span&gt;

&lt;span class="c"&gt;# Verify application starts with NFS PVC&lt;/span&gt;
kubectl delete pod authelia-xxxxx  &lt;span class="c"&gt;# force reschedule&lt;/span&gt;
&lt;span class="c"&gt;# → Pod starts, mounts NFS, passes health checks&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After the migration, Longhorn was uninstalled. The cluster went from three storage replicas on one disk to a single NFS server on the same disk — simpler, more predictable, and without the Multi-Attach false alarm.&lt;/p&gt;




&lt;p&gt;Distributed storage on a single host is the same anti-pattern as Azure Zone-Redundant Storage across availability zones that share the same power source. If the underlying infrastructure isn't actually distributed, the replication layer adds complexity without adding resilience. The fix in both cases: match the storage topology to the actual infrastructure topology. Single host = single NFS server. Multi-host = distributed storage.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/go/4aUtkCs"&gt;Designing Data-Intensive Applications*&lt;/a&gt; has the clearest explanation I've read of why replication only buys you resilience when the replicas are actually independent - which is the whole lesson of this migration in one sentence.&lt;/p&gt;



</description>
      <category>kubernetes</category>
      <category>storage</category>
      <category>homelab</category>
      <category>debugging</category>
    </item>
    <item>
      <title>The Operating Model: What Should Auto-Update vs. What Needs a Human</title>
      <dc:creator>david</dc:creator>
      <pubDate>Fri, 25 Sep 2026 11:00:27 +0000</pubDate>
      <link>https://dev.to/dwoitzik/the-operating-model-what-should-auto-update-vs-what-needs-a-human-1md6</link>
      <guid>https://dev.to/dwoitzik/the-operating-model-what-should-auto-update-vs-what-needs-a-human-1md6</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/operating-model-auto-update-human/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A homelab that requires daily attention isn't a homelab — it's a job. The goal: a cluster that runs for weeks without intervention, alerts when something breaks, self-heals what it can, and waits for a human only when a human is actually needed.&lt;/p&gt;

&lt;p&gt;This is the operating model that makes that work. Not aspirational — documented from what actually runs in production, including the deliberate gaps and the non-goals that keep the system honest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Auto-Updates
&lt;/h2&gt;

&lt;p&gt;Renovate runs every 2 hours via CronJob, tracking container images, Helm charts, and Terraform providers. But not everything gets the same treatment:&lt;/p&gt;

&lt;h3&gt;
  
  
  Auto-Merge (Patch/Minor/Digest, 3-Day Soak)
&lt;/h3&gt;

&lt;p&gt;Stateless workloads that can be rolled back by ArgoCD's &lt;code&gt;selfHeal&lt;/code&gt; if they break:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dashboard widgets (Homepage, Uptime Kuma display)&lt;/li&gt;
&lt;li&gt;Development tools (pre-commit hooks, linters)&lt;/li&gt;
&lt;li&gt;Non-critical utilities (SearXNG, Mealie, MySpeed)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Renovate opens a PR, CI runs, and after 3 days of green the PR auto-merges. If the new version breaks, ArgoCD's &lt;code&gt;syncPolicy.automated.selfHeal: true&lt;/code&gt; reverts to the previous version on the next sync.&lt;/p&gt;

&lt;h3&gt;
  
  
  PR-Only, Always (Manual Review Required)
&lt;/h3&gt;

&lt;p&gt;Stateful and critical services where a bad upgrade can cause data loss or auth outage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Databases:&lt;/strong&gt; CNPG Postgres, Redis/Valkey&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth:&lt;/strong&gt; &lt;a href="https://dev.to/blog/k3s-authelia-proxmox-homelab/"&gt;Authelia&lt;/a&gt;, &lt;a href="https://dev.to/blog/vault-auto-unseal-polling-sidecar/"&gt;Vault&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage:&lt;/strong&gt; &lt;a href="https://dev.to/blog/velero-garage-k3s-backup/"&gt;Garage S3&lt;/a&gt;, Velero&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apps with data migrations:&lt;/strong&gt; Nextcloud, Paperless, Gitea, Immich&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every update — patch, minor, or digest — gets a PR that requires manual review. The failure mode for these isn't "pod restarts" — it's "database schema mismatch" or "OIDC keys rotated."&lt;/p&gt;

&lt;h3&gt;
  
  
  PR-Only, Always (Major Version Bumps)
&lt;/h3&gt;

&lt;p&gt;Every major version bump, any package, regardless of tier. Major versions break APIs, change configuration formats, and introduce migration steps that no automated tool can reliably handle.&lt;/p&gt;

&lt;h3&gt;
  
  
  Terraform: Never Auto-Applied
&lt;/h3&gt;

&lt;p&gt;Every Terraform change goes through Atlantis: PR → plan → &lt;code&gt;atlantis apply&lt;/code&gt; comment → apply. No exceptions. Terraform changes physical and VM state on a single-point-of-failure host — that always gets a human in the loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Alerts (and Where)
&lt;/h2&gt;

&lt;p&gt;Prometheus + Alertmanager route to Discord via webhook. ArgoCD's notifications-controller sends app-state events through a separate path.&lt;/p&gt;

&lt;h3&gt;
  
  
  Critical Alerts (Always Discord)
&lt;/h3&gt;

&lt;p&gt;Everything with &lt;code&gt;severity: critical&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Proxmox host temperature&lt;/li&gt;
&lt;li&gt;k3s control-plane down&lt;/li&gt;
&lt;li&gt;Storage &amp;gt;93% full&lt;/li&gt;
&lt;li&gt;Postgres replication broken&lt;/li&gt;
&lt;li&gt;Certificate expiry soon&lt;/li&gt;
&lt;li&gt;Velero backup failed&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Warning Alerts (Selective)
&lt;/h3&gt;

&lt;p&gt;Only explicitly named warnings that matter at steady state:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;KubePodCrashLooping&lt;/code&gt; — a pod is crash-looping&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;KubeJobFailed&lt;/code&gt; — a CronJob failed (covers Renovate itself)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;KubeNodeNotReady&lt;/code&gt; — a node is unreachable&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;VeleroBackupPartialFailure&lt;/code&gt; — backup completed with warnings&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Deliberately NOT Alerted
&lt;/h3&gt;

&lt;p&gt;Routine warnings that would make Discord noisy without being actionable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;KubeMemoryQuotaOvercommit&lt;/code&gt; — LXC soft ceilings, not real pressure&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;KubeCPUThrottling&lt;/code&gt; — normal behavior for bursty workloads&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;KubePodNotReady&lt;/code&gt; during rolling updates — temporary, self-resolving&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The principle: &lt;strong&gt;an alert nobody acts on trains people to ignore Discord.&lt;/strong&gt; Every alert must have a corresponding action in the runbook. If there's no action, there's no alert.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Self-Heals
&lt;/h2&gt;

&lt;h3&gt;
  
  
  ArgoCD: Automatic Sync + Self-Heal
&lt;/h3&gt;

&lt;p&gt;All 41 live Applications have &lt;code&gt;syncPolicy.automated.selfHeal: true&lt;/code&gt;. A merged manifest change deploys itself. Manual &lt;code&gt;kubectl&lt;/code&gt; drift on anything ArgoCD tracks gets reverted automatically within minutes.&lt;/p&gt;

&lt;p&gt;This is the primary self-healing mechanism. If someone manually patches a Deployment (during debugging, for example), ArgoCD reverts it on the next sync cycle. The cluster always converges to the git state.&lt;/p&gt;

&lt;h3&gt;
  
  
  Docker Containers: restart: unless-stopped
&lt;/h3&gt;

&lt;p&gt;Media stack, Minecraft, AdGuard, Unbound — all run with &lt;code&gt;restart: unless-stopped&lt;/code&gt; in Docker Compose. Crashes and host reboots self-recover the process. The data recovery is handled by NFS/ZFS, not Docker.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kyverno: Audit-Only (Deliberately)
&lt;/h3&gt;

&lt;p&gt;Kyverno runs in Audit mode — it reports policy violations in PolicyReports but does not block or auto-remediate. This is deliberate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A bad policy in Enforce mode silently blocks legitimate deploys&lt;/li&gt;
&lt;li&gt;On a single-operator homelab, there's nobody to notice the block quickly&lt;/li&gt;
&lt;li&gt;Audit mode provides the visibility without the blast radius&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The policies exist (&lt;code&gt;require-resource-limits&lt;/code&gt;, &lt;code&gt;disallow-latest-tag&lt;/code&gt;, &lt;code&gt;disallow-privileged-containers&lt;/code&gt;) but they're watching, not enforcing. When I'm confident they won't cause false positives, they'll flip to Enforce.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Stays Manual (On Purpose)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Terraform Apply
&lt;/h3&gt;

&lt;p&gt;Always through Atlantis. A human comments &lt;code&gt;atlantis apply&lt;/code&gt;. Never auto-applied, regardless of what changed. This is the one rule with no exceptions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stateful Service Bumps
&lt;/h3&gt;

&lt;p&gt;Authelia, Vault, CNPG, Garage — reviewed before merge. The failure mode is data loss or auth outage, not "reroll the pod."&lt;/p&gt;

&lt;h3&gt;
  
  
  Snapshots Before State-Affecting Changes
&lt;/h3&gt;

&lt;p&gt;Proxmox VM/CT snapshot or Velero PV snapshot, taken manually before applying anything that touches running state. The judgment of "is this change state-affecting" doesn't belong to a script.&lt;/p&gt;

&lt;h3&gt;
  
  
  Velero Restore Testing
&lt;/h3&gt;

&lt;p&gt;Running, tested successfully at least once per namespace. But not yet exercised across every service. A backup that's only partly restore-tested is closer to hope than guarantee.&lt;/p&gt;

&lt;h3&gt;
  
  
  Full Cluster Rebuild
&lt;/h3&gt;

&lt;p&gt;Documented in &lt;code&gt;DISASTER-RECOVERY.md&lt;/code&gt;. Not automated. The target is fast, well-documented recovery — not zero-touch failover, because true HA isn't achievable on one physical host.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Non-Goals
&lt;/h2&gt;

&lt;p&gt;These are explicitly &lt;strong&gt;not&lt;/strong&gt; targets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero-downtime HA.&lt;/strong&gt; Not achievable with one physical host. Not attempted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fully unattended major-version upgrades.&lt;/strong&gt; Stateful service upgrades have repeatedly needed human judgment mid-migration (Postgres mount-point gotcha, capability-drop regression). Automating past that trades a 10-minute manual step for a much longer unattended-failure cleanup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alerting on everything.&lt;/strong&gt; Deliberately tuned to steady-state-relevant signals only.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Contract
&lt;/h2&gt;

&lt;p&gt;This document is the contract between the infrastructure and the operator. Everything not listed here is assumed to be working. If it's not working and it's not in this document, it's a gap — not an expected manual task.&lt;/p&gt;

&lt;p&gt;The operating model is a living document. As the cluster evolves (Cilium CNI, R2 offsite backup, Kyverno enforcement), the categories shift. But the principle stays: &lt;strong&gt;automate what's safe, alert what's important, and leave everything else to a human with context.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;This operating model maps directly to enterprise SRE practices: SLO-based alerting replaces threshold alerting, runbooks define the human response for each alert type, and change management gates (like Atlantis's &lt;code&gt;apply&lt;/code&gt; requirement) prevent unreviewed infrastructure changes. The difference is that enterprise environments have teams; a homelab has one person who needs to sleep through the night without Discord notifications.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/go/4f4RHQf"&gt;The Phoenix Project*&lt;/a&gt; is the book that most shaped this document - it's fiction, but the underlying argument (know exactly what's automated, what's gated, and why) is the same one this operating model tries to make explicit instead of leaving implicit.&lt;/p&gt;



</description>
      <category>kubernetes</category>
      <category>homelab</category>
      <category>gitops</category>
      <category>operations</category>
    </item>
    <item>
      <title>CPU Scheduling for etcd: Why Proxmox cpu.units Matters</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 21 Sep 2026 12:35:09 +0000</pubDate>
      <link>https://dev.to/dwoitzik/cpu-scheduling-for-etcd-why-proxmox-cpuunits-matters-5fe9</link>
      <guid>https://dev.to/dwoitzik/cpu-scheduling-for-etcd-why-proxmox-cpuunits-matters-5fe9</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/cpu-scheduling-etcd-proxmox-units/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Update:&lt;/strong&gt; this cluster has since moved off etcd entirely — a single-server k3s setup uses the&lt;br&gt;
embedded SQLite datastore by default, and the multi-server HA path that would have used etcd&lt;br&gt;
was never turned on here. The CPU-contention problem and the &lt;code&gt;cpu.units&lt;/code&gt; fix below were real&lt;br&gt;
while etcd was in play, and the underlying lesson (latency-sensitive, write-heavy workloads need&lt;br&gt;
scheduling priority when they share a host with bursty ones like LLM inference or game servers)&lt;br&gt;
still holds for anyone running real multi-server etcd today. Left as a historical/general&lt;br&gt;
reference rather than a description of this cluster's current architecture.&lt;/p&gt;

&lt;p&gt;etcd is the brain of Kubernetes. Every API call, every configmap update, every pod scheduling decision goes through etcd. When etcd is slow, everything is slow. When etcd times out, the API server becomes unreachable.&lt;/p&gt;

&lt;p&gt;On a single Proxmox host running k3s VMs alongside Ollama LLM inference, Minecraft game servers, and Docker media workloads, etcd shares CPU with everything else. Without scheduling priority, a Minecraft player join and an Ollama model load can delay etcd's fdatasync calls enough to trigger leader-election timeouts.&lt;/p&gt;

&lt;p&gt;The fix: &lt;code&gt;cpu.units = 2048&lt;/code&gt; on k3s VMs, giving them 2x the CPU scheduling priority over every other workload on the host — the same VMs that get &lt;a href="https://dev.to/blog/staggered-vm-boot-load-average-147/"&gt;staggered boot ordering&lt;/a&gt; to prevent I/O storms in the first place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  How Proxmox CPU Scheduling Works
&lt;/h2&gt;

&lt;p&gt;Proxmox uses the Linux CFS (Completely Fair Scheduler) with added weight controls. Each VM gets a &lt;code&gt;cpu.units&lt;/code&gt; value that determines its share of CPU time when multiple VMs compete for the same physical cores.&lt;/p&gt;

&lt;p&gt;The default is &lt;code&gt;units = 1024&lt;/code&gt;. When two VMs with equal units compete for a single core, they each get 50% of the CPU time. When one VM has &lt;code&gt;units = 2048&lt;/code&gt; and another has &lt;code&gt;units = 1024&lt;/code&gt;, the first gets 2/3 and the second gets 1/3.&lt;/p&gt;

&lt;p&gt;This is not CPU pinning (which restricts a VM to specific cores). It's CPU weighting (which determines priority when cores are shared). On a host with 16 threads and 12+ VMs/LXCs, most cores are shared between multiple workloads.&lt;/p&gt;
&lt;h2&gt;
  
  
  The etcd Problem
&lt;/h2&gt;

&lt;p&gt;etcd's performance depends on write latency. Every key-value operation (lease renewal, configmap update, secret sync) requires an fdatasync to the WAL (Write-Ahead Log) on disk. etcd's leader-election timeout is 5 seconds — if the leader can't renew its lease within that window, controller-runtime terminates the process.&lt;/p&gt;

&lt;p&gt;On a single NVMe shared by all VMs, etcd's fdatasync competes with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ollama reading model weights from disk (14GB qwen2.5 model)&lt;/li&gt;
&lt;li&gt;Minecraft world writes from the DMZ game server&lt;/li&gt;
&lt;li&gt;NFS serving PVC data to k3s pods&lt;/li&gt;
&lt;li&gt;ZFS txg commits flushing dirty data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Under normal load, this is fine — NVMe IOPS are high enough to handle all of them. But when Ollama starts loading a 26B model (Gemma 4) and Minecraft generates terrain simultaneously, the NVMe queue depth spikes, fdatasync latency increases, and etcd's lease renewal window shrinks.&lt;/p&gt;

&lt;p&gt;Without CPU priority, etcd's fdatasync also competes for CPU time with these workloads. The kernel's CFS scheduler doesn't know that etcd's fdatasync is more important than Ollama's matrix multiplication — it just sees two processes requesting CPU time and allocates it equally.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Fix
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# terraform/stacks/proxmox/vm.tf&lt;/span&gt;

&lt;span class="c1"&gt;# k3s VMs — 2x scheduling priority&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"vm_srv_k3s_11"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;cpu&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;cores&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;
    &lt;span class="nx"&gt;units&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2048&lt;/span&gt;  &lt;span class="c1"&gt;# 2x default&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Ollama LXC — default priority&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"ct_srv_ai_01"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;cpu&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;cores&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;
    &lt;span class="nx"&gt;units&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;  &lt;span class="c1"&gt;# default&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Minecraft LXC — default priority&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"ct_dmz_games_01"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;cpu&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;cores&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="nx"&gt;units&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;  &lt;span class="c1"&gt;# default&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The k3s VMs get &lt;code&gt;units = 2048&lt;/code&gt;. Everything else stays at the default &lt;code&gt;1024&lt;/code&gt;. When etcd and Ollama compete for the same CPU cycle, etcd gets 2/3 of the time and Ollama gets 1/3.&lt;/p&gt;

&lt;p&gt;This doesn't reduce Ollama's throughput under normal conditions — when there's no contention, Ollama still gets 100% of the CPU it requests. The priority only kicks in when multiple workloads compete for the same cores simultaneously.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Evidence
&lt;/h2&gt;

&lt;p&gt;Before &lt;code&gt;cpu.units = 2048&lt;/code&gt;, the CNPG operator was restarting due to leader-election timeouts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get pods &lt;span class="nt"&gt;-n&lt;/span&gt; cnpg-system
&lt;span class="c"&gt;# NAME                                    RESTARTS&lt;/span&gt;
&lt;span class="c"&gt;# cnpg-cloudnative-pg-5f8b9c4d6-xk2p4   310&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;310 restarts in 21 days. Each restart's log showed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Leader election retry deadline exceeded
context deadline exceeded
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After applying &lt;code&gt;cpu.units = 2048&lt;/code&gt; and the corresponding leader-election timeout increase (&lt;code&gt;--leader-renew-deadline=50&lt;/code&gt;), restarts dropped to zero.&lt;/p&gt;

&lt;p&gt;The CPU priority alone didn't fix it — the leader-election timeout increase was also necessary. But the CPU priority reduced the frequency of fdatasync delays enough that the 50-second deadline is never challenged under normal load.&lt;/p&gt;

&lt;h2&gt;
  
  
  The RAM Priority Interaction
&lt;/h2&gt;

&lt;p&gt;CPU priority and memory priority are separate in Proxmox. &lt;code&gt;cpu.units&lt;/code&gt; affects CPU scheduling; &lt;code&gt;memory.dedicated&lt;/code&gt; and &lt;code&gt;memory.floating&lt;/code&gt; affect RAM allocation.&lt;/p&gt;

&lt;p&gt;On this host, both matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;etcd needs low-latency CPU for fdatasync (solved by &lt;code&gt;cpu.units&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;k3s control-plane needs guaranteed RAM for API server and scheduler (solved by &lt;code&gt;memory.dedicated = 12284&lt;/code&gt;, the same VM-vs-LXC memory model covered in &lt;a href="https://dev.to/blog/overcommit-guard-python-host-freeze/"&gt;the overcommit guard writeup&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The RAM allocation is a hard reservation — 12 GB is always available for the k3s-11 VM. The CPU priority is a soft weighting — it only matters when cores are shared. Both are necessary: without RAM priority, the k3s VM could be ballooned down to 4 GB under pressure; without CPU priority, etcd could lose the fdatasync race to Ollama.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd Change
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pin etcd to specific cores.&lt;/strong&gt; CPU pinning would guarantee etcd always has CPU available, rather than just having priority. But pinning reduces overall CPU utilization — a pinned core can't be used by other VMs even when etcd is idle. On a 16-thread host with 12+ workloads, the utilization loss isn't worth it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Monitor etcd fdatasync latency directly.&lt;/strong&gt; Prometheus can scrape etcd's &lt;code&gt;etcd_disk_wal_fsync_duration_seconds&lt;/code&gt; metric. An alert on p99 &amp;gt; 100ms would catch CPU contention before it triggers leader-election timeouts.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;CPU scheduling priority is the same concept as Azure VM series selection: E-series VMs are memory-optimized, F-series are compute-optimized, and Dv5-series offer balanced resources. Choosing the wrong series for etcd (a latency-sensitive, write-heavy workload) produces the same performance degradation as running it without &lt;code&gt;cpu.units&lt;/code&gt; on Proxmox. The difference is that Azure makes the choice at VM creation time, while Proxmox lets you adjust it dynamically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/go/4aUtkCs"&gt;Designing Data-Intensive Applications*&lt;/a&gt; is the best resource I know for understanding why a consensus system like etcd is so much more latency-sensitive than an ordinary stateless workload in the first place.&lt;/p&gt;



</description>
      <category>proxmox</category>
      <category>kubernetes</category>
      <category>performance</category>
      <category>homelab</category>
    </item>
    <item>
      <title>3-2-1 Backup in Practice: Velero + PBS + Offsite That Doesn't Work Yet</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 21 Sep 2026 12:24:26 +0000</pubDate>
      <link>https://dev.to/dwoitzik/3-2-1-backup-in-practice-velero-pbs-offsite-that-doesnt-work-yet-52nc</link>
      <guid>https://dev.to/dwoitzik/3-2-1-backup-in-practice-velero-pbs-offsite-that-doesnt-work-yet-52nc</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/321-backup-velero-pbs-offsite/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The 3-2-1 backup rule is simple: 3 copies of your data, on 2 different media types, with 1 copy offsite. My homelab implements all three layers. In theory. In practice, each layer has its own failure mode that I discovered only when I needed it — starting with &lt;a href="https://dev.to/blog/velero-garage-k3s-backup/"&gt;the Garage S3 target Velero backs into&lt;/a&gt;, which is where layer 1 lives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three Layers
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Layer 1: Velero → Garage S3 (in-cluster, daily 05:00 UTC)
    ↓
Layer 2: PBS → External HDD (local, daily 03:00 UTC)
    ↓
Layer 3: rclone → Google Drive (offsite, DISABLED)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Layer 1: Velero → Garage S3
&lt;/h3&gt;

&lt;p&gt;Velero backs up k3s namespaces (apps, vault, database, argocd) to Garage S3 — a self-hosted S3-compatible object store running inside the same cluster.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kubernetes/system/velero/schedule.yml&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;defaultVolumesToFsBackup&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;includedNamespaces&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apps"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vault"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;database"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;argocd"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;excludedResources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;events"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;events.events.k8s.io"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;ttl&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;720h&lt;/span&gt;  &lt;span class="c1"&gt;# 30 days&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;defaultVolumesToFsBackup: true&lt;/code&gt; was the critical fix. Without it, Velero only captured Kubernetes manifests — not the actual PVC data. The backups "completed" for weeks with zero data.&lt;/p&gt;

&lt;p&gt;After the fix, Kopia sidecars run alongside each pod, copying PVC contents to the Garage S3 &lt;code&gt;velero&lt;/code&gt; bucket. Daily backups, 30-day retention.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The catch:&lt;/strong&gt; Garage runs inside the cluster. If the cluster dies, both Velero and Garage are gone. Layer 1 is useful for recovering individual PVCs or namespaces within a running cluster — not for full cluster recovery.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 2: PBS → External HDD
&lt;/h3&gt;

&lt;p&gt;Proxmox Backup Server (PBS) backs up VM and LXC disk images to a &lt;a href="https://dev.to/go/4vxwLpY"&gt;Seagate 2TB External HDD*&lt;/a&gt; connected via USB.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# PBS backup job — runs daily at 03:00 UTC&lt;/span&gt;
&lt;span class="c"&gt;# Backs up all VMs and LXCs to /mnt/backup (USB HDD)&lt;/span&gt;
proxmox-backup-client backup &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--repository&lt;/span&gt; &lt;span class="nb"&gt;local&lt;/span&gt;:/mnt/backup &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ns&lt;/span&gt; homelab &lt;span class="se"&gt;\&lt;/span&gt;
  vm/110/pct/200/pct/210/...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PBS handles deduplication, compression, and incremental backups. The USB HDD provides local, offline backup that survives cluster failures.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The gotcha:&lt;/strong&gt; PBS and the k3s VMs share the same physical host. A host-level failure (PSU, NVMe death) takes out both the primary data and the PBS backup. Layer 2 is protection against software failure (corruption, accidental deletion), not hardware failure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 3: rclone → Google Drive (DISABLED)
&lt;/h3&gt;

&lt;p&gt;The offsite layer was supposed to be rclone syncing from Garage S3 to Google Drive. In practice, Google's API throttled the sync to 1.6 KiB/s:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;Transferred:   1.6 KiB / 4.2 GiB,  0.00%
Elapsed time:  2h 30m
Transfer rate: 1.6 KiB/s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Google Drive's API rate limiting for server-to-server transfers (no user interaction) is aggressive. For 4 GB of data, the sync would take days. Combined with Google's 750 GB/day upload limit for personal accounts, and the fact that the sync would need to run continuously to keep up with daily backups, the offsite layer was disabled.&lt;/p&gt;

&lt;p&gt;The replacement (Cloudflare R2) is scaffolded but not active:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kubernetes/system/velero/offsite-schedule.yml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;velero.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Schedule&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;daily-offsite&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;4&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;includedNamespaces&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apps"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vault"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;database"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;argocd"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;storageLocation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;r2-offsite&lt;/span&gt;
    &lt;span class="na"&gt;ttl&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;168h&lt;/span&gt;  &lt;span class="c1"&gt;# 7 days&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Waiting on a Cloudflare R2 account with real API credentials. R2 has no egress fees and no API rate limiting for the volume this backup produces (~4 GB/day). Activating it also fixes a second problem beyond throttling: Garage runs inside the same cluster Velero is protecting, so an in-cluster-only backup target is circular regardless of how fast it uploads.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Restore Gotcha
&lt;/h2&gt;

&lt;p&gt;The first time I tested a full restore, PBS hit an IP/MAC conflict:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: VM 211 is running on a different node (10.0.20.11)
Proxmox cannot start the VM because the MAC address is already in use
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The issue: PBS restores the VM's network configuration exactly as it was, including the MAC address. If the original VM is still running (or its MAC is cached in the bridge), the restored VM can't start because Proxmox detects a MAC conflict on the virtual bridge.&lt;/p&gt;

&lt;p&gt;The fix: restore to a temporary VMID, verify the restore is complete, then shut down the original and rename the restored VM. This is a manual multi-step process — not something you want to do under time pressure during a real disaster.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Works
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Protects Against&lt;/th&gt;
&lt;th&gt;Doesn't Protect Against&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Velero → Garage&lt;/td&gt;
&lt;td&gt;PVC corruption, namespace deletion, accidental kubectl delete&lt;/td&gt;
&lt;td&gt;Cluster-wide failure (Garage is in-cluster)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PBS → USB HDD&lt;/td&gt;
&lt;td&gt;Software corruption, accidental VM deletion, ZFS pool issues&lt;/td&gt;
&lt;td&gt;Host hardware failure (PBS is on the same host)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rclone → Google Drive&lt;/td&gt;
&lt;td&gt;Host hardware failure, theft, fire&lt;/td&gt;
&lt;td&gt;Nothing yet — disabled due to throttling&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The honest assessment: I have two functional backup layers, both on the same physical host. True offsite backup doesn't exist yet. The 3-2-1 rule is aspirational, not achieved.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Verification Gap
&lt;/h2&gt;

&lt;p&gt;The most important lesson: &lt;strong&gt;a backup you haven't restored is not a backup.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Velero backups run daily, PBS runs daily, and until recently neither had been fully restored. The Velero &lt;code&gt;defaultVolumesToFsBackup&lt;/code&gt; gap existed for weeks because nobody tested a restore. The PBS IP conflict was discovered during the first full restore test.&lt;/p&gt;

&lt;p&gt;After these discoveries, I added:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Monthly Velero restore test&lt;/strong&gt; — restore the &lt;code&gt;database&lt;/code&gt; namespace to a temporary namespace, verify Postgres starts and contains expected data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quarterly PBS restore test&lt;/strong&gt; — restore one VM to a temporary VMID, verify it boots and services are functional&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Post-backup verification&lt;/strong&gt; — &lt;code&gt;velero backup describe --details | grep "Pod Volume Backups"&lt;/code&gt; after every scheduled backup&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;Backup verification is the same compliance requirement in Azure: ISO 27001 and NIS2 both require documented, tested restore procedures. Azure Backup reports "Completed" for VM snapshots, but the snapshot might not include the data disk if the backup policy was misconfigured. The only way to verify is to actually restore a test VM and confirm its contents — the same monthly drill I now run against Velero and PBS.&lt;/p&gt;



</description>
      <category>kubernetes</category>
      <category>backup</category>
      <category>homelab</category>
      <category>disasterrecovery</category>
    </item>
    <item>
      <title>NFS vs local-path: The SQLite Trap That Corrupted My S3 Metadata</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 14 Sep 2026 12:43:37 +0000</pubDate>
      <link>https://dev.to/dwoitzik/nfs-vs-local-path-the-sqlite-trap-that-corrupted-my-s3-metadata-h26</link>
      <guid>https://dev.to/dwoitzik/nfs-vs-local-path-the-sqlite-trap-that-corrupted-my-s3-metadata-h26</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/nfs-vs-local-path-sqlite-trap/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/blog/velero-garage-k3s-backup/"&gt;Garage S3&lt;/a&gt;, my self-hosted S3-compatible object store, uses SQLite for its metadata database. One morning, Terraform state operations started failing with &lt;code&gt;SQLITE_CORRUPT&lt;/code&gt;. The bucket metadata was gone. The &lt;code&gt;terraform-state&lt;/code&gt; bucket, the Atlantis lock table, the Velero backup index — all of it stored in a SQLite file that was now corrupted.&lt;/p&gt;

&lt;p&gt;The root cause: Garage was running on an NFS-backed PersistentVolume. NFS doesn't support the file-locking primitives that SQLite's WAL (Write-Ahead Logging) mode requires. Under concurrent access — Garage's metadata writer and a Velero backup reading the same database — the NFS lock delegation failed silently, and SQLite wrote to overlapping pages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Storage Classes
&lt;/h2&gt;

&lt;p&gt;My k3s cluster has two StorageClasses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# NFS — for most workloads&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;storage.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;StorageClass&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nfs-client&lt;/span&gt;
&lt;span class="na"&gt;provisioner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nfs-subdir-external-provisioner&lt;/span&gt;
&lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10.0.20.100&lt;/span&gt;
  &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/archive&lt;/span&gt;
&lt;span class="na"&gt;reclaimPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Retain&lt;/span&gt;

&lt;span class="c1"&gt;# Local-path — for embedded databases&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;storage.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;StorageClass&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;local-path&lt;/span&gt;
&lt;span class="na"&gt;provisioner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;rancher.io/local-path&lt;/span&gt;
&lt;span class="na"&gt;reclaimPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Retain&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NFS (&lt;code&gt;nfs-client&lt;/code&gt;) is the default. It's backed by a dedicated LXC running an NFS server on ZFS, providing storage that survives pod rescheduling — a pod on k3s-12 can access the same PVC as a pod on k3s-13 because the NFS server is independent of any specific node.&lt;/p&gt;

&lt;p&gt;Local-path (&lt;code&gt;local-path&lt;/code&gt;) pins the PV to whichever node created it. If the pod reschedules to a different node, the PVC is inaccessible until the pod returns to the original node. This is a limitation, but it's the right trade-off for certain workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why SQLite and NFS Don't Work
&lt;/h2&gt;

&lt;p&gt;SQLite's WAL mode requires &lt;code&gt;fcntl()&lt;/code&gt; file locks — specifically, &lt;code&gt;F_SETLK&lt;/code&gt; (non-blocking lock) and &lt;code&gt;F_SETLKW&lt;/code&gt; (blocking lock). These locks coordinate access between concurrent processes writing to the same database file.&lt;/p&gt;

&lt;p&gt;NFS handles file locks differently:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;NFSv3&lt;/strong&gt;: No native lock support. &lt;code&gt;fcntl()&lt;/code&gt; calls return success but locks are local to the client — two NFS clients can both acquire an "exclusive" lock on the same file simultaneously.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;NFSv4&lt;/strong&gt;: Has &lt;code&gt;LOCK&lt;/code&gt; operations, but the lock delegation model introduces latency and failure modes that SQLite's tight locking loop doesn't tolerate. If the NFS server is slow to respond to a lock request, SQLite's default 5-second busy timeout can expire, causing the application to retry — and the retry can conflict with the lock held by another client.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Kubernetes NFS provisioner&lt;/strong&gt;: The &lt;code&gt;nfs-subdir-external-provisioner&lt;/code&gt; uses NFSv4, but the lock delegation is handled by the NFS server's &lt;code&gt;rpc.lockd&lt;/code&gt; daemon, which runs in a separate process space. Under concurrent load, &lt;code&gt;lockd&lt;/code&gt; can lose track of which client holds which lock.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The result: SQLite thinks it has an exclusive lock, but another process (or the same process on a different connection) also has a lock. Both write to the database file. Pages overlap. The database corrupts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Garage Incident
&lt;/h2&gt;

&lt;p&gt;Garage runs with two storage mounts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kubernetes/apps/garage/garage.yml&lt;/span&gt;
&lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;data&lt;/span&gt;
    &lt;span class="na"&gt;persistentVolumeClaim&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;claimName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;garage-data&lt;/span&gt;      &lt;span class="c1"&gt;# NFS — bucket objects&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;meta&lt;/span&gt;
    &lt;span class="na"&gt;persistentVolumeClaim&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;claimName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;garage-meta&lt;/span&gt;      &lt;span class="c1"&gt;# local-path — SQLite metadata&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;data&lt;/code&gt; volume (bucket objects) is on NFS — fine, because S3 object storage doesn't use file locks for individual files. The &lt;code&gt;meta&lt;/code&gt; volume (SQLite database) was &lt;em&gt;also&lt;/em&gt; on NFS initially. This worked until a Velero backup and a Terraform state write happened simultaneously.&lt;/p&gt;

&lt;p&gt;Velero reads Garage's S3 API to enumerate backup objects. Terraform reads the same database to verify state file existence. Both hit SQLite through Garage's metadata layer. Under NFS, the concurrent reads triggered the lock delegation failure, and SQLite corrupted pages 169–184 of &lt;code&gt;db.sqlite&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Recovery
&lt;/h2&gt;

&lt;p&gt;SQLite has a &lt;code&gt;.recover&lt;/code&gt; command that can extract data from a corrupted database:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Dump recoverable data from corrupted SQLite&lt;/span&gt;
sqlite3 db.sqlite &lt;span class="s2"&gt;".recover"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; recovered.sql

&lt;span class="c"&gt;# Recreate the database from the dump&lt;/span&gt;
sqlite3 db_clean.sqlite &amp;lt; recovered.sql
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;.recover&lt;/code&gt; command scanned every page of the corrupted file and extracted whatever data it could read. For pages 169–184 (the corrupted range), it found partial data — enough to reconstruct the bucket and key metadata, but not enough to guarantee referential integrity.&lt;/p&gt;

&lt;p&gt;After recovery, the missing objects (terraform-state bucket and Atlantis lock key) had to be re-inserted manually via Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;msgpack&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;db_clean.sqlite&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# Garage uses msgpack-encoded metadata with 'G2key'/'G2bkt' prefixes
# Re-insert the terraform-state bucket
&lt;/span&gt;&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;INSERT INTO buckets (name, ...) VALUES (...)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The data was restored, but the trust was gone. The database could corrupt again under the same conditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix
&lt;/h2&gt;

&lt;p&gt;Move every embedded database to &lt;code&gt;local-path&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Garage — meta volume moved to local-path&lt;/span&gt;
&lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;meta&lt;/span&gt;
    &lt;span class="na"&gt;persistentVolumeClaim&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;claimName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;garage-meta&lt;/span&gt;      &lt;span class="c1"&gt;# NOW: local-path (was: nfs-client)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The apps that need &lt;code&gt;local-path&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;App&lt;/th&gt;
&lt;th&gt;Database&lt;/th&gt;
&lt;th&gt;Why local-path&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Garage S3&lt;/td&gt;
&lt;td&gt;SQLite (metadata)&lt;/td&gt;
&lt;td&gt;File-locking requirements&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mealie&lt;/td&gt;
&lt;td&gt;SQLite (recipes)&lt;/td&gt;
&lt;td&gt;WAL mode + concurrent access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home Assistant&lt;/td&gt;
&lt;td&gt;SQLite (state)&lt;/td&gt;
&lt;td&gt;Inotify-based DB writes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authelia&lt;/td&gt;
&lt;td&gt;PostgreSQL (CNPG)&lt;/td&gt;
&lt;td&gt;CNPG manages its own PV&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;PostgreSQL (via CNPG) doesn't have the SQLite lock problem because it uses its own file locking, but CNPG requires &lt;code&gt;local-path&lt;/code&gt; or a CSI driver that supports &lt;code&gt;ReadWriteOnce&lt;/code&gt; — NFS's &lt;code&gt;ReadWriteMany&lt;/code&gt; semantics can confuse CNPG's WAL archiving.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-off
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;local-path&lt;/code&gt; means the PVC is pinned to one node. If the pod reschedules, it loses access to the data. For Garage, this is acceptable: Garage runs on a single node, and if that node goes down, the S3 data is unavailable regardless (it's on the same host).&lt;/p&gt;

&lt;p&gt;For databases that need HA (Postgres, Redis), the solution isn't &lt;code&gt;local-path&lt;/code&gt; or NFS — it's a managed operator (CNPG for Postgres) that handles replication and failover independently of the storage layer.&lt;/p&gt;

&lt;p&gt;The principle: &lt;strong&gt;if the application uses file-level locking (SQLite, BoltDB, LMDB), it goes on local-path. If it uses network-level locking (PostgreSQL, MySQL), it goes on NFS or a managed operator.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;SQLite on NFS is the same failure mode as running SQLite on an SMB share in a Windows domain: the file-locking semantics are fundamentally incompatible. In Azure, this maps to Azure Files (SMB-backed) vs. Azure Disk (block storage). Azure Files supports SMB locks but has the same delegation latency issues under concurrent access — any application that needs tight file-level locking should use Azure Disks, not Azure Files. The principle is identical: embedded databases need local, low-latency storage with native file-locking support.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/go/4aUtkCs"&gt;Designing Data-Intensive Applications*&lt;/a&gt; covers exactly this class of correctness assumption — what a storage layer actually guarantees about concurrent access versus what an application silently assumes it guarantees — in far more depth than a corrupted &lt;code&gt;db.sqlite&lt;/code&gt; file teaches you in the moment.&lt;/p&gt;



</description>
      <category>kubernetes</category>
      <category>storage</category>
      <category>homelab</category>
      <category>debugging</category>
    </item>
    <item>
      <title>My Media Stack Lives in Two Containers and a Python CronJob</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 14 Sep 2026 12:37:38 +0000</pubDate>
      <link>https://dev.to/dwoitzik/my-media-stack-lives-in-two-containers-and-a-python-cronjob-j2k</link>
      <guid>https://dev.to/dwoitzik/my-media-stack-lives-in-two-containers-and-a-python-cronjob-j2k</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/media-stack-two-containers-python-cronjob/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;My media acquisition stack — SABnzbd for Usenet downloads, Sonarr for TV, Radarr for movies, Bazarr for subtitles, NZBHydra2 for indexer search — used to run as k3s Deployments. It worked, but three problems made it a bad fit for Kubernetes: GPU passthrough for Jellyfin transcoding, per-flow traffic isolation for indexer queries versus the actual download path, and the NFS file-locking trap for media libraries.&lt;/p&gt;

&lt;p&gt;The solution: move the media stack out of k3s entirely. Jellyfin runs in its own GPU-passthrough LXC. The acquisition stack runs in a second LXC with Docker Compose. A Python CronJob in k3s bridges the two via Traefik Service+Endpoints.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Update:&lt;/strong&gt; the first version of this stack wrapped the whole acquisition LXC in a Mullvad WireGuard tunnel via &lt;a href="https://github.com/qdm12/gluetun" rel="noopener noreferrer"&gt;gluetun&lt;/a&gt;, on the assumption that every flow out of that box needed VPN protection equally. Two days later I tore that back out — see "The Traffic-Isolation Rethink" below for why a single blanket tunnel was the wrong model for what these five apps actually do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Kubernetes Was Wrong for Media
&lt;/h2&gt;

&lt;h3&gt;
  
  
  GPU Passthrough
&lt;/h3&gt;

&lt;p&gt;Jellyfin needs GPU access for hardware video transcoding. The &lt;a href="https://dev.to/go/4bv3yF1"&gt;BMAX Mini PC*&lt;/a&gt; has an AMD Radeon Vega iGPU that supports VAAPI hardware transcoding. Proxmox GPU passthrough requires IOMMU group isolation — the GPU is passed to a single container or VM exclusively.&lt;/p&gt;

&lt;p&gt;Kubernetes doesn't natively support GPU passthrough for LXCs. The &lt;code&gt;nvidia-device-plugin&lt;/code&gt; works for NVIDIA GPUs on specific cloud providers, but for AMD iGPU passthrough on bare-metal Proxmox, you need a dedicated LXC with &lt;code&gt;/dev/dri/renderD128&lt;/code&gt; mapped directly.&lt;/p&gt;

&lt;p&gt;Jellyfin runs in &lt;code&gt;ct-srv-jellyfin-01&lt;/code&gt; with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# terraform/stacks/proxmox/lxc.tf&lt;/span&gt;
&lt;span class="nx"&gt;lxc_conf&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;desc&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"Jellyfin - GPU passthrough"&lt;/span&gt;
  &lt;span class="c1"&gt;# GPU device mapped via pct set&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;k3s's own nodes here (&lt;code&gt;vm-srv-k3s-11/12/13&lt;/code&gt;) are Proxmox &lt;strong&gt;VMs&lt;/strong&gt;, not LXCs — a VM doesn't share the host kernel, so there's no "just bind-mount the device node" path the way there is for an LXC. Getting a k3s pod real GPU access on this hardware would mean classic VFIO passthrough of the iGPU to one specific VM: unbinding &lt;code&gt;amdgpu&lt;/code&gt; from the Proxmox host and handing the whole device to &lt;code&gt;vfio-pci&lt;/code&gt; instead. I looked at this seriously before ruling it out, because "GPU-in-Kubernetes" device plugins exist and I wanted to know if they'd apply.&lt;/p&gt;

&lt;p&gt;They don't, for this hardware. AMD's own &lt;code&gt;rocm/k8s-device-plugin&lt;/code&gt; targets ROCm compute (HIP/OpenCL) — a materially heavier stack than what VAAPI hardware transcoding actually needs, which is just &lt;code&gt;/dev/dri&lt;/code&gt; visibility. And a Kubernetes device plugin can't manufacture GPU access a node's kernel doesn't already have — for a VM, that access only exists after real hypervisor-level VFIO passthrough, which a plugin doesn't do. Worse, this is a single consumer Ryzen APU, not a data-center part with SR-IOV or mediated-device support for splitting one GPU across VMs — passthrough would hand the &lt;em&gt;entire&lt;/em&gt; GPU to exactly one of the three k3s VMs, and the Proxmox host itself (which uses &lt;code&gt;amdgpu&lt;/code&gt; for its own display/telemetry) would permanently lose access to it. There's also no portability payoff to offset that cost: k3s's scheduler can't move a pod needing a passed-through device to a &lt;em&gt;different&lt;/em&gt; node than the one VFIO was bound to, so the usual "GPU follows the pod" reason people move transcoding into Kubernetes never materializes on a single-host, single-iGPU homelab. It would just be Jellyfin running on a VM instead of an LXC, at the permanent cost of the GPU being unavailable to anything else on the host.&lt;/p&gt;

&lt;p&gt;The LXC path avoids all of that: an LXC shares the host's kernel, so the host keeps the &lt;code&gt;amdgpu&lt;/code&gt; driver bound and simply grants the container access to the resulting &lt;code&gt;/dev/dri/renderD128&lt;/code&gt; device node. Non-exclusive from the host's perspective, already proven working, no PCI device binding to get wrong. If this box ever gets a GPU with real SR-IOV support, this is worth revisiting — the constraint here is the specific hardware, not a principled objection to GPU workloads in Kubernetes.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Traffic-Isolation Rethink
&lt;/h3&gt;

&lt;p&gt;The acquisition LXC has two genuinely different outbound flows, and my first pass treated them as one problem. SABnzbd connects to Eweka (my Usenet provider) over NNTPS on port 563 — already encrypted end-to-end, and Usenet copyright enforcement works exclusively via BitTorrent peer-list monitoring, so there's no mechanism by which an ISP or rights-holder observes or reports Usenet downloads in the first place. NZBHydra2's indexer search queries are a completely different flow: plain HTTP/HTTPS lookups against third-party indexer sites, which &lt;em&gt;does&lt;/em&gt; expose the home IP to whoever's on the other end, the same as browsing any site directly.&lt;/p&gt;

&lt;p&gt;Wrapping the whole LXC in gluetun/Mullvad "solved" both at once, but it was the wrong tool for either: VPN on the download path halves throughput for no privacy benefit Eweka's own SSL doesn't already provide, and it risks Eweka flagging the account for apparent multi-subscriber IP sharing. I tore gluetun out two days after standing it up and replaced it with a model that actually matches the two flows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SABnzbd → Eweka, direct, over SSL.&lt;/strong&gt; No VPN. The ISP sees "connected to news.eweka.nl," never content — that's the whole job done by NNTPS alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NZBHydra2 → indexers, through a Tor SOCKS5 proxy&lt;/strong&gt;, &lt;code&gt;tor&lt;/code&gt; container at &lt;code&gt;172.28.1.10:9050&lt;/code&gt;, configured with &lt;strong&gt;no direct-connection fallback&lt;/strong&gt;. Search queries are small and latency-tolerant, a good fit for Tor's limited bandwidth, and Tor is a better fit than a commercial VPN for metadata-only queries — no single operator to trust, no throughput to throttle. If Tor is down, indexer queries fail outright instead of silently leaking the home IP.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I verified this configuration directly rather than trusting the design intent on paper: &lt;code&gt;sabnzbd.ini&lt;/code&gt; shows &lt;code&gt;socks5_proxy_url = ""&lt;/code&gt; (empty — SABnzbd was never routed through Tor) with &lt;code&gt;ssl = 1&lt;/code&gt;, &lt;code&gt;ssl_verify = 2&lt;/code&gt; confirming the direct-to-Eweka SSL path is real; &lt;code&gt;nzbhydra.yml&lt;/code&gt; shows &lt;code&gt;proxyType: SOCKS&lt;/code&gt; pointed at the Tor container with no fallback option enabled. One live exception I noticed and haven't chased down yet: one specific indexer bypasses Tor and connects directly (&lt;code&gt;proxyIgnoreDomains&lt;/code&gt;) — possibly a site that blocks Tor exit nodes, possibly a leftover exception from before I understood this stack properly. Worth revisiting.&lt;/p&gt;

&lt;p&gt;Tor's bandwidth genuinely can't handle bulk transfers, and routing downloads through it would be abusive to a network that exists for people who need anonymity for safety — never route the actual download path through Tor, only small metadata lookups.&lt;/p&gt;

&lt;p&gt;The acquisition LXC (&lt;code&gt;ct-srv-media-acq-01&lt;/code&gt;) runs all five apps this way — no blanket tunnel, no kill-switch sidecar to maintain, no shared failure mode between "is Eweka's SSL up" and "is the VPN provider's WireGuard endpoint reachable today."&lt;/p&gt;

&lt;h3&gt;
  
  
  NFS File Locking
&lt;/h3&gt;

&lt;p&gt;Media libraries (downloaded files, metadata databases) live on NFS. As I wrote about in &lt;a href="https://dev.to/blog/nfs-vs-local-path-sqlite-trap/"&gt;the SQLite trap article&lt;/a&gt;, NFS file-locking semantics don't work with embedded databases. Sonarr and Radarr use SQLite internally for their media databases — and those databases corrupt on NFS under concurrent access.&lt;/p&gt;

&lt;p&gt;Moving the acquisition stack to a local LXC with local storage eliminates the NFS lock problem. The media files themselves (downloaded episodes, movies) still live on NFS for sharing, but the application databases stay local.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────────────────────────────────────┐
│  k3s Cluster                                │
│  ┌─────────────────────────────────────┐    │
│  │ Python CronJob (every 10 min)       │    │
│  │ - Checks Sonarr/Radarr queue        │    │
│  │ - Clears stuck items                │    │
│  │ - Triggers Jellyfin library scan    │    │
│  └──────────┬──────────────────────────┘    │
│             │ HTTP via Traefik              │
│  ┌──────────▼──────────────────────────┐    │
│  │ Traefik IngressRoutes               │    │
│  │ - sabnzbd.woitzik.dev               │    │
│  │ - sonarr.woitzik.dev                │    │
│  │ - radarr.woitzik.dev                │    │
│  │ - bazarr.woitzik.dev                │    │
│  └─────────────────────────────────────┘    │
└──────────────────┬──────────────────────────┘
                   │ Traefik Service+Endpoints
┌──────────────────▼──────────────────────────┐
│  ct-srv-media-acq-01 (LXC, Tor for indexers)│
│  ┌─────────┐ ┌────────┐ ┌────────┐         │
│  │ SABnzbd │ │ Sonarr │ │ Radarr │         │
│  └─────────┘ └────────┘ └────────┘         │
│  ┌─────────┐ ┌──────────────┐              │
│  │ Bazarr  │ │ NZBHydra2    │              │
│  └─────────┘ └──────────────┘              │
└─────────────────────────────────────────────┘
┌─────────────────────────────────────────────┐
│  ct-srv-jellyfin-01 (LXC, GPU passthrough)  │
│  ┌─────────┐ ┌──────────────┐              │
│  │ Jellyfin│ │ /dev/dri/    │              │
│  │         │ │ renderD128   │              │
│  └─────────┘ └──────────────┘              │
└─────────────────────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Traefik Bridge
&lt;/h2&gt;

&lt;p&gt;k3s Traefik routes external traffic to services inside the cluster. But the media stack isn't in the cluster — it's in LXCs. The bridge: Traefik IngressRoutes point at Kubernetes Services, which use &lt;code&gt;Endpoints&lt;/code&gt; objects with hardcoded IP addresses pointing at the LXC containers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kubernetes/apps/jellyfin/jellyfin.yml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Service&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;jellyfin&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8096&lt;/span&gt;
      &lt;span class="na"&gt;targetPort&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8096&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Endpoints&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;jellyfin&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps&lt;/span&gt;
&lt;span class="na"&gt;subsets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;addresses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;ip&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10.0.20.254&lt;/span&gt;  &lt;span class="c1"&gt;# ct-srv-jellyfin-01&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8096&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the same pattern for all external services. The Service has no &lt;code&gt;selector&lt;/code&gt; — it's a "headless" Service where the Endpoints are manually maintained. Traefik doesn't know or care that the backend is an LXC instead of a pod.&lt;/p&gt;

&lt;p&gt;The IngressRoute for Jellyfin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;traefik.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;IngressRoute&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;jellyfin&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;entryPoints&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;websecure&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;routes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Host(`media.woitzik.dev`)&lt;/span&gt;
      &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Rule&lt;/span&gt;
      &lt;span class="na"&gt;middlewares&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[{&lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;authelia&lt;/span&gt;&lt;span class="pi"&gt;}]&lt;/span&gt;
      &lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;jellyfin&lt;/span&gt;
          &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8096&lt;/span&gt;
  &lt;span class="na"&gt;tls&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;secretName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;wildcard-woitzik-dev-tls&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The media stack gets &lt;a href="https://dev.to/blog/k3s-authelia-proxmox-homelab/"&gt;Authelia&lt;/a&gt; protection, wildcard TLS, and &lt;a href="https://dev.to/blog/cloudflare-tunnel-zero-inbound-ports/"&gt;Cloudflare Tunnel&lt;/a&gt; external access — the same as every in-cluster service. The only difference is the backend IP.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Python CronJob
&lt;/h2&gt;

&lt;p&gt;A Python CronJob runs every 10 minutes inside k3s, bridging the gap between the media stack and the cluster:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# CronJob that monitors media acquisition&lt;/span&gt;
&lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*/10&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;
&lt;span class="na"&gt;jobTemplate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;media-watchdog&lt;/span&gt;
            &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;python:3.12-slim&lt;/span&gt;
            &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;python&lt;/span&gt;
              &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/scripts/media-watchdog.py&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The script:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Queries Sonarr/Radarr APIs for stuck queue items (downloads stuck for &amp;gt;30 minutes)&lt;/li&gt;
&lt;li&gt;Clears stuck items and re-triggers the download&lt;/li&gt;
&lt;li&gt;Checks Jellyfin's library scan status&lt;/li&gt;
&lt;li&gt;Posts status to Discord via webhook&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Without this CronJob, stuck downloads sit indefinitely. Sonarr and Radarr don't have built-in queue monitoring — they trust the download client to report status, and SABnzbd sometimes silently fails without notifying the *arr stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd Change
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Use a proper service mesh for cross-boundary traffic.&lt;/strong&gt; The manual Endpoints pattern works but is fragile — if the LXC IP changes, the Endpoints must be updated manually. A DNS-based service discovery (Headscale DNS entries, for example) would be more resilient.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Move the CronJob to a native LXC cron.&lt;/strong&gt; The Python script runs in a k3s Pod but talks to LXC-hosted services. It has no business being in the cluster. A systemd timer on the media acquisition LXC would be simpler.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;The media stack's departure from Kubernetes is the same pattern as running stateful workloads outside AKS: GPU workloads go to dedicated VMs with GPU passthrough, traffic-sensitive workloads get network-level isolation instead of a sidecar, and file-locking workloads go to local SSDs. Kubernetes excels at stateless, horizontally-scalable workloads. Media transcoding and acquisition are neither — they're stateful, single-instance, and hardware-dependent. The right platform for them is the bare metal underneath, not the orchestration layer on top.&lt;/p&gt;



</description>
      <category>homelab</category>
      <category>proxmox</category>
      <category>networking</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>Private AKS on Azure: No Public API Server, No Public Node IPs, No Default Outbound</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 07 Sep 2026 19:17:12 +0000</pubDate>
      <link>https://dev.to/dwoitzik/private-aks-on-azure-no-public-api-server-no-public-node-ips-no-default-outbound-5ga3</link>
      <guid>https://dev.to/dwoitzik/private-aks-on-azure-no-public-api-server-no-public-node-ips-no-default-outbound-5ga3</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/azure-private-aks-zero-trust/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Run &lt;code&gt;az aks create&lt;/code&gt; with the defaults and you get a public API server FQDN, a Standard Load Balancer giving every node its own path to the internet, and — unless you go out of your way to stop it — an admin kubeconfig that works from anywhere with the right token. None of that is a bug. It's the fast path, and for a demo cluster it's fine.&lt;/p&gt;

&lt;p&gt;For a cluster inside a governed Hub &amp;amp; Spoke, it's three separate audit findings. This is the Terraform to close all three at once: a private API server, forced-tunneled egress through a firewall you already control, and private-endpoint-only ACR and Key Vault.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/terraform-azurerm-private-aks" rel="noopener noreferrer"&gt;Get the base template free on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Target Architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                         ┌─────────────────────┐
                         │   Azure Firewall     │  ← you supply the IP
                         │  (existing / NVA)     │
                         └──────────┬───────────┘
                                    │ UDR: 0.0.0.0/0
                         ┌──────────┴───────────┐
                         │   vnet-aks            │
                         │                       │
              ┌──────────┴──────────┐  ┌─────────┴──────────┐
              │  snet-aks-nodes     │  │ snet-private-       │
              │  (default-deny NSG) │  │ endpoints            │
              │                     │  │  - ACR (Premium)     │
              │  Private AKS        │  │  - Key Vault         │
              │  (no public FQDN)   │  │                      │
              └─────────────────────┘  └─────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three engineering decisions drive this, same as the &lt;a href="https://dev.to/blog/azure-terraform-hub-spoke-zero-trust"&gt;Hub &amp;amp; Spoke module&lt;/a&gt; this one is designed to sit behind:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No public control-plane endpoint&lt;/strong&gt; — reachable only from inside the VNet, a peered network, or a VPN&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forced tunneling, not default outbound&lt;/strong&gt; — every node packet leaves through a firewall you control&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Private-endpoint-only dependencies&lt;/strong&gt; — ACR and Key Vault never touch the public internet either&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 1: The API Server Has No Public FQDN
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"azurerm_kubernetes_cluster"&lt;/span&gt; &lt;span class="s2"&gt;"this"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;# ...&lt;/span&gt;
  &lt;span class="nx"&gt;private_cluster_enabled&lt;/span&gt;             &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="nx"&gt;private_dns_zone_id&lt;/span&gt;                 &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;azurerm_private_dns_zone&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;private_cluster_public_fqdn_enabled&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;

  &lt;span class="nx"&gt;network_profile&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;network_plugin&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"azure"&lt;/span&gt;
    &lt;span class="nx"&gt;network_policy&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"azure"&lt;/span&gt;
    &lt;span class="nx"&gt;outbound_type&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"userDefinedRouting"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;private_cluster_public_fqdn_enabled = false&lt;/code&gt; is the setting people miss. &lt;code&gt;private_cluster_enabled = true&lt;/code&gt; alone still publishes a public FQDN that resolves to nothing reachable — a fingerprintable artifact for no benefit. Turning it off means &lt;code&gt;az aks show&lt;/code&gt; doesn't leak a hostname at all.&lt;/p&gt;

&lt;p&gt;The cluster needs its own Private DNS Zone for the API server (&lt;code&gt;privatelink.&amp;lt;region&amp;gt;.azmk8s.io&lt;/code&gt;), linked to the VNet — same DINE-policy-safe &lt;code&gt;lifecycle.ignore_changes&lt;/code&gt; pattern as the &lt;a href="https://dev.to/blog/azure-terraform-hub-spoke-zero-trust"&gt;Hub &amp;amp; Spoke module's&lt;/a&gt; centralized zones, because Azure Policy tagging fights this resource too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: &lt;code&gt;userDefinedRouting&lt;/code&gt; — the Setting That Actually Matters
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;network_plugin = "azure"&lt;/code&gt; and &lt;code&gt;network_policy = "azure"&lt;/code&gt; get most of the attention in AKS networking guides. The setting that actually determines whether traffic is inspectable is &lt;code&gt;outbound_type&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Left at its default (&lt;code&gt;loadBalancer&lt;/code&gt;), AKS provisions its own Standard Load Balancer outbound rule and every node gets a direct, un-inspected path to the internet — a private API server with fully public node egress is a common half-measure that looks locked down in the portal and isn't.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"azurerm_route_table"&lt;/span&gt; &lt;span class="s2"&gt;"aks_egress"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;route&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;name&lt;/span&gt;                   &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"default-via-firewall"&lt;/span&gt;
    &lt;span class="nx"&gt;address_prefix&lt;/span&gt;         &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"0.0.0.0/0"&lt;/span&gt;
    &lt;span class="nx"&gt;next_hop_type&lt;/span&gt;          &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"VirtualAppliance"&lt;/span&gt;
    &lt;span class="nx"&gt;next_hop_in_ip_address&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;var&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;firewall_private_ip&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"azurerm_subnet_route_table_association"&lt;/span&gt; &lt;span class="s2"&gt;"aks_nodes"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;subnet_id&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;azurerm_subnet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aks_nodes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;route_table_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;azurerm_route_table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aks_egress&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;outbound_type = "userDefinedRouting"&lt;/code&gt; tells AKS to trust this route table instead of provisioning its own path. The Route Table has to exist and be associated to the node subnet &lt;em&gt;before&lt;/em&gt; the cluster is created — &lt;code&gt;depends_on&lt;/code&gt; enforces the ordering, because AKS validates the UDR is actually in place at cluster-creation time and fails otherwise.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Terraform module (free, MIT):&lt;/strong&gt; &lt;a href="https://github.com/dwoitzik/terraform-azurerm-private-aks" rel="noopener noreferrer"&gt;Private AKS - Zero-Trust Edition →&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 3: ACR and Key Vault Without a Public Endpoint
&lt;/h2&gt;

&lt;p&gt;Nodes still need to pull images and (often) read secrets. The naive fix — &lt;code&gt;admin_enabled = true&lt;/code&gt; on the registry, a shared pull secret in a Kubernetes Secret — is a credential that outlives any single deployment and shows up in &lt;code&gt;kubectl get secret -o yaml&lt;/code&gt; for anyone with namespace read access.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"azurerm_container_registry"&lt;/span&gt; &lt;span class="s2"&gt;"this"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;sku&lt;/span&gt;                            &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"Premium"&lt;/span&gt; &lt;span class="c1"&gt;# required for Private Link&lt;/span&gt;
  &lt;span class="nx"&gt;admin_enabled&lt;/span&gt;                  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="nx"&gt;public_network_access_enabled&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"azurerm_role_assignment"&lt;/span&gt; &lt;span class="s2"&gt;"aks_acr_pull"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;scope&lt;/span&gt;                            &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;azurerm_container_registry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;this&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;role_definition_name&lt;/span&gt;             &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"AcrPull"&lt;/span&gt;
  &lt;span class="nx"&gt;principal_id&lt;/span&gt;                     &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;azurerm_kubernetes_cluster&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;kubelet_identity&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;object_id&lt;/span&gt;
  &lt;span class="nx"&gt;skip_service_principal_aad_check&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The kubelet identity — the identity that actually pulls images, not the cluster's control-plane identity — gets &lt;code&gt;AcrPull&lt;/code&gt; directly. No credential to rotate, nothing in a Secret, nothing that outlives the node it's bound to. Key Vault follows the identical shape: RBAC-authorized, &lt;code&gt;public_network_access_enabled = false&lt;/code&gt;, kubelet identity gets &lt;code&gt;Key Vault Secrets User&lt;/code&gt;, ready for the CSI Secrets Store driver to mount without ever touching a static key.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Doesn't Do
&lt;/h2&gt;

&lt;p&gt;It's a starting cluster, not a finished platform. No Entra ID–integrated &lt;code&gt;kubectl&lt;/code&gt; auth (add &lt;code&gt;azure_active_directory_role_based_access_control&lt;/code&gt; yourself), no multi-pool topology, no cluster autoscaler wiring, &lt;code&gt;sku_tier&lt;/code&gt; defaults to &lt;code&gt;Free&lt;/code&gt; (no uptime SLA — set &lt;code&gt;Standard&lt;/code&gt; for production). And it assumes you already have a firewall to hand it a private IP — pair it with the &lt;a href="https://github.com/dwoitzik/terraform-azurerm-firewall-forced-tunneling" rel="noopener noreferrer"&gt;Azure Firewall Forced Tunneling module&lt;/a&gt; if you don't.&lt;/p&gt;

&lt;p&gt;For the operational reality of running production Kubernetes past the "it deployed" stage — the workflows and failure modes that don't show up in a getting-started guide — &lt;a href="https://dev.to/go/gitops-kubernetes-book"&gt;GitOps and Kubernetes*&lt;/a&gt; is the reference I've actually kept open while doing this.&lt;/p&gt;




&lt;h2&gt;
  
  
  🐙 Private AKS - Zero-Trust Edition — free, MIT licensed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;private_cluster_enabled with no public FQDN at all&lt;/li&gt;
&lt;li&gt;Forced tunneling via userDefinedRouting - no default outbound path&lt;/li&gt;
&lt;li&gt;Private-endpoint-only ACR (Premium) and Key Vault, kubelet identity RBAC only&lt;/li&gt;
&lt;li&gt;Zero-Trust node subnet NSG, DINE-policy-safe Private DNS Zones&lt;/li&gt;
&lt;li&gt;Full source code - no lock-in, no black box&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Get it: &lt;a href="https://github.com/dwoitzik/terraform-azurerm-private-aks" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Full source code · free forever · no account, no checkout&lt;/em&gt;&lt;/p&gt;




</description>
      <category>azure</category>
      <category>terraform</category>
      <category>kubernetes</category>
      <category>zerotrust</category>
    </item>
    <item>
      <title>Staggered VM Boot: How I Prevented a Load Average of 147</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 07 Sep 2026 12:29:57 +0000</pubDate>
      <link>https://dev.to/dwoitzik/staggered-vm-boot-how-i-prevented-a-load-average-of-147-18cd</link>
      <guid>https://dev.to/dwoitzik/staggered-vm-boot-how-i-prevented-a-load-average-of-147-18cd</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/staggered-vm-boot-load-average-147/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;After a host reboot, every VM and LXC on the Proxmox host started simultaneously. Twelve containers and VMs, all booting at once, all hitting the same NVMe for root filesystem reads, service starts, and NFS mounts. The host load average hit 147.&lt;/p&gt;

&lt;p&gt;For context: load average represents the number of processes in the run queue or waiting for I/O. On a 16-thread CPU, a load average of 147 means 147 processes are competing for CPU or disk time. Every service was slow to start, k3s took minutes to become ready, and DNS didn't resolve for the first 90 seconds because the Raspberry Pi DNS nodes were waiting for services that hadn't booted yet.&lt;/p&gt;

&lt;p&gt;The fix wasn't more resources — it was boot order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Boot Storm
&lt;/h2&gt;

&lt;p&gt;When the Proxmox host starts, all VMs and LXCs configured with &lt;code&gt;onboot=1&lt;/code&gt; start simultaneously. The host's NVMe handles root filesystem reads for every container, plus the ZFS txg commits from the NFS server, plus etcd writes from the k3s control-plane.&lt;/p&gt;

&lt;p&gt;The simultaneous startup creates a thundering herd: 12 processes all requesting I/O at the same time, the NVMe queue depth maxes out, I/O latency spikes, and services that depend on each other (k3s needs NFS, k3s apps need DNS, DNS needs k3s services) enter a cascading wait state.&lt;/p&gt;

&lt;p&gt;The load average doesn't just spike and recover — it compounds. Services that fail to start within their timeout window retry, adding more processes to the queue. k3s control-plane tries to mount NFS volumes, NFS is slow because it's competing with 11 other containers for I/O, k3s retries the mount, adding more load.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix: Staggered Boot Order
&lt;/h2&gt;

&lt;p&gt;Proxmox supports &lt;code&gt;startup&lt;/code&gt; order with delays. The configuration in Terraform:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# terraform/stacks/proxmox/lxc.tf&lt;/span&gt;

&lt;span class="c1"&gt;# NFS first — k3s depends on it&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"ct_srv_nfs_01"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;# ...&lt;/span&gt;
  &lt;span class="nx"&gt;startup&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="nx"&gt;up&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;  &lt;span class="c1"&gt;# wait 30s after boot before starting next&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# k3s control-plane — waits for NFS&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"vm_srv_k3s_11"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;# ...&lt;/span&gt;
  &lt;span class="nx"&gt;startup&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="nx"&gt;up&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;  &lt;span class="c1"&gt;# 30s after NFS is up&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# k3s workers — 30s apart from each other&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"vm_srv_k3s_12"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;startup&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
    &lt;span class="nx"&gt;up&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"proxmox_virtual_machine"&lt;/span&gt; &lt;span class="s2"&gt;"vm_srv_k3s_13"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;startup&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;
    &lt;span class="nx"&gt;up&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;NFS&lt;/strong&gt; (order 1) — boots first, 30s head start. k3s PVCs mount from NFS, so NFS must be ready before k3s starts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;k3s-11&lt;/strong&gt; (order 2) — control-plane + etcd. Boots 30s after NFS. Needs NFS for system PVCs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;k3s-12&lt;/strong&gt; (order 3) — worker. Boots 30s after control-plane. Needs API server ready.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;k3s-13&lt;/strong&gt; (order 4) — worker. Boots 30s after k3s-12.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LXCs&lt;/strong&gt; (order 5+) — everything else. Docker workloads, media stack, DMZ.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Total boot sequence: ~3 minutes for everything to be online. Previously: all at once, 147 load average, 5+ minutes to stability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why NFS First
&lt;/h2&gt;

&lt;p&gt;NFS is the foundation of the storage layer. Every k3s PVC (Authelia, Vaultwarden, Paperless, Nextcloud) mounts from the NFS server at &lt;code&gt;10.0.20.100&lt;/code&gt;. If NFS isn't ready when k3s starts, the pod mount attempts fail, Kubernetes retries with exponential backoff, and the pods sit in &lt;code&gt;ContainerCreating&lt;/code&gt; for minutes.&lt;/p&gt;

&lt;p&gt;NFS itself depends on ZFS — the NFS export directory lives on the ZFS pool. ZFS needs a few seconds after boot to complete any pending txg commits and mount the pool. The 30-second head start gives ZFS and NFS time to stabilize before k3s starts hammering them with mount requests.&lt;/p&gt;

&lt;h2&gt;
  
  
  The CPU Scheduling Priority
&lt;/h2&gt;

&lt;p&gt;Beyond boot order, k3s VMs get CPU scheduling priority via &lt;code&gt;cpu.units&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# k3s VMs get 2x scheduling priority&lt;/span&gt;
&lt;span class="nx"&gt;cpu&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;cores&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;
  &lt;span class="nx"&gt;units&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2048&lt;/span&gt;  &lt;span class="c1"&gt;# default is 1024&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# LXCs stay at default priority&lt;/span&gt;
&lt;span class="nx"&gt;cpu&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;cores&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
  &lt;span class="nx"&gt;units&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;  &lt;span class="c1"&gt;# default&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;cpu.units&lt;/code&gt; tells the Proxmox scheduler how to weight CPU time when multiple VMs compete for the same physical cores. &lt;code&gt;units=2048&lt;/code&gt; means k3s VMs get twice the CPU scheduling priority over LXCs. When Ollama is running LLM inference on the AI LXC and k3s needs CPU for etcd fdatasync, etcd wins.&lt;/p&gt;

&lt;p&gt;This matters during boot too: even with staggered starts, there's overlap between late-booting LXCs and already-running k3s workloads. The CPU priority ensures k3s gets scheduling preference.&lt;/p&gt;

&lt;h2&gt;
  
  
  The &lt;code&gt;onboot&lt;/code&gt; Gotcha
&lt;/h2&gt;

&lt;p&gt;Proxmox's &lt;code&gt;bpg/proxmox&lt;/code&gt; Terraform provider doesn't reliably manage the &lt;code&gt;onboot&lt;/code&gt; attribute. &lt;code&gt;terraform plan&lt;/code&gt; always shows "No changes" regardless of the live value — a known limitation of the provider.&lt;/p&gt;

&lt;p&gt;This means &lt;code&gt;onboot&lt;/code&gt; must be set manually after any LXC recreate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pct &lt;span class="nb"&gt;set&lt;/span&gt; &amp;lt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nt"&gt;-onboot&lt;/span&gt; 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I discovered this the hard way: after recreating a container via Terraform, &lt;code&gt;onboot&lt;/code&gt; defaulted to &lt;code&gt;0&lt;/code&gt;. The next host reboot silently skipped that container. The k3s control-plane node came up without its NFS mount, and half the cluster was in CrashLoopBackOff until I noticed.&lt;/p&gt;

&lt;p&gt;The workaround: a manual &lt;code&gt;pct set&lt;/code&gt; step in the operations runbook, applied after every LXC creation. Not ideal, but documented and repeatable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before vs. After
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before (simultaneous)&lt;/th&gt;
&lt;th&gt;After (staggered)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Peak load average&lt;/td&gt;
&lt;td&gt;147&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to k3s ready&lt;/td&gt;
&lt;td&gt;5+ minutes&lt;/td&gt;
&lt;td&gt;90 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to DNS functional&lt;/td&gt;
&lt;td&gt;90 seconds&lt;/td&gt;
&lt;td&gt;30 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;I/O wait %&lt;/td&gt;
&lt;td&gt;85%&lt;/td&gt;
&lt;td&gt;15%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failed mount attempts&lt;/td&gt;
&lt;td&gt;12-15&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The staggering didn't add total boot time — it redistributed the I/O load over 3 minutes instead of concentrating it in 30 seconds. Services come up later individually but the cluster as a whole is stable sooner because nothing is fighting for I/O. The same host also runs &lt;a href="https://dev.to/blog/overcommit-guard-python-host-freeze/"&gt;a memory overcommit guard&lt;/a&gt; for the RAM side of this problem — boot order fixes the I/O storm, the guard fixes the memory one.&lt;/p&gt;




&lt;p&gt;Boot storm mitigation is the same problem in Azure: when you scale out a VMSS from 0 to 50 instances, all 50 hit the Azure fabric simultaneously. Azure handles this with staggered placement and shared disks, but the principle is identical — spread the I/O load over time instead of concentrating it. In a homelab, you do it yourself with &lt;code&gt;startup.order&lt;/code&gt; and &lt;code&gt;startup.up&lt;/code&gt; delays.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/go/4aUtkCs"&gt;Designing Data-Intensive Applications*&lt;/a&gt; has a genuinely useful framing for this kind of thundering-herd problem, even though it's written about databases rather than hypervisors - the queueing math behind "everything wants the same resource at the same instant" doesn't care what the resource actually is.&lt;/p&gt;



</description>
      <category>proxmox</category>
      <category>homelab</category>
      <category>performance</category>
      <category>debugging</category>
    </item>
    <item>
      <title>The Overcommit Guard: How a Python Script Prevents Host Freezes</title>
      <dc:creator>david</dc:creator>
      <pubDate>Thu, 03 Sep 2026 19:30:32 +0000</pubDate>
      <link>https://dev.to/dwoitzik/the-overcommit-guard-how-a-python-script-prevents-host-freezes-48h0</link>
      <guid>https://dev.to/dwoitzik/the-overcommit-guard-how-a-python-script-prevents-host-freezes-48h0</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/overcommit-guard-python-host-freeze/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;My Proxmox host froze three times in one week. The root cause was RAM pressure pushing ZFS into an I/O stall on a single shared NVMe. After fixing the freeze (ZFS tuning, cache=none+aio=native), I needed a way to prevent it from ever happening again — a CI guard that catches memory overcommit before it hits the host. It's the same host where &lt;a href="https://dev.to/blog/staggered-vm-boot-load-average-147/"&gt;uncontrolled boot storms once pushed load average to 147&lt;/a&gt;; memory pressure and I/O pressure are two versions of the same single-host bottleneck.&lt;/p&gt;

&lt;p&gt;The first version of the guard was wrong. It summed VM and LXC memory allocations together against one ceiling, and flagged 85 GB as "allocated" on a host with 62 GB of physical RAM. A live check showed only 41 GB actually in use — 21 GB available. The guard was overstating pressure by 40 GB because it treated two different memory physics as the same thing.&lt;/p&gt;

&lt;p&gt;The fix was a two-tier model that understands the difference between a VM reservation and an LXC ceiling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Two Memory Physics
&lt;/h2&gt;

&lt;h3&gt;
  
  
  VM &lt;code&gt;dedicated&lt;/code&gt; — Real Reservation
&lt;/h3&gt;

&lt;p&gt;When you set &lt;code&gt;dedicated = 12288&lt;/code&gt; (12 GB) on a Proxmox VM, QEMU pre-allocates that memory as a host process. The moment the VM starts, 12 GB of physical RAM is gone — reserved for the QEMU process, not available for anything else. The host's &lt;code&gt;free -h&lt;/code&gt; reflects this immediately.&lt;/p&gt;

&lt;p&gt;This is a hard reservation. If the sum of all VM &lt;code&gt;dedicated&lt;/code&gt; values plus ZFS ARC plus host overhead exceeds physical RAM, the host is overcommitted. The kernel will start swapping, ZFS will stall waiting for I/O, and the freeze pattern repeats.&lt;/p&gt;

&lt;h3&gt;
  
  
  LXC &lt;code&gt;dedicated&lt;/code&gt; — Soft Ceiling
&lt;/h3&gt;

&lt;p&gt;When you set &lt;code&gt;dedicated = 4096&lt;/code&gt; (4 GB) on a Proxmox LXC, you're setting &lt;code&gt;memory.max&lt;/code&gt; in the cgroup. This is a ceiling the kernel enforces &lt;em&gt;only if the container actually tries to use that much&lt;/em&gt;. It reserves nothing on the host up front.&lt;/p&gt;

&lt;p&gt;A container with &lt;code&gt;dedicated = 4096&lt;/code&gt; might be using 800 MB. The host sees 800 MB, not 4 GB. The remaining 3.2 GB is available for other workloads. This is why the original guard was wrong: summing LXC &lt;code&gt;dedicated&lt;/code&gt; values counts memory that isn't actually consumed.&lt;/p&gt;

&lt;p&gt;A live check on the host confirmed this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Check actual LXC memory usage vs configured limits&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;ct &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;pct list | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'NR&amp;gt;1 {print $1}'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nv"&gt;limit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;pct config &lt;span class="nv"&gt;$ct&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"dedicated"&lt;/span&gt; | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $2}'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="nv"&gt;actual&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;pct &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nv"&gt;$ct&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nb"&gt;cat&lt;/span&gt; /sys/fs/cgroup/memory.current 2&amp;gt;/dev/null&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"CT &lt;/span&gt;&lt;span class="nv"&gt;$ct&lt;/span&gt;&lt;span class="s2"&gt;: limit=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;limit&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;MB actual=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt;actual/1024/1024&lt;span class="k"&gt;))&lt;/span&gt;&lt;span class="s2"&gt;MB"&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;span class="c"&gt;# → Most CTs using 20-40% of their configured limit&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Two-Tier Guard
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Hard Gate (CI-Fail)
&lt;/h3&gt;

&lt;p&gt;The hard gate catches real overcommit risk. It sums only VM &lt;code&gt;dedicated&lt;/code&gt; values (real reservations) plus ZFS ARC max plus a fixed host reserve:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Hard gate: VM reservations + ARC + host reserve
&lt;/span&gt;&lt;span class="n"&gt;ZFS_ARC_MAX_GB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;       &lt;span class="c1"&gt;# /sys/module/zfs/parameters/zfs_arc_max
&lt;/span&gt;&lt;span class="n"&gt;HOST_RESERVE_GB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;       &lt;span class="c1"&gt;# kernel + QEMU overhead
&lt;/span&gt;&lt;span class="n"&gt;HARD_GATE_CEILING_GB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;44&lt;/span&gt; &lt;span class="c1"&gt;# safe ceiling for 62 GB physical
&lt;/span&gt;
&lt;span class="n"&gt;hard_gate_mb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vm_dedicated_total&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ZFS_ARC_MAX_GB&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;HOST_RESERVE_GB&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the hard gate exceeds 44 GB, the CI build fails. The PR cannot be merged. This is intentional: a VM &lt;code&gt;dedicated&lt;/code&gt; change that pushes past the ceiling is exactly the kind of change that caused the original freeze.&lt;/p&gt;

&lt;p&gt;The ceiling (44 GB on a 62 GB host) leaves 18 GB of headroom for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LXC actual usage (typically 8-12 GB across all containers)&lt;/li&gt;
&lt;li&gt;Kernel and system overhead not captured in the reserve&lt;/li&gt;
&lt;li&gt;Burst spikes from Ollama inference or Paperless OCR&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Soft Check (Warn-Only)
&lt;/h3&gt;

&lt;p&gt;The soft check sums LXC &lt;code&gt;dedicated&lt;/code&gt; values and compares against physical RAM. If it exceeds 62 GB, a warning is printed — but the build does &lt;strong&gt;not&lt;/strong&gt; fail. LXC ceilings are soft; summing them overstates real pressure.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Soft check: LXC CT limits vs physical RAM (visibility only)
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;lxc_dedicated_total&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;PHYSICAL_RAM_GB&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;WARN: LXC CT limits sum to &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;lxc_dedicated_total&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; GB, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
          &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;over physical RAM (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;PHYSICAL_RAM_GB&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; GB). &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
          &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Check actual usage: pct exec &amp;lt;id&amp;gt; -- cat /sys/fs/cgroup/memory.current&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The soft check exists for visibility. If someone adds a new LXC with 32 GB &lt;code&gt;dedicated&lt;/code&gt; and the sum crosses 100 GB, the warning fires — but it doesn't block the merge, because the actual usage is probably a fraction of the configured limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why &lt;code&gt;floating&lt;/code&gt; Is Ignored
&lt;/h2&gt;

&lt;p&gt;VMs can have both &lt;code&gt;dedicated&lt;/code&gt; (ceiling) and &lt;code&gt;floating&lt;/code&gt; (balloon-driven floor). Under host memory pressure, Proxmox can deflate the balloon down to the &lt;code&gt;floating&lt;/code&gt; minimum, freeing memory for other workloads.&lt;/p&gt;

&lt;p&gt;The guard conservatively uses &lt;code&gt;dedicated&lt;/code&gt;, not &lt;code&gt;floating&lt;/code&gt;, because:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The balloon only deflates &lt;em&gt;after&lt;/em&gt; pressure is detected — it doesn't prevent the pressure&lt;/li&gt;
&lt;li&gt;A sudden allocation spike (Ollama loading a 26B model) can't wait for the balloon to deflate&lt;/li&gt;
&lt;li&gt;The guard exists to catch &lt;em&gt;pending&lt;/em&gt; risk, not to model &lt;em&gt;steady-state&lt;/em&gt; usage&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;code&gt;floating&lt;/code&gt; values are parsed and printed for visibility, never counted in the hard gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The CI Integration
&lt;/h2&gt;

&lt;p&gt;The script runs in pre-commit hooks and GitHub Actions CI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# .github/workflows/ci.yml&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Memory overcommit guard&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;python scripts/check-host-memory-overcommit.py&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A Terraform change that adds a new VM or increases a VM's &lt;code&gt;dedicated&lt;/code&gt; value is checked against the ceiling before merge. If it pushes past 44 GB, the PR is blocked with a clear error message explaining why.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;python scripts/check-host-memory-overcommit.py
REL-035 memory overcommit guard &lt;span class="o"&gt;(&lt;/span&gt;two-tier model&lt;span class="o"&gt;)&lt;/span&gt;

&lt;span class="nt"&gt;--&lt;/span&gt; Hard gate &lt;span class="o"&gt;(&lt;/span&gt;CI-fail&lt;span class="o"&gt;)&lt;/span&gt;: VM reservations + ARC + host reserve &lt;span class="nt"&gt;--&lt;/span&gt;
  VM &lt;span class="sb"&gt;`&lt;/span&gt;dedicated&lt;span class="sb"&gt;`&lt;/span&gt; &lt;span class="nb"&gt;sum&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;3 VMs&lt;span class="o"&gt;)&lt;/span&gt;: 36864 MB &lt;span class="o"&gt;(&lt;/span&gt;36.0 GB&lt;span class="o"&gt;)&lt;/span&gt;
  VM &lt;span class="sb"&gt;`&lt;/span&gt;floating&lt;span class="sb"&gt;`&lt;/span&gt; &lt;span class="nb"&gt;sum&lt;/span&gt;  &lt;span class="o"&gt;(&lt;/span&gt;3 VMs&lt;span class="o"&gt;)&lt;/span&gt;: 40960 MB &lt;span class="o"&gt;(&lt;/span&gt;40.0 GB&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; informational only, NOT counted
  + ZFS ARC max: 4 GB
  + host/hypervisor reserve: 6 GB
  &lt;span class="o"&gt;=&lt;/span&gt; hard gate total: 47104 MB &lt;span class="o"&gt;(&lt;/span&gt;46.0 GB&lt;span class="o"&gt;)&lt;/span&gt;
  Ceiling: 44 GB

FAIL: hard gate total &lt;span class="o"&gt;(&lt;/span&gt;46.0 GB&lt;span class="o"&gt;)&lt;/span&gt; exceeds the 44 GB ceiling by 2.0 GB.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What It Prevents
&lt;/h2&gt;

&lt;p&gt;The guard exists because &lt;code&gt;mini&lt;/code&gt; — the single Proxmox host running this entire homelab — has no failover. A host freeze means every service is down simultaneously: k3s, databases, DNS, monitoring, backups. The only recovery is a hard power cycle.&lt;/p&gt;

&lt;p&gt;REL-016 (the ZFS freeze) happened because RAM pressure pushed ZFS into a stall-wait state on the shared NVMe. The hard gate catches the most common path to that state: a VM &lt;code&gt;dedicated&lt;/code&gt; change that leaves insufficient headroom for the kernel, ZFS, and LXC workloads.&lt;/p&gt;

&lt;p&gt;It doesn't prevent every possible freeze — a runaway process inside a VM can still consume all its allocated memory and cause host pressure. But it prevents the &lt;em&gt;planned&lt;/em&gt; overcommit: the Terraform change that accidentally pushes past the ceiling because someone added a new VM without checking the math. CPU has the same two-tier problem; see &lt;a href="https://dev.to/blog/cpu-scheduling-etcd-proxmox-units/"&gt;how &lt;code&gt;cpu.units&lt;/code&gt; scheduling priority solved it for etcd&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Memory overcommit modeling is the same problem in Azure: Reserved VM instances guarantee physical memory allocation, while Burstable VMs share host memory and can be throttled. Mixing both in the same Availability Set without understanding the difference produces the same false-sense-of-security that my original guard had. The fix is the same: separate hard reservations from soft ceilings, and never sum them together.&lt;/p&gt;



</description>
      <category>proxmox</category>
      <category>homelab</category>
      <category>debugging</category>
      <category>cicd</category>
    </item>
    <item>
      <title>GitOps for Firewall Rules: MikroTik + Terraform + Atlantis</title>
      <dc:creator>david</dc:creator>
      <pubDate>Mon, 31 Aug 2026 13:31:54 +0000</pubDate>
      <link>https://dev.to/dwoitzik/gitops-for-firewall-rules-mikrotik-terraform-atlantis-1i9l</link>
      <guid>https://dev.to/dwoitzik/gitops-for-firewall-rules-mikrotik-terraform-atlantis-1i9l</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://woitzik.dev/blog/mikrotik-terraform-gitops-firewall-management/" rel="noopener noreferrer"&gt;woitzik.dev&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Disclosure: This post contains Amazon affiliate links (marked with *). If you buy through them, I earn a small commission at no extra cost to you. I only link gear I actually own and use daily.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;My MikroTik RB5009 has no manual RouterOS configuration. Every firewall rule, every VLAN interface, every DHCP static lease, every NAT port-forward — all of it lives in Terraform, reviewed via Atlantis pull requests, and applied through the same GitOps workflow that manages the Kubernetes cluster. This is the delivery pipeline underneath &lt;a href="https://dev.to/blog/mikrotik-zero-trust-firewall-terraform/"&gt;the zero-trust default-deny firewall policy&lt;/a&gt; and &lt;a href="https://dev.to/blog/mikrotik-vlan-filtering-terraform-proxmox/"&gt;the VLAN matrix&lt;/a&gt; — this article covers the Atlantis plumbing, not the rules themselves.&lt;/p&gt;

&lt;p&gt;This article covers the full implementation: the &lt;code&gt;locals&lt;/code&gt; block that drives the VLAN matrix, the &lt;code&gt;place_before&lt;/code&gt; pattern for deterministic firewall ordering, and how a router's LED schedule ended up in Terraform.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/dwoitzik/homelab-infrastructure" rel="noopener noreferrer"&gt;View the complete homelab infrastructure source on GitHub 🐙&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Provider Setup
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;routeros&lt;/code&gt; provider (v1.80+) connects to the MikroTik via the REST API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# terraform/stacks/network/providers.tf&lt;/span&gt;
&lt;span class="nx"&gt;terraform&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;required_version&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"&amp;gt;= 1.5"&lt;/span&gt;
  &lt;span class="nx"&gt;required_providers&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;routeros&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;source&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"terraform-routeros/routeros"&lt;/span&gt;
      &lt;span class="nx"&gt;version&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"~&amp;gt; 1.80"&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nx"&gt;backend&lt;/span&gt; &lt;span class="s2"&gt;"s3"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;bucket&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"terraform-state"&lt;/span&gt;
    &lt;span class="nx"&gt;key&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"network/terraform.tfstate"&lt;/span&gt;
    &lt;span class="nx"&gt;endpoints&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;s3&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"https://s3.woitzik.dev"&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nx"&gt;skip_credentials_validation&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="nx"&gt;skip_metadata_api_check&lt;/span&gt;     &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="nx"&gt;skip_requesting_account_id&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The S3 backend is Garage, the self-hosted S3-compatible storage running inside k3s. The Terraform state for the network stack lives on the same cluster that depends on the network — another circular dependency that's acceptable because a full cluster failure means the network is the least of your problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  The VLAN Matrix
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;locals&lt;/code&gt; block is the single source of truth for the entire network:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# terraform/stacks/network/main.tf&lt;/span&gt;
&lt;span class="nx"&gt;locals&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;homelab_vlans&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"vlan20-srv"&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;
    &lt;span class="s2"&gt;"vlan30-dmz"&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;
    &lt;span class="s2"&gt;"vlan40-iot"&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;
    &lt;span class="s2"&gt;"vlan100-admin"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;rpi_port_mapping&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"ether6"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# RPi 4B #1 (Keepalived Node A)&lt;/span&gt;
    &lt;span class="s2"&gt;"ether7"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# RPi 4B #2 (Keepalived Node B)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;proxmox_port&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"ether5"&lt;/span&gt;

  &lt;span class="c1"&gt;# Port mapping for all VLAN-tagged trunks&lt;/span&gt;
  &lt;span class="nx"&gt;port_mapping&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"ether1"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;   &lt;span class="c1"&gt;# WAN (FritzBox)&lt;/span&gt;
    &lt;span class="s2"&gt;"ether2"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# k3s-11&lt;/span&gt;
    &lt;span class="s2"&gt;"ether3"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# k3s-12&lt;/span&gt;
    &lt;span class="s2"&gt;"ether4"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# k3s-13&lt;/span&gt;
    &lt;span class="s2"&gt;"ether5"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# Proxmox host&lt;/span&gt;
    &lt;span class="s2"&gt;"ether6"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# RPi #1&lt;/span&gt;
    &lt;span class="s2"&gt;"ether7"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;  &lt;span class="c1"&gt;# RPi #2&lt;/span&gt;
    &lt;span class="s2"&gt;"sfp-sfpplus1"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;  &lt;span class="c1"&gt;# DMZ switch&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adding a new VLAN means one entry in &lt;code&gt;homelab_vlans&lt;/code&gt;. Terraform auto-generates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Bridge VLAN entries&lt;/li&gt;
&lt;li&gt;DHCP server per VLAN&lt;/li&gt;
&lt;li&gt;Firewall rules per VLAN&lt;/li&gt;
&lt;li&gt;VLAN interface on the bridge&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Moving a device between VLANs means changing one number in &lt;code&gt;port_mapping&lt;/code&gt;. The firewall rules, DHCP leases, and bridge entries update automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deterministic Firewall Ordering
&lt;/h2&gt;

&lt;p&gt;MikroTik evaluates firewall rules in order and stops at the first match. Terraform manages rule ordering via &lt;code&gt;place_before&lt;/code&gt; references:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# terraform/stacks/network/firewall_deterministic.tf&lt;/span&gt;

&lt;span class="c1"&gt;# Final drop rule — must be last in the forward chain&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"routeros_ip_firewall_filter"&lt;/span&gt; &lt;span class="s2"&gt;"fwd_99_drop_all"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;action&lt;/span&gt;       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"drop"&lt;/span&gt;
  &lt;span class="nx"&gt;chain&lt;/span&gt;        &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"forward"&lt;/span&gt;
  &lt;span class="nx"&gt;place_before&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;  &lt;span class="c1"&gt;# explicit: this is the last rule&lt;/span&gt;
  &lt;span class="nx"&gt;comment&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"99: Global - Final Drop (Zero Trust Policy)"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Anti-spoofing — must come before any accept rules&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"routeros_ip_firewall_filter"&lt;/span&gt; &lt;span class="s2"&gt;"fwd_00b_anti_spoof"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;action&lt;/span&gt;       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"drop"&lt;/span&gt;
  &lt;span class="nx"&gt;chain&lt;/span&gt;        &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"forward"&lt;/span&gt;
  &lt;span class="nx"&gt;src_address&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"10.0.0.0/8"&lt;/span&gt;
  &lt;span class="nx"&gt;in_interface_list&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"WAN"&lt;/span&gt;
  &lt;span class="nx"&gt;place_before&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;routeros_ip_firewall_filter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;fwd_01_established&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;comment&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"00b: Anti-Spoof - Drop internal src from WAN"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Established/related — allows return traffic&lt;/span&gt;
&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"routeros_ip_firewall_filter"&lt;/span&gt; &lt;span class="s2"&gt;"fwd_01_established"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;action&lt;/span&gt;       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"accept"&lt;/span&gt;
  &lt;span class="nx"&gt;chain&lt;/span&gt;        &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"forward"&lt;/span&gt;
  &lt;span class="nx"&gt;connection_state&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"established"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"related"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="nx"&gt;place_before&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;routeros_ip_firewall_filter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;fwd_99_drop_all&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;comment&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"01: Allow Established/Related"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;place_before&lt;/code&gt; attribute creates explicit ordering dependencies. Rule 00b must come before rule 01, which must come before rule 99. Terraform resolves these dependencies during plan, so the apply order matches the intended evaluation order.&lt;/p&gt;

&lt;p&gt;Without &lt;code&gt;place_before&lt;/code&gt;, Terraform would apply rules in dependency-graph order, which is not guaranteed to match the intended firewall evaluation order. A rule that should be evaluated first might be created last, allowing a broader rule to match traffic before the narrower, more specific rule is evaluated.&lt;/p&gt;

&lt;p&gt;The full ruleset in &lt;code&gt;firewall_deterministic.tf&lt;/code&gt; has 22 rules with explicit ordering. The &lt;code&gt;firewall_extra.tf&lt;/code&gt; file has 19 additional rules for VPN access tiers, monitoring, and application-specific port forwards — all with &lt;code&gt;place_before&lt;/code&gt; references to maintain correct evaluation order.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Atlantis Flow
&lt;/h2&gt;

&lt;p&gt;Every change to the network stack goes through Atlantis:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Edit &lt;code&gt;terraform/stacks/network/*.tf&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Push to a feature branch&lt;/li&gt;
&lt;li&gt;Open a PR against &lt;code&gt;main&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Atlantis comments with the &lt;code&gt;terraform plan&lt;/code&gt; output&lt;/li&gt;
&lt;li&gt;Review the plan&lt;/li&gt;
&lt;li&gt;Comment &lt;code&gt;atlantis apply&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Atlantis applies the changes to the MikroTik&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;code&gt;atlantis.yaml&lt;/code&gt; at the repo root defines the project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;
&lt;span class="na"&gt;projects&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;network&lt;/span&gt;
    &lt;span class="na"&gt;dir&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;terraform/stacks/network&lt;/span&gt;
    &lt;span class="na"&gt;workspace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
    &lt;span class="na"&gt;autoplan&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;when_modified&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.tf"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.tfvars"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The plan output shows exactly what will change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;# routeros_ip_firewall_filter.fwd_04a_monitoring will be updated in-place&lt;/span&gt;
&lt;span class="err"&gt;~&lt;/span&gt; &lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"routeros_ip_firewall_filter"&lt;/span&gt; &lt;span class="s2"&gt;"fwd_04a_monitoring"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;action&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"accept"&lt;/span&gt;
      &lt;span class="nx"&gt;chain&lt;/span&gt;       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"forward"&lt;/span&gt;
  &lt;span class="err"&gt;~&lt;/span&gt;   &lt;span class="nx"&gt;dst_port&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"9100"&lt;/span&gt; &lt;span class="nx"&gt;-&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"9100,9090"&lt;/span&gt;
      &lt;span class="nx"&gt;src_address&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"10.0.20.0/24"&lt;/span&gt;
      &lt;span class="nx"&gt;comment&lt;/span&gt;     &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"04a: SRV - Prometheus scrape to MGMT"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No surprise changes. Every rule modification is visible in the PR before it touches the live router.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Goes in Terraform vs. What Doesn't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;In Terraform:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All firewall filter rules (input + forward chains)&lt;/li&gt;
&lt;li&gt;VLAN interfaces and bridge matrix&lt;/li&gt;
&lt;li&gt;DHCP static leases&lt;/li&gt;
&lt;li&gt;NAT port-forwards&lt;/li&gt;
&lt;li&gt;QoS traffic shaping&lt;/li&gt;
&lt;li&gt;SNMP community configuration&lt;/li&gt;
&lt;li&gt;Power LED scheduling (yes, really)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Not in Terraform:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;RouterOS user management (done once, rarely changes)&lt;/li&gt;
&lt;li&gt;Certificate renewals (handled by the router's own scheduler)&lt;/li&gt;
&lt;li&gt;Custom scripts for monitoring (run via SNMP, not Terraform)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The principle: if it affects traffic flow or network behavior, it's in Terraform. If it's one-time configuration or the router's own maintenance, it stays in RouterOS.&lt;/p&gt;

&lt;h2&gt;
  
  
  The LED Night-Mode
&lt;/h2&gt;

&lt;p&gt;The MikroTik RB5009 has a power LED that's bright enough to be annoying in a dark room. RouterOS supports LED scheduling via the system scheduler. This ended up in Terraform:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"routeros_system_scheduler"&lt;/span&gt; &lt;span class="s2"&gt;"led_night_mode_off"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;     &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"led-off"&lt;/span&gt;
  &lt;span class="nx"&gt;start_time&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"23:00:00"&lt;/span&gt;
  &lt;span class="nx"&gt;interval&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"00:24:00"&lt;/span&gt;
  &lt;span class="nx"&gt;on_event&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/system led set led1 state=off"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"routeros_system_scheduler"&lt;/span&gt; &lt;span class="s2"&gt;"led_night_mode_on"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;     &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"led-on"&lt;/span&gt;
  &lt;span class="nx"&gt;start_time&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"07:00:00"&lt;/span&gt;
  &lt;span class="nx"&gt;interval&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"00:24:00"&lt;/span&gt;
  &lt;span class="nx"&gt;on_event&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/system led set led1 state=on"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two scheduler entries. LED off at 23:00, on at 07:00. Managed via Terraform, reviewed via PR, applied via Atlantis. It's the most trivial thing in the entire network stack, and it's also the change that made my partner the happiest.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Anti-Patterns
&lt;/h2&gt;

&lt;p&gt;Two things I deliberately avoid:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Never create resources directly via RouterOS API.&lt;/strong&gt; Every resource must be created through Terraform. If I need to debug a live issue and make a direct API change, I immediately add a &lt;code&gt;moved {}&lt;/code&gt; block or &lt;code&gt;import {}&lt;/code&gt; block to bring that resource under Terraform management. This prevents the "77 rules, Terraform knows about 22" problem I wrote about earlier.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Never &lt;code&gt;terraform apply&lt;/code&gt; locally.&lt;/strong&gt; All applies go through Atlantis PRs. This creates an audit trail, forces me to review every change, and prevents accidental modifications during debugging sessions.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;Firewall-as-code via Terraform is the same pattern in Azure: NSG rules, Azure Firewall Policy rule collections, and Application Security Groups are all Terraform resources. The difference is that Azure's NSG API handles rule ordering automatically based on priority numbers, while MikroTik requires explicit &lt;code&gt;place_before&lt;/code&gt; dependencies. Both approaches achieve the same goal: firewall rules that are reviewed, version-controlled, and auditable.&lt;/p&gt;



</description>
      <category>terraform</category>
      <category>mikrotik</category>
      <category>gitops</category>
      <category>networking</category>
    </item>
  </channel>
</rss>
