DEV Community

Cover image for Firewall Zone Persistence: Why firewalld Runtime Changes Vanish and How to Pin Them
Guatu
Guatu

Posted on Originally published at guatulabs.dev

Firewall Zone Persistence: Why firewalld Runtime Changes Vanish and How to Pin Them

You run firewall-cmd --zone=public --change-interface=wlp2s0, firewalld answers success, and --get-active-zones shows the change. After the next reboot, a firewall-cmd --reload, or a Wi-Fi reconnect, the interface is back in whatever zone it was in before. firewalld doesn't log anything about it.

This isn't a bug. firewalld keeps two separate configurations, and on a NetworkManager system a third component also gets a say in which zone an interface belongs to. If you don't know about all three, any change you make by hand will eventually be undone.

What I Expected

My mental model was the iptables one: you run a command, the rule exists, and if you want it to survive a reboot you save it. I knew firewalld had a --permanent flag, so I assumed the fix was to tack it on and move on.

There are two problems with that assumption. First, --permanent on its own doesn't change the running firewall at all. Second, for interfaces managed by NetworkManager, the zone binding isn't really firewalld's to keep. NetworkManager pushes a zone to firewalld every time a connection activates. If you've only told firewalld, your binding gets replaced the next time the connection comes up.

What Actually Happens

Runtime and permanent are two separate configs

firewalld keeps a runtime configuration (what the kernel is enforcing right now) and a permanent configuration (XML files in /etc/firewalld/ that get loaded at startup and on reload). Every firewall-cmd call writes to exactly one of them:

# Runtime only: takes effect now, gone on reload/restart/reboot
sudo firewall-cmd --zone=public --add-service=ssh

# Permanent only: written to disk, NOT active until reload
sudo firewall-cmd --permanent --zone=public --add-service=ssh

# Both: the boring, correct pattern
sudo firewall-cmd --permanent --zone=public --add-service=ssh
sudo firewall-cmd --reload
Enter fullscreen mode Exit fullscreen mode

That gives you two ways to fail, and they mirror each other. If you use runtime only, the rule works until the next reload, and a reload is exactly what someone else's Ansible run or package update will trigger. If you use permanent only, the rule doesn't exist yet, so you test it, see it fail, and then "fix" it with a runtime command. Now the two configs disagree and nothing tells you.

The --reload itself is the destructive step that people overlook. A reload throws away the runtime state and rebuilds it from the permanent config. Every runtime-only change you made since the last reload is gone, including the one you made five minutes ago to debug something.

NetworkManager owns the zone of the interfaces it manages

This part is less well known. On Fedora, RHEL, and most desktop distros, NetworkManager manages your Ethernet and Wi-Fi interfaces. Each NM connection profile has a connection.zone property. When the connection activates, NetworkManager calls firewalld over D-Bus and assigns the interface to that zone. If connection.zone is empty, which is the default for every profile, NM puts the interface in firewalld's default zone.

So the sequence goes like this:

  1. You run firewall-cmd --zone=drop --change-interface=wlp2s0. firewalld moves the interface. Runtime only.
  2. Your laptop sleeps, wakes up, and reconnects to Wi-Fi.
  3. NetworkManager activates the connection and tells firewalld to put wlp2s0 in the zone from the profile (empty, so the default zone).
  4. Your interface is back in public. Your drop binding is gone.

Writing the binding into the permanent zone XML (--permanent --zone=drop --add-interface=wlp2s0) doesn't reliably fix this, because NM reasserts its own zone on every activation. Newer firewall-cmd builds detect NM-managed interfaces and print a message about the interface being under NetworkManager's control, and in some cases they forward the change to NM for you. That behavior has changed across versions, so I don't count on it. For an NM-managed interface, I treat NM as the source of truth and set the zone there.

Zone bindings are per connection, not per interface

There's one more consequence. The zone is attached to the connection profile, not to the device. wlp2s0 on your home network and wlp2s0 at an airport are two separate NM connections with two separate connection.zone values. Pinning the zone for the home SSID tells NM nothing about the next network you join. A new SSID creates a new profile with an empty zone, so it lands in the default zone.

That turns out to be the most useful behavior in the whole stack, as long as the default zone is set correctly.

The Fix

1. Set zones on the NetworkManager connection, not in firewalld

For any NM-managed interface, put the zone in the profile:

# See which connection is active on which device
nmcli -f NAME,UUID,DEVICE connection show --active

# Pin the zone on the connection profile (persists in the NM keyfile)
sudo nmcli connection modify "HomeWiFi" connection.zone home

# Re-activate so NM pushes the new zone to firewalld now
sudo nmcli connection up "HomeWiFi"

# Verify from firewalld's side
firewall-cmd --get-active-zones
Enter fullscreen mode Exit fullscreen mode

The setting is stored in the connection's keyfile under /etc/NetworkManager/system-connections/, so it survives reboots and reconnects. It also survives firewall-cmd --reload, because NM reapplies it the next time the connection activates. Since it's stored per connection, it applies only to that SSID.

If you'd rather not bounce the connection, sudo nmcli device reapply wlp2s0 usually applies the change without a full disconnect. I still run connection up when I'm making a change by hand, because the result is unambiguous.

2. Make the default zone the untrusted one

Every new network falls back to firewalld's default zone. That makes the default zone the most important security decision on a laptop, and most people never change it from what the distro shipped. Fedora Workstation ships FedoraWorkstation as the default zone, and that zone allows a wide high-port range. That's a reasonable choice for a desktop on a home LAN. It's a bad one for a café.

# Check what new networks will land in
firewall-cmd --get-default-zone

# Make unknown networks untrusted by default
sudo firewall-cmd --set-default-zone=public
Enter fullscreen mode Exit fullscreen mode

--set-default-zone is one of the few firewalld commands that changes runtime and permanent configuration together. It writes DefaultZone= to /etc/firewalld/firewalld.conf and applies the change immediately, so you don't need --permanent or a reload.

You end up with an allowlist built on inheritance. Every network is untrusted unless its NM profile explicitly says otherwise. You don't have to remember to lock down a new network. You only have to remember to unlock the networks you trust. If you forget, the failure is that something like a file share doesn't work at home. That's annoying, but it's a much safer way to fail than accidentally running a home-level ruleset on hotel Wi-Fi.

I build the rest of the mobile hardening on this rule (exit-node activation, DNS lockdown, captive portal handling), and I covered how the dispatcher scripts fit around it in Trust the Tailnet, Not the Network. The firewalld side is the foundation for all of it: if the zone doesn't persist, nothing built on top of it holds.

3. Bind interfaces NM doesn't manage in firewalld's permanent config

Not every interface goes through NetworkManager. tailscale0 is the one that matters to me. Check how NM sees it first:

nmcli device status | grep -E 'DEVICE|tailscale'
Enter fullscreen mode Exit fullscreen mode

If it shows as unmanaged, NM won't assert a zone for it, and firewalld's permanent config is the right place for the binding:

sudo firewall-cmd --permanent --zone=trusted --add-interface=tailscale0
sudo firewall-cmd --reload
firewall-cmd --get-zone-of-interface=tailscale0
Enter fullscreen mode Exit fullscreen mode

Bind the interface, not the source range. A common suggestion is --add-source=100.64.0.0/10 on the trusted zone. That trusts any packet carrying a CGNAT-range source address arriving on any interface, including the café Wi-Fi, where someone on the same L2 segment can put whatever source IP they like on a packet. Binding to tailscale0 means the trust applies only to traffic that came out of the WireGuard tunnel and was authenticated by the tailnet. It's the same least-privilege thinking as scoping Kubernetes service accounts, applied to a network interface instead. It's also what lets subnet routing into a homelab work without opening the LAN-facing side.

4. If you experimented in runtime, promote it on purpose

Sometimes the right workflow really is to try things in runtime and keep what works. firewalld has a command for that:

# Snapshot the current runtime state into permanent config
sudo firewall-cmd --runtime-to-permanent
Enter fullscreen mode Exit fullscreen mode

It copies everything in runtime to permanent, including the throwaway rule you added while debugging and meant to remove. Diff the two before you run it:

zone=public
diff <(firewall-cmd --zone=$zone --list-all) \
     <(firewall-cmd --permanent --zone=$zone --list-all)
Enter fullscreen mode Exit fullscreen mode

No output means the runtime and permanent configs agree. Any output shows exactly what a reload would throw away, or what --runtime-to-permanent would make permanent. I run this diff before every reload on a machine I didn't configure myself.

Runtime does have one feature I like: --timeout. When you're testing a rule over SSH and might lock yourself out, a time-limited runtime rule removes itself after the timeout:

# Rule expires automatically after 5 minutes
sudo firewall-cmd --zone=public --add-port=8443/tcp --timeout=300
Enter fullscreen mode Exit fullscreen mode

It's the firewall equivalent of a dead-man's switch. If the rule works, add it with --permanent and reload.

The Shadowing Problem You Create Along the Way

Even when you use --permanent correctly, it has a side effect people don't notice. Zone definitions ship in /usr/lib/firewalld/zones/. The first time you permanently modify a zone, firewalld writes a full copy of that zone to /etc/firewalld/zones/<zone>.xml. From then on, the /etc copy completely replaces the vendor version. If a later firewalld update changes the shipped definition of public, you won't get that change, because your local copy wins.

This is the same trap as editing /etc/systemd/resolved.conf directly instead of using a drop-in under resolved.conf.d/: one local edit and you've forked the whole vendor config. firewalld zones don't support drop-ins, so you can't avoid the copy. You can keep an eye on it, though:

# Which zones have you forked from vendor defaults?
ls /etc/firewalld/zones/

# Compare your copy against what the package ships
diff /usr/lib/firewalld/zones/public.xml /etc/firewalld/zones/public.xml

# Reset a zone back to the vendor definition
sudo firewall-cmd --permanent --load-zone-defaults=public
sudo firewall-cmd --reload
Enter fullscreen mode Exit fullscreen mode

I keep the list of forked zones short and track the /etc/firewalld/zones/ files in configuration management. If a zone XML sits there unmanaged, the next person who looks at the box has no way of knowing whether it was a deliberate choice or someone's debugging session that got saved.

Why This Matters

You'll run into this in a few predictable places.

Laptops and anything that roams. Every reconnect is a NetworkManager activation, and every activation reapplies the zone from the profile. Runtime zone changes on Wi-Fi don't last long. If your security model depends on "I moved the interface to drop before I connected," it lasts until the first sleep/wake cycle.

Servers managed by config management. An Ansible role that runs firewall-cmd --reload wipes out every runtime-only change since the last reload. Hand-applied emergency rules disappear at the next scheduled run, and the incident comes back. Whatever your config management writes should be the only permanent state on the box, and emergency changes should go into that config management afterward, not just into runtime.

Debugging sessions that turn into production config. The opposite mistake is just as common. Someone runs --runtime-to-permanent to "save the fix" and also saves three experimental port openings. Diff before you promote.

The habits I'd recommend:

  • Pair every --permanent with a --reload, or use commands like --set-default-zone that write both configs.
  • On NM-managed interfaces, set zones with nmcli connection modify ... connection.zone, never with firewall-cmd --change-interface.
  • Treat the default zone as your policy for unknown networks, and set it to something restrictive.
  • Bind trust to interfaces like tailscale0, not to source ranges that can be spoofed from the local segment.
  • Diff runtime against permanent before every reload and every promotion.
  • Know which zones you've forked into /etc/firewalld/zones/, and keep those files under version control.

The underlying issue is that firewalld, NetworkManager, and the kernel's nftables ruleset each hold their own piece of state, and firewall-cmd reports success while updating only one of them. Once you know which component owns which setting, persistence becomes a matter of writing each setting to its owner. If you're dealing with this across a fleet of edge boxes or mixed Linux hosts and want a second set of eyes on the design, that's the kind of infrastructure work I take on.

Top comments (0)