DEV Community

Dal
Dal

Posted on Edited on

Recovering a Dead Ubuntu VPS After an IP Change: VNC Recovery, /32 On-Link Routing, and systemd-networkd

Incident Post-Mortem: Recovering a "Dead" Ubuntu VPS After an IP Change

A technical walkthrough of recovering a production Ubuntu 26.04 minimal server after a provider IP reallocation caused total SSH isolation.


1. Incident Overview & Environment

  • Server: germany-play (KVM Virtual Machine)
  • Base OS: Ubuntu 26.04 LTS (Minimal cloud template)
  • Previous IP: 2.26.48.117/32
  • Reallocated IP: 193.222.99.171/32
  • Assigned Gateway: 10.0.0.1

Failure State

Immediately following the IP address change in the hosting control panel:

  • Inbound SSH sessions timed out unconditionally.
  • Inbound ping tests showed 100% packet loss.
  • Outbound traffic from the VM could not resolve external names or reach internet destinations.

2. Gaining Console Access via Out-of-Band VNC

Because network-level access (SSH) was non-functional, management had to be performed out-of-band via the provider's hypervisor web console (VNC).

Inspecting the network interface from the console showed:

root@germany-play:~# ip addr show ens3
2: ens3: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc fq_codel state UP group default qlen 1000
    link/ether 52:54:00:ca:68:fd brd ff:ff:ff:ff:ff:ff
    altname enp0s3
    altname enx525400ca68fd
    inet 2.26.48.117/32 scope global ens3
Enter fullscreen mode Exit fullscreen mode

The hypervisor bridge was routing traffic toward 193.222.99.171, but the guest operating system interface (ens3) remained bound to the decommissioned IP 2.26.48.117.


3. Root Cause Analysis

Minimal Ubuntu cloud templates often omit dynamic configuration layers:

  1. cloud-init was not configured to re-probe network metadata dynamically.
  2. netplan was absent from the installation image.
  3. NetworkManager was not installed.
  4. systemd-networkd was present but disabled (inactive (dead)).

The Point-to-Point /32 Subnet Constraint

In public cloud setups utilizing /32 host routes, the assigned gateway (10.0.0.1) does not reside within the broadcast subnet of the assigned IP (193.222.99.171/32).

Attempting to add a default route using traditional syntax produces an error:

RTNETLINK answers: Network is unreachable
Enter fullscreen mode Exit fullscreen mode

To instruct the Linux kernel that the gateway is reachable directly over the link layer despite differing subnets, the route must explicitly declare the onlink attribute.


4. Troubleshooting Obstacles

Ephemeral Commands vs. Persistence

Executing runtime commands:

ip addr add 193.222.99.171/32 dev ens3
ip route add default via 10.0.0.1 dev ens3 onlink
Enter fullscreen mode Exit fullscreen mode

Restored instant connectivity (sub-millisecond ping to 8.8.8.8 and active SSH). However, these settings exist solely in kernel runtime state and disappear upon system reboot.

Syntax Placement in systemd-networkd

When writing /etc/systemd/network/10-ens3.network, placing the GatewayOnLink=yes directive inside the [Network] section resulted in systemd ignoring the key:

10-ens3.network:9: Unknown key 'GatewayOnLink' in section [Network], ignoring.
Enter fullscreen mode Exit fullscreen mode

In systemd-networkd, GatewayOnLink is valid exclusively within a dedicated [Route] block.

Stale File Definitions

Residual references in /etc/hosts and /etc/network/interfaces caused dual-IP binding upon reboot, leaving the decommissioned IP active alongside the new address.


5. Verified Permanent Configuration

The following configuration provides persistent network initialization across reboots without requiring third-party network managers.

Step 1: Create the Network Unit

Write /etc/systemd/network/10-ens3.network:

[Match]
Name=ens3

[Network]
Address=193.222.99.171/32
DNS=8.8.8.8
DNS=1.1.1.1

[Route]
Gateway=10.0.0.1
GatewayOnLink=yes
Enter fullscreen mode Exit fullscreen mode

Step 2: Synchronize Legacy Files & Purge Stale Addresses

# Update host definition
sed -i 's/2.26.48.117/193.222.99.171/g' /etc/hosts

# Update legacy interface file
sed -i 's/2.26.48.117/193.222.99.171/g' /etc/network/interfaces

# Remove decommissioned IP from active runtime
ip addr del 2.26.48.117/32 dev ens3
Enter fullscreen mode Exit fullscreen mode

Step 3: Enable the Network Service

systemctl enable systemd-networkd
systemctl restart systemd-networkd

# Verify kernel routing table
ip route
# default via 10.0.0.1 dev ens3 proto static onlink
Enter fullscreen mode Exit fullscreen mode

6. Explaining the External ICMP Timeout

Following recovery, global ICMP ping diagnostics (such as check-host.net) continued to report packet drops, even while SSH and outbound internet functioned normally.

Examining firewall policies:

root@germany-play:~# iptables -L -n
Chain INPUT (policy ACCEPT)
Chain FORWARD (policy ACCEPT)
Chain OUTPUT (policy ACCEPT)
Enter fullscreen mode Exit fullscreen mode

Reason: Inbound ICMP echo requests are dropped upstream by the datacenter anti-DDoS scrubbing layer. Testing TCP connectivity on service ports (port 22) confirmed complete reachability with sub-millisecond round-trip response and zero dropped connections.


7. Operational Takeaways

  • Minimal cloud templates require verifying the active network management subsystem before performing changes.
  • Out-of-band hypervisor access (VNC/Serial) must be confirmed before reallocating production IP addresses.
  • Point-to-point /32 setups require explicit onlink flags for off-subnet gateways.
  • External ICMP timeouts do not indicate service downtime; always validate through TCP port handshakes.

Top comments (0)