Stop blindly restarting services, killing processes, and rebooting nodes. Here is the non-destructive triage sequence senior engineers run when systems degrade.
An alert fires in Slack at 3:15 AM. Latency on your production API gateway jumps from eighteen milliseconds to four seconds. PagerDuty starts triggering escalation policies, and error rates climb toward thirty percent.
What happens next in too many incident bridges?
Panic.
Someone opens an SSH session, runs systemctl restart nginx, watches the restart hang, slams Ctrl+C, immediately executes killall -9 nginx, runs echo 3 > /proc/sys/vm/drop_caches, and when the shell still feels sluggish, pulls the emergency brake: sudo reboot.
In sixty seconds of frantic typing, they destroyed volatile RAM evidence, severed open TCP connection states, wiped kernel ring buffers, aborted inflight database transactions, and flushed every breadcrumb required to identify what actually happened.
Worse yet, the underlying bottleneck remains untouched. If the root cause was a saturated network connection queue, a frozen NFS storage mount, or an exhausted inode table, the machine reboots and falls flat on its face five minutes later.
Troubleshooting under fire is not about executing loud commands to prove you are doing something. It is about disciplined observation before intervention.
Think of diagnostic triage as a ladder. You begin at the bottom rung with zero-impact, read-only telemetry. You measure pressure before you touch processes. You inspect wait states before you send signals. You only climb to the intervention rungs once you know the exact mechanism causing the degradation.
Here is the systematic troubleshooting ladder to run through before you make a production outage worse.
The "Do No Harm" Rules of Production Triage
Before touching the keyboard during a live outage, establish a perimeter. The goal during the first three minutes is stabilization and fact-finding, not heroic guesswork.
Four actions do almost permanent damage during initial triage:
First, never reboot a degrading machine. A reboot clears /proc, flushes volatile socket buffers, deletes temporary logs in memory-backed filesystems (/run and /tmp), and resets process runtimes. Unless the kernel is in a hard panic where the console is completely dead, rebooting is an admission that you gave up on finding the cause.
Second, do not blindly issue kill -9 (SIGKILL). When you send SIGKILL, the kernel immediately unmaps process memory without giving the application runtime an opportunity to flush write buffers, close socket descriptors, or release locks on shared storage. You risk database corruption and leave orphaned downstream state across your cluster.
Third, never run echo 3 > /proc/sys/vm/drop_caches. Junior admins often run this when free -m looks low. The Linux kernel uses unused memory for disk page cache to keep your storage reads fast. Dropping clean page caches forces the operating system to re-read binaries, shared libraries, and configuration files directly from physical disk, creating an immediate, massive storage I/O spike that can push a struggling server into a complete lockup.
Fourth, avoid recursive permission changes like chmod -R 777 or chown -R. If an application suddenly fails with permission errors, inspect the exact file and parent directory permissions. Running blanket recursive permissions destroys security baselines, breaks setuid binaries like sudo, and introduces security vulnerabilities you will spend weeks auditing later.
Rung 0: Verify Ground Truth and Blast Radius
When an alert triggers, start by verifying that the operating system itself agrees with your external monitoring. Monitoring tools and load balancers can report false outages if health-check endpoints time out due to external network jitter.
Log in and run two baseline commands:
uptime
Sample output:
03:15:22 up 142 days, 6:41, 2 users, load average: 28.45, 14.10, 6.20
Load average represents the number of tasks in an R (runnable) or D (uninterruptible sleep waiting for disk or I/O) state, averaged over one, five, and fifteen minutes.
To interpret load numbers, compare them against your available CPU cores:
nproc
If nproc returns 4 and your load average is 28.45, your run queues are backed up. Seven tasks are competing for every single core, or processes are queuing on storage operations. If your 1-minute load is 28 while your 15-minute load is 6, the problem is recent and escalating fast.
Immediately follow this by checking the kernel ring buffer for critical hardware, filesystem, or memory panics:
dmesg -T --level=alert,crit,err | tail -n 25
If the kernel killed a process due to out-of-memory pressure, dropped a network interface, or remounted an ext4 filesystem read-only because of journal corruption, dmesg tells you in plain text. You do not need to guess.
Inspect Pressure Stall Information (PSI)
Traditional load average is a blunt instrument. It mixes CPU demand with disk wait times, making it difficult to distinguish between compute saturation and storage bottlenecks.
Modern Linux kernels (version 4.20 and newer) provide Pressure Stall Information (PSI) via /proc/pressure/. These files show the exact percentage of wall-clock time that tasks spent waiting for resources over 10-second, 60-second, and 300-second windows.
Inspect all three subsystems in seconds:
head -n 2 /proc/pressure/{cpu,memory,io}
Sample output:
==> /proc/pressure/cpu <==
some avg10=42.15 avg60=18.40 avg300=8.10 total=3145820
==> /proc/pressure/memory <==
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0
==> /proc/pressure/io <==
some avg10=68.90 avg60=34.12 avg300=12.05 total=8492014
full avg10=54.20 avg60=22.80 avg300=7.40 total=5109400
Pay attention to the difference between some and full.
The some metric means at least one task was stalled on that resource. The full metric means all runnable tasks were blocked simultaneously, meaning the system made zero productive forward progress during that window.
In the snippet above, CPU has some pressure at 42 percent, but I/O pressure shows full stalls at 54 percent over the last 10 seconds. That single readout eliminates memory as a suspect and tells you that storage latency is choking every runnable thread on the system.
Rung 1: The Four Resource Pillars
Once you verify system pressure, inspect the four fundamental hardware resources: Memory, CPU, Storage Space, and Inodes.
1. Memory: Available vs Free
Open memory metrics using human-readable units:
free -h
Sample output:
total used free shared buff/cache available
Mem: 31Gi 18Gi 850Mi 1.2Gi 12Gi 11Gi
Swap: 4.0Gi 1.8Gi 2.2Gi
Do not make the classic mistake of looking at free (850Mi) and concluding the host is running out of RAM.
The free column represents memory that the kernel has zero use for at this moment. Linux intentionally allocates unused RAM to buff/cache for filesystem caching. If an application suddenly demands memory, the kernel evicts clean cache pages instantly.
The only metric that dictates immediate headroom is available (11Gi). This calculates how much memory can be allocated without causing the host to thrash swap space.
To check if memory pressure is actively destroying performance, verify whether the system is thrashing anonymous memory in and out of swap:
vmstat 1 5
Focus specifically on the si (swap in) and so (swap out) columns. If so consistently exceeds zero over several seconds, the kernel is desperately pushing memory pages to disk to prevent an out-of-memory crash. Memory access drops from nanosecond speeds to millisecond disk latency, stalling everything running on the machine.
2. CPU: Identify Steal Time and Wait States
A high CPU utilization percentage does not automatically mean your application code is running slowly. You must inspect where the CPU cycles are being spent:
mpstat -P ALL 1 1
Sample output:
15:18:01 CPU %usr %nice %sys %iowait %irq %soft %steal %guest %idle
15:18:02 all 12.50 0.00 14.20 45.10 0.00 2.20 26.00 0.00 0.00
15:18:02 0 14.00 0.00 16.00 42.00 0.00 3.00 25.00 0.00 0.00
15:18:02 1 11.00 0.00 12.40 48.20 0.00 1.40 27.00 0.00 0.00
Notice two critical fields here:
%iowait (45.10 percent): The CPU cores are completely idle, but they are idle because every runnable process is blocked waiting for physical disk transactions to return. Tuning code or adding worker threads will only increase contention.
%steal (26.00 percent): If you run on virtualized cloud infrastructure (AWS EC2, Google Compute Engine, or Azure VMs), steal time indicates that the physical hypervisor has descheduled your virtual CPU to serve another tenant on the same physical host. When steal time exceeds five to ten percent, you have a noisy-neighbor problem at the hypervisor layer. No application-level fix will resolve it. Your only recourse is migrating or resizing the instance.
3. Storage Space vs Inode Exhaustion
Every engineer knows how to check disk capacity:
df -h
If a partition reaches 100 percent, services fail to write temporary files, logs get dropped, and databases stop accepting transactions.
However, there is an insidious failure mode that causes applications to crash with No space left on device even when df -h shows hundreds of gigabytes of free disk space:
df -i
Sample output:
Filesystem Inodes IUsed IFree IUse% Mounted on
/dev/nvme0n1p1 1966080 1966080 0 100% /
An inode is a data structure on Unix filesystems that stores metadata about a file (ownership, access permissions, file size, block locations), excluding the file content itself.
When a filesystem is formatted, a fixed number of inodes are allocated. If an application (like a broken session handler, an email spool, or an unrotated micro-logger) generates millions of zero-byte files, every inode gets consumed.
Once inodes hit 100 percent (IFree = 0), the filesystem cannot create a single new file, directory, or socket descriptor, even if 95 percent of physical disk gigabytes are empty.
Find the directory hoarding inodes with this one-liner:
find /var -xdev -printf '%h\n' 2>/dev/null | sort | uniq -c | sort -k1 -rn | head -n 10
4. The Ghost File Trap: Open Deleted Descriptors
You log into a box with a full root volume (/var is at 100%). You find a massive 40GB log file at /var/log/app/trace.log and delete it using rm -f /var/log/app/trace.log.
You run df -h. The disk is still at 100 percent full.
Why?
In Unix semantics, removing a file with rm only removes the directory link. If a running process still holds an open file descriptor pointing to that file, the kernel keeps the allocated disk blocks active on storage until the process closes the descriptor or terminates.
Identify open deleted files holding disk space hostage:
lsof +L1 2>/dev/null | head -n 15
Alternatively, query /proc directly without external utilities:
find /proc/*/fd -ls 2>/dev/null | grep '(deleted)' | sort -k7 -rn | head -n 10
Sample output:
184920 lrwx------ 1 app app 64 Mar 15 03:20 /proc/4182/fd/7 -> /var/log/app/trace.log (deleted)
Process 4182 still has file descriptor 7 open. The disk space will not free up until you address that process.
Do not kill the process yet. In Rung 6, you will see how to reclaim that space safely in one second without restarting the daemon.
Rung 2: Process State and the Uninterruptible Sleep Trap
When a server becomes unresponsive, inexperienced responders often grep through ps aux and start firing kill -9 at whatever process shows the highest CPU or longest runtime.
Check process states first:
ps -eo pid,user,state,wchan:20,comm | grep -E ' (D|Z) '
Sample output:
4182 app D nfs_wait_client indexer
4205 app D sync_inodes_sb flush_daemon
9812 root Z - cleanup.sh <defunct>
Pay attention to process state D (Uninterruptible Sleep).
A process enters state D when it makes a system call that requires hardware interaction, usually reading or writing to a storage controller, waiting for a locked page in kernel memory, or awaiting a response from an NFS network share.
To maintain filesystem and hardware integrity, the Linux kernel puts the process to sleep and marks it as uninterruptible. That means the process cannot receive signals.
If you run kill -9 4182 on a process in state D, nothing happens. The process does not die.
Why? Because SIGKILL cannot be delivered until the process returns from the kernel syscall into user space. If the NFS server has hung, or if a block device driver is wedged waiting for a failing SCSI controller, the system call will never finish.
If you try to terminate multiple D-state processes, your shell commands will start queuing behind the same locked kernel mutexes, turning a slow server into an unresponsive brick.
Peeking Inside the Blocked Kernel Stack
Instead of attacking the process, read its kernel call stack to identify what resource it is waiting for:
cat /proc/4182/stack
Sample output:
[<0>] nfs_wait_client+0x9e/0x130 [nfs]
[<0>] nfs4_proc_lookup+0x18b/0x290 [nfsv4]
[<0>] __lookup_slow+0x7b/0x130
[<0>] walk_component+0x12a/0x1b0
[<0>] path_lookupat+0x6d/0x1a0
[<0>] filename_lookup+0xbb/0x190
[<0>] vfs_statx+0x72/0x120
[<0>] __do_sys_newstat+0x39/0x70
[<0>] do_syscall_64+0x5b/0x120
[<0>] entry_SYSCALL_64_after_hwframe+0x44/0xa9
This stack trace gives you instant clarity. The application is running stat() on a file located on a mounted NFS export, and the remote NFS client is unresponsive (nfs_wait_client).
You now know that killing the local application is useless. The actual failure is downstream on the storage network.
Rung 3: Network Sockets, Backlogs, and Connection Limits
If CPU, memory, and disk show normal numbers, the bottleneck is almost always in the network stack: socket exhaustion, saturated backlog queues, or connection table limits.
Start by inspecting socket totals:
ss -s
Sample output:
Total: 34120
TCP: 48210 (estab 1420, closed 45200, orphaned 80, timewait 42100)
Transport Total IP IPv6
RAW 2 1 1
UDP 14 8 6
TCP 3010 2410 600
Notice the timewait count: 42,100 sockets.
When a service makes thousands of short-lived outbound TCP connections (for instance, a microservice making HTTP requests without persistent connection pooling), each closed connection stays in TIME_WAIT for 60 seconds (two times the Maximum Segment Lifetime, or 2MSL).
If outbound requests happen faster than sockets expire from TIME_WAIT, the operating system exhausts its ephemeral port range (cat /proc/sys/net/ipv4/ip_local_port_range), and new outbound calls begin failing with Cannot assign requested address.
Decode Recv-Q and Send-Q on Listening Sockets
Run ss to inspect your active listening daemons:
ss -lnt
Sample output:
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 512 128 0.0.0.0:8080 0.0.0.0:*
LISTEN 0 511 0.0.0.0:80 0.0.0.0:*
The Recv-Q and Send-Q columns have a completely different meaning on listening sockets compared to active, established connections:
On an established connection, Recv-Q is bytes waiting to be read by the local app, and Send-Q is unacknowledged bytes waiting in the outbound network buffer.
On a listening socket, Send-Q represents the maximum listen backlog queue size (the maximum number of fully established TCP handshakes the kernel will hold before dropping new ones). Recv-Q represents the number of established connections currently waiting in that queue for the application to execute accept().
Look closely at port 8080 above:
Recv-Q is 512, while Send-Q is 128.
The queue is overflowing. The application's worker threads or event loop are blocked or saturated, failing to pull incoming TCP handshakes off the socket queue fast enough.
Incoming client connections reaching port 8080 are now getting silently dropped or rejected with TCP reset (RST) packets.
Verify connection drops in the kernel network statistics:
nstat -az TcpExtListenOverflows TcpExtListenDrops
If TcpExtListenOverflows is ticking up rapidly, your application runtime is choked, not the network card.
The Netfilter Connection Tracking Trap
If your Linux host runs Docker, Kubernetes, or complex iptables / nftables firewall rules, the kernel tracks connection states using the nf_conntrack module.
Every single TCP, UDP, and ICMP stream passing through the host consumes an entry in an in-memory hash table.
Check current tracking table usage against its absolute ceiling:
cat /proc/sys/net/netfilter/nf_conntrack_count
cat /proc/sys/net/netfilter/nf_conntrack_max
If nf_conntrack_count reaches nf_conntrack_max, the Linux kernel drops all new incoming and outgoing connections instantly.
Your CPU can sit at five percent idle, memory can have twenty gigabytes free, and disk I/O can be zero, but your server will drop packets indiscriminately.
Check for this silent killer in the kernel log:
dmesg -T | grep -i "table full, dropping packet"
Rung 4: Storage Subsystem Latency and I/O Bottlenecks
Throughput in megabytes per second rarely kills application servers. Latency and queued transactions are what bring databases and microservices down.
Inspect active disk latency across every physical block device:
iostat -xz 1 3
Sample output:
Device r/s w/s rkB/s wkB/s aqu-sz r_await w_await %util
nvme0n1 12.00 1280.00 48.0 84500.0 14.20 2.10 48.50 98.40
Three columns tell the entire story:
r_await and w_await: The average time in milliseconds that read and write requests took from when they were queued in the block layer to when the physical disk finished servicing them. Solid-state NVMe drives should rarely exceed 0.5 to 2.0 milliseconds. If w_await climbs to 48.50 milliseconds, write requests are queuing up forty times slower than normal.
aqu-sz (Average Queue Size): The number of requests waiting in the device driver queue. If this number climbs well above the number of concurrent channels supported by your hardware, disk contention is saturated.
%util: The percentage of time the device was busy servicing requests.
A warning on modern hardware: %util can be misleading on high-end NVMe drives and cloud storage volumes (like AWS EBS gp3 or io2) because they process dozens of I/O queues concurrently. A device can show 100 percent utilization while still having bandwidth available.
Always correlate %util with r_await, w_await, and aqu-sz. High wait times combined with high queue size confirm real physical storage saturation.
Detecting Silent Read-Only Remounts
When block storage controllers encounter hardware communication timeouts or unrecoverable journal errors, the kernel triggers an emergency safety mechanism: it remounts the filesystem as read-only (ro) to prevent structural filesystem corruption.
Check if any mounted filesystem flipped to read-only mode:
grep -E ' (ro|read-only)[, ]' /proc/mounts
If your root partition or database mount shows ro, write operations immediately fail with Read-only file system.
Do not try to force a read-write remount (mount -o remount,rw /) without checking dmesg. If the underlying disk or virtual volume experienced hardware corruption, forcing writes can permanently destroy filesystem metadata.
Rung 5: DNS and Local Resolution Traps
Many production outages diagnosed as "application crashes" or "API failures" are nothing more than local DNS resolution failures.
When an application attempts to connect to an external payment processor, a database cluster, or a cloud metadata service, it makes system calls to resolve hostnames via the glibc resolver.
If DNS lookups take two to five seconds to time out, application worker threads block. Connection pools exhaust themselves, thread pools reach max capacity, and the application stops answering requests on its health-check endpoints.
Inspect your current resolver configuration:
cat /etc/resolv.conf
Test lookup latency directly against the configured nameserver:
time dig @127.0.0.53 api.internal.service.net +time=2 +tries=1
If your server uses systemd-resolved (standard on modern Ubuntu and Debian installations), check its runtime health and dropped cache queries:
resolvectl status
resolvectl statistics
Look for transaction timeouts or failing upstream DNS links. If systemd-resolved loses connectivity to an upstream nameserver configured via DHCP, it can sit in a retry loop while your applications wait and time out.
Rung 6: Controlled Intervention (When You Finally Take Action)
You completed your inspection. You identified the bottleneck without destroying evidence or restarting nodes.
Now, and only now, you move to remediation.
Climbing to the intervention rung requires discipline. Execute changes in order of increasing impact:
1. Freeze a Runaway Process Before Killing It
If an errant process or runaway batch script is hammering CPU or disk I/O, do not immediately send SIGKILL.
Pause it first:
kill -SIGSTOP 4182
SIGSTOP halts the process execution instantly without terminating it. CPU consumption drops to zero. Disk I/O ceases immediately.
Because the process remains in memory, its file descriptors, network connections, memory mappings, and thread stacks stay completely intact.
You can now safely attach gdb or strace to extract evidence, verify open connections using lsof -p 4182, or examine /proc/4182/ at your own pace while the server recovers.
Once you have documented the cause, either resume it cleanly:
kill -SIGCONT 4182
Or terminate it gracefully:
kill -SIGTERM 4182
Only if it refuses to terminate after ten to fifteen seconds should you reach for kill -SIGKILL 4182.
2. Safely Reclaim Space from Deleted Open Files
Remember process 4182 holding open /var/log/app/trace.log (deleted) and consuming 40GB of storage?
You do not need to restart the application or kill the process to get that disk space back.
Truncate the file descriptor directly through /proc:
truncate -s 0 /proc/4182/fd/7
This single command tells the filesystem to instantly free every disk block assigned to that unlinked file, dropping disk utilization from 100 percent to normal levels immediately.
The application process keeps running without interruption, and the production crisis is averted without bouncing the daemon.
3. Graceful Reloads vs Hard Restarts
If a configuration change or memory leak requires refreshing an application daemon, always check if the service supports graceful reloads:
systemctl reload nginx
A graceful reload signals the master process to parse the new configuration, spin up clean worker processes, and allow existing workers to finish servicing inflight HTTP requests before shutting down.
A hard systemctl restart forcibly drops active connections, spikes upstream error rates, and causes connection retries that can trigger a thundering herd on downstream databases.
An Interesting Fact About Linux Diagnostics
The /proc virtual filesystem, which powers almost every non-invasive diagnostic tool mentioned on this ladder, was not originally a Linux invention.
Tom Killian first implemented /proc in Version 8 Unix at Bell Labs in 1984. His original implementation contained only process IDs as files, allowing debuggers to inspect process memory without using arcane kernel debugging system calls.
When Linus Torvalds and early Linux contributors expanded /proc in the early 1990s, they transformed it from a simple process directory into an entire synthetic window into kernel internals. None of the files in /proc exist on physical storage. When you run cat /proc/pressure/io or inspect /proc/<pid>/stack, the kernel intercepts your VFS read request and generates text representation on the fly from active kernel memory structures.
This architectural decision is what allows modern engineers to inspect kernel locks, socket states, and hardware pressure in production with zero disk overhead and near-zero CPU footprint.
The 60-Second Triage Cheatsheet
When an alert pages you in the middle of the night, keep this sequence pinned to your monitor:
- Rung 0: Verify Ground Truth
- Check load versus cores:
uptimeandnproc - Check kernel ring buffer:
dmesg -T --level=alert,crit,err | tail -n 25 - Check resource stalls:
head -n 2 /proc/pressure/{cpu,memory,io}
- Check load versus cores:
- Rung 1: Check Resource Exhaustion
- Check memory headroom:
free -h(look atavailable) - Check swap thrashing:
vmstat 1 5(look atsiandso) - Check CPU breakdown:
mpstat -P ALL 1 1(watch%iowaitand%steal) - Check storage space and inodes:
df -handdf -i - Check open deleted files:
find /proc/*/fd -ls 2>/dev/null | grep '(deleted)'
- Check memory headroom:
- Rung 2: Check Process States
- Find stuck processes:
ps -eo pid,user,state,wchan:20,comm | grep -E ' (D|Z) ' - Inspect blocked kernel stack:
cat /proc/<pid>/stack
- Find stuck processes:
- Rung 3: Check Sockets and Queues
- Check socket totals:
ss -s - Check listen backlogs:
ss -lnt(watch ifRecv-Q > Send-Q) - Check connection tracker limits:
cat /proc/sys/net/netfilter/nf_conntrack_count
- Check socket totals:
- Rung 4: Check Disk Latency
- Measure real await times:
iostat -xz 1 3(watchr_await,w_await,aqu-sz) - Check for read-only remounts:
grep -E ' (ro|read-only)[, ]' /proc/mounts
- Measure real await times:
- Rung 5: Check DNS Health
- Test resolution latency:
time dig @127.0.0.53 <hostname> +time=2
- Test resolution latency:
- Rung 6: Controlled Action
- Freeze before killing:
kill -SIGSTOP <pid> - Reclaim deleted file space:
truncate -s 0 /proc/<pid>/fd/<fd> - Prefer reloads over restarts:
systemctl reload <service>
- Freeze before killing:
Wrapping Up
Production incidents test discipline far more than memory.
The temptation to type fast, restart services, and reboot machines comes from anxiety, not analysis. Every time you restart a service without diagnosing it, you buy twenty minutes of relief at the expense of another outage tomorrow.
Climb the ladder one step at a time. Start with non-invasive read-only metrics, understand the kernel wait states, and intervene with surgical precision.
What is the single worst command you have ever seen someone execute during a live production outage?
If this breakdown saved you hours of debugging or gave you something practical to use in production, consider buying me a coffee. Your support directly fuels independent, zero-fluff Linux and DevOps technical guides.
About the Author
Asep Sayyad is a Linux and DevOps engineer passionate about Linux administration, automation, cloud technologies, containers, and open-source software. He enjoys solving real-world infrastructure challenges and sharing practical knowledge through in-depth technical articles, tutorials, and hands-on guides.
His goal is to help aspiring and experienced engineers build stronger Linux and DevOps skills with content focused on real production scenarios rather than theory alone.
Connect with Me
- Portfolio: asepsayyad007.in
- Blog: asepsayyad007.in/blogs
- GitHub: github.com/asepsayyad007
- LinkedIn: linkedin.com/in/asepsayyad
- Medium: asepsayyad007.medium.com
- Support: buymeacoffee.com/asepsayyad007
Enjoyed this article?
If this guide saved you hours of debugging or gave you something practical for production, consider:
- Buying me a coffee: Your support directly fuels independent, zero-fluff Linux and DevOps engineering breakdowns.
- Starring my open-source projects on GitHub.
- Sharing this article with fellow Linux and DevOps engineers.
You can also follow me for more practical content on Linux, DevOps, Cloud, Containers, Automation, and Open Source. Thanks for reading, and enjoy your learning!
© 2026 Asep Sayyad
Top comments (0)