The DevOps bootcamp · Chapter 2 of 14 · DevOps · new chapter every Thursday morning
By the end of this chapter: Work confidently with processes, files, permissions and networking.
The problem
Say you are a junior engineer on your first backend project. You need to test a branch on the team's shared staging server. You SSH in, pull your code, and run the start command. The terminal spits back Error: listen EADDRINUSE: address already in use :::8080.
You try changing the port, but the proxy routing traffic to your application expects 8080. You assume the previous version of the application is still running, but you do not know how to find it. You try running your new application as the superuser to force it through, and it starts, but immediately crashes with Permission denied: /var/log/staging-api/error.log.
Because you do not know how to ask Linux what is holding that port, or how to read the permissions on that file, you do the only thing you know: you reboot the staging server. It takes ten minutes to come back. When it does, your application works, but you have just taken down three other services your team was actively testing, costing four engineers an hour of lost work and earning a sharp message from your tech lead.
Before you start
You need a Linux environment. Ubuntu 24.04 LTS is the current standard for server deployments. You can use Windows Subsystem for Linux (WSL2), a local virtual machine, or a cloud instance.
Verify your environment by checking the operating system release:
cat /etc/os-release
The output must include PRETTY_NAME="Ubuntu 24.04 LTS".
Verify you have standard utilities installed:
python3 --version
The output should be Python 3.12.x.
How do you find what is holding a port?
When an application claims a network port, the Linux kernel prevents any other process from claiming it. If a deployment script fails, if an application crashes poorly, or if a process is sent to the background and forgotten, the process remains alive and the port remains locked.
You cannot ask the port to release the process. You must ask the kernel which process ID (PID) owns the port, and then terminate that PID.
The tool for querying network sockets is ss (socket statistics). By asking it to list listening TCP ports and the processes attached to them, you get the exact PID blocking your deployment. You must run this command as the superuser (sudo); if you run it as a standard user, the kernel will hide the PIDs of processes owned by other users, leaving the process column blank and you guessing.
Once you have the PID, you use the kill command. Despite the name, kill does not kill processes; it sends signals to them. A standard kill <PID> sends SIGTERM (signal 15), which politely asks the application to shut down. A well-written application will catch this, finish writing its logs, release its port, and exit. A hung application will ignore it. When that happens, you use kill -9 <PID>, which sends SIGKILL. This signal goes directly to the kernel, which instantly destroys the process without giving the application a chance to react.
Why does root own the log file?
Linux treats almost everything as a file, and every file has an owner, a group, and a strict set of permissions. Permissions are divided into read (r), write (w), and execute (x). They are applied to three categories: the user who owns the file, the group assigned to the file, and everyone else.
When you run an application with sudo, it executes as root (user ID 0). Any file it creates is owned by root. If you realise your mistake and later try to run that same application as a standard user, the application will attempt to append to its existing log file, discover it does not have write permissions to a root-owned file, and crash.
You view these permissions using ls -l. The output begins with a ten-character string, such as -rw-r--r--.
- The first character is the type (
-for a file,dfor a directory). - The next three are the owner's permissions (
rw-means read and write, but not execute). - The next three are the group's permissions (
r--means read only). - The final three are everyone else's permissions (
r--means read only).
You change who owns a file with chown (change owner). You change the permissions themselves with chmod (change mode). While you can use letters to change permissions, engineers use octal numbers because they are exact and overwrite the previous state entirely. Read is worth 4, write is worth 2, and execute is worth 1. You add them together for each category.
A permission of 755 means the owner gets 7 (4+2+1: read, write, execute), the group gets 5 (4+1: read, execute), and everyone else gets 5. A permission of 644 means the owner gets 6 (4+2: read, write), and everyone else gets 4 (read only).
Build it
We will recreate the staging server failure on your machine, diagnose the port collision, kill the ghost process, and fix the resulting permission trap.
Step 1: Create the ghost process
Start a dummy web server in the background, bound to port 8080. The & at the end tells the shell to run this in the background, returning control of the terminal to you.
python3 -m http.server 8080 > /dev/null 2>&1 &
Why: This simulates the previous version of your application that failed to shut down.
How to confirm: Run curl http://localhost:8080. It will output a directory listing of your current folder.
Step 2: Identify the blockage
Query the kernel for all listening TCP ports and the processes holding them.
sudo ss -lptn | grep :8080
Why: l shows listening sockets, p shows processes, t shows TCP only, and n forces numeric output (preventing Linux from trying to resolve 8080 to a service name, which is slow). We pipe (|) the output to grep to filter for our specific port.
How to confirm: You will see output resembling:
LISTEN 0 5 0.0.0.0:8080 0.0.0.0:* users:(("python3",pid=14235,fd=3))
The number after pid= (in this case, 14235) is what you need.
Step 3: Terminate the ghost process
Send a kill signal to the PID you found in Step 2. Replace 14235 with your actual PID.
kill -9 14235
Why: We use -9 to force the kernel to destroy the process immediately, simulating clearing a hung application.
How to confirm: Run sudo ss -lptn | grep :8080 again. The output will be completely empty. The port is free.
Step 4: Create the permissions trap
Create a log directory and file as the root user.
sudo mkdir -p /var/log/staging-api
sudo touch /var/log/staging-api/error.log
Why: This simulates the moment you tried to force your application to start using sudo, leaving behind root-owned state.
How to confirm: Run ls -l /var/log/staging-api/error.log. The output will show root root as the owner and group.
Step 5: Trigger the failure
Attempt to write to the log file as your standard user.
echo "Starting application..." >> /var/log/staging-api/error.log
Why: This simulates your application trying to start up and write its first log line.
How to confirm: The terminal will output exactly: bash: /var/log/staging-api/error.log: Permission denied.
Step 6: Fix the ownership and run successfully
Change the ownership of the directory and its contents to your current user. $USER is an environment variable containing your username.
sudo chown -R $USER:$USER /var/log/staging-api
Why: The -R flag applies the change recursively to the directory and everything inside it. The $USER:$USER syntax sets both the owner and the group to your user.
How to confirm: Run the application write command again:
echo "Starting application..." >> /var/log/staging-api/error.log
cat /var/log/staging-api/error.log
The output will successfully print Starting application.... You have cleared the port, fixed the permissions, and started the service without rebooting the machine.
When this breaks
Symptom: ss: command not found
Cause: You are running a minimal Linux distribution, often inside a container, that has stripped out network utilities to save space.
Fix: Install the package that provides ss. On Ubuntu/Debian, run sudo apt-get update && sudo apt-get install iproute2.
Symptom: kill -9 returns silently, but the process still appears in ss or ps.
Cause: The process is a zombie (listed as Z state in ps) or in uninterruptible sleep (listed as D state). A zombie is already dead, but its parent process has not acknowledged its death. Uninterruptible sleep means the process is waiting on hardware, like a hung network drive, and the kernel will not deliver the kill signal until the hardware responds.
Fix: For a zombie, you must find and kill its parent. Find the parent PID (PPID) with ps -o ppid= -p <PID>, then kill the parent. For a process in uninterruptible sleep waiting on dead hardware, you cannot kill it. You must reboot the machine.
Symptom: bash: /var/log/staging-api/error.log: Permission denied even when you run sudo echo "test" >> /var/log/staging-api/error.log.
Cause: The shell parses the command line before executing it. It sees the redirect (>>) and attempts to open the file using your standard user permissions before it runs sudo echo. The sudo only applies to the echo command, not the file redirection.
Fix: Pass the string to the tee command, running tee as root.
echo "test" | sudo tee -a /var/log/staging-api/error.log
What it costs
Manual Linux administration is imperative: you are issuing commands that change the state of the machine. The cost of this approach is configuration drift.
If you fix a permissions error on a staging server by running chown, you have fixed it once, on one machine. The next time the server is rebuilt, or when you attempt to deploy to production, the error will return. The server's true configuration now exists only in the command history of your terminal, not in source control.
This commits you to remembering what you typed. It scales poorly, guarantees failures during high-stress deployments, and turns servers into fragile pets that engineers are afraid to replace. We learn these commands so we can diagnose failures and understand what the operating system requires, but we do not use them to deploy software.
In the interview
A standard infrastructure interview question is: "You are deploying a new version of a backend service to a Linux machine, and it fails to start. Walk me through how you debug it."
A weak answer is: "I would check the application logs, and if that doesn't work, I would restart the server." This is weak because it assumes the application lived long enough to write logs (which it cannot do if the port is blocked or the log directory is restricted), and it treats restarting a production server as a valid diagnostic step.
A strong answer names the specific sequence of investigation. You check process status with ps or systemctl, check port bindings with ss -lptn, check file permissions with ls -la, and read the application's standard error output directly.
The follow-up questions will change based on the level you are interviewing for. A junior candidate is expected to know the commands (ss, kill, chmod) and what they do. A senior candidate is expected to explain the lifecycle: why the old version did not release the port (perhaps the deployment script failed to send a SIGTERM, or the application does not handle signals gracefully) and why the permissions are wrong. An engineering manager or staff-level loop will probe the systemic failure: why is a human SSHing into a machine to deploy in the first place, and how do we move this process to an automated, immutable pipeline?
Your tasks
-
Find the SSH daemon. Find the PID of the process listening on port 22.
Done looks like: You have a number. Running
ps -p <number>outputs a line ending insshd. -
Create a read-only secret. Create a file named
secret.txt. Change its permissions so that only your user can read it, and absolutely nobody (including you) can write to it or execute it. Done looks like: Runningls -l secret.txtshows-r--------. Runningecho "test" > secret.txtfails withPermission denied. -
Hunt a disconnected process. Start a background process by running
sleep 3600 &. Close your terminal completely to log out of the shell. Open a new terminal, log back in, find that specific sleep process, and kill it. Done looks like: Runningps aux | grep sleepshows nothing after you have killed it.
Your tasks this week
Do the exercises above before the next chapter. Reading a tutorial and doing
one are different activities and only one of them changes what you can build.
Stuck on any of them? Say so — describe what you tried and what happened:
tell me where you got stuck. I read every one, and the questions
that come back more than twice get answered in the next chapter.
The DevOps bootcamp
Chapter 2 of 14. New chapter every Thursday morning.
Next: The shell: pipes, scripts and automating your own machine.
· The full syllabus and every chapter so far
· Subscribers also get the condensed notes for this chapter, the running
recap of everything the series has covered, and the extended guidance:
subscribe
Written by Amit Chakraborty — founding engineer and senior architect: React Native, AI and RAG systems, production architecture. Portfolio · LinkedIn · GitHub.
Need this built, reviewed or taught to your team? Get in touch or email amit@devamit.co.in. Available for senior and founding engineering roles, consulting and training, remote worldwide.
Top comments (0)