Your server works fine until it doesn't
Every infrastructure engineer knows this moment: the server hums along fine in staging, then a traffic spike hits production and everything falls over. Requests queue, workers pin at max, response times spike from 200ms to 8 seconds. The usual reaction is to throw more CPU at it. That's rarely the actual problem.
Most servers that collapse under load aren't under-resourced, they're misconfigured. The default settings for PHP-FPM, Nginx, and your database connections were never meant for production traffic. Here's the exact sequence of changes that fixes this, in the order that actually matters.
Step 0: measure before you touch anything
You can't prove a fix worked without a baseline. Run this first:
ab -n 1000 -c 50 https://yourdomain.com/
Record requests per second, mean response time, and failed request count. Check uptime and vmstat 1 5 during the test too. Write it all down; you'll compare against it later.
Fix 1: PHP-FPM is probably misconfigured
The default pool config is not production-ready. Open it:
sudo nano /etc/php/8.3/fpm/pool.d/www.conf
Size your workers based on actual memory usage, not guesswork. Check real footprint per worker with ps aux | grep php-fpm (typically 30-60MB for WordPress/Laravel), then divide available RAM by that number.
pm = dynamic
pm.max_children = 40
pm.start_servers = 10
pm.min_spare_servers = 5
pm.max_spare_servers = 15
pm.max_requests = 500
pm.max_requests is underrated. It recycles workers periodically, which stops slow memory leaks from degrading your server over hours of uptime.
Fix 2: Nginx is dropping connections it doesn't need to
Reduce TCP handshake overhead with proper keepalive and connection settings:
worker_processes auto;
worker_rlimit_nofile 65535;
events {
worker_connections 4096;
use epoll;
multi_accept on;
}
http {
keepalive_timeout 65;
keepalive_requests 1000;
sendfile on;
tcp_nopush on;
tcp_nodelay on;
}
If you're proxying to PHP-FPM over a Unix socket, bump the backlog to match expected concurrency:
listen = /run/php/php8.3-fpm.sock
listen.backlog = 1024
Fix 3: your database connections are too expensive
Each unpooled connection costs a handshake, auth round trip, and memory allocation. At scale, this is a silent killer. For Postgres, PgBouncer is the standard fix:
sudo apt install pgbouncer
[databases]
yourapp = host=127.0.0.1 port=5432 dbname=yourapp
[pgbouncer]
listen_port = 6432
auth_type = md5
pool_mode = transaction
max_client_conn = 500
default_pool_size = 25
Point your app at port 6432 instead of 5432. This typically cuts connection overhead by 60-80% under concurrent load. MySQL users: look at ProxySQL or mysqlnd_mux.
Fix 4: cache what shouldn't hit the database every time
Install Redis:
sudo apt install redis-server
sudo systemctl enable redis-server
Set eviction policy so it doesn't just run out of memory:
maxmemory 512mb
maxmemory-policy allkeys-lru
For custom apps, wrap expensive queries in a cache-aside pattern:
$cacheKey = "product:{$id}";
$product = $redis->get($cacheKey);
if (!$product) {
$product = $db->query("SELECT * FROM products WHERE id = ?", [$id]);
$redis->setex($cacheKey, 300, serialize($product));
}
Fix 5: stop static assets from hitting your app server
location ~* \.(jpg|jpeg|png|gif|css|js|woff2)$ {
expires 30d;
add_header Cache-Control "public, immutable";
}
If there's a CDN in front, confirm it respects your origin headers instead of overriding with its own default TTL.
Verifying it actually worked
Re-run the same load test, same concurrency:
ab -n 1000 -c 50 https://yourdomain.com/
Compare against baseline. RPS should climb, often 30-100%. Response time variance between p50 and p95 should tighten. Failed requests should hit zero at your target concurrency.
Watch PHP-FPM workers scale during load:
watch -n 1 'ps aux | grep php-fpm | wc -l'
Workers should scale up smoothly and settle, not spike immediately to max_children and stay there. If they pin instantly, either your sizing is wrong or a slow query is holding connections open too long.
Check PgBouncer is actually in the path:
psql -h 127.0.0.1 -p 6432 -U youruser yourapp -c "SHOW POOLS;"
And check Redis hit rate:
redis-cli info stats | grep keyspace
Healthy cache-aside logic should show above 80% hit ratio within minutes.
Mistakes that undo all of this
- Oversizing max_children. More workers than RAM supports causes swapping, which is worse than queued requests.
- Caching without invalidation. Stale cached data is a harder bug to catch than a slow endpoint. Always set a TTL or invalidate on write.
- Pooling without adjusting app-level assumptions. Transaction-mode pooling breaks session-level features like advisory locks if your app still expects a persistent connection.
- Skipping the baseline. No before/after numbers means no proof anything worked.
This isn't a full re-architecture. It's five configuration changes applied in order, with measurement at each step. That's usually enough to take a server from falling over at the first spike to handling concurrency predictably.
Full walkthrough with more context: How to set up website server performance that holds under real traffic
Originally published on binadit.com
Top comments (0)