DEV Community

Cover image for He Put 5,000 Users on a 12-Dollar Server. The Database Was Fine.
Chizee
Chizee

Posted on Originally published at omenabyte.com AI-assisted

He Put 5,000 Users on a 12-Dollar Server. The Database Was Fine.

The server cost twelve dollars a month. One CPU core, two gigabytes of RAM.

On it, a developer built a Twitter-style social API and loaded it with 50,000 users, 500,000 posts, and two million likes. The whole thing ran on that one machine: nginx, Node, and Postgres, no separate database box.

Then he rented a second server for four hours and used it to attack the first one with k6, spinning up virtual users by the thousand. Each one ran a loop that looks roughly like a person. Load the feed, wait a second or two, like something, maybe post something, load the feed again.

Ten users did nothing. So did fifty, a hundred, two hundred, a thousand.

At two thousand, the tail latency started moving. The average was still clean. P95 and P99 were climbing.

At three thousand he'd crossed his own failure threshold and backed down to 2,500. That was the ceiling: roughly 232 requests per second, P95 around 288 milliseconds.

Here's the part worth sitting with. RAM was still under a gigabyte. Postgres was completely fine. The CPU was at 90%.

The bottleneck wasn't where he expected it

Most people building a social app spend their early architecture budget on the database. Replicas, connection poolers, a caching layer, maybe a queue. The assumption is that the data tier buckles first under real concurrency.

Not here. At the limit, the database was sitting idle-ish and the single CPU core was the wall.

That makes sense once you look at what each request costs. Assembling a feed means parsing a request, authenticating it, querying, building objects, and serializing JSON. All of that is CPU work. Every request takes a fixed slice of that one core. Divide the load test's throughput into its CPU budget and you land at roughly four milliseconds of CPU per request.

Four milliseconds is the number that sets the ceiling. Nothing about it involves Postgres.

It also explains why the averages looked fine while the tail didn't. A saturated CPU queues work. Most requests arrive when the queue is short and get served quickly. The ones that land behind a backlog wait proportionally longer. Averages hide that, because the queue is only deep some of the time.

Every request costs a fixed slice of CPU: parsing, auth, query building, JSON serialization. At the ceiling th

Every request costs a fixed slice of CPU: parsing, auth, query building, JSON serialization. At the ceiling that worked out to roughly four milliseconds per request. Photo: Martijn Boer / Public domain

And nobody experiences your average response time. They experience the worst request in their session. If a page fires twenty requests and your P99 is three seconds, that user has a real chance of feeling the app stall on every load. P95 and P99 aren't vanity metrics for dashboards nobody reads.

Two changes, roughly double the capacity

He resisted the obvious move. No Redis, no migration to a bigger box.

First he cached the feed response inside Node, with a one-second TTL. That's the whole change. His reasoning was practical: if you're running one server anyway, you already have the memory, and using it is free. You don't need to introduce a separate service to hold a value for one second.

It took him from 2,500 to about 4,000 concurrent users.

Why one second works so well comes down to traffic concentration. The feed is the same response for everyone during that window. If most of your 232 requests per second are asking for the same thing, collapsing thousands of identical requests into a handful of database queries and render passes removes almost all of that CPU cost at once. The gain is proportional to how much your traffic asks for the same thing. A global feed caches beautifully. Per-user data caches far worse.

Turns out you do not need much machine. The whole API stack ran on one core and two gigs of RAM while serving

Turns out you do not need much machine. The whole API stack ran on one core and two gigs of RAM while serving thousands of concurrent users. Photo: Jwrodgers / CC BY-SA 3.0

Then he moved that cache out of Node and into nginx. This added less than the first change did, which is worth noting: the easy win was the one inside the application, and the second change bought the remaining thousand. When nginx answers from its own cache the request never enters the Node process at all, so there's no request parsing, no middleware chain, no JSON serialization, and no garbage collection pressure from thousands of response objects. The CPU that was pegged at 90% is only busy on cache misses now.

That pushed him past 5,000 concurrent users on the same twelve-dollar machine.

Moving the cache out of Node and into the reverse proxy meant cached requests never entered the application pr

Moving the cache out of Node and into the reverse proxy meant cached requests never entered the application process at all. Photo: RAIL P (RAIL.PHOTOGRAPHY) / CC0

His rule for this is the useful part: move the cache as close to your users as you can reasonably get it. It sounds like a slogan until you've watched it happen. Every hop you remove from the request path is CPU you don't spend, and CPU was the entire problem.

Amber is where the server started. The first cache bought 1,500 users; moving it into nginx bought the next 1,

Amber is where the server started. The first cache bought 1,500 users; moving it into nginx bought the next 1,000. The fading tail on the last bar is deliberate: the test stopped at 5,000, so that end is a floor and not a ceiling.

Chart by Omenabyte Intelligence. Figures as measured by Arjaythedev.

The extrapolation, and where it gets shaky

Five thousand concurrent users doesn't mean five thousand customers. Real people don't all use an app at the same instant. Based on typical usage patterns he estimates that concurrency level would correspond to somewhere around 25,000 to 30,000 daily active users.

That's the honest version of the headline, and it's worth being careful with. "A twelve-dollar server handles 5,000 users" is true for this application, this load profile, and this definition of handling. Swap the workload for video transcoding, per-request image resizing, or LLM inference and the number moves by orders of magnitude, because those burn vastly more CPU per request. The transferable part of this experiment isn't the number. It's the method, and the two fixes.

The method deserves more attention than it usually gets. He measured before he optimized. He found the actual bottleneck instead of the assumed one, and he fixed that specific thing. Most teams never do the first step, and then spend real money on the second.

What the video skips

A single-machine setup that passes a load test still has problems the test doesn't cover.

When a one-second cache entry expires, every concurrent request misses at the same moment and they all stampede the database together. nginx has a real fix for this: proxy_cache_lock on. When an entry expires, exactly one request goes upstream while the rest wait for it and then get served from the newly filled cache. Hand-rolled in-process caches rarely bother, which is why a one-second TTL on a hot key can produce a nasty sawtooth in your database load. The cache helps on average and hurts in bursts.

In-process caches also don't survive a restart. A deploy empties them, and if a cold cache means every request hits the database at once, your deploy just created a self-inflicted traffic spike. Run Node in cluster mode or more than one container and each process holds its own copy with its own idea of the truth. Fine for one global feed. Actively wrong for sessions, rate-limit counters, or anything per-user.

The migration this experiment argues against. Big iron solves the problem money can solve, not the one measure

The migration this experiment argues against. Big iron solves the problem money can solve, not the one measurement would have found. Photo: Christopher Bowns / CC BY-SA 2.0

And one machine is one point of failure. The load test measured capacity. It measured nothing about availability. A deploy is downtime, a hardware fault is an outage, and there's no redundancy anywhere in that stack. Sometimes that's an acceptable trade for a product at this stage. It should be a decision you made on purpose rather than something you discover at 2am.

There's also the load model itself. k6 virtual users follow a fixed loop, which makes them far more polite than real people. Real traffic arrives in bursts. A single scheduled notification or one post that gets picked up can spike you to a multiple of your baseline within seconds. Capacity planning has to plan for peak rather than mean, which is the standard guidance for sizing load tests: test at your projected peak, then push past it to find your margin.

The instinct this should push back on

Reaching for distributed infrastructure is usually an anxiety response, not an engineering one. The reasoning goes: this is a real product now, real products have scale problems, and scale problems need Kubernetes plus a queue plus a managed cache tier plus a service mesh. Money gets spent on a topology diagram before anyone has measured a single bottleneck.

That twelve-dollar server had one CPU core and it served over 5,000 concurrent users. It got there with two changes that cost nothing, because the person running the experiment had the discipline to find out what was actually limited before deciding how to fix it.

You might not need a distributed system. You might need to know what your CPU is doing.

The uncomfortable part is that the only way to find out which one you need is to run the test. Until you have, the migration you're planning is a guess about a number you don't know.


Source: load-test figures are from How many users can a $12 server support? by Arjaythedev, a hands-on k6 experiment on a $12/month VPS (1 vCPU, 2 GB RAM).

Top comments (0)