DEV Community

Cover image for I Tried to Break a $12 Server. It Took 5,000+ Concurrent Users.
Gaurav Talesara
Gaurav Talesara

Posted on

I Tried to Break a $12 Server. It Took 5,000+ Concurrent Users.

There is a common assumption in backend engineering:

If you expect traffic to grow, you need to start thinking about scaling infrastructure early.

More servers.
More CPU.
A dedicated database.
Redis.
Load balancers.
Containers.
Kubernetes.

Sometimes you do.

But sometimes the application isn't actually asking for more infrastructure yet.

I wanted to see how far a very small server could go before reaching a real bottleneck.

So I rented a server for $12/month.

The machine had:

  • 1 CPU
  • 2 GB RAM
  • Nginx
  • Node.js
  • PostgreSQL
  • No dedicated database server
  • No Redis

Then I started increasing the load until something broke.

The interesting part wasn't how cheap the server was.

It was what became the bottleneck first.


The experiment

The goal was simple:

Find the point where the application starts showing meaningful performance degradation, identify the bottleneck, and improve it without upgrading the server.

I used a small API application running on a single machine.

The application had a database-backed feed/workflow-style endpoint that required reading data from PostgreSQL and generating a response for the client.

The database was running on the same machine as the application.

There was no separate cache layer.

No Redis.

No horizontally scaled application servers.

Just:

Internet
   |
 Nginx
   |
 Node.js
   |
PostgreSQL
Enter fullscreen mode Exit fullscreen mode

The server cost about $12/month.

That was the entire point.

I wanted to see what the application could actually handle before adding infrastructure.


Making the test more realistic

A simple benchmark that repeatedly calls one endpoint isn't particularly interesting.

Real users don't behave like that.

They open something.

They read.

They perform an action.

They come back.

They refresh.

They request more data.

So I used k6 to create virtual users that followed a more realistic pattern.

The virtual users would:

  1. Request the feed/workflow data.
  2. Perform an occasional write/action.
  3. Request the data again.
  4. Repeat the cycle.

I also populated the database with a meaningful amount of data rather than testing against an empty database.

The load generator itself ran from a second VM so that the machine being tested wasn't also responsible for generating all the traffic.

The goal was to put pressure on the application server while keeping the test environment simple.


Starting with a small load

I didn't jump directly to thousands of users.

The load increased progressively:

10
50
100
200
1,000
2,000
3,000
Enter fullscreen mode Exit fullscreen mode

This matters because performance problems aren't always obvious when you only look at average latency.

At low traffic, everything looked healthy.

As the number of virtual users increased, the system continued to respond normally.

Then something started changing around the higher loads.


2,000 users: the first warning sign

At around 2,000 virtual users, the average response time still looked reasonable.

But the tail latency started moving.

This is where metrics like p95 and p99 become much more useful than simply looking at average latency.

Imagine these two systems:

System A
Average: 100 ms
p95:     150 ms
p99:     200 ms
Enter fullscreen mode Exit fullscreen mode

versus:

System B
Average: 100 ms
p95:     700 ms
p99:     2,000 ms
Enter fullscreen mode Exit fullscreen mode

Both have the same average.

But the user experience is very different.

The second system has a tail-latency problem.

That was the first signal that the server was approaching a limit.


3,000 users: the test failed

I increased the load again.

At approximately 3,000 virtual users, the experiment crossed the failure criteria.

That gave me an important data point:

The server wasn't simply getting "a little slower."

There was a point where additional concurrency started pushing the system beyond the acceptable performance envelope.

So instead of continuing to increase the load blindly, I backed down.

I wanted to understand the bottleneck first.


2,500 concurrent users

At around 2,500 concurrent users, the system handled approximately:

~232 requests/sec
~288 ms p95 latency
Enter fullscreen mode Exit fullscreen mode

Those numbers were interesting.

But the most interesting information came from the resource utilization.

The server wasn't running out of memory.

RAM usage was still under 1 GB.

PostgreSQL was also behaving normally.

The CPU, however, was sitting around 90%.

That changed the direction of the investigation.

The first instinct could have been:

"The server is too small. Buy a bigger one."

But the measurements were telling a different story.


Finding the actual bottleneck

The resource picture looked roughly like this:

RAM       < 1 GB
Postgres  Healthy
CPU       ~90%
Enter fullscreen mode Exit fullscreen mode

The CPU was the constraint.

This is an important distinction.

If RAM had been nearly exhausted, adding memory would have been a reasonable direction.

If PostgreSQL had been saturated, database optimization would have been the next area to investigate.

But neither was the immediate problem.

The application was spending too much CPU doing work for requests.

So I asked a different question:

How much of this work actually needs to happen for every request?

That led directly to caching.


Optimization #1: Cache the response in Node.js

The feed/workflow response wasn't changing every millisecond.

Yet the application was repeatedly doing the same expensive work for requests that could safely receive a recently generated response.

So I introduced a very simple cache.

The idea was:

Request
   |
   v
Is cached response available?
   |
  Yes ----> Return cached response
   |
  No
   |
   v
Query database
   |
   v
Build response
   |
   v
Cache for 1 second
   |
   v
Return response
Enter fullscreen mode Exit fullscreen mode

The important detail was the TTL:

1 second.

This wasn't an attempt to build a complicated distributed caching architecture.

It was simply:

If another request arrives immediately after this one, don't repeat the same work unnecessarily.

That small change had a significant effect.

The system moved from roughly:

2,500 concurrent users
Enter fullscreen mode Exit fullscreen mode

to:

4,000 concurrent users
Enter fullscreen mode Exit fullscreen mode

before reaching the same kind of limitation.

And I hadn't changed the server.

No additional CPU.

No additional RAM.

No Redis cluster.

No second application server.

The workload had simply changed.


Why did a 1-second cache help so much?

Caching works because computation has a cost.

Consider a simplified request:

Request
   ↓
Node.js
   ↓
PostgreSQL query
   ↓
Process data
   ↓
Serialize response
   ↓
Send response
Enter fullscreen mode Exit fullscreen mode

If 100 requests arrive close together and they all need essentially the same data, doing the complete operation 100 times may be unnecessary.

With a short-lived cache:

Request 1
   ↓
Generate response
   ↓
Cache

Request 2 ─┐
Request 3 ─┤
Request 4 ─┤──> Cached response
Request 5 ─┘
Enter fullscreen mode Exit fullscreen mode

The expensive work can be shared across multiple requests.

That reduces CPU work.

It can also reduce database work.

And because the response is already available, the application has less work to perform per request.

The important point is that caching didn't make the CPU faster.

It reduced how much CPU work was required.


Optimization #2: Move the cache closer to the edge

The next question was:

Why should Node.js be responsible for serving a response that Nginx can serve directly?

The architecture originally looked like:

Client
  |
 Nginx
  |
 Node.js
  |
PostgreSQL
Enter fullscreen mode Exit fullscreen mode

After introducing application-level caching:

Client
  |
 Nginx
  |
 Node.js
  |
Cache
  |
PostgreSQL
Enter fullscreen mode Exit fullscreen mode

That was already better.

But there was another opportunity.

If Nginx could cache the response, a cache hit wouldn't need to travel through the Node.js application at all.

So the architecture became conceptually:

Client
  |
 Nginx
  |
  +---- Cache hit ----> Response
  |
  +---- Cache miss ---> Node.js
                           |
                       PostgreSQL
Enter fullscreen mode Exit fullscreen mode

Now the request could potentially be handled before reaching the application layer.

This is an important optimization because every request that Nginx can satisfy is a request that doesn't consume Node.js CPU.


The result

After moving the caching responsibility to Nginx, the system pushed past 5,000 concurrent users on the same $12/month server.

The progression was approximately:

~2,500 users
     |
     | 1-second application cache
     v
~4,000 users
     |
     | Move cache to Nginx
     v
5,000+ users
Enter fullscreen mode Exit fullscreen mode

The infrastructure didn't become more powerful.

The application simply stopped doing unnecessary work for every request.

That distinction is the main lesson from the experiment.


Concurrent users are not total users

One important clarification:

5,000 concurrent users does not mean the server can only support 5,000 users total.

Concurrency and total user count are different measurements.

If an application has 100,000 registered users but only 500 are actively making requests at a given moment, the server is dealing with roughly 500 concurrent users, not 100,000.

That's why saying:

"This server supports 5,000 users"

would be misleading.

The accurate statement from this experiment is:

The setup was able to push past 5,000 concurrent virtual users under this particular load-test workload.

That doesn't automatically translate into a production capacity number.

Real applications have different request patterns, payload sizes, database queries, background jobs, connection behavior, and traffic distributions.

Benchmarks are useful when we understand exactly what was benchmarked.


What I would not conclude from this experiment

This experiment doesn't prove that everyone should run production systems on a $12 server.

It doesn't.

A production system may need:

  • High availability
  • Backups
  • Replication
  • Monitoring
  • Failover
  • Security controls
  • Multiple application instances
  • Database replicas
  • Disaster recovery
  • Horizontal scaling
  • Background workers
  • Rate limiting

Those requirements are independent of whether a small server can handle a particular workload.

The experiment answers a narrower question:

How much performance can I get from simple infrastructure before infrastructure itself becomes the problem?

In this case, quite a lot.


The bigger lesson: measure before scaling

One of the easiest mistakes in backend engineering is solving a capacity problem before identifying the actual constraint.

You see latency increasing.

You add a bigger server.

But what if the application is wasting CPU?

You see database queries getting slower.

You add a larger database instance.

But what if the same expensive query is being executed thousands of times unnecessarily?

You see more traffic.

You add Redis.

But what if a simple Nginx cache would have handled the workload?

Infrastructure can absolutely solve performance problems.

But infrastructure should be guided by measurements.

In this experiment, the sequence was:

Load test
   ↓
Observe latency
   ↓
Measure resources
   ↓
Find CPU bottleneck
   ↓
Reduce repeated work
   ↓
Test again
   ↓
Find the next limit
Enter fullscreen mode Exit fullscreen mode

That loop is more valuable than any particular server size.


A simple scaling mindset

When an application starts struggling, I like thinking about the problem in this order:

1. What exactly is getting slower?

Don't stop at:

"The API is slow."
Enter fullscreen mode Exit fullscreen mode

Look at:

  • Average latency
  • p95
  • p99
  • Requests/sec
  • Error rate
  • CPU
  • Memory
  • Database utilization
  • Connection counts
  • Network
  • Disk I/O

2. What resource is actually saturated?

Find the constraint.

In this experiment:

RAM     → fine
Postgres → fine
CPU     → ~90%
Enter fullscreen mode Exit fullscreen mode

That pointed the investigation toward application work.

3. Can the work be eliminated?

Before making hardware bigger, ask:

  • Are we calculating the same thing repeatedly?
  • Can the result be cached?
  • Can a query be avoided?
  • Can data be precomputed?
  • Can work move to a cheaper layer?
  • Can Nginx/CDN handle something before it reaches the application?

4. Test the change

Don't assume the optimization worked.

Run the same workload again.

Compare the metrics.

5. Scale only when necessary

Eventually, the $12 server will hit another limit.

That's completely fine.

Scaling isn't the failure.

Scaling without understanding why you're scaling is the expensive part.


What surprised me most

The surprising part wasn't that a cheap server handled thousands of concurrent virtual users.

It was how quickly the result changed after removing unnecessary work.

At roughly 2,500 concurrent users, CPU was close to its limit.

A one-second cache increased the tested capacity to roughly 4,000.

Moving that cache to Nginx pushed the test past 5,000.

Same machine.

Same CPU.

Same RAM.

The biggest change was the amount of work the application had to perform.

That's a useful reminder for any backend system:

Before adding capacity, find out whether you can remove the work creating the demand.

Sometimes the cheapest performance optimization isn't a bigger server.

It's one less thing your server has to do.

Top comments (0)