DEV Community

Cover image for The 512 MB Ceiling: Load-Testing Using K6
Yar Khan
Yar Khan

Posted on

The 512 MB Ceiling: Load-Testing Using K6

How ten users appeared to sink a service — and what was really going on.

The setup

We had a backend written in Go, doing the ordinary things backends do: talk to a
SQL database, cache in Redis, check passwords, serve JSON to an app with
role-based access. It ran in a small managed container capped at 512 MB of
RAM
— the kind of tier you pick when you're being frugal and the traffic is
still modest.

Before opening the doors wider, we wanted an honest answer to a simple question:
how much can this thing actually take? So we wired up a load test with k6,
built two realistic user journeys — a regular user browsing their content, and an
admin working through management screens — and had each simulated user log in for
real, then move through the app with human-like pauses between clicks.

We started gently. A smoke test with a couple of users: flawless, responses in
~145 ms. Encouraged, we turned it up to a baseline of ten.

The crash

Ten users killed it.

The load tool reported 100% of requests failing. The platform's own logs told
the real story in one blunt line:

Ran out of memory (used over 512MB) while running your code.

The container had been killed and restarted mid-test. That's why everything
failed — not because the application returned errors, but because the process it
was running died.

This was genuinely surprising. Ten concurrent users is nothing. A healthy Go
service serving ten people should live comfortably in a few tens of megabytes.
Something was wrong — but the crash log only told us that it ran out of memory,
never why.

Chasing the wrong ghosts

The obvious suspects came first. Was a metrics system quietly hoarding memory,
one entry per unique URL, forever? We checked — no, it labeled by route
template, so its memory was bounded. Was the database connection pool
unbounded? No, it was capped at a sane number. No obvious leak, no runaway map.

We reached for the standard seatbelt for "Go app on a tiny container": we told the
Go runtime it had a memory budget (a soft limit a little under 512 MB) and asked
its garbage collector to work harder. This is usually the right move, and it did
change the symptom — but not in the way we hoped. Instead of a clean crash, logins
now hung for thirteen seconds and then failed with a gateway error, while
every other endpoint stayed fast. The runtime was fighting to stay under the
limit and losing.

That mismatch was the clue. If ordinary requests were quick and only login was
catastrophic, then login was doing something the others weren't.

The real culprit

Login checks a password. And this service, quite correctly, used a modern,
deliberately-expensive password hash — the kind security guides recommend
precisely because it's hard to brute-force. The way it earns that hardness is by
being memory-hard: every single password check allocates a large block of
RAM — here, 64 MB — and holds it for the entire time the hash is computing.

Suddenly the arithmetic was obvious, and a little alarming:

512 MB ÷ 64 MB per login ≈ 8 simultaneous logins to consume the entire
container — before counting anything else the app was doing.

Our "ten users" weren't ten people quietly reading pages. Because every simulated
user logged in at almost the same instant, they were ten password hashes firing
at once
, demanding ~640 MB between them. The box never stood a chance. And the
memory limit we'd added couldn't help: that memory wasn't garbage to collect, it
was in active use by hashes still running. So the runtime just thrashed instead
of crashing.

This wasn't only a testing curiosity. It described a real, plausible production
event: a group of users all signing in within the same minute — the start of a
session, a shift, a class — could knock the service over. The very feature meant
to keep accounts safe had become the thing most likely to take the service down.

The fix

We didn't weaken the password hashing. Weakening security to survive load is the
kind of trade you regret later.

Instead, we put a turnstile in front of it. A small counting semaphore now
limits how many password hashes may run at the same time — four. The rest wait
their turn. Because only four hashes are ever in flight, the memory they can
collectively demand is bounded to about 256 MB, comfortably inside the
512 MB ceiling, no matter how many people hit "sign in" at once. A login burst now
produces a short queue instead of a crash.

We re-ran the baseline. Zero failures. Every check passed. Responses back around
145 ms, logins around 1.6 seconds.
The wall was gone.

Finding the next wall

A fix that isn't measured is just a hope, so we kept pushing — up to fifty
concurrent users.

The result was quietly satisfying: no crash, no errors, and normal page
requests still returning in ~120 ms.
For everyday traffic, the little container
handled fifty users without complaint. Steady-state capacity was clearly much
higher than we'd feared.

But the turnstile revealed its own cost. With fifty people trying to sign in at
almost the same moment, and only four hashes allowed through at a time, the login
queue grew — the slowest sign-ins now waited up to forty seconds. The service
was no longer falling over; it was making people wait at the door. The
bottleneck had moved from "the whole thing dies" to "authentication throughput
during a burst," which is a far better problem to have and a far easier one to
reason about.

Where it stands

The service now survives bursts that used to kill it, and comfortably serves far
more concurrent users than the number that originally looked fatal. The remaining
limit — how fast a large crowd can sign in at once on a 512 MB box — is
understood, measured, and has a clear set of levers to widen when we need to:
make each password check a little cheaper, widen the turnstile, or give the
container more room. None of those are emergencies anymore. They're decisions.

And that, more than any single number, was the point of the exercise.

Top comments (0)