DEV Community

Timevolt
Timevolt

Posted on

Scaling Your App: Horizontal vs Vertical — A Lord of the Rings Quest

The Quest Begins (The “Why”)

Honestly, I remember the day our little side‑project started to feel like a dragon hoarding gold. We had a neat Express API that served JSON to a React frontend, and everything was smooth until our user base jumped from a few hundred to a few thousand requests per minute. The CPU on our single t2.medium instance spiked, latency crept up, and users started seeing those dreaded “502 Bad Gateway” pages.

I felt like Frodo staring at Mount Doom, wondering if I had enough stamina to keep going. The obvious answer was “make the server bigger,” but I’d heard whispers about adding more servers instead of just beefing up one. Was vertical scaling the easy way out, or was horizontal scaling the true path to glory? I decided to treat this as a quest: find the treasure of scalable architecture and bring it back to my team.

The Revelation (The Insight)

The treasure turned out to be a simple truth: vertical scaling makes a single instance stronger; horizontal scaling makes the system resilient by spreading the load.

Think of it like Neo realizing he can bend the Matrix. When you only upgrade the box (more CPU, more RAM), you’re still limited by that one box’s ceiling. Hit a hardware limit? You’re stuck. But when you clone the instance and put a load balancer in front, each request can be served by any copy. If one node fails, the others keep the party going.

The catch? Your app must be stateless (or at least externalize state). Anything stored in memory — like session objects, local caches, or in‑process queues — needs to live somewhere shared (Redis, a database, or a distributed cache). Once that’s true, adding more nodes is as easy as spinning up another container.

Wielding the Power (Code & Examples)

The Struggle: A Single‑Node Nightmare

Here’s what our original server looked like — simple, but not ready for traffic spikes:

// server.js (vertical‑only version)
const express = require('express');
const app = express();
const PORT = process.env.PORT || 3000;

// In‑memory “cache” – a red flag for horizontal scaling
const cache = new Map();

app.get('/data/:id', (req, res) => {
  const id = req.params.id;
  if (cache.has(id)) {
    return res.json({ source: 'cache', data: cache.get(id) });
  }

  // Simulate expensive work
  const result = expensiveComputation(id);
  cache.set(id, result);
  res.json({ source: 'computed', data: result });
});

function expensiveComputation(id) {
  // Pretend this is a CPU‑heavy task
  let sum = 0;
  for (let i = 0; i < 5e7; i++) sum += Math.sqrt(i);
  return sum * parseInt(id, 10);
}

app.listen(PORT, () => console.log(`🚀 Server listening on ${PORT}`));
Enter fullscreen mode Exit fullscreen mode

If you launch this on a bigger EC2 instance, you’ll see higher throughput — until you hit the CPU ceiling. Plus, if the process crashes, all cached data vanishes.

The Victory: Going Horizontal with Node’s Cluster Module

Node gives us a built‑in way to fork multiple workers that share the same port. Combined with an external Redis store, we can safely scale out:

// cluster.js (horizontal‑ready version)
const express = require('express');
const redis = require('redis');
const { fork } = require('cluster');
const numCPUs = require('os').cpus().length;

// Only the master process does this
if (cluster.isMaster) {
  console.log(`Master ${process.pid} is running`);

  // Fork workers.
  for (let i = 0; i < numCPUs; i++) {
    cluster.fork();
  }

  cluster.on('exit', (worker, code, signal) => {
    console.log(`Worker ${worker.process.pid} died. Restarting...`);
    cluster.fork();
  });
} else {
  // Workers can share any TCP connection.
  const app = express();
  const client = redis.createClient({ url: process.env.REDIS_URL });

  client.on('error', err => console.log('Redis error', err));
  client.connect().catch(console.error);

  app.get('/data/:id', async (req, res) => {
    const id = req.params.id;
    const cached = await client.get(`data:${id}`);
    if (cached) {
      return res.json({ source: 'cache', data: JSON.parse(cached) });
    }

    const result = expensiveComputation(id);
    await client.set(`data:${id}`, JSON.stringify(result), { EX: 60 }); // 1‑min TTL
    res.json({ source: 'computed', data: result });
  });

  const PORT = process.env.PORT || 3000;
  app.listen(PORT, () => {
    console.log(`🚀 Worker ${process.pid} listening on ${PORT}`);
  });
}

function expensiveComputation(id) {
  let sum = 0;
  for (let i = 0; i < 5e7; i++) sum += Math.sqrt(i);
  return sum * parseInt(id, 10);
}
Enter fullscreen mode Exit fullscreen mode

What changed?

  1. Statelessness – The in‑memory Map is gone. We now ask Redis for cached values, which all workers can reach.
  2. Clustering – The master process spins up a worker per CPU core, giving us true horizontal scaling on a single machine.
  3. Resilience – If a worker dies, the master instantly spawns a replacement; the load balancer (or just the OS’s round‑robin routing) sends traffic to the healthy ones.

If you want to go beyond a single host, replace the cluster module with a container orchestrator (Docker Swarm, Kubernetes, or even a managed service like AWS ECS). The same principles apply: stateless app + external state + load balancer = unlimited horizontal growth.

Traps to Avoid (The “Monsters” on the Path)

  • Sticky sessions without reason – Tying a user to a specific node defeats the purpose of horizontal scaling unless you absolutely need it (e.g., WebRTC).
  • Local file uploads – Writing to the server’s disk means another node can’t see the file. Use S3, GCS, or a shared volume.
  • In‑process queues – Libraries like BullMQ that store jobs in memory will lose work when a worker dies. Back them with Redis or a proper message broker.

Why This New Power Matters

Now that you’ve got the horizontal spell in your grimoire, you can:

  • Handle traffic spikes without over‑paying for a massive instance that sits idle most of the time.
  • Achieve zero‑downtime deploys – roll out a new version by gradually replacing workers.
  • Sleep better at night knowing the loss of a single node won’t bring your whole service down.

Imagine launching a feature that goes viral overnight. Instead of frantically scrambling for a bigger VM, you just hit “scale up” on your Kubernetes cluster and watch the autoscaler add pods. That’s the kind of power that turns a stressful on‑call into a confident victory lap.

Your Turn – The Challenge

Grab a small API you’ve built (maybe that todo app from last weekend).

  1. Externalize any in‑memory cache or session store to Redis.
  2. Wrap the entry point in Node’s cluster module (or Dockerize it and deploy with a replica count of 3).
  3. Hammer it with a loader like autocannon or hey and watch the latency stay flat as you increase the request rate.

How did it feel to see the load spread across cores? Did you notice any hidden state that crept back in? Drop a comment below — I’d love to hear about your scaling saga and the loot you collected along the way!

Happy scaling, fellow adventurer! 🚀

Top comments (0)