DEV Community

Cover image for Service discovery and load balancing in Node.js — without Consul or Kubernetes
Icebob
Icebob

Posted on

Service discovery and load balancing in Node.js — without Consul or Kubernetes

The moment you run a service on more than one machine, two questions appear that a monolith never had to answer: where is it (service discovery) and which copy do I call (load balancing). The standard answers are all infrastructure. A registry like Consul or etcd plus a client library in every service. Kubernetes Service objects plus kube-proxy. A service mesh sidecar next to every pod. An nginx or HAProxy in front of every pool.

Each of these works. Each is also a component you run, monitor, upgrade and get paged for, and none of them knows anything about your application — they see IP addresses and ports, not "the orders.create action, version 2, currently on three nodes".

This article shows the other approach: put the registry in the framework. Every Node.js process keeps a full copy of the service catalogue, learns about changes directly from its peers, and picks the target of each call locally. No registry server, no sidecar, no extra hop. I'll use Moleculer, which has done it this way since 2017, and I'll run every scenario for real — including the one where a node is killed with SIGKILL mid-traffic — so you can see exactly what happens and when.

How it works, in one paragraph

Each process runs a service broker (the Moleculer runtime) with a node ID. When a broker starts, it publishes an INFO packet on the transporter (the message bus — NATS here, but Redis, MQTT, AMQP, Kafka or plain TCP work the same) describing every service, action and event it hosts, plus free-form metadata. Every other broker receives it and updates its local registry. Brokers send a heartbeat every few seconds; a node that stops heartbeating is marked unavailable, and a node that shuts down cleanly says goodbye with a DISCONNECT packet so nobody waits for a timeout. When your code calls broker.call("greeter.hello", …), the caller's registry picks one live endpoint using a strategy and sends the request straight to that node. That is client-side discovery plus client-side load balancing, with the catalogue kept eventually consistent by gossip over the bus.

Every node has the full registry; the caller picks the endpoint. No registry server, no proxy hop.

The setup

One service, deliberately trivial, that reports which node answered:

// services/greeter.service.js — the only service in this article.
// Every response says which node served it, so load balancing is visible.
module.exports = {
  name: "greeter",
  actions: {
    hello: {
      params: { name: "string" },
      handler(ctx) {
        return { hello: ctx.params.name, servedBy: this.broker.nodeID };
      },
    },
  },
};
Enter fullscreen mode Exit fullscreen mode

One config, shared by every worker process and by the client. Everything the article varies is an environment variable; the service never changes:

// moleculer.config.js
const PreferZoneStrategy = require("./prefer-zone.strategy");   // custom strategy, shown later

// Built-in strategies are picked by name; a custom one is passed as a class.
const strategy = process.env.STRATEGY === "PreferZone" ? PreferZoneStrategy : process.env.STRATEGY || "RoundRobin";

module.exports = {
  nodeID: `worker-${process.pid}`,
  transporter: process.env.TRANSPORTER || "nats://localhost:4222",
  metadata: { zone: process.env.ZONE || "eu" }, // free-form node metadata, visible to every other node
  registry: {
    strategy,
    strategyOptions: { shardKey: "name" }, // only Shard reads this: route by ctx.params.name
    preferLocal: true,
  },
  heartbeatInterval: 2, // seconds; default 10
  heartbeatTimeout: 6,  // seconds; default 30 — shortened so the crash demo fits in a terminal
};
Enter fullscreen mode Exit fullscreen mode

Workers are started with moleculer-runner, the framework's CLI host: SERVICEDIR=services npx moleculer-runner --config moleculer.config.js. The client is a broker with no services, which is a perfectly normal thing to be — an API gateway or a cron job looks exactly like this:

// client.js (abridged) — a node with no services of its own; it only calls greeter.hello.
const broker = new ServiceBroker({ ...config, nodeID: `client-${process.pid}` });
await broker.start();
await broker.waitForServices("greeter");            // block until at least one instance is known

for (const name of ["ada", "linus", "grace", "dennis", "ken", "brian"]) {
  const res = await broker.call("greeter.hello", { name });
  console.log(`${name.padEnd(7)} → ${res.servedBy}`);
}
Enter fullscreen mode Exit fullscreen mode

1. Three workers, zero configuration

Start three workers and run the client:

$ node client.js calls 6
ada     → worker-3473401
linus   → worker-3473403
grace   → worker-3473402
dennis  → worker-3473401
ken     → worker-3473403
brian   → worker-3473402
Enter fullscreen mode Exit fullscreen mode

Round-robin across three processes. Nobody registered anything anywhere; the workers announced themselves on the bus and the client's registry did the rest. You can ask any node what it currently knows — the registry is exposed as a built-in $node service:

const nodes = await broker.call("$node.list", { withServices: true });
Enter fullscreen mode Exit fullscreen mode
$ node client.js nodes
client-3473436     available=true  zone=eu  services=-
worker-3473402     available=true  zone=eu  services=greeter
worker-3473401     available=true  zone=eu  services=greeter
worker-3473403     available=true  zone=eu  services=greeter
Enter fullscreen mode Exit fullscreen mode

Note the zone=eu column: that's the metadata from the config, carried in the INFO packet. It will matter in a minute.

2. Scaling down, crashing, scaling up — with traffic running

This is the part registry vendors don't put in the quick-start. The client calls greeter.hello every 500 ms for 16 seconds. While it runs, I stop one worker cleanly (SIGTERM), kill another (SIGKILL, no goodbye), and start a new one.

$ node client.js loop 16
t=  0.0s  → worker-3473034
t=  0.5s  → worker-3473036
t=  1.0s  → worker-3473035
t=  1.5s  → worker-3473034          ← SIGTERM sent to worker-3473034 right after this
t=  2.1s  → worker-3473036
t=  2.6s  → worker-3473036
t=  3.1s  → worker-3473035
t=  3.6s  → worker-3473036
t=  4.1s  → worker-3473035
t=  4.6s  → worker-3473036
t=  5.1s  → worker-3473035
t=  5.6s  → worker-3473036          ← SIGKILL sent to worker-3473035 here
[20:45:57.784Z] WARN  BROKER: Request 'greeter.hello' is timed out. { nodeID: 'worker-3473035', timeout: 2000 }
t=  6.2s  ✗ RequestTimeoutError: Request is timed out when call 'greeter.hello' action on 'worker-3473035' node.
t=  8.7s  → worker-3473036
[20:46:00.807Z] WARN  BROKER: Request 'greeter.hello' is timed out. { nodeID: 'worker-3473035', timeout: 2000 }
t=  9.2s  ✗ RequestTimeoutError: Request is timed out when call 'greeter.hello' action on 'worker-3473035' node.
t= 11.7s  → worker-3473036
[20:46:01.618Z] WARN  DISCOVERY: Heartbeat is not received from 'worker-3473035' node.
[20:46:01.619Z] WARN  REGISTRY: Node 'worker-3473035' disconnected unexpectedly.
t= 12.2s  → worker-3473036
t= 12.7s  → worker-3473036
t= 13.3s  → worker-3473036
t= 13.8s  → worker-3473098          ← new worker, first request within a second of starting
t= 14.3s  → worker-3473036
t= 14.8s  → worker-3473098
Enter fullscreen mode Exit fullscreen mode

Three different things happened, and it's worth being precise about each:

  • Graceful stop is instant. worker-3473034 got SIGTERM, the runner stopped the broker, the broker sent DISCONNECT, and the node was gone from every registry before the next call. Zero failed requests. This is what happens on every deploy, so it's the case that matters most.
  • A crash is detected by heartbeat, and until then calls to that node fail. worker-3473035 was killed at ~5.6 s. The registry only learned at ~11.6 s (heartbeatTimeout: 6), and in between, every request round-robined onto the dead node timed out. That's two failures out of eight calls. With the default heartbeatTimeout: 30 it would have been worse. No client-side registry can do better than its heartbeat window — Consul's agent has the same gap, Kubernetes' readiness probes have the same gap — the difference is only who notices and how you react. Reacting is what retries and circuit breakers are for; that's the next article. For now: tune heartbeatInterval/heartbeatTimeout to what your deployment can afford (2/6 is fine for a dozen nodes; the defaults are conservative for hundreds).
  • Scaling up is sub-second. worker-3473098 started, published INFO, and was serving traffic on the next tick. No registration step, no health-check grace period, no DNS TTL.

3. Choosing who gets the call: strategies

Round-robin is the default, but the strategy is a one-word option. Built in:

  • RoundRobin — the default; equal share, in order.
  • Random — equal share, no order; useful when many callers would otherwise stay in lockstep.
  • CpuUsage — samples the CPU load nodes report in their heartbeat, picks a lightly loaded one. Good for uneven hardware.
  • Latency — measures round-trip time to each node and prefers the fast ones. Good for multi-region.
  • Shard — consistent hashing on a field from the request, so the same key always lands on the same node. This is how you get in-memory per-user caches or ordered processing per key without a separate sticky-session layer.

Shard, keyed on name, same three workers:

$ STRATEGY=Shard node client.js calls 12
ada     → worker-3473403
linus   → worker-3473403
grace   → worker-3473471
dennis  → worker-3473471
ken     → worker-3473471
brian   → worker-3473471
ada     → worker-3473403
linus   → worker-3473403
grace   → worker-3473471
dennis  → worker-3473471
ken     → worker-3473471
brian   → worker-3473471
Enter fullscreen mode Exit fullscreen mode

Same name, same node, every time — with a consistent-hash ring, so adding a node only moves the keys that land on it. Two things I learned while writing this: the key must be given as registry.strategyOptions.shardKey (I first tried strategy: { type: "Shard", options: {…} }, which is silently ignored and falls back to random selection — check the output, not the config), and the key can be read from ctx.meta instead of params with a # prefix ("#userId"), which is usually what you want since the user ID rides in meta. Strategies can also be set per action (strategy/strategyOptions on the action definition) if one hot action needs sharding and the rest don't.

Writing your own in fifteen lines

The strategy interface is one method: select(endpoints, ctx). Here's one that keeps traffic inside the caller's own zone when it can, and falls back to everything when it can't — the classic "same-AZ first" rule, without a mesh:

// prefer-zone.strategy.js — a custom load-balancing strategy in ~15 lines.
// Picks an instance in the caller's own zone when one is alive; falls back to any instance.
const { Strategies } = require("moleculer");

class PreferZoneStrategy extends Strategies.Base {
  constructor(registry, broker, opts) {
    super(registry, broker, opts);
    this.zone = broker.metadata.zone;
    this.fallback = new Strategies.RoundRobin(registry, broker, opts);
  }

  select(endpoints, ctx) {
    const local = endpoints.filter((ep) => ep.node.metadata?.zone === this.zone);
    return this.fallback.select(local.length ? local : endpoints, ctx);
  }
}

module.exports = PreferZoneStrategy;
Enter fullscreen mode Exit fullscreen mode

Add a fourth worker with ZONE=us, then call from three different places:

$ node client.js nodes
worker-3473488     available=true  zone=us  services=greeter
worker-3473471     available=true  zone=eu  services=greeter
worker-3473403     available=true  zone=eu  services=greeter

$ STRATEGY=PreferZone ZONE=eu node client.js calls 6
ada     → worker-3473471
linus   → worker-3473403
grace   → worker-3473471
dennis  → worker-3473403
ken     → worker-3473471
brian   → worker-3473403

$ STRATEGY=PreferZone ZONE=us node client.js calls 6
ada     → worker-3473488
linus   → worker-3473488
grace   → worker-3473488
dennis  → worker-3473488
ken     → worker-3473488
brian   → worker-3473488

$ STRATEGY=PreferZone ZONE=asia node client.js calls 6      # no local instance → fall back to all
ada     → worker-3473403
linus   → worker-3473471
grace   → worker-3473488
dennis  → worker-3473403
ken     → worker-3473471
brian   → worker-3473488
Enter fullscreen mode Exit fullscreen mode

The eu caller round-robins over the two eu workers, the us caller pins to its one us worker, and a caller from a zone with no workers uses all three. The metadata came from the registry — the strategy never made a network call. You can put anything in metadata: version, hardware class, canary flag, tenant — and route on it the same way.

One more rule you get for free: preferLocal

If the calling node itself hosts an instance of the service, preferLocal: true (the default) calls it in-process and skips the network entirely. That's what lets the same code run as a modular monolith on a laptop and as a distributed system in production — I showed that end-to-end in the previous article.

4. No message broker either: discovery over plain TCP

Everything above used NATS as the bus. If you don't want to run even that — a small deployment, an on-prem box, an edge device — Moleculer's TCP transporter has discovery built in: nodes find each other by UDP multicast and then talk over direct TCP connections. Same two workers, same client, one variable changed:

$ TRANSPORTER=TCP SERVICEDIR=services npx moleculer-runner --config moleculer.config.js
[20:47:29.966Z] INFO  TRANSPORTER: TCP server is listening on port 42799
[20:47:29.971Z] INFO  TRANSPORTER: UDP Multicast Server is listening on 192.168.1.20:4445. Membership: 239.0.0.0
[20:47:29.972Z] INFO  TRANSPORTER: UDP discovery started.

$ TRANSPORTER=TCP node client.js calls 4
ada     → worker-3473535
linus   → worker-3473536
grace   → worker-3473535
dennis  → worker-3473536
Enter fullscreen mode Exit fullscreen mode

(IP address replaced; everything else is the actual log.) Service discovery and load balancing with no infrastructure process at all — not a registry, not a bus. The honest limitation: most cloud networks block multicast between VMs, so across hosts on AWS/GCP/Azure you either give the TCP transporter a static urls list of peers, or you use NATS/Redis, which is what I'd do anyway once there's more than one host.

When you still want Consul or Kubernetes

Being fair about the boundaries:

  • Polyglot systems. The registry lives in the Node.js processes. A Go or Python service can't join it (there are Java and Go ports that speak the same protocol, but they're separate community projects). If half your services aren't Node, an external registry is the lingua franca.
  • You're already on Kubernetes. Then keep it — for scheduling, restarts, secrets and ingress, it's excellent. It doesn't conflict: Kubernetes decides which pods exist, the broker's registry decides which action call goes to which pod and does so with application-level knowledge (versions, metadata, per-action strategies) that a Service object doesn't have. Many teams run exactly this combination, and it means the same code also runs on a plain VM.
  • Very large clusters. Every node hearing every heartbeat is O(n²) chatter. Past a few hundred nodes, switch the discoverer from the default gossip to the Redis- or etcd-backed one (registry.discoverer: "Redis") — heartbeats go to a shared store instead of the bus. That's one config line, and the rest of the article is unchanged.

For a Node.js system of, say, five to fifty services on a handful of machines — which is most systems — the registry in the framework covers service discovery, load balancing, zone affinity, sharding and node health, and the only infrastructure you ran for it in this article was one NATS container. Or, in the last demo, nothing.


Code in this article was run on Moleculer 0.15.2 and Node.js 22, with NATS 2 as the bus except in the TCP section. The moleculer-examples repository has a run.sh that reproduces all five demos, including the crash-and-recover timeline.

Top comments (0)