Imagine a busy café with one master chef and four other chefs. Customers keep placing orders. If the master chef prepares every dish alone, customers will wait too long.
Instead, the master chef distributes orders among the four chefs. In system design, this is what a load balancer does: it receives incoming requests and decides which server should handle each one.
The Café Team
- Master chef: Load balancer
- Four chefs: Servers
- Customer orders: Incoming requests
- Preparing dishes: Processing requests
1. Round Robin
The master chef assigns orders to chefs one by one in a repeating order.
Café example
- Order 1 goes to Chef 1.
- Order 2 goes to Chef 2.
- Order 3 goes to Chef 3.
- Order 4 goes to Chef 4.
- Order 5 goes back to Chef 1.
In system design
Requests are distributed sequentially across servers.
Best used when
Servers have similar capacity and requests take roughly the same amount of time.
Limitation
Round Robin does not check whether a chef is already busy. A chef working on a complicated dish may receive another order even when another chef is free.
2. Weighted Round Robin
Now imagine Chef 1 is more experienced and can prepare twice as many dishes as the other chefs. The master chef gives Chef 1 more orders.
Café example
A possible distribution is:
- Chef 1 receives 2 orders.
- Chef 2 receives 1 order.
- Chef 3 receives 1 order.
- Chef 4 receives 1 order.
The exact distribution depends on the configured weights.
In system design
Servers receive requests according to assigned weights. A server with a higher weight generally receives a larger share of requests.
Best used when
Servers have different processing capacities.
Limitation
Weights do not necessarily reflect how busy a server is at that moment.
3. Least Connections
Imagine Chef 1 is preparing three dishes, Chef 2 is preparing one, and Chefs 3 and 4 are free. The master chef sends the next order to a chef with the fewest active orders.
Café example
The next order could go to Chef 3 or Chef 4 because both have no active orders.
In system design
The load balancer selects a server with the fewest active connections.
Best used when
Connections or requests can remain active for different lengths of time.
Limitation
The number of active connections does not always represent the real workload. One request may require much more processing than another.
4. Weighted Least Connections
Now combine the previous two ideas. Chef 1 is twice as capable as the other chefs, but the master chef also checks how many active orders each chef has.
The master chef considers both capacity and current connections when choosing a chef.
In system design
The algorithm uses server weights along with active connection counts to help distribute requests.
Best used when
Servers have different capacities and workloads vary.
Limitation
It still relies on connection counts and weights, which may not perfectly represent actual processing effort.
5. Least Response Time
Imagine Chef 1 has only one order but is taking a long time to finish it. Chef 2 has two orders but usually finishes dishes quickly. The master chef considers how quickly each chef is responding and sends the next order to a suitable, faster option.
In system design
Requests are routed based on response time, often combined with active connections. The exact calculation depends on the load balancer.
Best used when
Fast responses and low latency are important.
Limitation
Response times can fluctuate, so a single slow or fast response may not represent long-term performance.
6. IP Hash
Imagine the master chef uses each customer's membership number to consistently assign them to a particular chef. Customer A may usually be assigned to Chef 1, while Customer B may usually be assigned to Chef 3.
In system design
The load balancer hashes the client's IP address to determine which server should receive the request.
Best used when
Consistent client-to-server routing is useful, such as in some session-persistence setups.
Limitation
Traffic can become uneven, and adding or removing servers can change which server a client reaches. IP-based routing may also be less reliable when many users share an IP address.
Quick Comparison
| Algorithm | Café rule | Best use |
|---|---|---|
| Round Robin | Take turns | Similar servers and requests |
| Weighted Round Robin | Stronger chef gets more orders | Different server capacities |
| Least Connections | Choose a chef with fewer active connections | Variable-duration connections |
| Weighted Least Connections | Consider capacity and active connections | Different capacities and workloads |
| Least Response Time | Prefer faster-responding chefs | Latency-sensitive applications |
| IP Hash | Assign clients consistently | Session persistence in suitable setups |
Interview Question
Your café has four chefs. Chef 1 is powerful but currently busy. Chef 2 is less powerful but completely free. Which load-balancing algorithm could be a good fit?
- A. Round Robin
- B. Weighted Round Robin
- C. Least Connections
- D. IP Hash
Answer
Weighted Least Connections is usually the best fit when both server capacity and current connections matter. Least Connections alone may choose Chef 2 because they have fewer active connections, but it does not account for the difference in capacity.
Key Takeaway
- The master chef is the load balancer.
- The four chefs are servers.
- Customer orders are incoming requests.
- The load-balancing algorithm decides which server receives each request.
There is no single best algorithm for every application. Choose based on server capacity, request duration, response time, and whether consistent client routing is needed.
Top comments (0)