Many engineers use these terms interchangeably, but they measure three different things. Understanding the distinction helps you design better systems and reason clearly about networking and distributed systems.
1️⃣ Latency
Definition: The time it takes for data to travel from sender to receiver — "how long does a single request take?"
Examples:
- API response time
- Database query time
- Network round-trip time (RTT)
If a request takes 40 ms to reach its destination, the latency is 40 ms.
Units: ms, μs
Lower is better.
2️⃣ Throughput
Definition: The amount of data/work successfully delivered per unit time — "how much work gets done every second?"
Examples:
- Requests per second (RPS)
- Messages processed per second
- MB transferred per second
A system processing 10,000 requests per second has high throughput.
Units: RPS, Mbps, Gbps, Transactions/sec
Higher is better.
3️⃣ Bandwidth
Definition: The maximum capacity of a network link — "the width of the highway."
Examples:
- 100 Mbps internet connection
- 1 Gbps network interface
- 10 Gbps data center link
Bandwidth = theoretical maximum. Throughput = what you actually achieve.
Units: Mbps, Gbps
Higher is better.
🎯 The Highway Analogy
| Concept | Highway Equivalent |
|---|---|
| Latency | How long it takes one car to travel from A to B |
| Throughput | How many cars reach the destination per hour |
| Bandwidth | Number of lanes available on the highway |
A highway can have:
- ✅ High bandwidth, ❌ High latency (many lanes, but the road is very long / has tolls)
- ✅ Low latency, ❌ Low throughput (short road, but only 1 lane — a car gets there fast, but few cars fit)
These metrics are related but not identical — you can have low latency with low throughput, or the reverse, and designing for one doesn't automatically improve the other.
❓ Q&A
How are latency and throughput related, and can a system have one without the other?
Yes. Latency is the time for a single unit of work to complete; throughput is the rate at which many units of work complete over time. A system can have low latency but low throughput (a single request is fast, but the system can't handle many concurrently), or higher latency but high throughput (each request is slower, but thousands are processed in parallel via pipelining/batching).
Can a network have high bandwidth but low throughput?
Yes. Bandwidth is the theoretical max capacity; throughput is what's actually achieved. Common causes: protocol overhead, packet loss/retransmission, congestion, small TCP window sizes, server-side processing bottlenecks, inefficient application code, or too few concurrent connections to saturate the link.
For example, a 100 Mbps link achieving only 60 Mbps throughput is usually explained by real-world losses — TCP/IP header overhead, retransmissions from packet loss, congestion control backing off, the latency/RTT limiting how much data can be "in flight" (bandwidth-delay product), competing traffic sharing the link, or hardware/software processing limits.
What factors contribute to network latency, and how do you reduce it?
Latency is made up of several components:
- Propagation delay (physical distance, speed of light in fiber)
- Transmission delay (time to push bits onto the wire, depends on bandwidth)
- Processing delay (routers, servers, load balancers, serialization/deserialization)
- Queuing delay (congestion, buffering at routers/servers)
- DNS resolution, TLS handshake, connection setup (TCP 3-way handshake)
- Application-level delays (DB queries, GC pauses, lock contention)
Common ways to reduce it:
- Cache close to the user (CDN, in-memory, edge caching)
- Reduce network hops / use geo-distributed regions closer to users
- Use connection pooling & keep-alive to avoid repeated handshakes
- Use async/non-blocking I/O to avoid queuing delays
- Batch or pipeline requests where possible
- Optimize serialization (protobuf/avro vs JSON)
- Use faster transport protocols (HTTP/2, QUIC/HTTP/3, gRPC)
- Reduce database query time (indexes, denormalization, read replicas)
- Load balance to avoid hotspots/queuing on a single node
- Precompute / prefetch data
How does caching impact latency?
Caching reduces latency by serving data from a faster, closer store (memory, CDN edge, local cache) instead of round-tripping to a database or origin server. It removes the need for expensive computation/disk/network I/O on repeated requests. It primarily helps latency, and indirectly throughput too, since the backend capacity freed up can serve more requests.
How does horizontal scaling impact throughput?
Adding more instances/nodes behind a load balancer lets a system handle more concurrent requests, directly increasing throughput (more RPS). It typically does not improve — and may slightly worsen — the latency of an individual request (still bounded by network RTT + per-request processing time), unless it also reduces contention/queueing delay.
Why doesn't increasing bandwidth always reduce latency?
Bandwidth affects how much data can be sent per second, not how fast a single bit/packet travels. Latency is dominated by propagation delay (distance/speed of light) and processing/queuing delays — a wider highway (more lanes) doesn't make a single car arrive faster if the road length and speed limit stay the same. Bandwidth helps throughput; it only marginally helps latency by reducing transmission delay for large payloads.
Does latency or throughput/bandwidth matter more for a given use case?
It depends on the workload:
- Online gaming: latency matters most. Games need near real-time responsiveness (low RTT) for actions to feel instant; packet sizes are small, so throughput is rarely the bottleneck. High latency causes lag even with plenty of bandwidth.
- Video streaming: bandwidth and sustained throughput matter most. Streaming needs enough sustained data rate to deliver frames without buffering; a few hundred ms of extra latency (buffering) is imperceptible, but insufficient bandwidth causes stuttering and quality drops. (Exception: live, interactive video — e.g., video calls — behaves more like gaming and is latency-sensitive.)
💡 Summary
Latency measures how long a request takes, throughput measures how much work is completed per unit time, and bandwidth measures the maximum capacity of the communication channel. Treating these as the same thing is one of the most common mistakes in reasoning about system performance.
Top comments (0)