DEV Community

M TOQEER ZIA
M TOQEER ZIA

Posted on

Five System Design Concepts Every Backend Engineer Should Understand

System design becomes easier when you understand how requests move through a system.

A user sends data across a network. A transport protocol moves the data. NAT rewrites addresses when traffic crosses network boundaries. A web server accepts the request. Hashing helps route data or requests. Throughput and latency tell you whether the whole system performs well under load.

These ideas connect closely. Learning them together gives you a stronger view of backend systems.

  1. Web servers sit at the request boundary

A web server receives HTTP or HTTPS requests and returns content. The response might contain HTML, JavaScript, an image, a PDF, or JSON from an API.

HTTP commonly runs over TCP. Before an HTTP exchange starts, the client and server establish a TCP connection. HTTPS adds a TLS handshake. Each connection also consumes server resources, including memory for a socket.

Repeated connection setup wastes time and resources. Keep alive keeps a connection open for a short period, which lets several requests reuse the same connection.

After receiving a request, the server parses the request, performs the required work, then sends a response. The work might involve reading a file, querying a database, or checking a cache.

Static content stays the same across requests. Dynamic content depends on request context, such as the logged in user or a database result. Dynamic content gives applications flexibility, but caching becomes harder because responses differ across users.

Concurrency matters as traffic grows. A simple blocking server handles one request at a time. Other requests wait. Production systems use threads, processes, or several server instances behind a load balancer. More concurrency also consumes more resources, so limits on connections and workers matter.

  1. TCP and UDP make different trade offs

TCP and UDP both operate at the transport layer. The IP address identifies the machine. The port identifies the application on the machine.

TCP focuses on reliable and ordered delivery. A three way handshake creates a connection. Acknowledgments confirm delivery. Retransmission replaces missing data. Sequence numbers restore packet order. Congestion control reduces the sending rate when the network becomes busy.

Those features add overhead and latency. TCP also keeps connection state on both endpoints. Every open connection consumes memory and file descriptors on the server.

For web traffic, databases, email, and file transfer, reliability often matters more than lower transport overhead.

UDP takes a different approach. UDP sends datagrams without a handshake, persistent connection state, acknowledgment, retransmission, packet ordering, or built in congestion control.

This gives UDP lower overhead. The trade off is weaker delivery guarantees.

Streaming media, voice traffic, gaming, and many DNS requests favor UDP because waiting for retransmission often hurts more than losing a small amount of data. Some applications add selected reliability features above UDP when needed.

Your protocol choice should follow application requirements. Use TCP when correctness and ordered delivery matter. Use UDP when low transport overhead and fast delivery matter more than guaranteed arrival.

  1. NAT changes how private systems reach public networks

Private IPv4 addresses do not travel across the public internet. NAT solves this by rewriting packet headers as traffic crosses a router or gateway.

Suppose a client uses 192.168.1.10 on port 49152. The NAT router owns a public address. The router replaces the private source address with its public address and often replaces the source port as well. The router stores the mapping in a translation table.

When the reply returns, the router checks the table, restores the private destination address and port, then forwards the packet to the original client.

The translation table makes NAT stateful. Losing the table through a reboot, timeout, or failure interrupts active mappings.

NAT also supports port forwarding. An application might listen on port 8080 while a DNAT rule rewrites incoming port 80 traffic toward port 8080. Layer 4 load balancers use a related approach with Virtual IP addresses. Incoming traffic reaches one Virtual IP, then the load balancer rewrites the destination toward a selected backend server.

NAT helped conserve IPv4 addresses by letting many private devices share one public address. The same mechanism also introduces trade offs. NAT breaks direct end to end addressing, complicates peer to peer traffic, adds state, and sometimes becomes a bottleneck. NAT also does not replace a firewall.

  1. Hashing decides where data belongs

A hash function maps an input, such as a user ID or cache key, into a fixed size output. System design uses the result for fast lookups, sharding, caching, deduplication, routing, and set membership checks.

A useful system design hash should be deterministic, fast, and evenly distributed.

Simple modulo sharding uses a rule such as hash of key modulo N. This works well while the server count stays fixed. Changing N remaps a large share of keys, which creates expensive movement across the cluster.

Consistent hashing reduces this movement.

Consistent hashing places both servers and keys on a circular hash space. A key belongs to the next server on the ring. When a server joins or leaves, only nearby keys move. Most keys stay with their existing owners.

Virtual nodes improve balance. Instead of placing each physical server at one point, the system places each server at many points around the ring. This spreads ownership more evenly.

Real systems use these ideas for storage, caches, messaging, and request routing. DynamoDB and Cassandra use consistent hashing concepts for partitioning. Kafka uses hash based partitioning by key. Memcached clients use consistent hashing so independent clients agree on key ownership.

Bloom filters use hashing for a different job. A Bloom filter answers whether an item might exist in a set. If one required bit is zero, the item is definitely absent. If all required bits are one, the item might exist. False positives are possible. False negatives are not.

This property helps systems skip expensive disk or network lookups for items known to be absent.

  1. Throughput and latency tell you whether the design works

Throughput measures completed work per unit of time. Examples include requests per second, transactions per second, or megabytes per second.

Latency measures the time required for one unit of work to finish.

A system with high throughput still might give users slow responses. A system with low latency still might support only a small amount of traffic.

For interactive APIs and web applications, latency matters heavily. For batch processing, throughput often carries more weight.

Average latency gives an incomplete picture. Percentiles show the distribution more clearly. p50 describes the median. p95 and p99 expose slower requests near the tail.

At 10,000 requests per second, a p99 slow rate means about 100 slow responses every second. Tail latency therefore matters even when slow requests form a small percentage of traffic.

Little's Law connects concurrency, throughput, and latency through L equals lambda times W. L represents average requests inside the system. Lambda represents arrival rate. W represents average time inside the system.

Queueing also changes performance under heavy load. As utilization approaches full capacity, waiting time rises sharply. The guide shows a simple queueing factor of about 2 times baseline at 50 percent utilization and about 10 times at 90 percent utilization.

This explains why production systems keep spare capacity. Running near maximum capacity leaves little room for traffic spikes.

Batching often raises throughput while increasing latency for individual items. Concurrency raises throughput until coordination and resource contention reduce the gains. Pipelining raises throughput by overlapping work. Caching improves both metrics when requests hit the cache.

The slowest stage sets the throughput ceiling for the full pipeline. Before tuning every component, find the bottleneck.

How these concepts connect

Think about one request to a backend API.

Your client sends the request toward a public address. NAT might rewrite the source address while traffic leaves a private network. TCP establishes reliable transport for the HTTP request. The web server accepts the connection and processes the request.

The application might hash a user ID to select a cache node or database shard. A load balancer might spread requests across several server instances. Each server consumes CPU, memory, network bandwidth, connections, and database capacity.

As traffic increases, queues start to form. Throughput approaches the limit of the slowest component. Tail latency rises. More servers, better caching, improved partitioning, or lower per request work might relieve the bottleneck.

System design is easier when you follow the request from the network edge to storage, then measure the result.

Learn where the request goes. Learn what state each layer keeps. Learn which guarantees each layer provides. Measure latency and throughput under load. Then change the part limiting the system.

Top comments (0)