DEV Community

Peakblick
Peakblick

Posted on Originally published at peakblick.com

System Design Interview Questions (With Model Answers)

These are the system design questions that actually come up — grouped by topic, each with a short model answer and a note on what the interviewer is really checking. System design rewards judgement over facts: there's rarely one right answer, only trade-offs you can name and defend.

The trap in a system design round is treating it like trivia. The interviewer isn't checking whether you've memorised what a load balancer is — they're checking whether you can reason about scale, spot the bottleneck, and make a call while naming what it costs. For each answer below, notice the pattern: name the mechanism, give the one trade-off that matters, then say when you'd actually reach for it.

Scaling fundamentals

What's the difference between vertical and horizontal scaling?

Vertical scaling (scaling up) means a bigger machine — more CPU, RAM, disk. It's simple and needs no code changes, but it has a hard ceiling and a single point of failure. Horizontal scaling (scaling out) means more machines behind a load balancer. It scales much further and adds redundancy, but it forces you to deal with distributed-system problems: shared state, consistency, and coordination. Short version: scale up until it's too expensive or too risky, then scale out.

What they're testing: that you know scaling out buys you headroom at the cost of complexity, not for free.

Latency vs throughput — what's the difference?

Latency is how long one request takes; throughput is how many requests you handle per second. They're related but not the same — you can raise throughput by adding workers while a single request stays just as slow, and you can cut latency without touching throughput. Know which one the problem is about: a slow page is a latency problem, a system falling over under load is a throughput problem.

What they're testing: that you reach for the right lever instead of "make it faster" in general.

A service is getting too much traffic. How do you think about handling it?

Measure first — find the actual bottleneck instead of guessing. Then work the usual order: cache what's read often and changes rarely; scale the stateless parts horizontally behind a load balancer; move slow or spiky work off the request path into a queue; and only then scale the database, which is usually the hardest piece to scale. The instinct interviewers want is "find the bottleneck, then apply the cheapest fix that moves it," not "add servers everywhere."

What they're testing: a systematic approach, and knowing the database is usually the real constraint.

Caching & data

What is caching, and where would you add it?

A cache stores the result of expensive work close to where it's needed, so you skip the work next time. You add it wherever reads dominate and the data tolerates being slightly stale — a database query layer, rendered pages, a CDN for static assets, or in-memory (Redis/Memcached) for hot keys. The win is latency and load; the cost is a second copy of the truth that can drift.

What they're testing: that you see caching as a read optimisation with a staleness price, not a free speed-up.

Why is cache invalidation hard?

Because a cache is a copy, and the moment the source changes, the copy is wrong — you have to decide when to throw it away without knowing in advance when the data will change. The common strategies are TTL (expire after N seconds — simple, but serves stale data in that window) and explicit invalidation on write (fresh, but easy to miss a path and leak stale data forever). There's no clean answer, which is why it's the classic "two hard things in computer science" joke.

What they're testing: that you understand the real cost of a cache is correctness, not just the extra moving part.

SQL vs NoSQL — when would you choose each?

Reach for SQL (a relational database) by default: you get strong consistency, transactions, flexible queries and joins, and decades of tooling — ideal when your data is related and correctness matters, like anything with money or accounts. Reach for NoSQL when you have a specific scale or shape problem SQL handles badly: massive write throughput, a flexible/changing schema, or simple key-value access patterns at scale. The honest answer in most interviews is "SQL unless I can name the specific reason it won't work."

What they're testing: that you don't pick NoSQL for hype — you justify it with a concrete constraint.

What is sharding, and what's the hard part?

Sharding splits one dataset across multiple databases by some key (user ID, region) so no single machine holds — or serves — all of it. That's how you scale writes past one box. The hard part is choosing the shard key: a bad one creates hot shards (one machine does most of the work) and makes cross-shard queries and joins painful or impossible. Resharding later is a major operation, so the key is a decision you want to get right early.

What they're testing: that you know sharding scales writes but the shard key is where it lives or dies.

What does replication give you, and what's the trade-off?

Replication keeps copies of your data on multiple machines. It buys you two things: read scaling (spread reads across replicas) and availability (if one node dies, another has the data). The trade-off is consistency: with asynchronous replication, a replica can lag behind the primary, so a read right after a write might return stale data. You're trading "always perfectly up to date" for "more reads and survives failures."

What they're testing: that you connect replication to the consistency lag it introduces, not just "backups."

What does the CAP theorem actually force you to choose?

CAP says that when a network partition happens — some nodes can't talk to others — a distributed system has to pick: stay consistent (reject or block requests it can't confirm, so no one sees wrong data) or stay available (keep answering, accepting that different nodes might briefly disagree). You don't get both during a partition. The useful framing isn't "CA vs CP vs AP" trivia — it's "when the network splits, would I rather return an error or risk a stale answer?", and that depends on whether you're moving money or showing a feed.

What they're testing: that you can turn CAP into a real product decision, not recite three letters.

Communication & scale

What does a load balancer do?

It sits in front of a pool of identical servers and spreads incoming requests across them, so no single server is overwhelmed and the pool can grow or shrink freely. It also does health checks — stop sending traffic to a server that's down — which is half the reason it exists. For it to work, your servers generally need to be stateless (any server can handle any request), which is why "where does the session live?" is the follow-up question.

What they're testing: that you connect load balancing to statelessness, not just "it shares the load."

When would you use a message queue?

When work doesn't need to happen inside the request. A queue lets you accept a request, drop a job on the queue, respond immediately, and process it later with separate workers — think sending email, generating a report, processing an upload. It buys you three things: a faster response, resilience to spikes (the queue absorbs the burst), and decoupling (the producer doesn't care who consumes). The cost is that the work is now asynchronous, so you design for "eventually done" and for retries.

What they're testing: that you move slow or spiky work off the request path on purpose, and handle the async consequences.

How would you rate-limit an API?

Track requests per client (by API key or IP) over a time window and reject once they exceed the limit, usually with a 429 Too Many Requests. The common algorithm is a token bucket — each client has a bucket that refills at a steady rate and each request spends a token — which allows short bursts while capping the sustained rate. In a multi-server setup the counter has to live somewhere shared, like Redis, so the limit holds across the whole fleet rather than per server.

What they're testing: that you remember the limit has to be enforced globally, not per instance.

What's a CDN for?

A CDN caches your static content — images, CSS, JS, video — on servers spread around the world, so a user in Tokyo is served from a nearby edge instead of your origin in Virginia. That cuts latency dramatically and takes load off your servers. The trade-off is the usual cache one: you have to handle invalidation (cache-busting filenames or purges) when the content changes.

What they're testing: that you push static, geographically-shared content to the edge and know how you'd update it.

Reliability & approach

What is idempotency, and why does it matter in a distributed system?

An operation is idempotent if doing it twice has the same effect as doing it once. It matters because in a distributed system, requests get retried — a network timeout doesn't tell you whether the first attempt succeeded. If "charge this card" isn't idempotent, a retry double-charges. The fix is an idempotency key: the client sends a unique ID with the request, and the server records it so a repeat with the same key is recognised and ignored. Any operation that changes state and can be retried needs this.

What they're testing: that you design for retries and at-least-once delivery, not a perfect network.

How do you approach an open-ended "design X" question?

Don't start drawing boxes. Start by scoping: ask about the requirements and scale — how many users, read-heavy or write-heavy, what has to be consistent, what can be eventual. Sketch the core data model and the main API, then the high-level components (clients, load balancer, services, database, cache, queue). Then find the bottleneck for the given scale and go deep on one or two pieces — how you'd cache, shard, or queue — naming the trade-off at each step. The interviewer is grading the conversation, not a finished diagram.

What they're testing: that you drive the problem — scope, structure, then depth — instead of memorising one architecture.

How to actually answer these

System design is the round where candidates most often sink themselves by saying too much. The strong move is the opposite: scope before you build, name the trade-off instead of claiming a choice is just "better," and admit what you don't know rather than bluff — "I'd need to check the real numbers, but my instinct is to shard by user ID because our access pattern is per-user" reads far better than a confident wrong architecture.

Three habits that quietly raise your score: always ask about scale first (the right design for 1,000 users is wrong for 100 million); say the trade-off out loud for every choice; and drive the conversation — a good system design answer is a discussion you lead, not a monologue. Prepping the backend stack too? Pair this with the SQL, Python backend, and Django questions.


Rehearse these out loud, scored. Peakblick asks you system design questions like these on a timer and scores every answer 1–10 with feedback — so you find the weak spots before the real interview. Free to try, no card.

Top comments (1)

Collapse
 
anh_nguynvn_0478e614ba profile image
Anh Nguyễn Văn •

Việc phân loại các câu hỏi theo chủ đề như thế này giúp ích rất nhiều cho việc ôn tập có hệ thống thay vì học vẹt từng case riêng lẻ. Một kinh nghiệm xương máu khi đi phỏng vấn là đừng quá sa đà vào việc chọn ngay một database cụ thể ngay từ đầu. Thay vào đó, hãy tập trung giải thích được trade-off giữa consistency và availability dựa trên yêu cầu cụ thể của bài toán. Nếu bạn chọn NoSQL, hãy sẵn sàng giải thích tại sao mô hình eventual consistency lại phù hợp với use case đó hơn là ACID của SQL. Việc chứng minh được tư duy đánh đổi luôn quan trọng hơn là đưa ra một đáp án "đúng" duy nhất.