SysDesign Studio is an open-source, interactive way to practice system design interviews. Evolve architectures step by step, trace failure paths, and defend every decision.
Most system design prep teaches you to draw boxes: load balancer, cache, queue, done.
Then the interviewer asks, "Why a cache and not read replicas?" or "What happens when that Redis node dies?" A memorized diagram can't answer either question.
That gap is what interviews actually test. So I built SysDesign Studio, an open-source app for practicing the reasoning behind the boxes.
⭐ Source:
SysDesign Studio
Learn to reason about systems, not just draw them.
A polished learning app for system design interviews. It helps engineers reason through architecture decisions, trade-offs, bottlenecks, and failure modes instead of memorizing buzzwords.
Built with Next.js, TypeScript, Tailwind CSS, and a content-driven problem model so each design lesson is easy to extend without touching the UI.
Author
This project is created and maintained by Santosh Goteti.
About
SysDesign Studio is built for engineers who want to strengthen the mental models behind real systems:
- why a component exists in the first place
- what problem it solves under realistic load
- which alternatives were considered and rejected
- what breaks first as scale increases
- how the design evolves from a simple starting point to a resilient system
This repository turns those ideas into a guided interview-style learning journey.
Why this project?
System design interviews are not about drawing fancy diagrams — they…
The problem with how we prep
Most resources show you the final architecture of a system. That's like learning chess by staring at checkmate positions. You never see:
- Why each component showed up, and what number forced it in
- Which alternatives were on the table, and why they lost
- What breaks next: every fix introduces a new failure mode
- How the design evolved from one server to global scale
Interviewers aren't grading your diagram. They're grading whether you can defend it.
How it works: a 7-stage interview flow
Every problem follows the shape of a real interview:
| Stage | What you practice |
|---|---|
| 1. Understand | Clarifying questions, plus how each answer changes the design |
| 2. Size | Back-of-envelope math that always ends in a "so what" for the design |
| 3. Start simple | APIs, data model, sync vs. async, and the simplest design that works |
| 4. Evolve | v1 → v7, one bottleneck at a time |
| 5. Flows | Animated write, read, and failure paths through the diagram |
| 6. Deep dives | The hard sub-problems, with options and a recommendation |
| 7. Defend | Trade-offs, follow-up questions, and common mistakes |
Every evolution step uses the same strict template:
Problem → Evidence (numbers) → Options (each with a verdict) → Decision → New risks + mitigations → "Say it"
The "Say it" line is what you'd literally say out loud to the interviewer. It's the piece I wished I'd had when I was preparing.
Four problems are fully written, with more on the way!
| Problem | Difficulty | The core tension |
|---|---|---|
| 🔗 URL Shortener | Medium | A read-heavy key-value system where ID generation is the hard part |
| 🚦 Rate Limiter | Medium | A ~1 ms allow/deny decision on every request, with counters shared across a fleet |
| 📞 Phone Directory | Medium | "Who's calling?": a lookup that must answer while the phone is still ringing |
| 🕷️ Web Crawler | Hard | Politeness, dedupe, and a frontier that feeds thousands of fetchers without hammering anyone |
Each one has 7 evolution steps, 7–8 flows (including failure scenarios), 3 deep dives, trade-offs, follow-ups, and mistakes to avoid.
Here's a taste of each.
🔗 URL Shortener, v3: "Cache the read path"
Problem: every redirect hits the database. That's 350K reads/s at peak against one primary.
Evidence: a single Postgres primary handles roughly 10–20K primary-key lookups/s, so we're 20–30× over. Reads outnumber writes 100:1, and the traffic is heavily skewed.
Options:
- ❌ Read replicas: you'd need 25+ full copies, replication lag breaks the "creator clicks immediately" case, and it ignores the skew. It's the expensive way to buy throughput.
- ❌ Shard now: sharding is coming for storage reasons, but for reads alone it's slower to roll out and costs more.
- ✅ Cache-aside with Redis: the hot set is about 100 GB, and you get a hit rate above 90%.
The insight that makes the cache win is that short-code → URL mappings never change. The classic cache headache, invalidation, mostly disappears; the only exception is takedowns.
New risks the step introduces: a cache stampede on a viral link (fixed with single-flight request coalescing), losing a cache node (Redis Cluster with replicas), and scanners probing random codes (negative caching plus per-IP limits).
💬 Say it: "Short-code-to-URL mappings never change, so caching them is unusually safe. With 100:1 reads and a skewed distribution I expect a 90%+ hit rate, which brings DB load from 350K/s down to about 35K/s."
🚦 Rate Limiter: the race condition nobody draws
The naive version is GET count → compare → SET count. With two gateways, both can read 99 and both allow request #100. In a load test with 1,000 concurrent checks against a 100-token bucket, about 140 got through.
The step compares three fixes. GET/SET from the gateway is racy and slow. Optimistic locking (WATCH/MULTI) is correct, but its retries pile up on hot keys exactly when you need speed. The winner is a token bucket in a single atomic Lua script: one round trip, no race, and it uses Redis's clock so gateway clock skew doesn't matter.
The deep dives then cover algorithm choice, how exact the count really needs to be, and the classic question: fail open or fail closed?
📞 Phone Directory: the fastest lookup is the one you never make
Caller ID has a hard deadline: the answer must arrive before the user picks up. The server path is already about 15 ms, but mobile p99 round trips run 300–800 ms on weak networks.
Two tricks from this problem:
- A Bloom filter for unknown numbers. About 30% of lookups are numbers we've never seen, so a cache can't help. A 3.6 GB Bloom filter answers "definitely not here" in memory.
- An on-device snapshot. The top 5M numbers cover about 50% of lookups but weigh only about 100 MB. They ship nightly through a CDN with daily deltas, so half of all lookups never leave the phone and keep working offline.
It also covers the messy parts real systems face: number normalization, merging conflicting crowd-sourced names, and privacy-by-design deletion.
🕷️ Web Crawler: why the frontier is not "just a Kafka topic"
The estimate that changes everything: politeness limits you to about 1 request/s per host. To reach 2,000 pages/s, you must be fetching from at least 2,000 different hosts at every moment. A 500M-URL site would take about 16 years at that rate.
So the frontier can't be one global FIFO. A broker can't express "not before time T, per host, with priority." The design uses a purpose-built scheduler instead: priority front queues, a back queue per host, and a min-heap keyed on next_allowed_time. Kafka still shows up, between fetch and parse, where ordering and timing don't matter.
The failure flows cover slow and hostile sites, fetchers dying mid-fetch, crawler traps (infinite URL spaces), and a frontier partition losing its politeness state during failover.
Failure paths are first-class
Every problem includes animated failure flows that run on the final architecture, with failing hops drawn in red:
- A cache node dies mid-spam-wave
- A viral link triggers a stampede
- A whole region goes down
- An event gets processed twice
Each flow can also be viewed as an auto-generated sequence diagram.

Under the hood
- Next.js 15, React 19, TypeScript, Tailwind v4
- @xyflow/react for the architecture diagrams, with custom floating edges
- Packets animate along edges with SVG
animateMotion - Mermaid sequence diagrams are generated from the same flow data
- Framer Motion for stage transitions
The design decision I'm happiest with is that content is pure data. Each problem is a single TypeScript file, and each evolution step is a delta on the previous diagram:
{
version: "v3",
title: "Cache the read path",
problem: "Every redirect hits the database: 350K reads/s at peak against one primary.",
evidence: "A single Postgres primary does roughly 10–20K simple PK lookups/s...",
options: [ /* each with pros, cons, and a verdict */ ],
decision: "Cache-aside: on GET, try Redis...",
newRisks: [ /* risk + mitigation pairs */ ],
sayIt: "Short-code-to-URL mappings never change, so caching them is unusually safe...",
delta: {
addNodes: [{ id: "cache", label: "Cache", sub: "Redis Cluster ~100 GB", kind: "cache", x: 960, y: 80 }],
addEdges: [{ id: "api-cache", from: "api", to: "cache", label: "GET / SET" }],
},
}
diagramAt(problem, n) replays the deltas to build the diagram at any step and tracks which nodes changed, which drives the NEW / CHANGED highlights. Flows are validated against the final diagram's edges, so a flow can't route a packet along a connection that doesn't exist.
As a result, adding a new problem never touches the UI.
What's next
- [ ] More problems: Video Streaming, Chat App, News Feed, Payment System, Distributed KV Store, and more
- [ ] Practice mode: reveal one stage at a time and sketch your own answer first
- [ ] Drag-and-drop canvas to build your own design and compare it to the reference
- [ ] Building-block guides, decision trees, and a scaling cheat sheet
Try it / contribute
git clone https://github.com/Santoshrt999/sysDesign-studio.git
cd sysDesign-studio
npm install
npm run dev
Want to add a problem? It's one TypeScript file plus one line in the registry. PRs and content reviews are very welcome, especially from people who've sat on the interviewer side of the table.
Which system design problem should I write next? Tell me in the comments 👇
If this helps your prep, a ⭐ on GitHub helps others find it.



Top comments (1)
Do the SVG animateMotion flows honor prefers-reduced-motion? For architecture diagrams, switching to static highlighted paths could keep failure flows readable without motion.