I built a 3D system design simulator to stop hand-waving architectures
System design interviews and architecture reviews share a problem: we draw boxes and arrows on a whiteboard and assert things like "the cache absorbs most of this" or "we'll just auto-scale." Nobody checks. I wanted a toy where you can actually watch those claims break — so I built one.
Give it a try: https://system-design-simulator-navy.vercel.app/
What it is
System Design 3D Simulator is a web app: you compose infrastructure out of components (DNS, CDN, load balancers, rate limiters, API gateways, app servers, WebSocket gateways, databases with shardable primaries, read replicas, caches, search clusters, queues, workers, object storage, monitoring), wire them together in a 3D scene, dial up traffic, and watch throughput, p50/p99 latency, error rate, availability, queue depth, and cost respond in real time.
It also ships one-click interview architectures — URL shortener, social feed with fan-out on write, chat with WebSockets, CDN-heavy video platform — plus demand patterns (steady, viral spikes, daily cycle, growth ramp) so you can capacity-plan for peaks instead of averages.
The stack is React 19 + Vite 8 + Three.js via React Three Fiber, with Zustand holding all simulation state and Lucide for icons. Four files do most of the work: store.js (881 lines — the whole simulation engine), Scene.jsx (980 lines — 3D models, failure visuals, packet animation), ControlPanel.jsx (388 lines), and componentInfo.js (458 lines).
One technical decision: a tick-based capacity model with magic numbers up front
The core of the simulation is a fixed-tick loop (TICK_S = 0.8 seconds simulated per tick) where every component type gets a per-node capacity constant:
const APP_RPS_PER_NODE = 150;
const LC_EFFICIENCY = 1.08; // least-connections spreads better
const CDN_OFFLOAD_PCT = 0.6;
const DB_LOAD_CAP = 220; // load units per primary before saturation
const CB_DROP_FRACTION = 0.35; // requests shed while breaker is open
const MQ_BUFFER_CAP = 8000; // buffered msgs per broker before drops
const WORKER_DRAIN_PER_NODE = 200;
The tradeoff is obvious: these numbers are invented, not measured. A real CDN offloads far more than 60% for cacheable content; a real app server's RPS depends on what the handler does. But the point of the tool was never absolute accuracy — it's relative behavior. Kill the only load balancer and you get a total outage. Add read replicas and the primary stops saturating. Switch the LB strategy from round robin to least-connections (that 1.08 efficiency factor) or consistent hashing and you can watch the hot-partition trap appear. The constants being rough doesn't matter as long as the relationships between components behave plausibly, and keeping them as named constants in one place means anyone can tune them toward realism later.
The same thinking applies to the cost model: every component has a monthly $/node figure, cost per 1M requests is derived live — and dead nodes still bill. That last detail is deliberate. Idle capacity costing money is the whole argument for auto-scaling, and the simulator lets you toggle auto-scaling and watch scaling events fire.
Chaos was the fun part
Each node accumulates damage under sustained overload and eventually crashes; auto-heal restarts it, modeled loosely on Kubernetes liveness probes. A circuit breaker opens when the primary DB saturates and sheds read pressure until recovery. You can kill any node by hand, or roll the chaos monkey. The failure I keep coming back to: search queries with no search cluster fall back to expensive scans on the app servers and DB, and p99 quietly degrades instead of erroring. Silent degradation is harder to notice than an outage, which is exactly why I kept it.
What broke / what's next
Honest status: the failure visuals and packet animation live in the same 980-line Scene.jsx as the component models, and that file is the hardest thing in the repo to change — it needs splitting before anything else. The 3D drag-and-drop interaction (added a few commits ago along with the info modals) still feels fiddly for precise topologies. And the capacity constants, as noted, are placeholders with opinions — calibrating even one path (say, queue drain behavior) against a real benchmark would make the whole thing more trustworthy.
Next up would be per-link latency visualization and a shareable-URL encoding of a topology so two people can argue about the same diagram.
Run it
npm install
npm run dev
Then pick the URL shortener scenario, switch demand to viral spike, kill a node, and watch what redundancy actually buys you. That's the whole pitch: fewer whiteboard assertions, more watching things break.
In action
Real captures from the running app — a healthy system at 850/850 rps, p99 at 302ms:
Built by Irfan Wani. Repo: https://github.com/Irfanwani/system-design-simulator



Top comments (0)