Most system design material I've seen ends at the diagram. You draw a load balancer, a cache, a queue, everyone nods, and nobody ever checks whether the design actually survives the traffic. I've spent years on backends with Kafka, Postgres and Kubernetes, and the lessons that stuck with me were always the ones where something broke in production.
So I built ArchPath, a free course where you learn system design by running traffic through your own architecture.
How it works
You drag servers, load balancers, caches, queues and databases onto a canvas, wire them up and press run. The simulator shows latency, errors, utilisation of every node and the monthly bill as the traffic grows.
In the clip above a small service gets twice its usual traffic. p99 latency climbs past 400 ms because the database sits at 86%. Add one cache in front of it and latency drops to 140 ms, the database relaxes to 30%, and the bill still fits the budget.
The simulator is deterministic and runs entirely in the browser. There's no real Redis or Kafka behind it, every component is modelled by capacity, latency, hit ratio and cost. Same design, same numbers every time, so you can change one thing and compare. Each challenge also has a gold medal for the cheapest design that still holds, because in real life "just add more servers" is rarely the answer.
What's inside
- 52 topics in 10 acts, from a single server to global systems, each with an animated lecture and three story challenges
- around 200 production incidents to fix: bad deploys, hot keys, retry storms, a region going down
- a sandbox where you shape the workload yourself
No account needed, progress stays in your browser unless you sign up. It works best on a laptop.
What I'd love to hear
I'm looking for honest feedback, especially from people who run system design interviews or operate this stuff in production. Where do the numbers look wrong to you? Which topic felt too easy or too confusing?

Top comments (0)