<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Irfan Wani</title>
    <description>The latest articles on DEV Community by Irfan Wani (@irfan_wani).</description>
    <link>https://dev.to/irfan_wani</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1788786%2F1b75ec39-af48-4a58-af19-a894d98ef8e3.jpeg</url>
      <title>DEV Community: Irfan Wani</title>
      <link>https://dev.to/irfan_wani</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/irfan_wani"/>
    <language>en</language>
    <item>
      <title>I built a 3D system design simulator to stop hand-waving architectures</title>
      <dc:creator>Irfan Wani</dc:creator>
      <pubDate>Tue, 29 Sep 2026 17:29:31 +0000</pubDate>
      <link>https://dev.to/irfan_wani/i-built-a-3d-system-design-simulator-to-stop-hand-waving-architectures-4nd6</link>
      <guid>https://dev.to/irfan_wani/i-built-a-3d-system-design-simulator-to-stop-hand-waving-architectures-4nd6</guid>
      <description>&lt;h2&gt;
  
  
  I built a 3D system design simulator to stop hand-waving architectures
&lt;/h2&gt;

&lt;p&gt;System design interviews and architecture reviews share a problem: we draw boxes and arrows on a whiteboard and assert things like "the cache absorbs most of this" or "we'll just auto-scale." Nobody checks. I wanted a toy where you can actually watch those claims break — so I built one.&lt;/p&gt;

&lt;p&gt;Give it a try at &lt;a href="https://system-design-simulator-navy.vercel.app/" rel="noopener noreferrer"&gt;https://system-design-simulator-navy.vercel.app/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;System Design 3D Simulator is a web app: you compose infrastructure out of components (DNS, CDN, load balancers, rate limiters, API gateways, app servers, WebSocket gateways, databases with shardable primaries, read replicas, caches, search clusters, queues, workers, object storage, monitoring), wire them together in a 3D scene, dial up traffic, and watch throughput, p50/p99 latency, error rate, availability, queue depth, and cost respond in real time.&lt;/p&gt;

&lt;p&gt;It also ships one-click interview architectures — URL shortener, social feed with fan-out on write, chat with WebSockets, CDN-heavy video platform — plus demand patterns (steady, viral spikes, daily cycle, growth ramp) so you can capacity-plan for peaks instead of averages.&lt;/p&gt;

&lt;p&gt;The stack is React 19 + Vite 8 + Three.js via React Three Fiber, with Zustand holding all simulation state and Lucide for icons. Four files do most of the work: &lt;code&gt;store.js&lt;/code&gt; (881 lines — the whole simulation engine), &lt;code&gt;Scene.jsx&lt;/code&gt; (980 lines — 3D models, failure visuals, packet animation), &lt;code&gt;ControlPanel.jsx&lt;/code&gt; (388 lines), and &lt;code&gt;componentInfo.js&lt;/code&gt; (458 lines).&lt;/p&gt;

&lt;h2&gt;
  
  
  One technical decision: a tick-based capacity model with magic numbers up front
&lt;/h2&gt;

&lt;p&gt;The core of the simulation is a fixed-tick loop (&lt;code&gt;TICK_S = 0.8&lt;/code&gt; seconds simulated per tick) where every component type gets a per-node capacity constant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;APP_RPS_PER_NODE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;150&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;LC_EFFICIENCY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.08&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;      &lt;span class="c1"&gt;// least-connections spreads better&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CDN_OFFLOAD_PCT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.6&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DB_LOAD_CAP&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;220&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;         &lt;span class="c1"&gt;// load units per primary before saturation&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CB_DROP_FRACTION&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.35&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// requests shed while breaker is open&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MQ_BUFFER_CAP&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;      &lt;span class="c1"&gt;// buffered msgs per broker before drops&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;WORKER_DRAIN_PER_NODE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tradeoff is obvious: these numbers are invented, not measured. A real CDN offloads far more than 60% for cacheable content; a real app server's RPS depends on what the handler does. But the point of the tool was never absolute accuracy — it's relative behavior. Kill the only load balancer and you get a total outage. Add read replicas and the primary stops saturating. Switch the LB strategy from round robin to least-connections (that 1.08 efficiency factor) or consistent hashing and you can watch the hot-partition trap appear. The constants being rough doesn't matter as long as the &lt;em&gt;relationships&lt;/em&gt; between components behave plausibly, and keeping them as named constants in one place means anyone can tune them toward realism later.&lt;/p&gt;

&lt;p&gt;The same thinking applies to the cost model: every component has a monthly $/node figure, cost per 1M requests is derived live — and dead nodes still bill. That last detail is deliberate. Idle capacity costing money is the whole argument for auto-scaling, and the simulator lets you toggle auto-scaling and watch scaling events fire.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chaos was the fun part
&lt;/h2&gt;

&lt;p&gt;Each node accumulates damage under sustained overload and eventually crashes; auto-heal restarts it, modeled loosely on Kubernetes liveness probes. A circuit breaker opens when the primary DB saturates and sheds read pressure until recovery. You can kill any node by hand, or roll the chaos monkey. The failure I keep coming back to: search queries with no search cluster fall back to expensive scans on the app servers and DB, and p99 quietly degrades instead of erroring. Silent degradation is harder to notice than an outage, which is exactly why I kept it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke / what's next
&lt;/h2&gt;

&lt;p&gt;Honest status: the failure visuals and packet animation live in the same 980-line &lt;code&gt;Scene.jsx&lt;/code&gt; as the component models, and that file is the hardest thing in the repo to change — it needs splitting before anything else. The 3D drag-and-drop interaction (added a few commits ago along with the info modals) still feels fiddly for precise topologies. And the capacity constants, as noted, are placeholders with opinions — calibrating even one path (say, queue drain behavior) against a real benchmark would make the whole thing more trustworthy.&lt;/p&gt;

&lt;p&gt;Next up would be per-link latency visualization and a shareable-URL encoding of a topology so two people can argue about the same diagram.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install
&lt;/span&gt;npm run dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then pick the URL shortener scenario, switch demand to viral spike, kill a node, and watch what redundancy actually buys you. That's the whole pitch: fewer whiteboard assertions, more watching things break.&lt;/p&gt;




&lt;p&gt;Built by Irfan Wani. Repo: &lt;a href="https://github.com/Irfanwani/system-design-simulator" rel="noopener noreferrer"&gt;https://github.com/Irfanwani/system-design-simulator&lt;/a&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>webdev</category>
      <category>react</category>
      <category>javascript</category>
    </item>
  </channel>
</rss>
