<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: kirandeepjassal-crypto</title>
    <description>The latest articles on DEV Community by kirandeepjassal-crypto (@kirandeepjassalcrypto).</description>
    <link>https://dev.to/kirandeepjassalcrypto</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3948965%2F92a8ebec-5c78-46dc-b19b-babd45c794b0.png</url>
      <title>DEV Community: kirandeepjassal-crypto</title>
      <link>https://dev.to/kirandeepjassalcrypto</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kirandeepjassalcrypto"/>
    <language>en</language>
    <item>
      <title>Angular Performance Optimization in 2026 — OnPush, trackBy, Lazy Loading, Signals, Standalone (Real Code, Production Metrics)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Wed, 30 Sep 2026 17:34:20 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/angular-performance-optimization-in-2026-onpush-trackby-lazy-loading-signals-standalone-real-3f8m</link>
      <guid>https://dev.to/kirandeepjassalcrypto/angular-performance-optimization-in-2026-onpush-trackby-lazy-loading-signals-standalone-real-3f8m</guid>
      <description>&lt;p&gt;Angular apps don't go slow because Angular is slow. They go slow because the defaults — &lt;code&gt;Default&lt;/code&gt; change detection, NgModules with eager imports, no &lt;code&gt;trackBy&lt;/code&gt;, BehaviorSubject everywhere, no lazy loading — were designed for a world where bundle size and per-keystroke render cost weren't the user-facing scoreboard. In 2026 they are: Core Web Vitals (LCP, INP) are part of the Google ranking signal &lt;em&gt;and&lt;/em&gt; the conversion funnel.&lt;/p&gt;

&lt;p&gt;This is the production playbook for Angular 19+. Five techniques compose into the biggest wins, each with before/after code and real metrics from &lt;strong&gt;Mattrx&lt;/strong&gt; — a multi-tenant marketing-analytics SaaS (Angular 16 → 19, 540+ components, 22k LOC, 110k MAU).&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR — the five, ranked by impact
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Technique&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;th&gt;Win&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Lazy loading&lt;/strong&gt; (routes + &lt;code&gt;@defer&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Initial JS 1.2 MB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;290 KB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;LCP −67%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;OnPush&lt;/strong&gt; change detection&lt;/td&gt;
&lt;td&gt;60 CD cycles/sec&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;3/sec&lt;/strong&gt; on dashboard&lt;/td&gt;
&lt;td&gt;INP −78%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Signals&lt;/strong&gt; (replace BehaviorSubject)&lt;/td&gt;
&lt;td&gt;manual &lt;code&gt;markForCheck&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;auto-react on read&lt;/td&gt;
&lt;td&gt;fewer bugs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Standalone components&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;NgModule barrels&lt;/td&gt;
&lt;td&gt;direct imports, tree-shake&lt;/td&gt;
&lt;td&gt;bundle −18%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;trackBy&lt;/code&gt; / &lt;code&gt;track&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4,000 nodes recreated&lt;/td&gt;
&lt;td&gt;only changed rows&lt;/td&gt;
&lt;td&gt;filter 280ms → 35ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Aggregate (Angular 16 → 19 + the 5 techniques, 4 weeks, no new features):&lt;/strong&gt; LCP 4.6s → &lt;strong&gt;1.5s&lt;/strong&gt;, INP 420ms → &lt;strong&gt;80ms&lt;/strong&gt;, initial JS 1.2 MB → &lt;strong&gt;290 KB&lt;/strong&gt;, CrUX "Good" CWV 38% → &lt;strong&gt;91%&lt;/strong&gt;, trial signup conversion &lt;strong&gt;+12%&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model you need: change detection
&lt;/h2&gt;

&lt;p&gt;Every time Zone.js patches an async API (click, XHR, &lt;code&gt;setTimeout&lt;/code&gt;), Angular schedules a CD pass and walks the tree top-down. &lt;strong&gt;Default CD re-checks every binding on every component on every pass&lt;/strong&gt; — even if nothing changed. A &lt;code&gt;setInterval(fn, 100)&lt;/code&gt; triggers a full-app CD pass every 100ms. &lt;code&gt;OnPush&lt;/code&gt; prunes the tree: a component is checked only when an &lt;code&gt;@Input&lt;/code&gt; reference changes, an event fires in it, an &lt;code&gt;async&lt;/code&gt; pipe emits, a signal it reads changes, or &lt;code&gt;markForCheck()&lt;/code&gt; is called.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Lazy loading — the biggest LCP win
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;routes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Routes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;loadComponent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;import&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./pages/login.component&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;LoginComponent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;dashboard&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;loadComponent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;import&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./pages/dashboard.component&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DashboardComponent&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="na"&gt;canMatch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;authGuard&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;reports&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;loadChildren&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;import&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./pages/reports/reports.routes&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;REPORTS_ROUTES&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="na"&gt;canMatch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;authGuard&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add background prefetch so subsequent navigation is instant too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;provideRouter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;routes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;withPreloading&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;PreloadAllModules&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then &lt;code&gt;@defer&lt;/code&gt; for anything heavy that isn't above-the-fold:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;@defer (on viewport) {
  &lt;span class="nt"&gt;&amp;lt;mx-revenue-chart&lt;/span&gt; &lt;span class="na"&gt;[data]=&lt;/span&gt;&lt;span class="s"&gt;"revenue()"&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
} @placeholder {
  &lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"chart-skeleton h-64"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
} @loading (minimum 200ms) {
  &lt;span class="nt"&gt;&amp;lt;mx-spinner&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
} @error {
  &lt;span class="nt"&gt;&amp;lt;mx-chart-load-error&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;@defer&lt;/code&gt; downloads the component &lt;strong&gt;and its transitive imports&lt;/strong&gt; as a separate chunk automatically — no webpack config. Triggers: &lt;code&gt;on idle&lt;/code&gt;, &lt;code&gt;on viewport&lt;/code&gt;, &lt;code&gt;on interaction&lt;/code&gt;, &lt;code&gt;on hover&lt;/code&gt;, &lt;code&gt;on timer(2s)&lt;/code&gt;, &lt;code&gt;when condition&lt;/code&gt;. &lt;strong&gt;Result: initial bundle 1.2 MB → 290 KB, LCP 4.6s → 1.5s.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  2. OnPush — the biggest CPU win
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Component&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;selector&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;mx-kpi-card&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;standalone&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;changeDetection&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ChangeDetectionStrategy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OnPush&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;`...&lt;/span&gt;&lt;span class="se"&gt;\`&lt;/span&gt;&lt;span class="s2"&gt;,
})
export class KpiCardComponent {
  @Input() label!: string;
  @Input() value!: number;
}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Mattrx dashboard did ~&lt;strong&gt;13,000 binding evaluations/sec at idle&lt;/strong&gt;. Blanket OnPush on leaf components: &lt;strong&gt;540/sec&lt;/strong&gt;. The contract: never mutate input objects in place — create new references (&lt;code&gt;this.user = {...this.user, name}&lt;/code&gt;), or use Signals to enforce it structurally. Convert leaf components first (cards, rows, badges) — highest fan-out.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Signals — replace BehaviorSubject for state
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;@&lt;/span&gt;&lt;span class="nd"&gt;Injectable&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;providedIn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;root&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;KpiService&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;kpis&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Kpis&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;revenuePerCampaign&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;computed&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;kpis&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;campaigns&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;revenue&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;campaigns&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;http&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;get&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Kpis&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;`/api/kpis?u=&lt;/span&gt;&lt;span class="se"&gt;\$&lt;/span&gt;&lt;span class="s2"&gt;{id}&lt;/span&gt;&lt;span class="se"&gt;\`&lt;/span&gt;&lt;span class="s2"&gt;).subscribe(v =&amp;gt; this.kpis.set(v)); }
}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reading &lt;code&gt;kpis()&lt;/code&gt; in a template auto-subscribes the component; &lt;code&gt;kpis.set(...)&lt;/code&gt; marks only readers for check. No async pipe, no &lt;code&gt;markForCheck&lt;/code&gt;, no &lt;code&gt;takeUntilDestroyed&lt;/code&gt;. Use Signals for component/service state; keep RxJS for genuine async pipelines (HTTP, WebSocket, debounced search) and bridge with &lt;code&gt;toSignal&lt;/code&gt; / &lt;code&gt;toObservable&lt;/code&gt;. &lt;strong&gt;/inbox: re-renders per WebSocket message 23 → 2, 62 &lt;code&gt;takeUntilDestroyed&lt;/code&gt; removed.&lt;/strong&gt; Once everything reads signals, you're on the path to zoneless (−28 KB bundle).&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Standalone components — tree-shaking + lazy loading
&lt;/h2&gt;

&lt;p&gt;Standalone (default in Angular 19) drops the NgModule ceremony and lets the bundler see the exact import graph — no barrel of unrelated junk. Run the schematic in chunks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ng generate @angular/core:standalone
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Mattrx: 540 components in ~3 days, ~8% needed manual tweaks (mostly &lt;code&gt;forwardRef&lt;/code&gt; cycles). &lt;strong&gt;Bundle −18% on top of lazy; compile time 42s → 28s.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5. trackBy / &lt;a class="mentioned-user" href="https://dev.to/for"&gt;@for&lt;/a&gt; track — the afternoon fix
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;@for (campaign of filteredCampaigns(); track campaign.id) {
  &lt;span class="nt"&gt;&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;&lt;/span&gt;{{ campaign.name }}&lt;span class="nt"&gt;&amp;lt;/td&amp;gt;&amp;lt;td&amp;gt;&lt;/span&gt;{{ campaign.spend | currency }}&lt;span class="nt"&gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;/span&gt;
} @empty {
  &lt;span class="nt"&gt;&amp;lt;tr&amp;gt;&amp;lt;td&lt;/span&gt; &lt;span class="na"&gt;colspan=&lt;/span&gt;&lt;span class="s"&gt;"2"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;No campaigns match.&lt;span class="nt"&gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;/span&gt;
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without a stable &lt;code&gt;track&lt;/code&gt;, filtering a 4,000-row table (new array of references) destroys and recreates &lt;strong&gt;all 4,000 DOM nodes&lt;/strong&gt; — and silently unchecks checkboxes. With &lt;code&gt;track campaign.id&lt;/code&gt;: only changed rows. &lt;strong&gt;Use a stable id, never &lt;code&gt;$index&lt;/code&gt; for dynamic lists.&lt;/strong&gt; Filter latency &lt;strong&gt;280ms → 35ms&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bonus: block control flow
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;@if&lt;/code&gt; / &lt;code&gt;@for&lt;/code&gt; / &lt;code&gt;@switch&lt;/code&gt; (Angular 17+) compile smaller and faster than &lt;code&gt;*ngIf&lt;/code&gt; / &lt;code&gt;*ngFor&lt;/code&gt;, support &lt;code&gt;@else if&lt;/code&gt; and &lt;code&gt;@empty&lt;/code&gt;. Use for all new code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The right mental model
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do less work, and do it later.&lt;/strong&gt; Lazy loading = do it later. OnPush = do less. &lt;code&gt;track&lt;/code&gt; = don't recreate what didn't change. Signals = do less, more precisely. Standalone = ship only what you use.&lt;/p&gt;

&lt;p&gt;Three habits: OnPush + standalone on every new component by default; lazy-load before you optimize; profile before you tune (Chrome Performance, &lt;code&gt;ng build --stats-json&lt;/code&gt;, Lighthouse).&lt;/p&gt;

&lt;p&gt;The full guide has every before/after, the Default-vs-OnPush tree-walk diagrams, the bundle-split diagram, the full &lt;code&gt;@defer&lt;/code&gt; trigger list, the Signals interop table, and the step-by-step diagnostic checklist for inheriting a slow app:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/angular-performance-optimization-guide-onpush-trackby-lazy-loading-signals-standalone-components" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/angular-performance-optimization-guide-onpush-trackby-lazy-loading-signals-standalone-components&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/angular-performance-optimization-guide-onpush-trackby-lazy-loading-signals-standalone-components" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>angular</category>
      <category>performance</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Angular Micro Frontends in 2026 — Webpack Module Federation vs Native Federation, Independent Deployments (Real Project, Production Metrics)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:06:26 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/angular-micro-frontends-in-2026-webpack-module-federation-vs-native-federation-independent-58h5</link>
      <guid>https://dev.to/kirandeepjassalcrypto/angular-micro-frontends-in-2026-webpack-module-federation-vs-native-federation-independent-58h5</guid>
      <description>&lt;p&gt;Micro frontends sound like an architecture pattern. They're actually a &lt;strong&gt;deployment&lt;/strong&gt; pattern that happens to have UI consequences. The question they answer isn't "how do we structure code?" — it's &lt;em&gt;"can team A ship without waiting for team B, even though both teams' code runs in the same browser tab?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The wrong reason to adopt them: "our app is big." (Big apps want an Nx monorepo with feature libraries.) The &lt;strong&gt;right&lt;/strong&gt; reason: independent deployments, different release cadences, separate teams that don't want to coordinate every push to production.&lt;/p&gt;

&lt;p&gt;This is the production playbook, built on the same real app throughout: &lt;strong&gt;Mattrx&lt;/strong&gt;, a multi-tenant marketing-analytics SaaS on Angular 19, where the partner-widgets team carved out into a separate remote in early 2026. Real numbers from that split are throughout.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one idea
&lt;/h2&gt;

&lt;p&gt;Independent deployment is the whole point. The shell's HTML never re-bundles the remote — it fetches it fresh at runtime (or from cache), evaluates it in place, and reuses the singleton shared deps. Update the remote on its CDN → the next navigation picks it up. &lt;strong&gt;The shell does not redeploy when the remote ships.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Three concepts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Host&lt;/strong&gt; (shell) — loads first. Owns routing, auth, layout, in-host features. Knows the &lt;em&gt;URLs&lt;/em&gt; of its remotes, not their &lt;em&gt;code&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remote&lt;/strong&gt; — a separately-built, separately-deployed Angular sub-app that exposes lazy-loadable components/routes via a manifest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contract&lt;/strong&gt; — the named exports each remote exposes. The host imports them &lt;em&gt;by name&lt;/em&gt;; the remote promises to keep them stable.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Two paths: Webpack MF vs Native Federation
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Webpack MF&lt;/th&gt;
&lt;th&gt;Native Federation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Basis&lt;/td&gt;
&lt;td&gt;Webpack 5 ModuleFederationPlugin&lt;/td&gt;
&lt;td&gt;Import Maps + ESM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Build&lt;/td&gt;
&lt;td&gt;webpack (slower)&lt;/td&gt;
&lt;td&gt;esbuild (faster)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Config&lt;/td&gt;
&lt;td&gt;custom webpack&lt;/td&gt;
&lt;td&gt;default Angular builder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New Angular 19 project&lt;/td&gt;
&lt;td&gt;only with webpack expertise&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Recommended&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Already on webpack&lt;/td&gt;
&lt;td&gt;stay&lt;/td&gt;
&lt;td&gt;migrate later if convenient&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For Mattrx we chose &lt;strong&gt;Native Federation&lt;/strong&gt; for the two new remotes. Honest reason: it didn't require swapping our build system back to webpack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Native Federation — the modern setup
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm i &lt;span class="nt"&gt;-D&lt;/span&gt; @angular-architects/native-federation
ng add @angular-architects/native-federation &lt;span class="nt"&gt;--project&lt;/span&gt; shell &lt;span class="nt"&gt;--type&lt;/span&gt; host &lt;span class="nt"&gt;--port&lt;/span&gt; 4200
ng add @angular-architects/native-federation &lt;span class="nt"&gt;--project&lt;/span&gt; partner-widgets &lt;span class="nt"&gt;--type&lt;/span&gt; remote &lt;span class="nt"&gt;--port&lt;/span&gt; 4201
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The shell points at remotes via a small JSON manifest (&lt;code&gt;.json&lt;/code&gt;, not &lt;code&gt;.js&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json-doc"&gt;&lt;code&gt;&lt;span class="c1"&gt;// apps/shell/src/assets/manifest.json&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"partner-widgets"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://partners.mattrx.io/remoteEntry.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"labs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="s2"&gt;"https://labs.mattrx.io/remoteEntry.json"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A lazy route pulls the remote:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;loadRemoteModule&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@angular-architects/native-federation&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;APP_ROUTES&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Routes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;dashboard&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;loadChildren&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;import&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@mattrx/features/dashboard&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DASHBOARD_ROUTES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;partners&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="na"&gt;loadChildren&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;loadRemoteModule&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;partner-widgets&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./Module&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;PARTNER_ROUTES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The remote declares what it exposes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// apps/partner-widgets/federation.config.js&lt;/span&gt;
&lt;span class="nx"&gt;module&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;exports&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;withNativeFederation&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;partner-widgets&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;exposes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./Module&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                 &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./apps/partner-widgets/src/app/partner.routes.ts&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./PartnerCampaignsWidget&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./apps/partner-widgets/src/app/widgets/partner-campaigns.component.ts&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nf"&gt;shareAll&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;singleton&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;strictVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;requiredVersion&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;auto&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The day-2 reality (where the real engineering lives)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Immutable chunks, mutable manifest.&lt;/strong&gt; &lt;code&gt;remoteEntry.json&lt;/code&gt; changes every deploy → &lt;code&gt;Cache-Control: no-cache&lt;/code&gt;. The content-hashed chunks it references → cache forever. Get this wrong and a user loads &lt;em&gt;yesterday's&lt;/em&gt; manifest pointing at &lt;em&gt;today's&lt;/em&gt; missing chunks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Atomic deploys.&lt;/strong&gt; Upload chunks first, verify they return 200 from the CDN, &lt;em&gt;then&lt;/em&gt; flip the manifest pointer. If the manifest goes live before the chunks exist, anyone loading in that window crashes. We learned this the hard way — two broken minutes in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared dependency version drift — the #1 source of pain.&lt;/strong&gt; If the shell is on &lt;code&gt;@angular/core@19.1.3&lt;/code&gt; and the remote was built against &lt;code&gt;19.0.7&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;singleton: true, strictVersion: true&lt;/code&gt; — boot-time error. Loud, deterministic, right. &lt;strong&gt;Use this for Angular packages.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;strictVersion: false&lt;/code&gt; — uses whichever loaded first. Silent breakage if APIs drift. OK only for a backwards-compatible design system.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;singleton: false&lt;/code&gt; — two Angulars. Don't.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When the shell bumps an Angular minor, every remote bumps the same week — a "version-bump-only" PR per remote.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The exposes contract is a SemVer surface.&lt;/strong&gt; The host imports &lt;code&gt;'./PartnerCampaignsWidget'&lt;/code&gt; by name — that string is a public contract. Add freely, coordinate on signature changes, deprecate-for-one-release before removing. Keep a literal &lt;code&gt;EXPOSES.md&lt;/code&gt; per remote. Contract-test it in CI against the remote's staging URL, on every shell PR &lt;em&gt;and&lt;/em&gt; every remote PR.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rollback is flipping the manifest back:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws s3 &lt;span class="nb"&gt;cp &lt;/span&gt;s3://mattrx-partners/remoteEntry.PREVIOUS.json s3://mattrx-partners/remoteEntry.json
aws cloudfront create-invalidation &lt;span class="nt"&gt;--distribution-id&lt;/span&gt; ABC &lt;span class="nt"&gt;--paths&lt;/span&gt; &lt;span class="s2"&gt;"/remoteEntry.json"&lt;/span&gt;
&lt;span class="c"&gt;# Live within ~30s. No build, no redeploy.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The production numbers (after carving out 2 remotes)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Partner team deploys/week&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;15–40&lt;/strong&gt; (3–8/day)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to ship a partner hotfix&lt;/td&gt;
&lt;td&gt;25 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cross-team coordination tickets/sprint&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prod incidents from cross-team merges/mo&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shell bundle (gzipped)&lt;/td&gt;
&lt;td&gt;290 KB&lt;/td&gt;
&lt;td&gt;290 KB (unchanged)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cold-start LCP (first remote hit)&lt;/td&gt;
&lt;td&gt;1.5s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1.7s&lt;/strong&gt; (+200ms — honest cost)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm-start LCP (cached remote)&lt;/td&gt;
&lt;td&gt;1.5s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.4s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rollback on a remote regression&lt;/td&gt;
&lt;td&gt;25 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;45 seconds&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The LCP cost is real and we own it. Idle-prefetch the manifest after the shell loads and cold-start drops back to ~+40ms:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;provideAppInitializer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;requestIdleCallback&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;r&lt;/span&gt;&lt;span class="p"&gt;()));&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;prefetchRemote&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;partner-widgets&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  When NOT to use micro frontends
&lt;/h2&gt;

&lt;p&gt;The most-skipped section of every MFE article. &lt;strong&gt;Most teams shouldn't.&lt;/strong&gt; Don't reach for MFEs if you're under ~25 frontend engineers in 1–2 teams, if you want smaller bundles (that's lazy loading), if you want faster CI (&lt;code&gt;nx affected&lt;/code&gt; is cheaper), or if your teams ship together anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do&lt;/strong&gt; consider them if a team ships to a different surface (partner CDN, extension), needs a conflicting deploy cadence, embeds the same code in third-party hosts, or is an acquired team merging in without a rewrite. For Mattrx, exactly one of those was true — so we carved out one remote, not ten.&lt;/p&gt;

&lt;h2&gt;
  
  
  The right mental model
&lt;/h2&gt;

&lt;p&gt;Micro frontends are &lt;strong&gt;an operational tool dressed in architecture clothes.&lt;/strong&gt; The architecture exists &lt;em&gt;so that&lt;/em&gt; the deployment property (independent ship) becomes possible. Without the deployment property, the architecture is overhead with no upside.&lt;/p&gt;

&lt;p&gt;Three habits: adopt one remote at a time for one reason at a time; treat the &lt;code&gt;exposes&lt;/code&gt; contract as a SemVer surface (document + test in CI); coordinate the boring stuff (Angular minor bumps, design-system versions, CDN caching).&lt;/p&gt;

&lt;p&gt;The full guide has the complete Webpack MF setup too, the request-flow and deploy-flow diagrams, the embeddable-widget use case, the 8-week adoption path, and the honest "what we'd do again / what we'd skip":&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/angular-micro-frontends-module-federation-webpack-native-federation-independent-deployments-guide" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/angular-micro-frontends-module-federation-webpack-native-federation-independent-deployments-guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/angular-micro-frontends-module-federation-webpack-native-federation-independent-deployments-guide" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>angular</category>
      <category>architecture</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Angular Signals vs RxJS in 2026 — When to Use Each, Performance Benchmarks, and the Interop Pattern (Real Code)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Sat, 26 Sep 2026 17:51:32 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/angular-signals-vs-rxjs-in-2026-when-to-use-each-performance-benchmarks-and-the-interop-pattern-28b0</link>
      <guid>https://dev.to/kirandeepjassalcrypto/angular-signals-vs-rxjs-in-2026-when-to-use-each-performance-benchmarks-and-the-interop-pattern-28b0</guid>
      <description>&lt;p&gt;"Should I use Signals or RxJS for this?" is the most-asked question in modern Angular code review. The bad answer — &lt;em&gt;"use Signals everywhere, RxJS is legacy"&lt;/em&gt; — produces fragile WebSocket handling and re-invented &lt;code&gt;switchMap&lt;/code&gt;. The other bad answer — &lt;em&gt;"stick with RxJS, Signals are toys"&lt;/em&gt; — produces 30 lines of &lt;code&gt;BehaviorSubject&lt;/code&gt; + &lt;code&gt;takeUntilDestroyed&lt;/code&gt; to track which tab is active.&lt;/p&gt;

&lt;p&gt;The right answer is mechanical, not philosophical: &lt;strong&gt;Signals are pull-based, synchronous, fine-grained reactive values; RxJS is push-based, asynchronous, time-aware event composition.&lt;/strong&gt; Each is excellent at exactly one half of the problem. Real code, benchmarks, and production numbers from migrating &lt;strong&gt;Mattrx&lt;/strong&gt; — a multi-tenant marketing-analytics SaaS (540 components, 22k LOC, 110k MAU).&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model (the one chart you need)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                        SIGNALS                  RxJS
Direction              PULL (read latest)       PUSH (subscribe, get next)
Time                   SYNCHRONOUS              ASYNCHRONOUS
Granularity            Per-value, fine-grained  Per-stream, coarse-grained
Cancellation           N/A (no async)           takeUntil / abort
Memoization            Built-in (computed)      Manual (shareReplay)
OnPush integration     Auto-marks for check     Needs async pipe / markForCheck
Conceptual fit         "the current value of X" "the stream of X events over time"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask one question:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"I need to know the latest X" -&amp;gt; &lt;strong&gt;Signal&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;"I need to react to every event of X over time" -&amp;gt; &lt;strong&gt;RxJS&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That resolves 90% of the real ambiguity.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Signals win (state)
&lt;/h2&gt;

&lt;p&gt;Local UI state, derived values, service-level shared state, side effects — all Signal territory. The service-level case is where the code shrinks most:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// BEFORE - BehaviorSubject ceremony (29 lines)&lt;/span&gt;
&lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="nx"&gt;dateRange$&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nx"&gt;BehaviorSubject&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;DateRange&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(...);&lt;/span&gt;
&lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="nx"&gt;status$&lt;/span&gt;    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nx"&gt;BehaviorSubject&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Status&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;active&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="c1"&gt;// ...asObservable(), setters, combineLatest, shareReplay...&lt;/span&gt;

&lt;span class="c1"&gt;// AFTER - Signals (9 lines)&lt;/span&gt;
&lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;dateRange&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;DateRange&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;lastWeek&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="na"&gt;to&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;today&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;status&lt;/span&gt;    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Status&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;active&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;filters&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;computed&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;dateRange&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dateRange&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;}));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;computed()&lt;/code&gt; &lt;strong&gt;is&lt;/strong&gt; the multicast — lazy, memoized, O(1) cached reads. No &lt;code&gt;shareReplay&lt;/code&gt;, no &lt;code&gt;combineLatest&lt;/code&gt;, no &lt;code&gt;next()&lt;/code&gt;. Mattrx had 240 BehaviorSubjects; the cleanup deleted ~1,400 lines of state code.&lt;/p&gt;

&lt;h2&gt;
  
  
  When RxJS wins (events over time) — and Signals genuinely can't
&lt;/h2&gt;

&lt;p&gt;Debounced search with race cancellation is the canonical example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;results$&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;valueChanges&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pipe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nf"&gt;debounceTime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;250&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="nf"&gt;distinctUntilChanged&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="nf"&gt;switchMap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;q&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;http&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`/api/search?q=&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;pipe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;catchError&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt;&lt;span class="p"&gt;([])))),&lt;/span&gt;
  &lt;span class="nf"&gt;shareReplay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="nf"&gt;takeUntilDestroyed&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To replicate this in pure Signals you'd hand-roll a &lt;code&gt;setTimeout&lt;/code&gt; debounce, an &lt;code&gt;AbortController&lt;/code&gt; you abort per keystroke, try/catch, and a loading flag — ~40 lines re-inventing &lt;code&gt;switchMap&lt;/code&gt;. Same story for &lt;strong&gt;WebSocket streams&lt;/strong&gt; (&lt;code&gt;webSocket()&lt;/code&gt; + &lt;code&gt;retry({delay})&lt;/code&gt; + &lt;code&gt;share()&lt;/code&gt;), &lt;strong&gt;drag-to-resize&lt;/strong&gt; (&lt;code&gt;mousedown -&amp;gt; switchMap(mousemove) -&amp;gt; takeUntil(mouseup)&lt;/code&gt;), and &lt;strong&gt;multi-source merge&lt;/strong&gt; (&lt;code&gt;combineLatest&lt;/code&gt; + &lt;code&gt;throttleTime&lt;/code&gt;). Anything with &lt;strong&gt;debounce / throttle / retry / timeout / switchMap / scan&lt;/strong&gt; in its description belongs in RxJS. &lt;strong&gt;Signals do not model time.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance (numbers, not vibes)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Operation&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;signal()&lt;/code&gt; read&lt;/td&gt;
&lt;td&gt;~12 ns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;signal.set(v)&lt;/code&gt; write (no listeners)&lt;/td&gt;
&lt;td&gt;~50 ns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;BehaviorSubject.next(v)&lt;/code&gt; (one sync sub)&lt;/td&gt;
&lt;td&gt;~700 ns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;computed()&lt;/code&gt; read (warm cache)&lt;/td&gt;
&lt;td&gt;~15 ns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observable &lt;code&gt;pipe(map).subscribe()&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;~3.4 us&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Signals are ~14x faster than BehaviorSubject for "publish a value, one sync reader." But the real win is structural, not micro: on Mattrx's &lt;code&gt;/inbox&lt;/code&gt;, re-renders per WebSocket message went &lt;strong&gt;23 -&amp;gt; 2&lt;/strong&gt; — because fine-grained dependency tracking re-renders only the &lt;em&gt;specific slice&lt;/em&gt; that changed, not 23 components subscribed via &lt;code&gt;| async&lt;/code&gt;. And with &lt;strong&gt;zoneless&lt;/strong&gt; + Signals, idle change-detection passes on &lt;code&gt;/dashboard&lt;/code&gt; dropped to &lt;strong&gt;0&lt;/strong&gt; (only run when a signal changes).&lt;/p&gt;

&lt;p&gt;The honest caveat: Signals don't make a slow app fast on their own — they make change-detection &lt;strong&gt;precise&lt;/strong&gt;. The big wins come combined with OnPush + standalone + &lt;code&gt;@for&lt;/code&gt; track.&lt;/p&gt;

&lt;h2&gt;
  
  
  Interop — you don't have to pick
&lt;/h2&gt;

&lt;p&gt;Bridge at the boundary and keep each half native:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// stream -&amp;gt; state (HTTP as an OnPush-aware Signal)&lt;/span&gt;
&lt;span class="nx"&gt;campaigns&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;toSignal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;http&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;get&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Campaign&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/campaigns&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;initialValue&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// state -&amp;gt; stream (debounce a Signal-driven input through RxJS, back to a Signal)&lt;/span&gt;
&lt;span class="nx"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;toSignal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nf"&gt;toObservable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;pipe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;debounceTime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;250&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;distinctUntilChanged&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="nf"&gt;switchMap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;q&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;http&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`/api/search?q=&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;))),&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;initialValue&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;toSignal()&lt;/code&gt; and &lt;code&gt;toObservable()&lt;/code&gt; aren't workarounds — they're the design. (One gotcha: &lt;code&gt;toSignal&lt;/code&gt; without &lt;code&gt;initialValue&lt;/code&gt; returns &lt;code&gt;Signal&amp;lt;T | undefined&amp;gt;&lt;/code&gt; — always supply &lt;code&gt;initialValue&lt;/code&gt; for HTTP.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision tree
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is this "the current value of X" (state)? -&amp;gt; SIGNAL
  ├── Local to a component -&amp;gt; signal()
  ├── Shared across the app -&amp;gt; signal() in a service
  └── Derived from other signals -&amp;gt; computed()
Otherwise it's an event stream / time-based / async:
  ├── HTTP -&amp;gt; HttpClient Observable; toSignal() if read as state
  ├── debounce/throttle/switchMap input -&amp;gt; RxJS
  ├── WebSocket / SSE -&amp;gt; RxJS (webSocket + retry + share)
  └── composing DOM events (drag, gesture) -&amp;gt; RxJS (fromEvent + operators)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The right mental model
&lt;/h2&gt;

&lt;p&gt;In one line: &lt;strong&gt;Signals are nouns. RxJS is verbs.&lt;/strong&gt; "What is the current set of selected campaigns?" -&amp;gt; Signal. "What's happening when the user types?" -&amp;gt; RxJS. Three habits: default to Signals for state, default to RxJS for streams, bridge with interop instead of fighting it.&lt;/p&gt;

&lt;p&gt;The full guide has all 10 code examples each way, the full benchmark tables, the end-to-end &lt;code&gt;/campaigns&lt;/code&gt; component using both, and the 4-week Mattrx migration path (including the bugs that bit us):&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/angular-signals-vs-rxjs-when-to-use-which-performance-comparison-guide" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/angular-signals-vs-rxjs-when-to-use-which-performance-comparison-guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/angular-signals-vs-rxjs-when-to-use-which-performance-comparison-guide" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>angular</category>
      <category>rxjs</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Enterprise Angular Architecture in 2026 — Core, Shared, Feature + Nx Monorepo (Real Project Layout, Production Metrics)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Wed, 23 Sep 2026 18:44:42 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/enterprise-angular-architecture-in-2026-core-shared-feature-nx-monorepo-real-project-layout-1kd4</link>
      <guid>https://dev.to/kirandeepjassalcrypto/enterprise-angular-architecture-in-2026-core-shared-feature-nx-monorepo-real-project-layout-1kd4</guid>
      <description>&lt;p&gt;Most Angular apps don't fail because the framework runs out of steam. They fail around the 18-month mark, when the directory tree turns into a swamp: &lt;code&gt;shared/utils/helpers/index.ts&lt;/code&gt; re-exports everything, three teams import each other's internals through it, CI takes 14 minutes because &lt;em&gt;one&lt;/em&gt; test change rebuilds the world, and nobody can ship a feature without breaking two others.&lt;/p&gt;

&lt;p&gt;The four patterns here — &lt;strong&gt;Feature, Core, Shared, and an Nx monorepo&lt;/strong&gt; with enforced library boundaries — are what stops that. Not 2017 NgModule dogma; the &lt;em&gt;directory-and-dependency-graph discipline&lt;/em&gt; underneath, which works just as cleanly with Angular 19 standalone components.&lt;/p&gt;

&lt;p&gt;Every section uses the same real app: &lt;strong&gt;Mattrx&lt;/strong&gt;, a multi-tenant marketing-analytics SaaS — Angular 19 standalone, 4 apps in one repo, 540+ components, 22k LOC, 5 product teams sharing a design system and data-access layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model: the dependency graph IS the architecture
&lt;/h2&gt;

&lt;p&gt;The single biggest decision in a large Angular app is &lt;strong&gt;what's allowed to import what&lt;/strong&gt;. The arrows go ONE way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Apps (shells) -&amp;gt; Features (lazy) -&amp;gt; Shared (ui, util)
                     |
                     v
                   Core (imported only by the app shell)

Never:
  Shared -&amp;gt; Features    (shared can't know about features)
  Features -&amp;gt; Features  (features can't reach across each other)
  Anything -&amp;gt; Core, except the app shell
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Owns&lt;/th&gt;
&lt;th&gt;Imported by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Core&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;auth, interceptors, error handler, config, logger&lt;/td&gt;
&lt;td&gt;only main.ts, once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Shared&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;buttons, tables, pipes, DTOs - no business logic&lt;/td&gt;
&lt;td&gt;any feature&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Feature&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;routes, components, services, state for one capability&lt;/td&gt;
&lt;td&gt;the app router (lazy)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Nx libs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;tagged boundaries, ESLint enforcement, nx affected&lt;/td&gt;
&lt;td&gt;the whole repo&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Core — the app-singleton layer (modern, no NgModule)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;provideCore&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nx"&gt;EnvironmentProviders&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;makeEnvironmentProviders&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
    &lt;span class="nx"&gt;AuthService&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ConfigService&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;LoggerService&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;provide&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ErrorHandler&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;useClass&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;GlobalErrorHandler&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="nf"&gt;provideHttpClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;withInterceptors&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nx"&gt;authInterceptor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;errorInterceptor&lt;/span&gt;&lt;span class="p"&gt;])),&lt;/span&gt;
    &lt;span class="nf"&gt;provideAppInitializer&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;inject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ConfigService&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;()),&lt;/span&gt;
  &lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// apps/customer/src/main.ts — called exactly once&lt;/span&gt;
&lt;span class="nf"&gt;bootstrapApplication&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;AppComponent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;providers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;provideCore&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nf"&gt;provideRouter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;routes&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;provideAnimationsAsync&lt;/span&gt;&lt;span class="p"&gt;()],&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;makeEnvironmentProviders&lt;/code&gt; is the modern equivalent of the "Core module can only be imported once" guard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shared — the cardinal rule
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Shared cannot import from Features. Ever.&lt;/strong&gt; If a "shared" component needs to know about campaigns, it isn't shared — it's a campaigns component.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Each Shared component is its own Nx library, imported at the leaf — &lt;strong&gt;no barrel &lt;code&gt;@mattrx/shared/ui&lt;/code&gt; that re-exports everything&lt;/strong&gt; (so the bundler tree-shakes per-library, not per-monolith):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;MxButton&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@mattrx/shared/ui/button&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;MxTable&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@mattrx/shared/ui/table&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="s2"&gt;```

Don't ship a `&lt;/span&gt;&lt;span class="nx"&gt;SharedModule&lt;/span&gt;&lt;span class="s2"&gt;` in 2026 — importing it to use `&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;mx&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;button&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="s2"&gt;` drags in `&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;mx&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;modal&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="s2"&gt;`, `&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;mx&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;toast&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="s2"&gt;`, `&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;mx&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;table&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="s2"&gt;`, and everything else. Use standalone leaves.

## Feature — one bounded business capability, lazy-loaded



```&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;
&lt;span class="c1"&gt;// app.routes.ts — the app router doesn't know the feature's internal structure&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;campaigns&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;canMatch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;authGuard&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="nx"&gt;loadChildren&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;import&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@mattrx/features/campaigns&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CAMPAIGNS_ROUTES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bundler emits one chunk per feature. Landing on /dashboard downloads main.js (~280 KB) + dashboard.chunk.js (~140 KB) — &lt;em&gt;not&lt;/em&gt; campaigns, inbox, reports. Feature state is Signals-first (toSignal at the HTTP boundary, computed for derived).&lt;/p&gt;

&lt;h2&gt;
  
  
  Nx — the structural firewall
&lt;/h2&gt;

&lt;p&gt;Tags declare what each library &lt;em&gt;is&lt;/em&gt; and &lt;em&gt;belongs to&lt;/em&gt;, and ESLint enforces the arrows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json-doc"&gt;&lt;code&gt;&lt;span class="nl"&gt;"@nx/enforce-module-boundaries"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"depConstraints"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sourceTag"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"type:feature"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;     &lt;/span&gt;&lt;span class="nl"&gt;"onlyDependOnLibsWithTags"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"type:ui"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"type:data-access"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"type:util"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sourceTag"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"type:ui"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;          &lt;/span&gt;&lt;span class="nl"&gt;"onlyDependOnLibsWithTags"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"type:ui"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"type:util"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sourceTag"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"type:util"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nl"&gt;"onlyDependOnLibsWithTags"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"type:util"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sourceTag"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"scope:customer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="nl"&gt;"onlyDependOnLibsWithTags"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"scope:customer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s2"&gt;"scope:shared"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;type:ui&lt;/code&gt; library that imports a &lt;code&gt;type:feature&lt;/code&gt; library is an ESLint error in CI. That's the difference between "we have a convention" and "the codebase enforces the convention."&lt;/p&gt;

&lt;p&gt;And &lt;code&gt;nx affected&lt;/code&gt; runs CI only for what the PR touched:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx nx affected --target=build --base=origin/main&lt;/span&gt;   &lt;span class="c1"&gt;# 2.5 min, not 12&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The before/after (after the structure landed, 4 weeks)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CI on a typical PR&lt;/td&gt;
&lt;td&gt;22 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;affected:build&lt;/td&gt;
&lt;td&gt;12 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.5 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;affected:test&lt;/td&gt;
&lt;td&gt;8 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;90 s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Initial JS (gzipped)&lt;/td&gt;
&lt;td&gt;1.2 MB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;290 KB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cross-team merge conflicts/wk&lt;/td&gt;
&lt;td&gt;~6&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~1&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scaffold a new feature&lt;/td&gt;
&lt;td&gt;~3 days&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~4 hours&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Where does this go?" PR comments/wk&lt;/td&gt;
&lt;td&gt;~12&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~1&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Velocity (features/sprint)&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The migration path (4 weeks, not one big PR)
&lt;/h2&gt;

&lt;p&gt;Week 1 — tease apart &lt;strong&gt;Core&lt;/strong&gt; (move Auth, interceptors, error handler, logger, config; replace &lt;code&gt;CoreModule&lt;/code&gt; with &lt;code&gt;provideCore()&lt;/code&gt;). Week 2 — tease apart &lt;strong&gt;Shared&lt;/strong&gt; (one lib per component, tag everything, turn on &lt;code&gt;@nx/enforce-module-boundaries&lt;/code&gt;, fix violations). Week 3 — extract &lt;strong&gt;Features&lt;/strong&gt; one at a time, most isolated first, ~1-2/day. Week 4 — enable &lt;code&gt;nx affected&lt;/code&gt; in CI + add any second frontend as a new app that reuses your shared libs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The right mental model
&lt;/h2&gt;

&lt;p&gt;Enterprise Angular at scale isn't NgModules vs standalone — it's &lt;strong&gt;the dependency graph being a layered DAG, and the codebase enforcing that picture so humans don't have to&lt;/strong&gt;. Core = app-singletons imported once. Shared = presentation + utilities that know no business concepts. Features = lazy-loaded leaves that can't reach each other. Nx = makes the graph visible, enforces it, makes CI fast.&lt;/p&gt;

&lt;p&gt;Three habits: generate every library (&lt;code&gt;nx g&lt;/code&gt;) with tags from day one; treat boundary violations as test failures; keep apps as thin shells.&lt;/p&gt;

&lt;p&gt;The full guide has the real Mattrx layout with file counts, the classic-vs-standalone code for every layer, the &lt;code&gt;nx graph&lt;/code&gt;, the full ESLint config, the decision tree for "where does this go?", and the week-by-week migration path:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/enterprise-angular-architecture-feature-core-shared-modules-nx-monorepo-guide" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/enterprise-angular-architecture-feature-core-shared-modules-nx-monorepo-guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/enterprise-angular-architecture-feature-core-shared-modules-nx-monorepo-guide" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>angular</category>
      <category>architecture</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>LLMs &amp; Transformers from First Principles (2026) — Tokenization, Attention, LoRA, and a Tiny GPT You Can Train (with PyTorch Code)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:18:08 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/llms-transformers-from-first-principles-2026-tokenization-attention-lora-and-a-tiny-gpt-h13</link>
      <guid>https://dev.to/kirandeepjassalcrypto/llms-transformers-from-first-principles-2026-tokenization-attention-lora-and-a-tiny-gpt-h13</guid>
      <description>&lt;p&gt;Everyone uses LLMs in 2026. Far fewer can explain what happens between &lt;code&gt;text in&lt;/code&gt; and &lt;code&gt;text out&lt;/code&gt;. The gap matters because &lt;em&gt;every&lt;/em&gt; LLM problem — bad outputs, high latency, wrong answers, costly fine-tunes — is solved by knowing which mechanism inside the model is responsible.&lt;/p&gt;

&lt;p&gt;This rebuilds the LLM stack piece by piece with real PyTorch and Hugging Face code: tokenization, embeddings, self-attention, the transformer block, a tiny GPT you can train on your laptop, pretraining vs fine-tuning vs LoRA vs DPO, and decoding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one idea
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;An LLM is a function from token sequences to a next-token probability distribution.&lt;/strong&gt; That's it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sequence&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;bos&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; sky&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; is&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;done&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;probs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LLM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sequence&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;next_token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;sequence&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;next_token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every chat API, code assistant, agent, and RAG system is a sampling loop around "predict the next token."&lt;/p&gt;

&lt;h2&gt;
  
  
  Tokenization — text becomes integers first
&lt;/h2&gt;

&lt;p&gt;A model can't read characters. Modern LLMs use &lt;strong&gt;subword tokenization&lt;/strong&gt; (BPE for GPT/Llama/Mistral, WordPiece for BERT, SentencePiece for T5).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;
&lt;span class="n"&gt;tok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;tok&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;convert_ids_to_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tok&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Mumbai and Indore are cities in India.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="c1"&gt;# ['Mumbai', 'and', 'Ind', 'ore', 'are', 'cities', 'in', 'India', '.']
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;"Mumbai" is &lt;strong&gt;1 token&lt;/strong&gt;; "Indore" is &lt;strong&gt;2&lt;/strong&gt;. Token counts drive API cost, context window, and latency — French text costs 2-3x more tokens than English. &lt;strong&gt;Tokenization is the source of most "why does the model do X?" bugs.&lt;/strong&gt; And always use &lt;code&gt;apply_chat_template&lt;/code&gt; — hand-concatenating "user: ... assistant: ..." is a classic bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-attention — the one idea that changed everything
&lt;/h2&gt;

&lt;p&gt;For each token, compute a weighted sum of every other token's &lt;strong&gt;value&lt;/strong&gt;, where the weights depend on relevance. Every token gets three roles: Query ("what am I looking for?"), Key ("what do I represent?"), Value ("what do I contribute?").&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Attention(Q, K, V) = softmax(Q @ K.T / sqrt(d)) @ V
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The whole revolution is one matrix multiply + a softmax:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Q&lt;/span&gt; &lt;span class="o"&gt;@&lt;/span&gt; &lt;span class="n"&gt;K&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;d_k&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;mask&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;masked_fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mask&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-inf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;softmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="o"&gt;=-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;@&lt;/span&gt; &lt;span class="n"&gt;V&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;strong&gt;causal mask&lt;/strong&gt; (lower-triangular) is why GPT can't see the future. &lt;strong&gt;Multi-head attention&lt;/strong&gt; runs several of these in parallel. In production, use Flash Attention via &lt;code&gt;F.scaled_dot_product_attention&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The transformer block
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;x -&amp;gt; LayerNorm -&amp;gt; MultiHeadAttention -&amp;gt; +residual -&amp;gt;
  -&amp;gt; LayerNorm -&amp;gt; FeedForward -&amp;gt; +residual -&amp;gt; output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Residuals&lt;/strong&gt; are the gradient highway — without them, deep transformers don't train.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LayerNorm&lt;/strong&gt; (Pre-LN) stabilizes training.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feed-forward&lt;/strong&gt; is a 2-layer MLP per token, usually &lt;code&gt;d_ff = 4 x d_model&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stack 6-96 of these and you have a model. Llama-3-8B is 32 blocks; GPT-3 is 96. The guide has a full runnable ~200-line GPT you can train on Shakespeare on a laptop — &lt;strong&gt;that's the entire architecture of GPT-3, just bigger and on more data.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Training, decoded
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pretraining&lt;/strong&gt; = predict the next token on 5-15 trillion tokens. The intelligence emerges from that one boring objective at scale ($10M-$100M for a frontier model).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning&lt;/strong&gt; — start with prompting/few-shot; if that's not enough, &lt;strong&gt;LoRA&lt;/strong&gt; trains ~0.1% of params at ~1000x less memory:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;lora_config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LoraConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lora_alpha&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;target_modules&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;q_proj&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;k_proj&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;v_proj&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;o_proj&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;lora_dropout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.05&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;task_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CAUSAL_LM&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_peft_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lora_config&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# trainable%: 0.4%, adapter ~130 MB
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;QLoRA&lt;/strong&gt; (LoRA on a 4-bit base) fine-tunes a 70B model on a single 24 GB GPU — the standard recipe for small teams in 2026.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SFT&lt;/strong&gt; trains on (instruction, response); &lt;strong&gt;DPO&lt;/strong&gt; on (prompt, chosen, rejected) — much simpler than classic RLHF.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Decoding
&lt;/h2&gt;

&lt;p&gt;Greedy is deterministic but repetitive. The production default is temperature + top-p:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;do_sample&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;top_p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                     &lt;span class="n"&gt;repetition_penalty&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.05&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_new_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Chat: &lt;code&gt;0.7&lt;/code&gt;. Code: &lt;code&gt;0.0-0.3&lt;/code&gt;. For production, swap &lt;code&gt;model.generate&lt;/code&gt; for &lt;strong&gt;vLLM&lt;/strong&gt; — paged attention + continuous batching, often 5-20x faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest stuff
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Context length costs scale &lt;strong&gt;quadratically&lt;/strong&gt; — 32k isn't 4x 8k, it's ~16x.&lt;/li&gt;
&lt;li&gt;Most prod LLM systems are 90% retrieval + 10% LLM.&lt;/li&gt;
&lt;li&gt;Fine-tuning is rarely the answer — try prompting, RAG, and a bigger model first.&lt;/li&gt;
&lt;li&gt;Evals beat vibes — build a golden set of 50-200 queries before you fine-tune.&lt;/li&gt;
&lt;li&gt;Inference cost dominates training cost. Optimize for serving.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The right mental model
&lt;/h2&gt;

&lt;p&gt;An LLM is not magic. It's: a tokenizer, an embedding table, a stack of transformer blocks, and a head that projects back to vocabulary. Pretraining gives general competence; fine-tuning teaches your task; decoding turns probabilities into text.&lt;/p&gt;

&lt;p&gt;Three habits: always inspect the tokens first; build the eval set before the model; reach for the smallest model that works.&lt;/p&gt;

&lt;p&gt;The full guide has the complete PyTorch — scaled dot-product + multi-head attention, the transformer block, the runnable tiny GPT + training loop, full/LoRA/QLoRA fine-tuning, DPO, the decoding recipes, RAG, tool-use, and the 2026 model menu:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/llms-transformers-tokenization-attention-embeddings-pretraining-finetuning-pytorch-guide" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/llms-transformers-tokenization-attention-embeddings-pretraining-finetuning-pytorch-guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/llms-transformers-tokenization-attention-embeddings-pretraining-finetuning-pytorch-guide" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>python</category>
      <category>llm</category>
    </item>
    <item>
      <title>AI Deployment &amp; MLOps in 2026 — Serving, Monitoring, Cloud, and a Full End-to-End Project (with Code)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Sat, 19 Sep 2026 08:16:31 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/ai-deployment-mlops-in-2026-serving-monitoring-cloud-and-a-full-end-to-end-project-with-3f3c</link>
      <guid>https://dev.to/kirandeepjassalcrypto/ai-deployment-mlops-in-2026-serving-monitoring-cloud-and-a-full-end-to-end-project-with-3f3c</guid>
      <description>&lt;p&gt;A model that gets 0.94 F1 in a notebook is worth nothing until it's serving real requests, staying healthy under load, and being retrained before it goes stale. &lt;strong&gt;The gap between "trained" and "in production" is where most AI projects quietly die&lt;/strong&gt; — not because the model was bad, but because nobody owned deployment, monitoring, and maintenance.&lt;/p&gt;

&lt;p&gt;This is the MLOps playbook: deploying models (REST/batch/streaming/serverless), best practices (versioning, CI/CD, reproducibility, the registry), monitoring live systems (drift, performance, business metrics), cloud strategies, and a complete end-to-end app — with real code throughout.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one idea
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Deployment is not a &lt;code&gt;model.pkl&lt;/code&gt; on a server.&lt;/strong&gt; It's a versioned artifact + reproducible serving image + an API + health checks + autoscaling + monitoring + a rollback path. MLOps = DevOps for ML &lt;em&gt;plus&lt;/em&gt; the things ML adds: data versioning, a model registry, experiment tracking, and &lt;strong&gt;retraining&lt;/strong&gt; (because models go stale in a way code never does).&lt;/p&gt;

&lt;h2&gt;
  
  
  Three serving patterns — pick by latency
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Real-time (REST/gRPC)&lt;/td&gt;
&lt;td&gt;ms&lt;/td&gt;
&lt;td&gt;per-request, now (fraud check, chatbot)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch&lt;/td&gt;
&lt;td&gt;minutes-hours&lt;/td&gt;
&lt;td&gt;scheduled bulk scoring (nightly churn)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Streaming&lt;/td&gt;
&lt;td&gt;sub-second&lt;/td&gt;
&lt;td&gt;react to events (anomaly on a sensor stream)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The bug that kills more models than any other: serving/training skew
&lt;/h2&gt;

&lt;p&gt;The serving code preprocesses input differently than training did. Fix it by saving the &lt;em&gt;entire pipeline&lt;/em&gt; — preprocessing + model — as one artifact:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Pipeline&lt;/span&gt;&lt;span class="p"&gt;([(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prep&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;preprocessor&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;LGBMClassifier&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;))])&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;joblib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dump&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model.joblib&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# preprocessing + model travel together. Cannot drift apart.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At serving time, &lt;code&gt;joblib.load("model.joblib").predict(raw_df)&lt;/code&gt; applies the exact training-time preprocessing. &lt;strong&gt;This single discipline prevents the most common production failure.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A real-time REST service (FastAPI) — three non-negotiables
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;joblib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model.joblib&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;# 1. load once at startup
&lt;/span&gt;&lt;span class="n"&gt;MODEL_VERSION&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1.4.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/predict&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;features&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;CustomerFeatures&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;features&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;model_dump&lt;/span&gt;&lt;span class="p"&gt;()])&lt;/span&gt;
    &lt;span class="n"&gt;proba&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict_proba&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;churn_probability&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;proba&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;will_churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;proba&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model_version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;MODEL_VERSION&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;# 2. ALWAYS return the version
&lt;/span&gt;
&lt;span class="nd"&gt;@app.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/health/ready&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ready&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ready&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;      &lt;span class="c1"&gt;# 3. readiness + liveness endpoints
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Version everything — the three artifacts
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;code (git SHA) + data (DVC / dataset hash) + model (registry version) = a reproducible prediction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With a registry (MLflow), serving loads by &lt;em&gt;stage&lt;/em&gt;, not a file path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mlflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pyfunc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;models:/churn-model/Production&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Swapping models becomes "promote version 6 to Production" — no redeploy, full audit trail, and rollback is just promoting the previous version.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ML-specific CI/CD piece: the quality gate
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Quality gate - block deploy if metrics regress&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;python - &amp;lt;&amp;lt;'PY'&lt;/span&gt;
    &lt;span class="s"&gt;import json; m = json.load(open("metrics.json"))&lt;/span&gt;
    &lt;span class="s"&gt;assert m["roc_auc"] &amp;gt;= 0.82, f"ROC-AUC {m['roc_auc']} below 0.82"&lt;/span&gt;
    &lt;span class="s"&gt;assert m["f1"] &amp;gt;= 0.70, f"F1 {m['f1']} below 0.70"&lt;/span&gt;
    &lt;span class="s"&gt;PY&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A new model only ships if it meets minimum metrics on a held-out set.&lt;/p&gt;

&lt;h2&gt;
  
  
  Monitor four layers (not just accuracy)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;System health&lt;/strong&gt; - latency (p50/p95/p99), error rate, throughput (Prometheus + Grafana).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data drift&lt;/strong&gt; - production inputs drift from training inputs; accuracy silently drops &lt;em&gt;with no error thrown&lt;/em&gt;. Detect with a KS-test per feature; tools: Evidently, NannyML, WhyLabs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concept drift&lt;/strong&gt; - inputs look the same but the input-&amp;gt;output relationship changed. Shows up as falling live AUC once true outcomes arrive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Business metrics&lt;/strong&gt; - revenue saved, fraud caught, tickets deflected. A model with great AUC that doesn't move the business metric is a science project, not a product.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Cloud strategy: containerize once, then choose
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;Ops&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Managed endpoints (SageMaker, Vertex, Azure ML)&lt;/td&gt;
&lt;td&gt;fastest to prod&lt;/td&gt;
&lt;td&gt;lowest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kubernetes (AKS/EKS/GKE)&lt;/td&gt;
&lt;td&gt;full control, multi-model, GPU sharing&lt;/td&gt;
&lt;td&gt;highest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Serverless (Cloud Run, Lambda)&lt;/td&gt;
&lt;td&gt;low/spiky traffic, scale-to-zero&lt;/td&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Start with a managed endpoint unless you have a reason not to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The end-to-end project
&lt;/h2&gt;

&lt;p&gt;A customer-churn system wired end to end: &lt;code&gt;train.py&lt;/code&gt; (train -&amp;gt; quality gate -&amp;gt; register to MLflow) -&amp;gt; FastAPI service in Docker (loads from the registry, exports Prometheus metrics) -&amp;gt; &lt;code&gt;monitor.py&lt;/code&gt; (KS-test drift + live-AUC performance) -&amp;gt; GitHub Actions CI/CD with a &lt;code&gt;schedule&lt;/code&gt; block that retrains every Monday, runs the same quality gate, and only ships if the new model clears the bar.&lt;/p&gt;

&lt;h2&gt;
  
  
  The MLOps maturity ladder
&lt;/h2&gt;

&lt;p&gt;Level 0 (notebook -&amp;gt; model.pkl -&amp;gt; manual upload) -&amp;gt; Level 1 (automated training) -&amp;gt; Level 2 (CI/CD + registry + quality gate) -&amp;gt; Level 3 (drift + performance + business monitoring) -&amp;gt; Level 4 (auto-retraining). &lt;strong&gt;Most teams should target Level 2-3.&lt;/strong&gt; Match the maturity to the stakes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The right mental model
&lt;/h2&gt;

&lt;p&gt;MLOps is treating a model as a &lt;strong&gt;living production system&lt;/strong&gt;, not a deliverable. Three habits:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;One artifact, one version, always returned.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gate every deploy on quality.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assume the model will rot - monitor and retrain.&lt;/strong&gt; Drift is the default, not an edge case. Build the feedback loop before you need it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The full guide has every file (train/serve/monitor/CI), the Dockerfile, the K8s manifest with HPA, the Azure ML managed-endpoint code, and the mental checklist for \"in production\":&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/ai-deployment-mlops-production-serving-monitoring-cloud-end-to-end-project" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/ai-deployment-mlops-production-serving-monitoring-cloud-end-to-end-project&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/ai-deployment-mlops-production-serving-monitoring-cloud-end-to-end-project" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>mlops</category>
      <category>python</category>
    </item>
    <item>
      <title>No Python, No PhD: Train Real ML Models in C# with ML.NET (Regression, Classification, Clustering)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Thu, 17 Sep 2026 19:08:04 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/no-python-no-phd-train-real-ml-models-in-c-with-mlnet-regression-classification-clustering-ml6</link>
      <guid>https://dev.to/kirandeepjassalcrypto/no-python-no-phd-train-real-ml-models-in-c-with-mlnet-regression-classification-clustering-ml6</guid>
      <description>&lt;p&gt;Mattrx ran its predictive features on a Python scikit-learn microservice for two years. We replaced it with &lt;strong&gt;ML.NET running in-process&lt;/strong&gt; inside the existing .NET 9 app — and trained real regression, classification, and clustering models in C# with no separate service, no second language in production, and no implementing gradient descent by hand.&lt;/p&gt;

&lt;p&gt;If you're a .NET team that "does ML" by shipping a Python sidecar, you're paying a tax most teams never question: a second runtime, a second deploy pipeline, a second on-call surface, a cross-process hop on every prediction, and a data contract that drifts between two languages. For &lt;strong&gt;classical&lt;/strong&gt; ML — regression, classification, clustering — you usually don't need any of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before vs after
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Before (Python sidecar)&lt;/th&gt;
&lt;th&gt;After (ML.NET in-process)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Languages in production&lt;/td&gt;
&lt;td&gt;C# &lt;strong&gt;and&lt;/strong&gt; Python&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;C# only&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deploy pipelines&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prediction path&lt;/td&gt;
&lt;td&gt;HTTP to Flask -&amp;gt; scikit-learn&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;in-memory call&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prediction p95&lt;/td&gt;
&lt;td&gt;45 ms&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.8 ms&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model artifact&lt;/td&gt;
&lt;td&gt;pickle in a container image&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2 MB .zip, loaded by the app&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;On-call surface&lt;/td&gt;
&lt;td&gt;app + ML service&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;app only&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infra cost&lt;/td&gt;
&lt;td&gt;+$160/mo for the ML service&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0 (decommissioned)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The one mental shift
&lt;/h2&gt;

&lt;p&gt;The reason .NET teams reach for Python isn't the models — it's a belief that "real ML needs Python and a math background." For deep-learning research, fair. For the bread-and-butter business ML that 90% of products ship — &lt;em&gt;predict a number, predict a category, group similar things&lt;/em&gt; — it's not true.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You don't implement algorithms; you compose a pipeline.&lt;/strong&gt; You describe &lt;code&gt;data -&amp;gt; transforms -&amp;gt; trainer -&amp;gt; metrics&lt;/code&gt;, call &lt;code&gt;Fit()&lt;/code&gt;, call &lt;code&gt;Evaluate()&lt;/code&gt;. You never write gradient descent, a tree split, or a k-means iteration. The skill that matters is &lt;strong&gt;data preparation and honest evaluation&lt;/strong&gt; — and that's language-agnostic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Regression — forecast campaign conversions
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;pipeline&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Transforms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Categorical&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;OneHotEncoding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"ChannelEnc"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"Channel"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Transforms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Categorical&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;OneHotEncoding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"VerticalEnc"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"Vertical"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Transforms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Concatenate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Features"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"Impressions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"Week1Clicks"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"AudienceSize"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"ChannelEnc"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"VerticalEnc"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Transforms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;NormalizeMinMax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Features"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Regression&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Trainers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FastTree&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;labelColumnName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Label"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;featureColumnName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Features"&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;

&lt;span class="n"&gt;ITransformer&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pipeline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TrainSet&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;metrics&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Regression&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Evaluate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TestSet&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;FastTree&lt;/code&gt; is the gradient-boosted tree trainer — you just &lt;em&gt;select&lt;/em&gt; it. &lt;strong&gt;Result:&lt;/strong&gt; R² &lt;strong&gt;0.78&lt;/strong&gt;, RMSE &lt;strong&gt;41&lt;/strong&gt; on campaigns averaging ~600 conversions. Same accuracy band as the old scikit-learn model, now with no service to call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Classification — predict tenant churn
&lt;/h2&gt;

&lt;p&gt;The catch every real churn model hits: &lt;strong&gt;class imbalance&lt;/strong&gt;. Most tenants don't churn, so a model that always predicts "no" looks 94% accurate and is useless. Evaluate on AUC / precision / recall — &lt;strong&gt;never&lt;/strong&gt; raw accuracy.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;BinaryClassification&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Evaluate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TestSet&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="c1"&gt;// AUC, PositivePrecision, PositiveRecall, F1&lt;/span&gt;

&lt;span class="c1"&gt;// CS can only call ~30 tenants/week -&amp;gt; tune the threshold for PRECISION:&lt;/span&gt;
&lt;span class="kt"&gt;bool&lt;/span&gt; &lt;span class="n"&gt;flag&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prediction&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Probability&lt;/span&gt; &lt;span class="p"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="m"&gt;0.62&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// from the PR curve, not 0.5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Result:&lt;/strong&gt; AUC &lt;strong&gt;0.86&lt;/strong&gt;, precision &lt;strong&gt;0.71&lt;/strong&gt; at the threshold CS actually works. The weekly at-risk list is model-ranked instead of a brittle &lt;code&gt;if&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Clustering — segment tenants (no labels)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;pipeline&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Transforms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Concatenate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Features"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"CampaignsPerMonth"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"AvgAudienceSize"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"ReportDownloads"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"SeatUtilization"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"ApiCallsPerDay"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Transforms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;NormalizeMinMax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Features"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;   &lt;span class="c1"&gt;// critical: k-means is scale-sensitive&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Clustering&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Trainers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;KMeans&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Features"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;numberOfClusters&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Result:&lt;/strong&gt; &lt;strong&gt;5 clusters&lt;/strong&gt;, silhouette &lt;strong&gt;0.52&lt;/strong&gt; — distinct enough the product team named them ("power users," "dormant SMBs," "report-only"). The old size-based SQL never surfaced "report-only," a high-churn group hiding inside "enterprise."&lt;/p&gt;

&lt;h2&gt;
  
  
  Serving in production (the part tutorials skip)
&lt;/h2&gt;

&lt;p&gt;A trained &lt;code&gt;ITransformer&lt;/code&gt; is not thread-safe to predict from directly. Use &lt;code&gt;PredictionEnginePool&lt;/code&gt; — thread-safe, fast, hot-reloadable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="n"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddPredictionEnginePool&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;ChurnInput&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ChurnPrediction&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FromFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;modelName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"churn"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;filePath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Models/churn.zip"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;watchForChanges&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The nightly retrain job trains, evaluates, and &lt;strong&gt;gates&lt;/strong&gt; on a metric floor before swapping the file — never ship a regression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;auc&lt;/span&gt; &lt;span class="p"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="m"&gt;0.80&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;ml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;trainSet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Schema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"Models/churn.zip"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// pool hot-reloads&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Result:&lt;/strong&gt; prediction p95 &lt;strong&gt;45 ms -&amp;gt; 2.8 ms&lt;/strong&gt;, zero-downtime model promotion.&lt;/p&gt;

&lt;h2&gt;
  
  
  The classic mistake: data leakage
&lt;/h2&gt;

&lt;p&gt;Catching one leaked feature (a &lt;code&gt;final_invoice_flag&lt;/code&gt; that only existed &lt;em&gt;after&lt;/em&gt; churn) dropped offline AUC from a too-good &lt;strong&gt;0.97&lt;/strong&gt; to an honest &lt;strong&gt;0.86&lt;/strong&gt; — and the honest model is the one that works on live tenants. Every feature must be knowable at prediction time. A suspiciously high AUC is a leak until proven otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  When ML.NET is the &lt;em&gt;wrong&lt;/em&gt; call
&lt;/h2&gt;

&lt;p&gt;Deep learning, transformers, LLMs, computer vision belong in Python (or an API) — ML.NET can consume an ONNX model but won't train a state-of-the-art net. If your data scientists live in Python daily, don't fight that. ML.NET wins when the &lt;strong&gt;engineering team&lt;/strong&gt; owns the model and the problem is classical.&lt;/p&gt;

&lt;h2&gt;
  
  
  The closing mental model
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Classical ML is data engineering with an evaluation step — pick the language your app is already in.&lt;/strong&gt; Regression, classification, and clustering are &lt;code&gt;data -&amp;gt; transforms -&amp;gt; trainer -&amp;gt; metrics&lt;/code&gt;. ML.NET gives a .NET team all four in C#, in-process, with no second runtime to operate.&lt;/p&gt;

&lt;p&gt;The full guide has the before/after architecture diagrams, every pipeline in full, the trainer cheat-sheet, the pre-ship checklist, and the aggregate Mattrx metrics:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/no-python-no-phd-train-ml-models-csharp-mlnet" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/no-python-no-phd-train-ml-models-csharp-mlnet&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/no-python-no-phd-train-ml-models-csharp-mlnet" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>dotnet</category>
      <category>csharp</category>
      <category>machinelearning</category>
      <category>ai</category>
    </item>
    <item>
      <title>Build a RAG Chatbot in C# with Semantic Kernel + Azure AI Search (2026) — Full Production Guide, Real Code, Metrics</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Wed, 16 Sep 2026 16:56:56 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/build-a-rag-chatbot-in-c-with-semantic-kernel-azure-ai-search-2026-full-production-guide-9o1</link>
      <guid>https://dev.to/kirandeepjassalcrypto/build-a-rag-chatbot-in-c-with-semantic-kernel-azure-ai-search-2026-full-production-guide-9o1</guid>
      <description>&lt;p&gt;Every "build a RAG chatbot in 30 minutes" tutorial ends at the demo. Then you ship to production and discover the index has no security trimming, hallucinations are 22% because there's no grounding instruction, the model burns tokens with no cache, partner A can read partner B's docs, and latency is 4 seconds because chunks are too big.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;production&lt;/strong&gt; RAG chatbot in C# has ten more pieces. This is the full build, with real metrics from &lt;strong&gt;Mattrx Help&lt;/strong&gt; — a multi-tenant marketing analytics SaaS (Angular 19 + .NET 9, 110k MAU). The in-app docs assistant indexes ~3,200 docs (~24k chunks), serves ~8,400 MAU, and runs at &lt;strong&gt;4% hallucination, p95 retrieval 95ms, $0.004/query&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Orchestrator&lt;/td&gt;
&lt;td&gt;Semantic Kernel&lt;/td&gt;
&lt;td&gt;First-class .NET, plugins, telemetry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vector index&lt;/td&gt;
&lt;td&gt;Azure AI Search&lt;/td&gt;
&lt;td&gt;Hybrid BM25+vector, security trimming, reranker&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM&lt;/td&gt;
&lt;td&gt;Azure OpenAI (gpt-4o-mini)&lt;/td&gt;
&lt;td&gt;Same tenant, managed identity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Embeddings&lt;/td&gt;
&lt;td&gt;text-embedding-3-small (1536d)&lt;/td&gt;
&lt;td&gt;Best price/quality in 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chunking&lt;/td&gt;
&lt;td&gt;400-600 tokens / 80 overlap&lt;/td&gt;
&lt;td&gt;Sweet spot for docs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retrieval&lt;/td&gt;
&lt;td&gt;Hybrid + semantic reranker&lt;/td&gt;
&lt;td&gt;Best recall at scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grounding&lt;/td&gt;
&lt;td&gt;System prompt + cited chunks&lt;/td&gt;
&lt;td&gt;Cuts hallucination 22% -&amp;gt; 4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Streaming&lt;/td&gt;
&lt;td&gt;Server-Sent Events&lt;/td&gt;
&lt;td&gt;TTFT &amp;lt; 500ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache&lt;/td&gt;
&lt;td&gt;Redis (question hash, 60s TTL)&lt;/td&gt;
&lt;td&gt;34% hit rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eval&lt;/td&gt;
&lt;td&gt;80-question golden set, nightly&lt;/td&gt;
&lt;td&gt;Catches regressions before users&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Azure AI Search vs pgvector — the honest call
&lt;/h2&gt;

&lt;p&gt;The deciding factor was the &lt;strong&gt;semantic reranker&lt;/strong&gt;. Hybrid search alone got 84% top-5 recall. Adding the L2 cross-encoder reranker pushed it to &lt;strong&gt;91%&lt;/strong&gt;. For a customer-facing assistant, that 7-point bump was worth $248/month.&lt;/p&gt;

&lt;p&gt;If you're not on Azure or don't need the reranker, &lt;strong&gt;pgvector is still the right answer&lt;/strong&gt;. Don't pick AI Search because it sounds enterprise — pick it if the constraints match: multi-tenant security + reranking + already on Azure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five things that actually moved the numbers
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The system prompt does most of the work.&lt;/strong&gt; Three rules dropped hallucination 22% -&amp;gt; 4%: "Use ONLY the excerpts" (grounding), "If not in excerpts, say so" (an escape hatch so the model isn't pressured to confabulate), and "Cite [n]" (makes hallucinations visible and reviewable in eval).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Security trimming is part of the query, not a post-filter.&lt;/strong&gt; In a multi-tenant app the most common production bug is showing tenant A's data to tenant B. The Azure AI Search filter (&lt;code&gt;partnerId eq '*' or partnerId eq '{partnerId}'&lt;/code&gt;) is enforced server-side — even a manipulated query can't get rows outside the partner's scope. Write an integration test that asserts it, every release.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The reranker is the most underrated feature.&lt;/strong&gt; A 7-point recall bump for one config flag + a tier upgrade.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Streaming is product UX.&lt;/strong&gt; A 2-second response feels broken; one that starts streaming in 380ms feels instant. SSE, end-to-end.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Tenant-scoped cache key.&lt;/strong&gt; 34% hit rate cuts cost by a third. NEVER a global key.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production metrics (after 4-week build + 2-week tuning)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Top-5 retrieval recall&lt;/td&gt;
&lt;td&gt;91% (84% hybrid -&amp;gt; 91% reranker)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answer cites a source&lt;/td&gt;
&lt;td&gt;96%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hallucination rate&lt;/td&gt;
&lt;td&gt;4% (22% before grounded prompt)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Query latency p95 (Azure Search)&lt;/td&gt;
&lt;td&gt;95 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TTFT p95&lt;/td&gt;
&lt;td&gt;380 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;End-to-end p95 (cache miss)&lt;/td&gt;
&lt;td&gt;2.1 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;End-to-end p95 (cache hit)&lt;/td&gt;
&lt;td&gt;120 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache hit rate&lt;/td&gt;
&lt;td&gt;34%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per query (cache incl.)&lt;/td&gt;
&lt;td&gt;$0.004&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total monthly infra&lt;/td&gt;
&lt;td&gt;~$1,688&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tickets deflected / month&lt;/td&gt;
&lt;td&gt;~520 (~$13,000 saved)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User-rated "helpful"&lt;/td&gt;
&lt;td&gt;84% (vs 61% human first-response)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The chatbot pays for itself ~8x in deflected support tickets. The architecture is mostly Microsoft pieces glued together with ~600 lines of careful C#.&lt;/p&gt;

&lt;h2&gt;
  
  
  The right mental model
&lt;/h2&gt;

&lt;p&gt;A production RAG chatbot is &lt;strong&gt;four boring pieces composed correctly&lt;/strong&gt;: a vector store, a retriever, a grounded prompt, and a streaming UI. Three habits that make it pay off:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Treat the system prompt + eval set as a &lt;strong&gt;unit&lt;/strong&gt; — never change the prompt without running the eval.&lt;/li&gt;
&lt;li&gt;Tenant-scope every cache key and security filter.&lt;/li&gt;
&lt;li&gt;Streaming + grounded prompt + cache + cost cap on day one — none are optimization, they're the minimum.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The full guide has the complete Semantic Kernel C# — index schema, batched ingestion, the hybrid retriever, the grounded prompt, SSE streaming, tool-calling plugins, the OpenTelemetry trace, the eval harness, and per-partner cost monitoring:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/build-rag-chatbot-csharp-semantic-kernel-azure-ai-search-step-by-step-guide" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/build-rag-chatbot-csharp-semantic-kernel-azure-ai-search-step-by-step-guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/build-rag-chatbot-csharp-semantic-kernel-azure-ai-search-step-by-step-guide" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>dotnet</category>
      <category>azure</category>
      <category>rag</category>
    </item>
    <item>
      <title>Agentic RAG vs Traditional RAG in .NET (2026) — When Each Wins, Semantic Kernel Code, Production Metrics</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Mon, 14 Sep 2026 12:24:13 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/agentic-rag-vs-traditional-rag-in-net-2026-when-each-wins-semantic-kernel-code-production-3k6</link>
      <guid>https://dev.to/kirandeepjassalcrypto/agentic-rag-vs-traditional-rag-in-net-2026-when-each-wins-semantic-kernel-code-production-3k6</guid>
      <description>&lt;p&gt;Traditional RAG is what every "ChatGPT for your docs" tutorial builds: embed the question, fetch top-k chunks, stuff them into a prompt, return the answer. It works beautifully for ~75% of the questions you'd ask a support assistant. Then someone asks &lt;em&gt;"My conversion rate dropped 18% last week — check my webhook logs, the dashboard error rate, related docs, and tell me what's wrong"&lt;/em&gt; and traditional RAG falls over. That question needs log inspection, a SQL query, a doc lookup, a recent-events check, a hypothesis, and validation. That's an &lt;strong&gt;agent&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This is the condensed walkthrough; the full guide (complete Semantic Kernel code for both, the router, and the full metrics table) is on my site 👇&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Full guide:&lt;/strong&gt; &lt;a href="https://prepstack.co.in/blog/agentic-rag-vs-traditional-rag-dotnet-comparison-guide" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/agentic-rag-vs-traditional-rag-dotnet-comparison-guide&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The decision matrix
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Traditional RAG&lt;/th&gt;
&lt;th&gt;Agentic RAG&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Steps per query&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (retrieve → generate)&lt;/td&gt;
&lt;td&gt;3–8 (plan → tools → critique → synth)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool calls&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2–6 on average&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost / query&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.004&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0.038 (~10×)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency p95&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.1 s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;8.2 s (~4×)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;FAQ, doc lookup, "where is X"&lt;/td&gt;
&lt;td&gt;Multi-step analysis, debugging, "why is X"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy — simple Qs&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;78%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy — complex Qs&lt;/td&gt;
&lt;td&gt;32% (hallucinates)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;84%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Right model&lt;/td&gt;
&lt;td&gt;gpt-4o-mini&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;gpt-4o&lt;/strong&gt; (mini struggles to plan)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The 2026 rule of thumb:&lt;/strong&gt; use a router. Traditional RAG by default; agentic RAG when the question requires multiple tools, multiple knowledge sources, or iteration.&lt;/p&gt;

&lt;h2&gt;
  
  
  What makes a RAG "agentic"
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Traditional RAG = &lt;code&gt;retrieve(question) → generate(prompt)&lt;/code&gt;. Agentic RAG = &lt;code&gt;agent(question)&lt;/code&gt;&lt;/strong&gt; — the agent decides what to retrieve, in what order, and stops only when it's confident.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TRADITIONAL RAG                 AGENTIC RAG
retrieve top-k                  plan -&amp;gt; pick tool -&amp;gt; execute
build prompt                      -&amp;gt; critique ("enough?")
generate                          -&amp;gt; loop until confident (cap at 6)
                                  -&amp;gt; synthesize with all context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The four things only agents can do: &lt;strong&gt;decompose&lt;/strong&gt; ("compare Q1 to last year and recommend") into sub-questions; &lt;strong&gt;iterate&lt;/strong&gt; (reformulate if the first retrieval returned junk); &lt;strong&gt;choose tools&lt;/strong&gt; (&lt;code&gt;searchDocs&lt;/code&gt; for definitions, &lt;code&gt;runSqlQuery&lt;/code&gt; for numbers, &lt;code&gt;getRecentLogs&lt;/code&gt; for debugging); &lt;strong&gt;self-critique&lt;/strong&gt; (judge whether the answer is grounded before returning). Need none of those four? Traditional RAG is the right choice.&lt;/p&gt;

&lt;p&gt;The four costs of going agentic: &lt;strong&gt;money&lt;/strong&gt; (4–8 LLM calls/query), &lt;strong&gt;latency&lt;/strong&gt; (sequential tool calls push p95 from 2s to 8s+), &lt;strong&gt;debuggability&lt;/strong&gt; (a wrong agent means reading 6 prompts + 6 tool results + a plan tree), and &lt;strong&gt;failure modes&lt;/strong&gt; that can't happen with traditional RAG (loop forever, stop too early, wrong tool).&lt;/p&gt;

&lt;h2&gt;
  
  
  Agentic RAG needs five pieces
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Tools&lt;/strong&gt; — typed functions the agent can call (side-effect aware).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System prompt&lt;/strong&gt; — role, instructions, "when to stop" rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loop&lt;/strong&gt; — orchestration that keeps calling the LLM until done (Semantic Kernel's auto function-calling).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Critic&lt;/strong&gt; — "you have enough info" vs "go look more."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget guards&lt;/strong&gt; — max iterations, max cost, max tool calls.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Tools follow three rules: &lt;strong&gt;server-side identity&lt;/strong&gt; (&lt;code&gt;user.TenantId&lt;/code&gt; from the JWT, never from the agent — the agent cannot access another tenant), &lt;strong&gt;read-only by default&lt;/strong&gt; (mutations need explicit user confirmation), and &lt;strong&gt;rich &lt;code&gt;Description&lt;/code&gt; attributes&lt;/strong&gt; (the LLM reads them to choose tools; bad descriptions = bad choices).&lt;/p&gt;

&lt;h2&gt;
  
  
  The router — the highest-ROI piece
&lt;/h2&gt;

&lt;p&gt;The router classifies each incoming query "traditional" or "agentic" and dispatches. It's a gpt-4o-mini classification call at temperature 0 with a JSON response — ~$0.0002/query, ~80ms p95.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Without router (everything agentic):
  13,800 queries/day x $0.038 = ~$15,700/month

With router (78% traditional, 22% agentic):
  10,800 x $0.004 + 3,000 x $0.038 + 13,800 x $0.0002 (router) = ~$4,800/month

Monthly savings: ~$10,900.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the single highest-ROI decision in the AI stack. Build the router first.&lt;/p&gt;

&lt;h2&gt;
  
  
  I run both in production (Mattrx)
&lt;/h2&gt;

&lt;p&gt;Mattrx Help is &lt;strong&gt;traditional RAG&lt;/strong&gt; (in-product docs assistant — "how do I X", "what does error 4012 mean"). Mattrx Insights is &lt;strong&gt;agentic RAG&lt;/strong&gt; (analytical assistant — "why did conversions drop", "debug my integration") with six tools (docs search, analytics query, recent events, log search, config status, period compare); the agent picks 2–4 per query. A router sits in front. After 4 weeks running both:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Traditional (Help)&lt;/th&gt;
&lt;th&gt;Agentic (Insights)&lt;/th&gt;
&lt;th&gt;Routed (blended)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Daily queries&lt;/td&gt;
&lt;td&gt;12,000&lt;/td&gt;
&lt;td&gt;1,800&lt;/td&gt;
&lt;td&gt;13,800&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg cost / query&lt;/td&gt;
&lt;td&gt;$0.004&lt;/td&gt;
&lt;td&gt;$0.038&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.012&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy — simple&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;78%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71%&lt;/td&gt;
&lt;td&gt;78%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy — complex&lt;/td&gt;
&lt;td&gt;32%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;84%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;84%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hallucination (complex)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;18%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6%&lt;/td&gt;
&lt;td&gt;overall 5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly OpenAI bill&lt;/td&gt;
&lt;td&gt;~$1,440&lt;/td&gt;
&lt;td&gt;~$2,050&lt;/td&gt;
&lt;td&gt;~$4,800 (vs $15,700 agentic-only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"This was helpful"&lt;/td&gt;
&lt;td&gt;84%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;86%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Routed mode is strictly better than either alone: better accuracy than traditional (complex queries get the agent), cheaper than agentic-only (simple queries skip the agent), acceptable latency (only the 22% that need agentic pay 8s).&lt;/p&gt;

&lt;h2&gt;
  
  
  The model to carry forward
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Traditional RAG is &lt;code&gt;retrieve → generate&lt;/code&gt;. Agentic RAG is &lt;code&gt;plan → loop(tool → critique) → synthesize&lt;/code&gt;. A router decides which to use.&lt;/strong&gt; Three habits prevent 90% of the pain: build the router first (cheaper, saves money day one, you'll need it forever); treat tools as a public API (server-side identity, read-only by default, rich descriptions, multi-tenant tested); hard caps on iterations + cost + daily volume (agents &lt;em&gt;will&lt;/em&gt; try to loop and &lt;em&gt;will&lt;/em&gt; try to spend $5 on a $0.04 question).&lt;/p&gt;

&lt;p&gt;The full guide has the complete Semantic Kernel C# — traditional service, agentic service with the auto-function-calling loop and budget guards, the tool plugins, and the router + classifier — plus a step-by-step trace of an agent debugging a real conversion drop, the architecture diagram, and the full metrics:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/agentic-rag-vs-traditional-rag-dotnet-comparison-guide" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/agentic-rag-vs-traditional-rag-dotnet-comparison-guide&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/agentic-rag-vs-traditional-rag-dotnet-comparison-guide" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>dotnet</category>
      <category>rag</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Design YouTube — Video Upload, Transcoding &amp; CDN Delivery at Scale (with Production .NET Code)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Sun, 13 Sep 2026 12:01:27 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/design-youtube-video-upload-transcoding-cdn-delivery-at-scale-with-production-net-code-3h19</link>
      <guid>https://dev.to/kirandeepjassalcrypto/design-youtube-video-upload-transcoding-cdn-delivery-at-scale-with-production-net-code-3h19</guid>
      <description>&lt;p&gt;"Design YouTube" is where two systems live in one product, pulling in opposite directions. The &lt;strong&gt;write&lt;/strong&gt; side is a brutal batch-processing problem: someone uploads a multi-gigabyte file, and you have to turn it into a dozen resolutions and formats without melting your servers. The &lt;strong&gt;read&lt;/strong&gt; side is a caching problem at planetary scale: billions of people press play and expect video to start in under a second, anywhere on Earth. The interview is about keeping those two apart — a &lt;strong&gt;queue-and-worker transcoding pipeline&lt;/strong&gt; for the write, and a &lt;strong&gt;CDN&lt;/strong&gt; for the read.&lt;/p&gt;

&lt;p&gt;This is the condensed walkthrough; the full guide (estimates, API, data model, and the full production .NET 9 code) is on my site 👇&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Full guide:&lt;/strong&gt; &lt;a href="https://prepstack.co.in/blog/design-youtube-system-design" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/design-youtube-system-design&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The design at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Upload&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Resumable, chunked&lt;/strong&gt; upload straight to blob storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processing&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Transcoding pipeline&lt;/strong&gt; — queue + stateless worker fleet, chunk → transcode → package&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Object/blob storage&lt;/strong&gt; for raw + renditions (petabyte scale, tiered)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delivery&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;CDN&lt;/strong&gt; at the edge — the origin sees only cache misses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Playback&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Adaptive bitrate&lt;/strong&gt; (HLS/DASH) — the player picks quality by bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;View counts&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Approximate + aggregated&lt;/strong&gt; — never a DB increment per view&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  It's two systems sharing a blob store
&lt;/h2&gt;

&lt;p&gt;The two sides scale very differently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Uploads:  ~1M videos/day, raw files tens of MB-GBs  -&amp;gt; petabytes of storage, huge transcode compute
Views:    billions/day  -&amp;gt;  ~tens of thousands of segment fetches/sec, served from the CDN edge
Read:write ratio: enormous  -&amp;gt;  the CDN, not the origin, carries the traffic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Storage and transcode compute dominate the write side; the CDN is not optional on the read side — it's the only way billions of views don't vaporize your origin.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Uploader --&amp;gt; [ Upload service ] == resumable, chunked ==&amp;gt; [ Blob: raw upload ]
                                                                 | event: uploaded
                                                                 v
                                     [ Transcoding pipeline - queue + worker fleet ]
                                     chunk -&amp;gt; transcode (144p...4K, multi-codec) -&amp;gt; package (HLS/DASH)
                                                                 |
                                                                 v
                                     [ Blob: renditions ] --&amp;gt; [ CDN edge ] --&amp;gt; Viewers
                                     [ Metadata DB ]   [ View counter (aggregated) ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The write pipeline (upload → transcode → store) and the read path (CDN → viewer) meet only at the blob store. Playback never waits on a transcoder; uploads never touch the CDN.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hard parts
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Resumable, chunked upload.&lt;/strong&gt; A multi-gigabyte upload over mobile &lt;em&gt;will&lt;/em&gt; drop mid-transfer; a single POST that restarts from zero is unusable. The client uploads fixed-size chunks (with offsets) directly to blob storage via a pre-signed URL, and a dropped connection &lt;strong&gt;resumes from the last good chunk&lt;/strong&gt;. The app server issues the URL and reacts to the completion event — it never proxies the bytes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The transcoding pipeline — the write core.&lt;/strong&gt; A raw upload must become many renditions (144p→4K) across codecs, packaged for streaming:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Chunk&lt;/strong&gt; the raw video into segments so they transcode independently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transcode&lt;/strong&gt; each segment to each target resolution/bitrate — an embarrassingly parallel fan-out across a stateless &lt;strong&gt;worker fleet&lt;/strong&gt; (this is where the compute goes).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Package&lt;/strong&gt; the transcoded segments into HLS/DASH + a manifest.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It's a &lt;strong&gt;DAG of jobs driven by a message queue&lt;/strong&gt;: workers pull segment-transcode tasks, scale horizontally with queue depth, and are idempotent and retried (a failed segment re-runs; the pipeline dead-letters what won't).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adaptive bitrate streaming (HLS/DASH).&lt;/strong&gt; The packaged video is a set of short segments at multiple bitrates plus a manifest. The &lt;strong&gt;player&lt;/strong&gt; downloads the manifest, starts at a modest bitrate, and switches up or down per segment based on measured bandwidth.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Player reads the manifest, then picks a rendition per segment by bandwidth:

   4K    ############   ~25 Mbps    &amp;lt;- fast wifi
  1080p  ########       ~8 Mbps
   720p  #####          ~5 Mbps
   480p  ###            ~2.5 Mbps
   240p  #              ~0.7 Mbps    &amp;lt;- weak mobile
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The crucial part: the &lt;strong&gt;server serves static segments&lt;/strong&gt; — all the adaptation logic lives in the player. That's what makes delivery a pure CDN problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CDN delivery — the read path.&lt;/strong&gt; Views dwarf uploads and viewers are global. Serve every segment from a &lt;strong&gt;CDN&lt;/strong&gt;: edge caches near the viewer hold popular content, so the origin (blob store) is hit only on a cache miss. This is what makes playback start fast worldwide and keeps billions of views from ever reaching your origin. The CDN &lt;em&gt;is&lt;/em&gt; the read architecture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;View counts at scale.&lt;/strong&gt; Billions of views can't each &lt;code&gt;UPDATE views SET count = count + 1&lt;/code&gt; — a hot-row meltdown. &lt;strong&gt;Aggregate&lt;/strong&gt;: buffer increments (in Redis or a stream), flush periodic batches to the metadata store, accept approximate, eventually consistent counts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scaling gotchas
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Transcode compute&lt;/strong&gt; is the write bottleneck — a big worker fleet autoscaled by &lt;strong&gt;queue depth&lt;/strong&gt;, often on cheap spot instances (jobs are idempotent, so eviction is fine).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CDN&lt;/strong&gt; absorbs the read load; pre-warm popular videos at the edge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage&lt;/strong&gt; is petabyte-scale — &lt;strong&gt;tier&lt;/strong&gt; cold, rarely-watched videos to cheap storage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pipeline resilience&lt;/strong&gt; — idempotent jobs, retries, a DLQ for un-transcodable uploads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never proxy the bytes through your app server&lt;/strong&gt; — the client talks to blob storage directly; your server orchestrates.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  I shipped this shape in production (Mattrx)
&lt;/h2&gt;

&lt;p&gt;Mattrx isn't a video site, but its &lt;strong&gt;report generation&lt;/strong&gt; runs on the exact same shape: a heavy artifact produced by an async worker fleet, stored in blob storage, and delivered from a CDN. Marketers export branded PDF reports of campaign performance — Mattrx renders about &lt;strong&gt;1.2 million every 48 hours&lt;/strong&gt; with PuppeteerSharp. V1 generated the PDF &lt;strong&gt;synchronously inside the request&lt;/strong&gt;: a big report took 30–60 seconds, tied up a web worker, timed out under load, and was streamed back through the app server. We rebuilt it as this design — enqueue on Azure Service Bus, render on a worker fleet, drop the PDF in Azure Blob Storage, serve it from &lt;strong&gt;Azure Front Door (CDN)&lt;/strong&gt; via a signed URL:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Report request&lt;/td&gt;
&lt;td&gt;Blocked 30–60s in the request (timeouts under load)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;202 Accepted&lt;/code&gt;, enqueue p95 ~90 ms&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generation&lt;/td&gt;
&lt;td&gt;Synchronous, tied up a web worker&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Async worker fleet&lt;/strong&gt; (PuppeteerSharp)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;Capped by the web tier&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~1.2M reports / 48h&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delivery&lt;/td&gt;
&lt;td&gt;Streamed through the app server&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Azure Front Door CDN&lt;/strong&gt; (signed URL, edge)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API write-path p95&lt;/td&gt;
&lt;td&gt;Dragged down by report load&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;120 ms&lt;/strong&gt; (fully decoupled)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worker cost&lt;/td&gt;
&lt;td&gt;Baseline&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;~$1,300/mo saved&lt;/strong&gt; (right-sized async fleet)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The API does no rendering — it writes a pending row, drops a command on Service Bus, and returns a job id in ~90 ms; the worker fleet does the expensive PuppeteerSharp render and stores the PDF in Blob Storage; and delivery is a signed URL fronted by Azure Front Door, so the finished report streams from the edge — the same decoupling that keeps YouTube's playback independent of its transcoders. (Full .NET 9 API + worker is in the post.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The model to carry forward
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A video platform is two decoupled systems sharing a blob store.&lt;/strong&gt; The write side is a queue-and-worker batch pipeline — resumable upload, parallel transcoding across a stateless fleet, packaged into adaptive segments — and the read side is a CDN in front of static files that carries essentially all the traffic. Keep them apart: playback depends only on the blob store and the CDN, never on the transcoders; uploads never touch the read path.&lt;/p&gt;

&lt;p&gt;Three habits it teaches: split the write pipeline from the read path (batch-process on workers, serve static from a CDN, let them meet only at the blob store); never move big bytes through your app server (clients talk to storage directly); approximate what's expensive to be exact (view counts are aggregated and eventually consistent).&lt;/p&gt;

&lt;p&gt;That wraps my System Design Interview series — ten "Design X" walkthroughs on one framework. The full guide has the estimates, API, data model, all the hard parts in depth, scaling, the complete production .NET 9 pipeline, and the "when it's overkill" section:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/design-youtube-system-design" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/design-youtube-system-design&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/design-youtube-system-design" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>interview</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Design Uber — Geospatial Matching, Live Location &amp; Surge Pricing (with Production .NET Code)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Sat, 12 Sep 2026 09:58:20 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/design-uber-geospatial-matching-live-location-surge-pricing-with-production-net-code-1ap3</link>
      <guid>https://dev.to/kirandeepjassalcrypto/design-uber-geospatial-matching-live-location-surge-pricing-with-production-net-code-1ap3</guid>
      <description>&lt;p&gt;"Design Uber" comes down to one question that sounds easy and isn't: &lt;em&gt;given a rider standing here, find the nearby available drivers — right now, among millions of them, while every one of those drivers is moving and pinging a new location every few seconds.&lt;/em&gt; Do it the obvious way — compute the distance from the rider to every driver — and you're doing millions of calculations per request. The entire interview is the data structure that turns "search everywhere" into "search the few blocks around the rider."&lt;/p&gt;

&lt;p&gt;This is the condensed walkthrough; the full guide (estimates, API, data model, and the full production .NET 9 code) is on my site 👇&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Full guide:&lt;/strong&gt; &lt;a href="https://prepstack.co.in/blog/design-uber-system-design" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/design-uber-system-design&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The design at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Find nearby drivers&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Geospatial index&lt;/strong&gt; (geohash / quadtree / H3) — search the rider's cell + neighbors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Driver locations&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;In-memory geo index&lt;/strong&gt; (Redis GEO), sharded by region — not a DB table&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Location ingestion&lt;/td&gt;
&lt;td&gt;Persistent connections; update the index, don't write each ping to a DB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Matching&lt;/td&gt;
&lt;td&gt;Rank nearby drivers by &lt;strong&gt;ETA&lt;/strong&gt;, offer, and &lt;strong&gt;lock&lt;/strong&gt; the driver to avoid double-dispatch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Surge&lt;/td&gt;
&lt;td&gt;Per-cell &lt;strong&gt;demand ÷ supply&lt;/strong&gt; ratio, updated continuously&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Live trip&lt;/td&gt;
&lt;td&gt;Stream driver location to the rider over a &lt;strong&gt;WebSocket&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Two forces: a write firehose and a proximity read
&lt;/h2&gt;

&lt;p&gt;The write side is the monster:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Drivers: ~5M active, each pings location every ~4s
  -&amp;gt; 5,000,000 / 4  ~  1.25M location updates/sec

Ride requests: millions/day  ~  hundreds/sec, each a nearby-driver search
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You cannot write 1.25M rows/sec to a relational database, so driver location lives in an &lt;strong&gt;in-memory, sharded geo index&lt;/strong&gt; updated in place; and the proximity search must be &lt;strong&gt;cell-bounded&lt;/strong&gt;, never a scan over millions of drivers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Drivers == location every ~4s ==&amp;gt; [ Location ingestion ] --&amp;gt; [ Geo index (Redis GEO), sharded by region ]
                                                                    ^
Rider request (pickup) --&amp;gt; [ Matching service ] -- search nearby cells
                                |  rank by ETA - offer - LOCK driver
                                v
                         [ Ride store ] - [ Trip tracking (WebSocket) ] - [ Surge (per cell) ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The hard parts
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Geospatial indexing — the heart of it.&lt;/strong&gt; The naive approach — distance from the rider to every driver — is O(N) per request and dies at millions of drivers. Instead, partition space into cells and bucket drivers by cell; a proximity search only examines the rider's cell and its neighbors.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Geohash:&lt;/strong&gt; interleave lat/lng bits into a string. Nearby points share a &lt;strong&gt;prefix&lt;/strong&gt;, so a cell is a prefix and "nearby" is "matching prefix + the 8 neighbor cells." Simple, and what Redis GEO uses under the hood.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quadtree:&lt;/strong&gt; recursively split space into four quadrants; dense areas (downtown) subdivide deeper — adaptive to density.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;S2 (Google) / H3 (Uber):&lt;/strong&gt; hierarchical global cell systems; Uber's H3 uses &lt;strong&gt;hexagons&lt;/strong&gt; (uniform neighbor distance, no corner problems).
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A geohash cell + its 8 neighbors — a "nearby" search scans these 9 cells:

      +-----+-----+-----+
      | NW  |  N  | NE  |
      +-----+-----+-----+
      |  W  |  *  |  E  |     * = the rider's cell
      +-----+-----+-----+
      | SW  |  S  | SE  |
      +-----+-----+-----+

You examine ~9 cells' worth of drivers, never all N drivers.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The location write firehose.&lt;/strong&gt; 1.25M updates/sec can't hit a database. Drivers stream location over &lt;strong&gt;persistent connections&lt;/strong&gt;; the ingestion layer &lt;strong&gt;updates the driver's position in the in-memory geo index in place&lt;/strong&gt; (one current position, not a history of rows). The index is &lt;strong&gt;sharded by region&lt;/strong&gt; (geohash prefix), so each shard owns its slice of the map and the write load spreads. Hot regions get finer shards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Matching — closest isn't the answer, best ETA is.&lt;/strong&gt; A nearby search returns candidates; ranking by &lt;strong&gt;straight-line distance is wrong&lt;/strong&gt; — a driver 200m away across a river is farther &lt;em&gt;by road&lt;/em&gt; than one 500m away on the same street. Rank by &lt;strong&gt;ETA&lt;/strong&gt; (road-network + traffic). Then &lt;strong&gt;offer&lt;/strong&gt; the ride to the top driver and &lt;strong&gt;lock&lt;/strong&gt; them (a short hold) so a second rider's search can't dispatch the same car; if they decline or time out, release and offer the next. Locking prevents double-dispatch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Surge pricing.&lt;/strong&gt; Per cell, compute &lt;strong&gt;demand ÷ supply&lt;/strong&gt; — open requests vs available drivers. When demand outstrips supply, a surge multiplier rises for that cell, pricing the scarcity and nudging drivers toward it. A continuously recomputed, per-region number — and a feedback loop, so tune it or it oscillates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live trip tracking.&lt;/strong&gt; Once matched, the driver's location streams to the rider in real time — the exact &lt;strong&gt;push-over-WebSocket&lt;/strong&gt; problem from the chat design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scaling gotchas
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ingestion firehose:&lt;/strong&gt; in-memory geo index, sharded by region; persistent connections, not per-ping HTTP.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hot cells (downtown, a stadium at closing):&lt;/strong&gt; split finer; the index adapts to density, or that cell melts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Proximity search:&lt;/strong&gt; cell-bounded and served from memory — the whole point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consistency:&lt;/strong&gt; driver location is eventually consistent (seconds stale is fine); dispatch uses a &lt;strong&gt;lock&lt;/strong&gt; where it matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Durability split:&lt;/strong&gt; rides persist relationally; location stays ephemeral in Redis (a lost ping just re-arrives in 4s).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy:&lt;/strong&gt; streaming precise movement of millions is a serious responsibility — minimize retention.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  I shipped this in production (Mattrx)
&lt;/h2&gt;

&lt;p&gt;Mattrx isn't ride-hailing, but its &lt;strong&gt;geo-analytics&lt;/strong&gt; run on exactly this proximity-search machinery. Conversions carry coordinates, and marketers ask location questions: "how many conversions within 5 km of this store?" (footfall attribution) and "render a heatmap of engagement." V1 answered those with a &lt;code&gt;Haversine&lt;/code&gt; distance filter scanned across the &lt;code&gt;CampaignEvents&lt;/code&gt; table (1.2B rows) — a full range scan per query, p95 ~1,800 ms that pinned the DB whenever a marketer dragged the map. We moved recent conversion events into per-tenant &lt;strong&gt;Redis GEO&lt;/strong&gt; sets on ingestion, so a radius question becomes a &lt;strong&gt;geohash-bounded &lt;code&gt;GEOSEARCH&lt;/code&gt;&lt;/strong&gt; over a few cells:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"Conversions within 5 km of a store" p95&lt;/td&gt;
&lt;td&gt;~1,800 ms (Haversine scan)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;~12 ms&lt;/strong&gt; (Redis &lt;code&gt;GEOSEARCH&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rows examined per query&lt;/td&gt;
&lt;td&gt;full &lt;code&gt;CampaignEvents&lt;/code&gt; range scan&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;a few geohash cells&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Store-attribution DB load&lt;/td&gt;
&lt;td&gt;heavy per query&lt;/td&gt;
&lt;td&gt;offloaded to the Redis geo index&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Geo-heatmap render&lt;/td&gt;
&lt;td&gt;seconds, laggy panning&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;sub-second, smooth&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;GeoAddAsync&lt;/code&gt; upserts a conversion's position by encoding it into a geohash score, and &lt;code&gt;GeoSearchAsync&lt;/code&gt; with a &lt;code&gt;GeoSearchCircle&lt;/code&gt; does the cell-bounded radius query natively — so a "near this store" question touches a handful of geohash cells in memory instead of Haversine-scanning a billion rows in Azure SQL, the whole difference between an 1,800 ms map drag and a 12 ms one. (Full .NET 9 geo index is in the post.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The model to carry forward
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Ride-hailing is a proximity-search problem wrapped in a location firehose.&lt;/strong&gt; You never search all drivers — you bucket the map into cells, drop drivers into cells, and a nearby search touches only the cells around the rider. Keep that index in memory because the write rate is enormous, shard it by region so it scales and adapts to density, match on ETA rather than distance, and lock a driver during an offer so you never book one car twice.&lt;/p&gt;

&lt;p&gt;Three habits it teaches: reach for a spatial index immediately ("distance to every driver" fails; "cell + neighbors" scales); keep the firehose out of your database (a current-location-per-driver index in memory beats a million writes/sec to disk); match on ETA, lock on dispatch (distance is a lie the map tells; the driver lock keeps matching correct).&lt;/p&gt;

&lt;p&gt;The full guide has the estimates, API, data model, all the hard parts in depth, scaling, the complete production .NET 9 Redis GEO code, and the "when it's overkill" section:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/design-uber-system-design" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/design-uber-system-design&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/design-uber-system-design" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>interview</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Design a Payment System — Idempotency, Ledgers &amp; Exactly-Once at Scale (with Production .NET Code)</title>
      <dc:creator>kirandeepjassal-crypto</dc:creator>
      <pubDate>Fri, 11 Sep 2026 16:24:43 +0000</pubDate>
      <link>https://dev.to/kirandeepjassalcrypto/design-a-payment-system-idempotency-ledgers-exactly-once-at-scale-with-production-net-code-1jfg</link>
      <guid>https://dev.to/kirandeepjassalcrypto/design-a-payment-system-idempotency-ledgers-exactly-once-at-scale-with-production-net-code-1jfg</guid>
      <description>&lt;p&gt;"Design a payment system" is the one where scale is &lt;em&gt;not&lt;/em&gt; the point. Nobody cares that it does ten thousand transactions a second; they care that it never does one transaction &lt;em&gt;twice&lt;/em&gt;, never loses one, and can prove — line by line, months later — exactly where every cent went. The interview lives in four words: &lt;strong&gt;idempotency, ledger, and reconciliation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the condensed walkthrough; the full guide (estimates, API, data model, and the full production .NET 9 code) is on my site 👇&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Full guide:&lt;/strong&gt; &lt;a href="https://prepstack.co.in/blog/design-a-payment-system-system-design" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/design-a-payment-system-system-design&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The design at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Double-charge safety&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Idempotency key&lt;/strong&gt; (unique) — a retry returns the first result, never re-charges&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Money records&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Double-entry ledger&lt;/strong&gt; — balanced entries; balances derived, not stored&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider events&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;At-least-once + dedup&lt;/strong&gt; by event id (exactly-once is a myth)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consistency&lt;/td&gt;
&lt;td&gt;Ledger writes &lt;strong&gt;atomic (DB transaction)&lt;/strong&gt;; external charge bridged by a state machine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Safety net&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Reconciliation&lt;/strong&gt; — compare ledger to provider daily; flag discrepancies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Card data&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Never store it&lt;/strong&gt; — tokenize via the provider (PCI)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  It's a correctness problem, not a scale problem
&lt;/h2&gt;

&lt;p&gt;Say the quiet part out loud: even a large processor does a few million transactions/day ≈ tens–hundreds/sec. That fits comfortably on a well-indexed relational database. So you spend the budget not on sharding and caching but on &lt;strong&gt;invariants&lt;/strong&gt;: uniqueness constraints for idempotency, a balanced ledger, and a reconciliation job. Correctness is the scarce resource here, not QPS.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client --&amp;gt; [ Payment API ] -- idempotency check --&amp;gt; [ Provider (Stripe) ]
                |  (pending)                              |
                v                                         | webhook (at-least-once)
          [ Ledger (double-entry, atomic) ] &amp;lt;-- dedup ---+  payment_intent.succeeded
                |
          [ Reconciliation job ] -- daily compare vs provider --&amp;gt; flag discrepancies
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The hard parts
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Idempotency — the whole ballgame.&lt;/strong&gt; A client charges, the request times out, the client retries. Did the first attempt succeed? You can't know from the client side — so the &lt;em&gt;server&lt;/em&gt; makes the retry safe. The client sends an &lt;strong&gt;idempotency key&lt;/strong&gt;; the server stores it under a &lt;strong&gt;unique constraint&lt;/strong&gt; &lt;em&gt;before&lt;/em&gt; doing anything expensive. On a retry the key already exists, so you return the stored result instead of charging again. Pass the &lt;strong&gt;same key to the provider&lt;/strong&gt; too (Stripe's &lt;code&gt;Idempotency-Key&lt;/code&gt; header) so even the external call dedups. One key, one charge, forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You can't make the charge and the ledger one transaction.&lt;/strong&gt; The provider is external — you cannot enroll "call Stripe" and "write my ledger" in one ACID transaction. Use a &lt;strong&gt;state machine&lt;/strong&gt;: record the payment as &lt;code&gt;pending&lt;/code&gt; (committed), call the provider, and let the &lt;strong&gt;webhook&lt;/strong&gt; drive the authoritative transition to &lt;code&gt;succeeded&lt;/code&gt;/&lt;code&gt;failed&lt;/code&gt;, at which point you post the ledger — all in one &lt;em&gt;local&lt;/em&gt; transaction. Internal ledger strongly consistent; external settlement eventually consistent; the webhook is the bridge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Double-entry ledger — the books must balance.&lt;/strong&gt; Never store a single mutable &lt;code&gt;balance&lt;/code&gt;. Every money movement posts &lt;strong&gt;balanced entries&lt;/strong&gt; — debits equal credits, netting to zero — into an immutable, append-only ledger. A successful charge might &lt;strong&gt;debit Cash&lt;/strong&gt; and &lt;strong&gt;credit&lt;/strong&gt; the customer's &lt;strong&gt;Receivable&lt;/strong&gt;. Balances are &lt;em&gt;derived&lt;/em&gt; (&lt;code&gt;SUM&lt;/code&gt; over an account), so the ledger is auditable, replayable, and self-checking: if a transaction's entries don't sum to zero, reject it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Webhooks are at-least-once — dedup them.&lt;/strong&gt; The provider will occasionally deliver the same &lt;code&gt;payment_intent.succeeded&lt;/code&gt; twice; applying it twice double-credits the ledger. Guard every handler with a &lt;strong&gt;&lt;code&gt;UNIQUE(eventId)&lt;/code&gt;&lt;/strong&gt; dedup log: record the event id first; if it's already there, the second delivery is a no-op.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reconciliation — the backstop.&lt;/strong&gt; What if a webhook is &lt;em&gt;lost&lt;/em&gt;? The charge succeeded at the provider but your ledger never recorded it — a silent gap. A scheduled job pulls the provider's transaction list and compares it to your ledger, flagging anything that exists on one side but not the other. Reconciliation turns "hope the webhook arrived" into "prove the books match."&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure &amp;amp; scaling gotchas
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Provider timeout (the classic):&lt;/strong&gt; you called Stripe and got no response. Don't blindly retry the charge — retry with the &lt;em&gt;same idempotency key&lt;/em&gt; (safe), or &lt;strong&gt;query the provider&lt;/strong&gt; for the intent's status. Reconciliation catches whatever slips through.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ledger scale:&lt;/strong&gt; append-only writes scale well; partition by account/time; balances via periodic snapshots + entries since.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Outbox for downstream events:&lt;/strong&gt; publish &lt;code&gt;payment.succeeded&lt;/code&gt; to other services via the outbox pattern so a crash can't drop the event.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refunds/reversals&lt;/strong&gt; are just &lt;em&gt;more balanced entries&lt;/em&gt; — never delete or mutate existing ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PCI:&lt;/strong&gt; never let raw card numbers touch your servers; tokenize through the provider and store only the token.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  I shipped this in production (Mattrx)
&lt;/h2&gt;

&lt;p&gt;Mattrx bills ~thousands of tenant workspaces through Stripe. V1 was dangerously naive: it charged inline, and a timeout on Stripe's response left it not knowing whether the charge landed — so a retry occasionally &lt;strong&gt;double-charged a customer&lt;/strong&gt;, cleaned up by hand. "Paid" was a boolean on the subscription row, so there was no audit trail and monthly reconciliation was a spreadsheet. We rebuilt it as exactly the design above:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Double-charges on retry&lt;/td&gt;
&lt;td&gt;A handful a month (manual refunds)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (idempotency key + Stripe &lt;code&gt;Idempotency-Key&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Money record&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;paid&lt;/code&gt; boolean on the subscription&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Immutable double-entry ledger&lt;/strong&gt; (always reconciles)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Webhook double-application&lt;/td&gt;
&lt;td&gt;Occasional double-credit&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (&lt;code&gt;UNIQUE(eventId)&lt;/code&gt; dedup)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reconciliation vs Stripe&lt;/td&gt;
&lt;td&gt;Manual spreadsheet (~hours/month)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Automated nightly&lt;/strong&gt;, discrepancies auto-flagged&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Billing disputes from double-charges&lt;/td&gt;
&lt;td&gt;Non-zero&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 in the last two quarters&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The charge path reserves the idempotency key &lt;em&gt;before&lt;/em&gt; calling Stripe and passes the same key onward, so a client retry and a provider retry both collapse to one charge; the webhook handler dedups on &lt;code&gt;event_id&lt;/code&gt; and posts a balanced ledger transaction inside one local DB transaction; and &lt;code&gt;Ledger.PostAsync&lt;/code&gt; refuses to write anything that doesn't net to zero — so the books are correct by construction, and the nightly reconciliation job catches the one thing code can't: a settlement Stripe recorded but never told us about. (Full .NET 9 billing service + webhook handler + ledger is in the post.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The model to carry forward
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A payment system is a machine for staying correct when you don't know what happened.&lt;/strong&gt; The timeout is the enemy, and every part of the design is a defense against it: an idempotency key so retries are safe, a state machine so an external charge and an internal record don't have to be one transaction, a double-entry ledger so the truth is auditable and self-checking, deduped webhooks so at-least-once behaves like once, and reconciliation so a lost message can't quietly corrupt the books.&lt;/p&gt;

&lt;p&gt;Three habits it teaches: lead with idempotency ("the client sends an idempotency key" frames the entire correctness story); make money a ledger, not a number (immutable balanced entries with derived balances is the only design an auditor or a bug can't break); assume the webhook can be lost (reconciliation is what turns hope into proof).&lt;/p&gt;

&lt;p&gt;The full guide has the estimates, API, data model, all the hard parts in depth, failure/scaling, the complete production .NET 9 code, and the "when it's overkill" section:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://prepstack.co.in/blog/design-a-payment-system-system-design" rel="noopener noreferrer"&gt;https://prepstack.co.in/blog/design-a-payment-system-system-design&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://prepstack.co.in/blog/design-a-payment-system-system-design" rel="noopener noreferrer"&gt;PrepStack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>interview</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
