<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Amit chakraborty</title>
    <description>The latest articles on DEV Community by Amit chakraborty (@techamit95ch).</description>
    <link>https://dev.to/techamit95ch</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F362727%2F44e6ad7c-ce73-4784-a4e6-965148cf26ee.jpg</url>
      <title>DEV Community: Amit chakraborty</title>
      <link>https://dev.to/techamit95ch</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/techamit95ch"/>
    <language>en</language>
    <item>
      <title>The Hermes Bytecode Trap: Why Your App is Fast in Development and Slow in Production</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:30:04 +0000</pubDate>
      <link>https://dev.to/techamit95ch/the-hermes-bytecode-trap-why-your-app-is-fast-in-development-and-slow-in-production-7n5</link>
      <guid>https://dev.to/techamit95ch/the-hermes-bytecode-trap-why-your-app-is-fast-in-development-and-slow-in-production-7n5</guid>
      <description>&lt;p&gt;I was the first engineering hire at Synapsis Medical Technologies, responsible for taking our HealthTech AI platform from zero to one. We were building a React Native application that integrated with clinical wearables and served real-time AI insights via a HIPAA-aligned RAG pipeline. On my local M2 MacBook, the app felt instantaneous. In our CI/CD pipeline—which I had just overhauled to cut release cycles from two days down to four hours—everything passed.&lt;/p&gt;

&lt;p&gt;Then we pushed to a group of beta users in a clinical setting.&lt;/p&gt;

&lt;p&gt;The symptom wasn't a crash; it was a "hang." Users on mid-range Android devices reported that the app would sit on the splash screen for six to eight seconds. Our TTI (Time to Interactive) metrics, which looked fine in the lab, spiked to unacceptable levels in the field. We were losing the "instant" feel required for a clinical tool.&lt;/p&gt;

&lt;p&gt;The culprit wasn't our React code. It was a failure to properly manage the Hermes bytecode pre-compilation during the transition from development to production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the "Hermes Boost" Fails Under Load
&lt;/h2&gt;

&lt;p&gt;In development, Hermes acts as a Just-In-Time (JIT) engine. It parses JavaScript on the fly. In production, it switches to Ahead-of-Time (AOT) compilation. The build process transforms your JavaScript into &lt;code&gt;.hbc&lt;/code&gt; (Hermes Bytecode) files. &lt;/p&gt;

&lt;p&gt;The promise is that the device doesn't have to parse the JS; it just maps the bytecode into memory and executes. However, three things break this in production:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Instruction Cache Misses:&lt;/strong&gt; If your bundle is massive (common in AI-heavy apps with large vendored libraries), the bytecode size can exceed the device's efficient memory mapping limits.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;The Bridge Bottleneck:&lt;/strong&gt; Even with Hermes, if you are initialising heavy native modules (like FHIR/HL7 parsers or encrypted local storage) on the main thread during the same tick as the Hermes VM startup, the UI thread will starve.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Bytecode Version Mismatch:&lt;/strong&gt; If your CI/CD pipeline doesn't strictly align the Hermes compiler version with the specific Hermes runtime version bundled in your &lt;code&gt;node_modules&lt;/code&gt;, the engine falls back to a slower execution mode or fails to load the pre-compiled bundle entirely.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Fix: Step-by-Step
&lt;/h2&gt;

&lt;p&gt;To solve this at Synapsis, I had to move beyond the default "enable Hermes" flag and implement a strict verification and optimisation pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Force Bytecode Alignment in CI
&lt;/h3&gt;

&lt;p&gt;Never assume your build server is using the correct compiler. If you are on React Native 0.73+, the Hermes version is tied to the React Native version. &lt;/p&gt;

&lt;p&gt;Verify your bytecode version by adding a check in your &lt;code&gt;__tests__&lt;/code&gt; or a pre-build script. Use the &lt;code&gt;hermesc&lt;/code&gt; command-line tool to inspect the generated bundle.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Locate the compiler in your node_modules&lt;/span&gt;
&lt;span class="nv"&gt;HERMESC_PATH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"./node_modules/react-native/sdks/hermesc/%OS%-bin/hermesc"&lt;/span&gt;

&lt;span class="c"&gt;# Compile a test file and check the bytecode version&lt;/span&gt;
&lt;span class="nv"&gt;$HERMESC_PATH&lt;/span&gt; &lt;span class="nt"&gt;-emit-binary&lt;/span&gt; &lt;span class="nt"&gt;-out&lt;/span&gt; ./test.hbc ./test.js
&lt;span class="nv"&gt;$HERMESC_PATH&lt;/span&gt; &lt;span class="nt"&gt;-dump-bytecode&lt;/span&gt; ./test.hbc | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"Bytecode version"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why:&lt;/strong&gt; If the version in your &lt;code&gt;package.json&lt;/code&gt; doesn't match the compiler used by your CI runner, you will see a &lt;code&gt;format version mismatch&lt;/code&gt; error in Logcat, and the app will default to the much slower JS parsing.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Implement "Dead Code" Pruning for the VM
&lt;/h3&gt;

&lt;p&gt;Hermes' memory footprint is directly tied to the number of functions it needs to register. In our AI platform, we had large JSON schemas for FHIR data. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Action:&lt;/strong&gt; Move static, large data structures out of the JS bundle and into native assets or fetch them on demand. &lt;br&gt;
&lt;strong&gt;Confirmation:&lt;/strong&gt; Monitor the &lt;code&gt;index.android.bundle&lt;/code&gt; size. If it exceeds 10MB, your TTI on a device like a Samsung A50 will increase by ~500ms per additional MB due to I/O overhead.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. Profile the "Quiet" Startup Failures
&lt;/h3&gt;

&lt;p&gt;Use the Chrome DevTools protocol with Hermes to capture a &lt;code&gt;.cpuprofile&lt;/code&gt;. &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; Open the app in release-mode-like conditions (but with debugging enabled).&lt;/li&gt;
&lt;li&gt; Capture the profile from the very first millisecond of the &lt;code&gt;Entry Point&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt; Look for &lt;code&gt;HermesVM::init&lt;/code&gt;. &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you see a long gap between &lt;code&gt;HermesVM::init&lt;/code&gt; and the first &lt;code&gt;FunctionCall&lt;/code&gt;, your native modules are blocking the VM. In my experience, this is usually caused by synchronous calls to &lt;code&gt;SharedPreferences&lt;/code&gt; or &lt;code&gt;Keychain&lt;/code&gt; during the &lt;code&gt;C++&lt;/code&gt; initialization phase of the React Native bridge.&lt;/p&gt;
&lt;h3&gt;
  
  
  4. Enable Memory-Mapped I/O (mmap)
&lt;/h3&gt;

&lt;p&gt;Ensure your Android build is actually using &lt;code&gt;mmap&lt;/code&gt; for the bytecode. In your &lt;code&gt;MainApplication.java&lt;/code&gt; (or &lt;code&gt;MainApplication.kt&lt;/code&gt;), verify the &lt;code&gt;JSExecutorFactory&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Ensure you are using the default HermesExecutor, &lt;/span&gt;
&lt;span class="c1"&gt;// but check if you have custom configurations that override the memory limit.&lt;/span&gt;
&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;factory&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;HermesExecutorFactory&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Confirmation:&lt;/strong&gt; Use &lt;code&gt;adb shell procrank&lt;/code&gt; while the app is starting. If &lt;code&gt;USS&lt;/code&gt; (Unique Set Size) jumps instantly to 100MB+, your bytecode isn't being memory-mapped efficiently; it's being copied into RAM.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cost of the Fix
&lt;/h2&gt;

&lt;p&gt;Optimising Hermes is not free. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Build Times:&lt;/strong&gt; AOT compilation adds roughly 2-5 minutes to a clean build, depending on bundle size. In our CI/CD overhaul, we had to use distributed caching for the &lt;code&gt;android/.gradle&lt;/code&gt; and &lt;code&gt;ios/build&lt;/code&gt; folders to keep our 4-hour release cycle intact.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Debugging Difficulty:&lt;/strong&gt; Bytecode stack traces are notoriously difficult to read. You must strictly manage your Source Maps. If you lose the source map for a specific bytecode build, production crashes will be impossible to trace.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The Trade-off:&lt;/strong&gt; You are trading build-time CPU cycles for end-user battery life and TTI. For a clinical app, this is always the right choice.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  At Your Level
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting Out
&lt;/h3&gt;

&lt;p&gt;Focus on ensuring Hermes is actually enabled. Check &lt;code&gt;global.HermesInternal&lt;/code&gt; in your &lt;code&gt;App.js&lt;/code&gt;. If it returns &lt;code&gt;null&lt;/code&gt;, your configuration is wrong, and you're running on the old, slower JSC engine.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working Engineer
&lt;/h3&gt;

&lt;p&gt;Stop using &lt;code&gt;console.log&lt;/code&gt; for performance timing. Use &lt;code&gt;Performance.now()&lt;/code&gt; and track the time from the first line of &lt;code&gt;index.js&lt;/code&gt; to the &lt;code&gt;useEffect&lt;/code&gt; in your root component. If this is over 1000ms on a real device, you have a bundle initialization problem that Hermes alone won't fix.&lt;/p&gt;

&lt;h3&gt;
  
  
  Senior or Staff
&lt;/h3&gt;

&lt;p&gt;You own the architecture. You must enforce a "bundle budget." If a new library adds 2MB to the bytecode, it must be justified by a 2MB removal elsewhere or a critical feature. Implement automated bundle size reporting in every Pull Request to catch regressions before they hit the 4-hour release window.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead or Director
&lt;/h3&gt;

&lt;p&gt;Understand that TTI is a retention metric. At Synapsis, a 2-second delay in clinical AI insights could mean a clinician stops using the tool. Budget time for "performance sprints" where the team does nothing but prune dependencies and optimize the bridge.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the Interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "We've enabled Hermes, but our Android startup time is still poor. What do you investigate first?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Weak Answer:&lt;/strong&gt; "I'd check the React code for heavy loops or unnecessary re-renders." &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Why it's weak:&lt;/strong&gt; Startup time (TTI) is rarely about React re-renders; it's about the time it takes to get the VM running and the first frame painted.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The Strong Answer:&lt;/strong&gt; A senior candidate should point to the &lt;strong&gt;Bridge and the Bytecode&lt;/strong&gt;. They should discuss the difference between JIT and AOT, the overhead of memory-mapping large &lt;code&gt;.hbc&lt;/code&gt; files, and the potential for synchronous native module initialization to block the UI thread.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Follow-up:&lt;/strong&gt; "How do you handle the fact that Hermes doesn't support &lt;code&gt;Intl&lt;/code&gt; on older versions of React Native?" &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  This separates those who have read the docs from those who have shipped. The answer involves polyfills that increase bundle size, directly impacting the TTI benefits Hermes provides. You have to weigh the cost of the polyfill against the speed of the engine.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=hermes-and-startup-time-what-actually-breaks-in-production" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>hermes</category>
      <category>android</category>
      <category>performance</category>
    </item>
    <item>
      <title>Defending the 16ms Frame: A Performance Budget for React Native Skia</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Fri, 09 Oct 2026 12:30:34 +0000</pubDate>
      <link>https://dev.to/techamit95ch/defending-the-16ms-frame-a-performance-budget-for-react-native-skia-3l30</link>
      <guid>https://dev.to/techamit95ch/defending-the-16ms-frame-a-performance-budget-for-react-native-skia-3l30</guid>
      <description>&lt;p&gt;I was the first engineering hire at Synapsis Medical Technologies, where I eventually scaled the team to 21 engineers. We were building a HealthTech AI platform that required real-time visualisations of wearable data—ECG waveforms and live heart-rate variability—layered over a React Native UI. &lt;/p&gt;

&lt;p&gt;We chose &lt;code&gt;@shopify/react-native-skia&lt;/code&gt; because the standard React Native &lt;code&gt;View&lt;/code&gt; hierarchy collapsed under the weight of 500+ data points updating at 60Hz. However, three weeks before a major clinical pilot, our "smooth" animations started stuttering. On an iPhone 15 Pro, it looked fine. On the mid-range Android devices our partner clinic used, the frame rate dropped to 22 FPS. The UI thread was clear, but the JavaScript thread was choking, and the GPU was waiting for data that arrived too late.&lt;/p&gt;

&lt;p&gt;The cost was immediate: we couldn't pass the clinical usability audit. Stuttering waveforms in a medical context aren't just a "bad UX"—they look like equipment failure. We had to move from a vague "make it faster" mandate to a rigid performance budget that every engineer on the team had to defend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Skia stutters in React Native
&lt;/h2&gt;

&lt;p&gt;In React Native 0.76 and earlier, the bottleneck is rarely Skia’s C++ engine. It is the bridge or the JSI (JavaScript Interface) overhead. &lt;/p&gt;

&lt;p&gt;When you use &lt;code&gt;useDrawSelection&lt;/code&gt; or &lt;code&gt;Canvas&lt;/code&gt;, every frame requires a trip from the JS thread to the UI thread. If you are calculating paths inside a standard &lt;code&gt;useMemo&lt;/code&gt; and passing them as props to a Skia component, you are serialising data across the bridge. Even with the New Architecture (TurboModules), if your JS thread is busy with a heavy JSON parse or a RAG pipeline update—as was common in our AI workflows—the rendering instruction arrives late.&lt;/p&gt;

&lt;p&gt;The "16ms budget" (for 60fps) or "8ms budget" (for 120fps) isn't just for drawing. It includes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;JS Execution:&lt;/strong&gt; Calculating the new coordinates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JSI Transfer:&lt;/strong&gt; Moving that data to the Skia host objects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rasterization:&lt;/strong&gt; Skia actually drawing pixels.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If JS takes 12ms, you have only 4ms left for everything else. On a low-end Android device, Skia might need 10ms to rasterize a complex blur or a heavy path, putting you at 22ms total—well over the 16ms limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix: Implementing the 16ms Ceiling
&lt;/h2&gt;

&lt;p&gt;We moved from "best effort" rendering to a "budget-first" architecture. Follow these steps to lock down your rendering performance.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Externalise State to Skia Values
&lt;/h3&gt;

&lt;p&gt;Stop using &lt;code&gt;useState&lt;/code&gt; or &lt;code&gt;useSharedValue&lt;/code&gt; from Reanimated for high-frequency Skia updates. Use &lt;code&gt;Skia.makeMutable&lt;/code&gt; or &lt;code&gt;useValue&lt;/code&gt; from the Skia library itself. This keeps the data within the Skia engine's reach without triggering React render cycles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to confirm:&lt;/strong&gt; Open the React DevTools profiler. If your Skia component shows a re-render count matching your frame rate, you have failed this step. The component should render &lt;em&gt;once&lt;/em&gt;, and the Skia &lt;code&gt;Value&lt;/code&gt; should update the drawing internally.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Offload Path Generation to a Worklet
&lt;/h3&gt;

&lt;p&gt;Generating a complex SVG path string in JS is expensive. For our ECG waveforms, we shifted path generation to &lt;code&gt;useComputedValue&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Do not do this in the main body of the component&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Skia&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Make&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nx"&gt;points&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forEach&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;moveTo&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="nx"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lineTo&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Do this instead&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useComputedValue&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Skia&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Make&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="c1"&gt;// ... logic here&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;pointsValue&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why:&lt;/strong&gt; This ensures the path is built as a host object. You are passing a reference across the JSI, not a massive string or array of coordinates.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Establish the "Frame-Drop" Metric
&lt;/h3&gt;

&lt;p&gt;We integrated &lt;code&gt;react-native-performance&lt;/code&gt; to track &lt;code&gt;js_fps&lt;/code&gt; and &lt;code&gt;ui_fps&lt;/code&gt;. I set a CI gate: if the 95th percentile of frame time exceeded 16ms on a Galaxy A54 (our baseline low-end device), the PR was blocked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Command:&lt;/strong&gt;&lt;br&gt;
Use the Flashlight CLI (a tool for mobile performance measurement) to get an actual score:&lt;br&gt;
&lt;code&gt;flashlight measure --bundleId com.synapsis.app --duration 10000&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Look for the &lt;strong&gt;"Frame Duration"&lt;/strong&gt; metric. If the mean is &amp;gt;16ms, you are dropping frames.&lt;/p&gt;
&lt;h3&gt;
  
  
  4. Use Picture Recording for Static Overlays
&lt;/h3&gt;

&lt;p&gt;If your UI has complex gradients or shadows that don't change every frame (like a background grid), do not redraw them. Use &lt;code&gt;SkPicture&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;picture&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useMemo&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;recorder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Skia&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;PictureRecorder&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;canvas&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;beginRecording&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rect&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="c1"&gt;// Draw complex, static background here&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;finishRecording&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[]);&lt;/span&gt;

&lt;span class="c1"&gt;// In your draw function&lt;/span&gt;
&lt;span class="nx"&gt;canvas&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drawPicture&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;picture&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;How to confirm it worked:&lt;/strong&gt; Use the Skia Overlay (&lt;code&gt;&amp;lt;SkiaDomView debug /&amp;gt;&lt;/code&gt;). You will see the "Draw Calls" count drop significantly. In our case, it dropped from 140 calls to 12 per frame.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cost of a Performance Budget
&lt;/h2&gt;

&lt;p&gt;Enforcing a 16ms budget has trade-offs that aren't always pleasant:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Development Velocity:&lt;/strong&gt; It took my team roughly 30% longer to ship new visual features because they couldn't just "drop a component in." They had to architect the data flow to avoid the bridge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code Complexity:&lt;/strong&gt; You end up with "Skia-specific" state management that sits parallel to your Redux or Zustand store. Synchronising these two can lead to bugs where the UI shows one value while the Skia canvas shows another.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hardware Ceiling:&lt;/strong&gt; There is a point where a device simply cannot render what you want. We had to implement "Feature Degradation"—detecting low-end GPUs and disabling Gaussian blurs or reducing the sample rate of the ECG waveform from 250Hz to 125Hz.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  At your level
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Starting out:&lt;/strong&gt; &lt;br&gt;
Stop using &lt;code&gt;setInterval&lt;/code&gt; or &lt;code&gt;requestAnimationFrame&lt;/code&gt; to drive animations. Use &lt;code&gt;useTiming&lt;/code&gt; or &lt;code&gt;useLoop&lt;/code&gt; from &lt;code&gt;@shopify/react-native-skia&lt;/code&gt;. The one thing to do next: move one &lt;code&gt;useState&lt;/code&gt; value that drives a Skia element into a &lt;code&gt;useValue&lt;/code&gt; hook and observe the reduction in React re-renders.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Working engineer:&lt;/strong&gt;&lt;br&gt;
Profile your app using the Flashlight CLI on a physical Android device, not an iOS simulator. The simulator uses your Mac's GPU and will lie to you about performance. Your goal is to keep the "JS Thread" usage below 40% during animations to leave headroom for background tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior or staff:&lt;/strong&gt;&lt;br&gt;
Architect a "Layered Rendering" strategy. Move static elements to the standard React Native View hierarchy (which is heavily optimised by the OS) and only use Skia for the volatile, high-frequency pixels. The one thing to do next: implement a &lt;code&gt;PerformanceMonitor&lt;/code&gt; provider that samples frame times and automatically downgrades graphical fidelity (e.g., removing &lt;code&gt;BlurMask&lt;/code&gt;) for users on older devices.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lead or director:&lt;/strong&gt;&lt;br&gt;
Define the "Baseline Device." If you don't name a specific $200 Android phone as your target, your engineers will develop on $1,200 iPhones and ship a product that fails in the real world. Budget for a "Performance Sprint" every quarter where no features are added, only frame-time regressions are fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "How do you handle a performance bottleneck in a React Native app with complex graphics?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Weak Answer:&lt;/strong&gt; "I would use &lt;code&gt;useMemo&lt;/code&gt; to cache calculations and make sure I'm not doing too much on the JS thread." (Too vague; doesn't acknowledge the specific constraints of the bridge or the GPU).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strong Answer:&lt;/strong&gt; A strong answer identifies the &lt;strong&gt;JSI bridge&lt;/strong&gt; as the primary bottleneck. You should discuss the trade-off between &lt;strong&gt;Declarative API&lt;/strong&gt; (easier to read, higher overhead) and the &lt;strong&gt;Imperative API&lt;/strong&gt; (manual drawing, lower overhead). Mention specific tools like &lt;strong&gt;Flashlight&lt;/strong&gt; or &lt;strong&gt;Perfetto&lt;/strong&gt; for tracing. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Senior Follow-up:&lt;/strong&gt; "How do you handle the synchronisation between a Skia animation and a React Native &lt;code&gt;FlatList&lt;/code&gt; scroll?"&lt;br&gt;
&lt;em&gt;The Answer:&lt;/em&gt; This probes for knowledge of thread synchronisation. A senior candidate should explain that since Skia rendering happens on the UI thread (if using the right hooks), it can be perfectly synced with &lt;code&gt;onScroll&lt;/code&gt; events using &lt;code&gt;useSharedValue&lt;/code&gt; and &lt;code&gt;useDerivedValue&lt;/code&gt;, provided you avoid the jump back to the JS thread. They should mention the risk of "checkerboarding" if the JS thread is too slow to feed the list while the GPU is busy drawing Skia layers.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=react-native-skia-rendering-a-performance-budget-approach" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>skia</category>
      <category>performance</category>
      <category>jsi</category>
    </item>
    <item>
      <title>Measuring the ROI of TurboModules before you commit to the rewrite</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Fri, 09 Oct 2026 08:30:03 +0000</pubDate>
      <link>https://dev.to/techamit95ch/measuring-the-roi-of-turbomodules-before-you-commit-to-the-rewrite-2hij</link>
      <guid>https://dev.to/techamit95ch/measuring-the-roi-of-turbomodules-before-you-commit-to-the-rewrite-2hij</guid>
      <description>&lt;p&gt;I was leading the architecture for a clinical AI platform at Synapsis Medical Technologies, where we were integrating real-time health data from wearables into a React Native environment. We were hitting a wall with the Bridge. As we scaled to handle high-frequency biometric streams, our UI thread started dropping frames during heavy data ingest. &lt;/p&gt;

&lt;p&gt;The immediate internal pressure was to migrate everything to TurboModules and the New Architecture (React Native 0.76+). The promise of synchronous execution and type-safe Codegen is seductive. However, I had a team of 21 engineers and a product roadmap that didn't include a month of "architectural refactoring" with no visible feature output.&lt;/p&gt;

&lt;p&gt;If you blindly migrate a stable Native Module to a TurboModule, you might spend 40 hours of engineering time to save 2ms of overhead that wasn't your bottleneck. I’ve shipped 18 production apps, and the most expensive mistake I see is optimising for architectural purity rather than measured latency.&lt;/p&gt;

&lt;p&gt;Before you touch &lt;code&gt;codegen&lt;/code&gt;, you need to know if the Bridge is actually your problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Bridge fails under load
&lt;/h2&gt;

&lt;p&gt;In the legacy architecture, every call between JavaScript and Native is asynchronous, serialised as JSON, and passed over a single queue. &lt;/p&gt;

&lt;p&gt;When we were streaming data from a wearable heart-rate monitor, each data point was a Bridge message. If the user was also interacting with a complex UI—say, a real-time graph—the Bridge became a bottleneck. The symptom wasn't just "slowness"; it was non-deterministic lag. The JSON serialisation overhead for a large array of health metrics can take 5–10ms. When you do that 60 times a second, your message queue backs up, and your &lt;code&gt;onPress&lt;/code&gt; events stay stuck behind data packets.&lt;/p&gt;

&lt;p&gt;TurboModules solve this by using JSI (JavaScript Interface), allowing JS to hold a reference to C++ host objects. You move from "sending a letter" (Bridge) to "calling a function" (JSI). But if your module only fires once every 10 seconds—like a battery level check—the Bridge overhead is irrelevant.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: Step-by-step instrumentation
&lt;/h2&gt;

&lt;p&gt;Do not start by changing your C++ or Java code. Start by profiling the actual cost of the Bridge in your current production build.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Trace the Bridge Traffic
&lt;/h3&gt;

&lt;p&gt;You need to see the volume and frequency of messages. Use the built-in MessageQueue spy. In your entry file (e.g., &lt;code&gt;App.js&lt;/code&gt;), add:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;MessageQueue&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-native/Libraries/BatchedBridge/MessageQueue&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;__DEV__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;MessageQueue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;spy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why:&lt;/strong&gt; This logs every call between JS and Native to your console.&lt;br&gt;
&lt;strong&gt;Confirm:&lt;/strong&gt; Watch the console during the specific user action that feels slow. If you see a flood of calls to the same module (e.g., &lt;code&gt;DeviceEventManager.emit&lt;/code&gt;), you have a high-frequency bottleneck. If you see very large JSON payloads, you have a serialisation bottleneck.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. Measure Bridge Latency with Systrace
&lt;/h3&gt;

&lt;p&gt;Standard console logs don't show you the time spent in the serialisation layer. Use the &lt;code&gt;react-native-performance&lt;/code&gt; library or the built-in Profiler in Flipper.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open Flipper and select the &lt;strong&gt;React Native Tracers&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Start a trace and perform the heavy action.&lt;/li&gt;
&lt;li&gt;Look for the &lt;code&gt;Bridge&lt;/code&gt; or &lt;code&gt;JS Call&lt;/code&gt; markers.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Confirm:&lt;/strong&gt; Look for gaps where the JS thread is idle but the Native thread is busy, or vice versa. If the "Bridge" bar is consistently wider than 5ms per call, the overhead is significant enough to justify a TurboModule.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. Define the Codegen Spec
&lt;/h3&gt;

&lt;p&gt;If the data proves you need to move, start with the Spec. This is the "contract" that ensures type safety. Create a file named &lt;code&gt;Native[YourModuleName].ts&lt;/code&gt; in a &lt;code&gt;specs&lt;/code&gt; folder.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;TurboModule&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-native&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;TurboModuleRegistry&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-native&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;Spec&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nx"&gt;TurboModule&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;getHealthData&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;streamData&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;values&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nx"&gt;TurboModuleRegistry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;getEnforcing&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Spec&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;YourModuleName&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why:&lt;/strong&gt; Codegen uses this TypeScript interface to generate the C++ glue code (JSI) and the Java/Objective-C boilerplate.&lt;br&gt;
&lt;strong&gt;Confirm:&lt;/strong&gt; Run &lt;code&gt;node node_modules/react-native/scripts/generate-codegen-artifacts.js&lt;/code&gt;. If it fails, your types aren't compatible with the restricted subset of TypeScript that Codegen supports (e.g., you cannot use generic &lt;code&gt;any&lt;/code&gt; or complex unions).&lt;/p&gt;
&lt;h3&gt;
  
  
  4. Implement the Synchronous Method
&lt;/h3&gt;

&lt;p&gt;One of the primary reasons I move to TurboModules is for synchronous methods. In the old Bridge, everything returned a &lt;code&gt;Promise&lt;/code&gt;. In TurboModules, you can return a value directly.&lt;/p&gt;

&lt;p&gt;In your native implementation (e.g., &lt;code&gt;YourModule.mm&lt;/code&gt; for iOS):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight objective_c"&gt;&lt;code&gt;&lt;span class="k"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;NSString&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;getHealthData&lt;/span&gt;&lt;span class="p"&gt;:(&lt;/span&gt;&lt;span class="n"&gt;NSString&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nv"&gt;id&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;@"Synchronous Data"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; 
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Confirm:&lt;/strong&gt; Call the method in JS: &lt;code&gt;const data = YourModule.getHealthData('123');&lt;/code&gt;. If it returns the value immediately without an &lt;code&gt;await&lt;/code&gt;, the JSI layer is working.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, and when not to do it
&lt;/h2&gt;

&lt;p&gt;The cost of TurboModules is primarily in build-time complexity and developer experience.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Build Times:&lt;/strong&gt; Codegen adds a significant overhead to your build process. In our CI/CD overhaul, where I cut release cycles from 2 days to 4 hours, we found that poorly managed native builds could add 10 minutes just to the compilation phase of a clean build.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;C++ Requirement:&lt;/strong&gt; To get the full performance benefit, you often end up writing C++ (via JSI). If your team is purely JS/React, you are introducing a massive bus-factor risk.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Library Compatibility:&lt;/strong&gt; Many third-party libraries still rely on the Bridge. Mixing the New Architecture with legacy modules requires the "Interop Layer," which can introduce its own performance regressions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Do not migrate if:&lt;/strong&gt; Your app is a standard CRUD app where the most complex thing you do is fetch a JSON API and display it in a list. The Bridge is more than fast enough for 90% of applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  At your level
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting out
&lt;/h3&gt;

&lt;p&gt;Focus on understanding why &lt;code&gt;MessageQueue.spy()&lt;/code&gt; exists. Before you learn how to write a TurboModule, learn how to read the Bridge traffic. If you can identify a "chatty" bridge, you've already provided more value than someone who just follows a tutorial.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working engineer
&lt;/h3&gt;

&lt;p&gt;Implement a performance monitoring suite. Don't wait for a bug report. Use &lt;code&gt;react-native-performance&lt;/code&gt; to track &lt;code&gt;bridge_transfer_time&lt;/code&gt; in your staging environment. The goal is to have a baseline so you can prove the ROI of a rewrite to your lead.&lt;/p&gt;

&lt;h3&gt;
  
  
  Senior or staff
&lt;/h3&gt;

&lt;p&gt;Your job is to prevent the migration until it is necessary. Evaluate the "Interoperability Layer" cost. If you have 50 legacy modules, moving one to TurboModules might actually slow down the app due to the overhead of the New Architecture's compatibility bridge. Own the decision of &lt;em&gt;when&lt;/em&gt; to flip the switch for the entire repo.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead or director
&lt;/h3&gt;

&lt;p&gt;Consider the hiring implications. Moving to a JSI-heavy architecture means you need engineers who understand the memory management differences between JS and C++. If you don't have that talent or the budget to hire it, stick to the Bridge and optimise your JSON payloads instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "We are seeing UI stutters when processing high-frequency data from a native sensor. How do you decide between optimising the Bridge or moving to TurboModules?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Weak Answer:&lt;/strong&gt; "TurboModules are faster because they use JSI and the New Architecture, so we should migrate to get better performance." This is weak because it assumes the architecture is the bottleneck without proof and ignores the massive engineering cost of migration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strong Answer:&lt;/strong&gt; A strong answer starts with measurement. You should discuss using the &lt;code&gt;MessageQueue&lt;/code&gt; spy to check for "chattiness" and measuring the serialisation overhead. You would mention that the Bridge overhead is usually around JSON stringification/parsing. If the frequency of calls is the issue, you might first try batching the data on the native side before it hits the Bridge. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Senior/Staff Follow-up:&lt;/strong&gt; "What happens if we have a mix of TurboModules and legacy modules?"&lt;br&gt;
A senior candidate must discuss the &lt;strong&gt;Interoperability Layer&lt;/strong&gt;. They should know that the New Architecture has to wrap legacy modules to make them compatible, which can actually increase memory pressure. They should be able to explain that the real win of TurboModules isn't just "speed," but the ability to perform &lt;strong&gt;synchronous&lt;/strong&gt; calls to native code, which eliminates the need for complex async state management in some UI interactions.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=turbomodules-and-codegen-measuring-before-optimising" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>turbomodules</category>
      <category>architecture</category>
      <category>performance</category>
    </item>
    <item>
      <title>Production failures: the things that page you · 2. Thundering herd: when the cache expires all at once</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Fri, 09 Oct 2026 04:30:11 +0000</pubDate>
      <link>https://dev.to/techamit95ch/production-failures-the-things-that-page-you-2-thundering-herd-when-the-cache-expires-all-at-2cjj</link>
      <guid>https://dev.to/techamit95ch/production-failures-the-things-that-page-you-2-thundering-herd-when-the-cache-expires-all-at-2cjj</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Production failures: the things that page you&lt;/strong&gt; · Chapter 2 of 24 · Reliability · new chapter every Friday morning&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;By the end of this chapter:&lt;/strong&gt; Recognise correlated cache expiry and apply jitter and request coalescing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Say you are an engineer on call for a content API. A celebrity mentions your site, and traffic spikes to ten times its normal volume. You open your dashboards, applying the structured reading method from Chapter 1, and the symptom is clear: the primary Postgres database is pegged at 100% CPU. Connections are maxed out, and the API is returning 503s. &lt;/p&gt;

&lt;p&gt;You check the cache hit rate. It is normally 95%, but right now it is oscillating wildly between 99% and 0%. You assume the cache time-to-live (TTL) is too short, so you deploy a change increasing it from five minutes to thirty minutes. The database recovers instantly. You close the incident. Thirty minutes later, the pager goes off again with the exact same database CPU spike. &lt;/p&gt;

&lt;p&gt;You have built a thundering herd. When a cache key expires under heavy load, hundreds of concurrent requests all miss the cache at the exact same millisecond. They all query the database simultaneously. The database stalls, connections queue up, and the API drops requests. Increasing the TTL did not fix the problem; it just delayed the next synchronised failure by thirty minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you start
&lt;/h2&gt;

&lt;p&gt;You need a local Redis server and Node.js. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Node.js 20.x or higher.&lt;/li&gt;
&lt;li&gt;  Redis 7.2 or higher.&lt;/li&gt;
&lt;li&gt;  The &lt;code&gt;redis&lt;/code&gt; npm package (v4.6.x).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Verify your environment by running this in your terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; redis-cli ping
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see &lt;code&gt;v20.x.x&lt;/code&gt; (or higher) and &lt;code&gt;PONG&lt;/code&gt;. If &lt;code&gt;redis-cli&lt;/code&gt; is not installed, ensure your Redis server is running on the default port (6379).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does a fixed TTL synchronise traffic?
&lt;/h2&gt;

&lt;p&gt;A cache with a fixed TTL acts as a traffic synchroniser. If a popular item is not in the cache, the first request fetches it from the database and writes it to Redis with a TTL of exactly 300 seconds. &lt;/p&gt;

&lt;p&gt;For the next 299 seconds, the database does no work for that item. The API serves thousands of requests directly from Redis. But exactly 300 seconds later, the key evaporates. If your API is handling 500 requests per second for that item, all 500 requests in that specific second will check Redis, find nothing, and query the database. &lt;/p&gt;

&lt;p&gt;The database does not see a smooth average of traffic. It sees zero requests for five minutes, followed by a wall of 500 concurrent queries, followed by zero requests for another five minutes. If those 500 queries require complex joins, the database CPU spikes, queries time out, and the API fails.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do we decouple the expiry?
&lt;/h2&gt;

&lt;p&gt;You prevent synchronised expiry by adding jitter. Jitter is randomness applied to a fixed interval. &lt;/p&gt;

&lt;p&gt;Instead of caching an item for exactly 300 seconds, you cache it for 300 seconds minus a random percentage—usually up to 20%. One request might set a TTL of 281 seconds, the next 245 seconds, the next 299 seconds. &lt;/p&gt;

&lt;p&gt;When you apply jitter to a large number of cached items, they no longer expire simultaneously. The database load smooths out. However, jitter only prevents different keys from expiring at the same time. It does not stop 500 concurrent requests from hitting the database when a single, highly popular key expires.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do we handle concurrent misses?
&lt;/h2&gt;

&lt;p&gt;To protect the database from a single popular key expiring, you must implement request coalescing, also known as single-flight.&lt;/p&gt;

&lt;p&gt;When 500 requests ask for the same missing key, only the first request should query the database. The other 499 requests should wait for the first request to finish, and then share its result. &lt;/p&gt;

&lt;p&gt;In an asynchronous language like JavaScript, you do this by caching the Promise of the database call in memory, rather than just the final result. When a request comes in, it checks the in-memory map. If a Promise exists for that key, the request awaits it. If not, it creates the Promise, puts it in the map, and queries the database. &lt;/p&gt;

&lt;h2&gt;
  
  
  Build it
&lt;/h2&gt;

&lt;p&gt;We will build a Node.js script that simulates 50 concurrent requests for a single item that is not in the cache. We will implement both request coalescing and TTL jitter so the database is only queried once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Initialize the project&lt;/strong&gt;&lt;br&gt;
Create a new directory and install the Redis client. This client handles the connection to your local Redis server.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm init &lt;span class="nt"&gt;-y&lt;/span&gt;
npm &lt;span class="nb"&gt;install &lt;/span&gt;redis@4.6.10
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2: Write the coalescing cache implementation&lt;/strong&gt;&lt;br&gt;
Create a file named &lt;code&gt;index.js&lt;/code&gt; and paste the following code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;redis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createClient&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;error&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Redis Client Error&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;

&lt;span class="c1"&gt;// The in-memory map that holds pending database calls&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;pendingRequests&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// A simulation of a slow database query&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;queryDatabase&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`[DB] Executing expensive query for &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;...`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;resolve&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; 
        &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Data for &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getWithCoalescingAndJitter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;baseTtlSeconds&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cacheKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`item:&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="c1"&gt;// 1. Check the distributed cache (Redis)&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// 2. Check for an in-flight request (Coalescing)&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pendingRequests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`[Cache] Coalescing request for &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;pendingRequests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// 3. Fetch from DB, cache it, and clean up the pending map&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;promise&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;queryDatabase&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

            &lt;span class="c1"&gt;// Calculate jitter: subtract up to 20% from the base TTL&lt;/span&gt;
            &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;jitter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;floor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;baseTtlSeconds&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
            &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;finalTtl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;baseTtlSeconds&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;jitter&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

            &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;EX&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;finalTtl&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
            &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`[Cache] Wrote to Redis with TTL &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;finalTtl&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;s`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="c1"&gt;// Always remove the promise, whether it resolved or rejected&lt;/span&gt;
            &lt;span class="nx"&gt;pendingRequests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;delete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;})();&lt;/span&gt;

    &lt;span class="c1"&gt;// Store the promise so subsequent requests can await it&lt;/span&gt;
    &lt;span class="nx"&gt;pendingRequests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cacheKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;promise&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;promise&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;flushAll&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Firing 50 concurrent requests...&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// Fire 50 requests at the exact same time&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;promises&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;length&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; 
        &lt;span class="nf"&gt;getWithCoalescingAndJitter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;user-1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;promises&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Completed &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; requests.`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`First result: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;disconnect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 3: Run the code and verify the output&lt;/strong&gt;&lt;br&gt;
Execute the script.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node index.js
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see output similar to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Firing 50 concurrent requests...
[DB] Executing expensive query for user-1...
[Cache] Coalescing request for user-1
[Cache] Coalescing request for user-1
... (repeated 49 times)
[Cache] Wrote to Redis with TTL 284s
Completed 50 requests.
First result: Data for user-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that despite 50 requests firing simultaneously, the &lt;code&gt;[DB]&lt;/code&gt; log only appears once. The other 49 requests found the pending Promise in the map and waited for it. The TTL was set to 284 seconds, demonstrating the jitter subtracting from the 300-second base.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this breaks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; Memory usage on the application servers grows linearly until the process crashes with &lt;code&gt;FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory&lt;/code&gt;.&lt;br&gt;
&lt;strong&gt;Cause:&lt;/strong&gt; The &lt;code&gt;pendingRequests&lt;/code&gt; map is leaking Promises. This happens if a database call hangs indefinitely without a timeout, or if the code that removes the Promise from the map is inside a &lt;code&gt;.then()&lt;/code&gt; block rather than a &lt;code&gt;finally&lt;/code&gt; block. If the database call rejects, the &lt;code&gt;.then()&lt;/code&gt; block is skipped, the Promise stays in the map forever, and all future requests for that key will await a rejected Promise.&lt;br&gt;
&lt;strong&gt;Fix:&lt;/strong&gt; Always use a &lt;code&gt;finally&lt;/code&gt; block to delete the key from the coalescing map, as shown in the implementation above. Enforce strict timeouts on the database queries themselves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; Redis throws &lt;code&gt;OOM command not allowed when used memory &amp;gt; 'maxmemory'&lt;/code&gt;.&lt;br&gt;
&lt;strong&gt;Cause:&lt;/strong&gt; You implemented jitter by &lt;em&gt;adding&lt;/em&gt; a random amount of time to the base TTL, rather than subtracting it. If your base TTL was tuned to exactly fit your available Redis memory, adding 20% to the TTL increases the total number of keys stored at any given moment by 20%, breaching your memory limit.&lt;br&gt;
&lt;strong&gt;Fix:&lt;/strong&gt; Always calculate jitter by subtracting from the maximum acceptable TTL. The base TTL should represent the absolute longest you are willing to serve stale data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; The database still sees a thundering herd, just a smaller one.&lt;br&gt;
&lt;strong&gt;Cause:&lt;/strong&gt; In-memory request coalescing only protects the database from concurrent requests hitting the &lt;em&gt;same application instance&lt;/em&gt;. If you run 100 Node.js containers, and a popular key expires, each container will allow one request through to the database. You will still see 100 concurrent database queries.&lt;br&gt;
&lt;strong&gt;Fix:&lt;/strong&gt; For moderate scale, 100 queries is acceptable and in-memory coalescing is sufficient. For massive scale, you must implement distributed coalescing using Redis locks, ensuring only one container globally is allowed to query the database for a missing key.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Request coalescing trades database CPU for application memory and open connections. &lt;/p&gt;

&lt;p&gt;When you coalesce 500 requests, 499 of them are sitting in your application server doing nothing but holding open HTTP connections to your clients, waiting for the single database query to resolve. If the database query takes five seconds, you are holding 500 connections open for five seconds. Under heavy load, this can easily exhaust your application server's connection limits or memory, causing it to drop traffic before it even reaches the coalescing logic.&lt;/p&gt;

&lt;p&gt;This commits you to keeping your database queries fast. Coalescing is a shield against concurrent load, not a band-aid for slow queries. If the underlying query is inherently slow, you should not be coalescing requests; you should be pre-computing the cache asynchronously in a background worker so the API never has to wait for the database at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;When an interviewer asks, "How do you protect a database from a sudden traffic spike?", they are probing your understanding of cache failure modes.&lt;/p&gt;

&lt;p&gt;A weak answer is, "I would put a Redis cache in front of it." This is weak because it assumes caches never expire and never start empty. It demonstrates that the candidate has read about caching but has not supported a high-traffic cache in production.&lt;/p&gt;

&lt;p&gt;A strong answer names the failure mode: "I would use a cache, but I'd need to protect against a cache stampede or thundering herd when keys expire." A senior candidate will explain how they use jitter to prevent correlated expiry windows, and request coalescing to handle concurrent misses. &lt;/p&gt;

&lt;p&gt;If you are interviewing for a Staff or Director role, the interviewer is probing your grasp of trade-offs. The follow-up will be: "What happens to the application servers while they wait for the coalesced request?" A successful answer acknowledges the cost: holding those requests consumes memory and connections. The candidate should then pivot to discussing stale-while-revalidate patterns or background cache warming as superior architectural alternatives that avoid holding client connections open entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your tasks
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Observe the herd:&lt;/strong&gt; Modify the &lt;code&gt;index.js&lt;/code&gt; script. Comment out the coalescing check (step 2 in the function). Run the script again. You should see 50 &lt;code&gt;[DB]&lt;/code&gt; logs, simulating 50 concurrent database queries. This is the failure mode you are preventing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Break the cleanup:&lt;/strong&gt; Restore the coalescing check. Remove the &lt;code&gt;finally&lt;/code&gt; block entirely. Modify the &lt;code&gt;queryDatabase&lt;/code&gt; function to throw an error for &lt;code&gt;user-2&lt;/code&gt;. Fire a request for &lt;code&gt;user-2&lt;/code&gt;, catch the error, and then fire a second request for &lt;code&gt;user-2&lt;/code&gt;. Verify that the second request hangs forever because the rejected Promise is permanently stuck in the &lt;code&gt;pendingRequests&lt;/code&gt; map.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Implement stale-while-revalidate:&lt;/strong&gt; Modify the implementation so that if the cache has expired, but the data is still in Redis (you will need to store the insertion timestamp inside the cached JSON), the function immediately returns the stale data to the user, but fires the database query in the background to update Redis for the next user.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Your tasks this week
&lt;/h2&gt;

&lt;p&gt;Do the exercises above before the next chapter. Reading a tutorial and doing&lt;br&gt;
one are different activities and only one of them changes what you can build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stuck on any of them?&lt;/strong&gt; Say so — describe what you tried and what happened:&lt;br&gt;
&lt;a href="https://www.amitchakraborty.dev/learn/production-failures#stuck" rel="noopener noreferrer"&gt;tell me where you got stuck&lt;/a&gt;. I read every one, and the questions&lt;br&gt;
that come back more than twice get answered in the next chapter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production failures: the things that page you
&lt;/h2&gt;

&lt;p&gt;Chapter 2 of 24. New chapter every Friday morning.&lt;br&gt;
Next: &lt;strong&gt;Retry storms&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;· &lt;a href="https://www.amitchakraborty.dev/learn/production-failures" rel="noopener noreferrer"&gt;The full syllabus and every chapter so far&lt;/a&gt;&lt;br&gt;
· Subscribers also get the condensed notes for this chapter, the running&lt;br&gt;
  recap of everything the series has covered, and the extended guidance:&lt;br&gt;
  &lt;a href="https://www.amitchakraborty.dev/#newsletter" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by Amit Chakraborty — founding engineer and senior architect: React Native, AI and RAG systems, production architecture. &lt;a href="https://www.amitchakraborty.dev?utm_source=curriculum&amp;amp;utm_medium=content&amp;amp;utm_campaign=production-failures" rel="noopener noreferrer"&gt;Portfolio&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Need this built, reviewed or taught to your team? &lt;a href="https://www.amitchakraborty.dev/#contact" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt; or email &lt;a href="mailto:amit@devamit.co.in"&gt;amit@devamit.co.in&lt;/a&gt;. Available for senior and founding engineering roles, consulting and training, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>reliability</category>
      <category>productionfailures</category>
      <category>howdowedecoupletheexpiry</category>
    </item>
    <item>
      <title>Why Your Lighthouse Score is Lying About Your Core Web Vitals Regression</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Fri, 09 Oct 2026 04:30:03 +0000</pubDate>
      <link>https://dev.to/techamit95ch/why-your-lighthouse-score-is-lying-about-your-core-web-vitals-regression-mg5</link>
      <guid>https://dev.to/techamit95ch/why-your-lighthouse-score-is-lying-about-your-core-web-vitals-regression-mg5</guid>
      <description>&lt;p&gt;I was monitoring the Vercel dashboard for a Next.js clinical AI platform I built at Synapsis Medical Technologies when the Interaction to Next Paint (INP) metric for our patient portal spiked from 180ms to 450ms. In a HIPAA-compliant environment where clinicians rely on real-time RAG (Retrieval-Augmented Generation) responses, a 300ms delay isn't just a "yellow" warning in a report—it is the difference between an interface feeling responsive and a user double-clicking a button, potentially triggering duplicate API calls to an LLM pipeline.&lt;/p&gt;

&lt;p&gt;I ran Lighthouse. The score was a perfect 100. I ran a manual Trace in Chrome DevTools on my M3 Max MacBook; the main thread was clear. Yet, the field data—the real-world experience of patients on three-year-old Android devices and spotty hospital Wi-Fi—showed a catastrophic regression.&lt;/p&gt;

&lt;p&gt;We spent two days chasing phantom layout shifts before identifying that the regression wasn't in our code, but in how we were measuring it. If you rely on lab data (Lighthouse/Synthetic) to debug production regressions, you are looking at a curated, sterile version of your app that does not exist for your users.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Gap Between Lab and Field
&lt;/h2&gt;

&lt;p&gt;Lab data is a snapshot in a controlled environment. Field data (Chrome User Experience Report or CrUX) is the messy reality. At Synapsis, I oversaw the architecture of five production systems, and the most expensive lesson I learned was that a "Performance" score in a CI/CD pipeline is a vanity metric.&lt;/p&gt;

&lt;p&gt;Field regressions usually stem from three causes that lab tests cannot simulate:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Device Heterogeneity:&lt;/strong&gt; Your CI runner has a fast CPU; your user is on a $200 Motorola.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Network Jitter:&lt;/strong&gt; Lab tests use simulated throttling, which is predictable. Real-world 4G has packet loss that wreaks havoc on hydration.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;User Interaction:&lt;/strong&gt; Lighthouse doesn't click buttons. INP, which replaced FID (First Input Delay), only exists when a user interacts.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step 1: Isolate the Population, Not the Page
&lt;/h2&gt;

&lt;p&gt;When a regression hits, do not look at the aggregate score. A 200ms jump in Largest Contentful Paint (LCP) across your entire site is often a false signal caused by a change in traffic mix (e.g., a new marketing campaign hitting lower-end devices).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to do:&lt;/strong&gt; Open your Real User Monitoring (RUM) tool—whether that is Vercel Speed Insights, Datadog RUM, or a custom &lt;code&gt;web-vitals&lt;/code&gt; library implementation. Filter the regression by device type and connection type.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to confirm:&lt;/strong&gt; If the LCP regression only appears on "4G" and "Mobile" but "Desktop" remains flat, your issue is likely image encoding or a bloated JavaScript bundle that is choking lower-end CPUs during decompression, not a server-side logic error.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Identify the "Attribution" Element
&lt;/h2&gt;

&lt;p&gt;The biggest mistake I see leads make is guessing which element caused a Cumulative Layout Shift (CLS). The documentation tells you what CLS is; it doesn't tell you that the "offending" element in the console is often the victim, not the cause.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to do:&lt;/strong&gt; Use the &lt;code&gt;web-vitals&lt;/code&gt; JavaScript library to log the &lt;code&gt;attribution&lt;/code&gt; object to your analytics provider.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;onCLS&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;web-vitals/attribution&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;onCLS&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;attribution&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;largestShiftTarget&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; 
  &lt;span class="c1"&gt;// Send this string to your logging service&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;How to confirm:&lt;/strong&gt; Look for the &lt;code&gt;largestShiftTarget&lt;/code&gt; in your logs. In one instance, our "regression" was actually a third-party cookie consent banner that loaded 2 seconds late. Lighthouse missed it because the headless browser didn't trigger the consent flow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Debugging INP with Long Animation Frames (LoAF)
&lt;/h2&gt;

&lt;p&gt;If your INP has spiked, the standard "Performance" tab in DevTools is often too noisy. You need to see what is blocking the main thread at the exact moment of interaction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to do:&lt;/strong&gt; Enable the "Long Animation Frames API" in your testing. If you are on Chrome 123 or higher, you can use the &lt;code&gt;Long Animation Frames&lt;/code&gt; (LoAF) entry to see exactly which script contributed to the delay.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; Open DevTools -&amp;gt; Performance.&lt;/li&gt;
&lt;li&gt; Click the "Interactions" track.&lt;/li&gt;
&lt;li&gt; Look for the red bars.&lt;/li&gt;
&lt;li&gt; Identify the "Presentation Delay."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;How to confirm:&lt;/strong&gt; If the "Input Delay" is high, your main thread was busy when the user clicked. If the "Processing Duration" is high, your event listener code is slow. If "Presentation Delay" is high, the browser was struggling to paint the frame (often due to excessive DOM size or complex CSS).&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cost of Over-Optimisation
&lt;/h2&gt;

&lt;p&gt;At Synapsis, I oversaw a CI/CD overhaul that cut release cycles from 2 days to 4 hours. Part of that was removing "Performance Gatekeeping" in CI that relied on Lighthouse scores. &lt;/p&gt;

&lt;p&gt;Why? Because forcing an engineer to fix a 2-point Lighthouse drop in a PR often leads to "hacks" like delaying the load of essential scripts (like analytics or support chats) until after the first interaction. This improves the lab score but destroys the field INP, as the main thread suddenly hits a wall the moment the user tries to use the app. &lt;/p&gt;

&lt;p&gt;The trade-off is simple: Lab data is for catching catastrophic failures before they merge; Field data is for making business decisions. Never roll back a release based solely on a Lighthouse score if your RUM data remains stable.&lt;/p&gt;

&lt;h2&gt;
  
  
  At Your Level
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting out
&lt;/h3&gt;

&lt;p&gt;Stop using "Fast 3G" throttling in DevTools as a proxy for reality. It is a linear slowdown. Instead, use the "Network" tab to simulate offline and slow speeds, but focus on understanding the "Main Thread" work in the Performance tab. Your goal is to see how your code execution blocks the browser.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working engineer
&lt;/h3&gt;

&lt;p&gt;Implement the &lt;code&gt;web-vitals&lt;/code&gt; library with attribution. You cannot fix what you cannot name. When a ticket comes in for "slowness," your first response should be to pull the &lt;code&gt;largestShiftTarget&lt;/code&gt; or the &lt;code&gt;processingDuration&lt;/code&gt; for that specific user session.&lt;/p&gt;

&lt;h3&gt;
  
  
  Senior or staff
&lt;/h3&gt;

&lt;p&gt;Own the RUM strategy. You should be looking at the 75th and 95th percentiles (P75/P95). A regression in P95 often signals a memory leak or a specific edge case (like a very large FHIR resource loading into a state manager), while a regression in P75 signals a general architectural bloat.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead or director
&lt;/h3&gt;

&lt;p&gt;Distinguish between "Performance" as a technical metric and "User Experience" as a business one. A 500ms LCP increase is acceptable if it comes from high-resolution product images that increase conversion by 10%. Do not let your team spend 40 hours chasing a green score if the business metric is moving in the right direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the Interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "Our LCP has regressed by 1 second in production, but Lighthouse shows no change. How do you find the cause?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Weak Answer:&lt;/strong&gt; "I would check the images, minify the CSS, and maybe try to implement a CDN or use &lt;code&gt;next/image&lt;/code&gt;." This is weak because it jumps to solutions without diagnosing why the tools are giving conflicting signals.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Strong Answer:&lt;/strong&gt; A strong candidate identifies that Lighthouse is a synthetic environment and the regression is likely environment-dependent. They will talk about &lt;strong&gt;Field vs. Lab data&lt;/strong&gt;. They will mention checking &lt;strong&gt;CrUX reports&lt;/strong&gt; or &lt;strong&gt;RUM data&lt;/strong&gt; to segment by device and geography. They will specifically name &lt;strong&gt;Interaction to Next Paint (INP)&lt;/strong&gt; as a metric that Lighthouse cannot accurately measure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Senior Follow-up:&lt;/strong&gt; "How do you handle a third-party script that is tanking your INP but is required by the marketing team?" &lt;br&gt;
This separates those who have read docs from those who have shipped. A senior engineer knows you can't always delete the script. They will discuss &lt;code&gt;requestIdleCallback&lt;/code&gt;, using &lt;code&gt;Partytown&lt;/code&gt; to move scripts to a Web Worker, or strategically yielding to the main thread using &lt;code&gt;setTimeout(..., 0)&lt;/code&gt; to break up long tasks. They understand the organizational friction between performance and third-party requirements.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=reading-a-core-web-vitals-regression-properly" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>frontend</category>
      <category>corewebvitals</category>
      <category>performance</category>
      <category>architecture</category>
    </item>
    <item>
      <title>MLOps: models that survive production · 2. Data versioning</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Thu, 08 Oct 2026 16:30:07 +0000</pubDate>
      <link>https://dev.to/techamit95ch/mlops-models-that-survive-production-2-data-versioning-p3k</link>
      <guid>https://dev.to/techamit95ch/mlops-models-that-survive-production-2-data-versioning-p3k</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;MLOps: models that survive production&lt;/strong&gt; · Chapter 2 of 12 · AI Engineering · new chapter every Thursday night&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;By the end of this chapter:&lt;/strong&gt; Reproduce a training run from six months ago.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Say you are an engineer who has just been asked to retrain the churn prediction model from Q1. A new enterprise customer needs it deployed in their own environment, but with a slight tweak to the output format. You check out the Git tag &lt;code&gt;release-q1&lt;/code&gt;, run the training script, and wait.&lt;/p&gt;

&lt;p&gt;The original run logged an F1 score of 0.88. Your new run logs 0.64.&lt;/p&gt;

&lt;p&gt;You have not changed a single line of code. The random seeds are pinned. The dependencies are locked in a requirements file. But the data in &lt;code&gt;s3://company-ml-data/churn/latest.csv&lt;/code&gt; is from today. Since Q1, the upstream engineering team has added three new columns, changed the definition of an 'active user', and dropped all rows from a deprecated region. &lt;/p&gt;

&lt;p&gt;You spend three days trying to manually reconstruct the Q1 dataset from database backups. You fail. You have to explain to your lead that the model currently running in production cannot be rebuilt, audited, or modified, because the data it was trained on no longer exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you start
&lt;/h2&gt;

&lt;p&gt;You need a terminal, Python, Git, and DVC (Data Version Control). &lt;/p&gt;

&lt;p&gt;Install DVC via pip. Do not use Homebrew or apt for this; keeping DVC in your Python environment ensures your CI pipeline and your local machine run the same version.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;&lt;span class="nv"&gt;dvc&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;3.48.4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Verify the setup. This command must return a version number for both tools:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git &lt;span class="nt"&gt;--version&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; dvc &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create a new directory for this exercise and move into it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;churn-model &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;churn-model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why Git cannot hold your data
&lt;/h2&gt;

&lt;p&gt;The instinct when faced with this problem is to put the data in Git alongside the code. If the code and data are in the same commit, checking out the commit restores both.&lt;/p&gt;

&lt;p&gt;Git is designed for text. It tracks changes line by line. If you commit a 5GB CSV file, Git will attempt to compress it and store the delta. It will fail to do this efficiently. Your &lt;code&gt;.git&lt;/code&gt; directory will balloon. Cloning the repository will take hours. Eventually, you will push to GitHub, which will reject any file larger than 100MB, and you will have to rewrite your Git history to remove the commit.&lt;/p&gt;

&lt;p&gt;Git Large File Storage (LFS) is the standard workaround in software engineering, but it fails in machine learning. Git LFS tightly couples your data storage to your Git provider. If you have 5TB of training data, storing it in GitHub LFS is prohibitively expensive compared to an AWS S3 bucket. Furthermore, Git LFS lacks the semantics for data pipelines; it cannot tell you if a dataset was generated by a specific script, only that the file changed.&lt;/p&gt;

&lt;p&gt;You need Git to track the exact state of your data, without Git ever touching the data itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pointer file pattern
&lt;/h2&gt;

&lt;p&gt;To solve this, we use the pointer file pattern. This is the mechanism underlying DVC.&lt;/p&gt;

&lt;p&gt;Instead of committing &lt;code&gt;dataset.csv&lt;/code&gt; to Git, you hash the file's contents. You store the heavy &lt;code&gt;dataset.csv&lt;/code&gt; in a dumb object store (like S3, or just a hidden folder on your laptop) renamed to its hash, for example &lt;code&gt;a1b2c3d4.csv&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Then, you create a tiny text file called &lt;code&gt;dataset.csv.dvc&lt;/code&gt;. This file contains nothing but the hash &lt;code&gt;a1b2c3d4&lt;/code&gt; and the file size. You commit this tiny text file to Git.&lt;/p&gt;

&lt;p&gt;Git tracks the code and the pointer. DVC tracks the heavy file. When you check out a Git commit from six months ago, Git updates &lt;code&gt;dataset.csv.dvc&lt;/code&gt; to contain the old hash. You then tell DVC to read that pointer, find the corresponding heavy file in the object store, and copy it back into your workspace. &lt;/p&gt;

&lt;p&gt;Code and data are synchronised, but Git only handles text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build it
&lt;/h2&gt;

&lt;p&gt;We are going to create a dataset, version it, overwrite it with new data, and then successfully time-travel back to the original state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Initialise the repositories&lt;/strong&gt;&lt;br&gt;
Initialise Git, then initialise DVC. DVC will create a &lt;code&gt;.dvc&lt;/code&gt; directory to act as your local object store.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git init
dvc init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;DVC creates some internal configuration files. Commit them to Git so your repository is clean.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"Initialise DVC"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2: Generate the Q1 data&lt;/strong&gt;&lt;br&gt;
Create a Python script named &lt;code&gt;generate_data.py&lt;/code&gt; to simulate our Q1 dataset.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# generate_data.py
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;csv&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;write_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;newline&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;writer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;csv&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writerow&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;active_days&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;churned&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writerows&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Q1 Data
&lt;/span&gt;&lt;span class="nf"&gt;write_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;dataset.csv&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wrote Q1 data to dataset.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python generate_data.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 3: Track the data with DVC&lt;/strong&gt;&lt;br&gt;
Tell DVC to track the CSV file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dvc add dataset.csv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at the output. DVC tells you exactly what to do next:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;To track the changes with git, run:
    git add dataset.csv.dvc .gitignore
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When you ran &lt;code&gt;dvc add&lt;/code&gt;, DVC did three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It calculated the MD5 hash of &lt;code&gt;dataset.csv&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;It moved &lt;code&gt;dataset.csv&lt;/code&gt; into &lt;code&gt;.dvc/cache/&lt;/code&gt; under its hash name, and put a hard link back in your workspace.&lt;/li&gt;
&lt;li&gt;It created &lt;code&gt;dataset.csv.dvc&lt;/code&gt; (the pointer) and updated &lt;code&gt;.gitignore&lt;/code&gt; so you do not accidentally commit the raw CSV to Git.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Step 4: Commit the pointer to Git&lt;/strong&gt;&lt;br&gt;
Commit the code, the pointer, and the gitignore file. This creates our Q1 snapshot.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git add generate_data.py dataset.csv.dvc .gitignore
git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"Train Q1 model"&lt;/span&gt;
git tag release-q1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 5: Generate the Q2 data&lt;/strong&gt;&lt;br&gt;
Time passes. The definition of the data changes. Modify &lt;code&gt;generate_data.py&lt;/code&gt; to output the Q2 data. Change the rows and add a new column.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# generate_data.py
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;csv&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;write_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;newline&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;writer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;csv&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="c1"&gt;# Schema change: added 'region'
&lt;/span&gt;        &lt;span class="n"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writerow&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;user_id&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;active_days&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;churned&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;writer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writerows&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Q2 Data
&lt;/span&gt;&lt;span class="nf"&gt;write_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;dataset.csv&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;EU&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;US&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;45&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;EU&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wrote Q2 data to dataset.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it to overwrite the CSV:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python generate_data.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 6: Track and commit the Q2 data&lt;/strong&gt;&lt;br&gt;
Update DVC with the new file, then update Git with the new pointer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dvc add dataset.csv
git add generate_data.py dataset.csv.dvc
git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"Train Q2 model"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you open &lt;code&gt;dataset.csv&lt;/code&gt; now, you will see the Q2 data with the &lt;code&gt;region&lt;/code&gt; column.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 7: Time travel&lt;/strong&gt;&lt;br&gt;
You are asked to reproduce the Q1 model. First, check out the old Git commit using the tag we made.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git checkout release-q1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at &lt;code&gt;dataset.csv&lt;/code&gt;. &lt;strong&gt;It has not changed.&lt;/strong&gt; It still contains the Q2 data. &lt;/p&gt;

&lt;p&gt;This is the most common point of confusion. Git only updated the text files it tracks. It updated &lt;code&gt;generate_data.py&lt;/code&gt; and it updated the pointer file &lt;code&gt;dataset.csv.dvc&lt;/code&gt;. It did not touch the actual CSV. &lt;/p&gt;

&lt;p&gt;To synchronise your workspace with the pointer file, run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dvc checkout
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;M       dataset.csv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open &lt;code&gt;dataset.csv&lt;/code&gt; again. The &lt;code&gt;region&lt;/code&gt; column is gone. The data is exactly as it was in Q1. You can now run your training script and get the exact 0.88 F1 score you got six months ago.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this breaks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The file is modified but &lt;code&gt;dvc checkout&lt;/code&gt; does nothing&lt;/strong&gt;&lt;br&gt;
You manually edited &lt;code&gt;dataset.csv&lt;/code&gt; to fix a typo. You realise you made a mistake, so you run &lt;code&gt;dvc checkout&lt;/code&gt; to restore the file to the state of the pointer. Nothing happens. The typo is still there.&lt;/p&gt;

&lt;p&gt;DVC uses file timestamps and sizes to optimise checkouts. If you edit a file in place without changing its size enough, DVC might not realise it has changed. To force DVC to verify the hashes and overwrite your local modifications, use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dvc checkout &lt;span class="nt"&gt;--force&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;ERROR: unexpected error - .dvc/cache is not tracked by git&lt;/strong&gt;&lt;br&gt;
You cloned your repository to a new machine, ran &lt;code&gt;dvc checkout&lt;/code&gt;, and got an error about a missing cache or a missing file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR: failed to pull data from the cloud - Access Denied
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;.dvc/cache&lt;/code&gt; directory is in your local &lt;code&gt;.gitignore&lt;/code&gt;. It does not push to GitHub. When you clone the repository, you get the pointers, but not the heavy files. To get the heavy files, you must configure a DVC remote (like an S3 bucket) using &lt;code&gt;dvc remote add&lt;/code&gt;, push the data there from the original machine (&lt;code&gt;dvc push&lt;/code&gt;), and pull it on the new machine (&lt;code&gt;dvc pull&lt;/code&gt;). If you get Access Denied, your AWS credentials in your terminal session do not have &lt;code&gt;s3:GetObject&lt;/code&gt; permissions for that bucket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Git merge conflicts in .dvc files&lt;/strong&gt;&lt;br&gt;
You and a colleague both updated the dataset on different branches. When you merge, Git throws a conflict in &lt;code&gt;dataset.csv.dvc&lt;/code&gt;.&lt;br&gt;
A &lt;code&gt;.dvc&lt;/code&gt; file is just YAML containing a hash. You cannot merge hashes. You must pick one. Accept either your changes or their changes in the Git merge tool. Then, whichever one you picked, run &lt;code&gt;dvc checkout&lt;/code&gt; to ensure your local CSV matches the hash you just committed to the merge.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;This approach costs storage space, and it scales poorly for append-only data.&lt;/p&gt;

&lt;p&gt;Because DVC hashes the entire file, it does not do delta compression for tabular data. If you have a 10GB CSV file and you append one row, &lt;code&gt;dvc add&lt;/code&gt; will hash the new file, copy the entire 10GB into &lt;code&gt;.dvc/cache&lt;/code&gt;, and update the pointer. You are now storing 20GB of data. If you do this every day for a month, you are storing 300GB of redundant data.&lt;/p&gt;

&lt;p&gt;For datasets that update frequently, flat-file versioning with DVC is the wrong tool. You are paying the storage cost of full duplication. In those scenarios, you must move to a time-travel database format like Apache Iceberg or Delta Lake, which version data at the row or block level, or rely on a Feature Store. &lt;/p&gt;

&lt;p&gt;DVC is best used for unstructured data (images, audio) where files are immutable and added discretely, or for tabular data that is generated as a distinct, immutable artifact for a specific training run.&lt;/p&gt;

&lt;p&gt;It also costs developer friction. Every data change now requires two commands (&lt;code&gt;dvc add&lt;/code&gt;, then &lt;code&gt;git add&lt;/code&gt;). If you forget &lt;code&gt;dvc add&lt;/code&gt;, you will commit a stale pointer to Git, and your CI pipeline will train on old data while your Git history claims otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;When an interviewer asks, "How do you ensure your training runs are reproducible?", they are checking if you understand that code is only half the state of an ML system.&lt;/p&gt;

&lt;p&gt;A weak answer focuses entirely on the environment: "I set the random seeds, pin the requirements, and use Docker." This is weak because it assumes the data is static. In production, data is a moving target. If you run the same Docker container against a database that has changed, the output changes.&lt;/p&gt;

&lt;p&gt;A strong answer explicitly names the pointer file pattern. "I version the code, the environment, and the data. I use a tool like DVC to hash the training artifacts, store them in an object store, and commit the pointer files to Git. This guarantees that checking out a Git commit restores the exact bytes of data used for that run."&lt;/p&gt;

&lt;p&gt;If you are interviewing for a Senior or Staff role, the interviewer will probe the limits of this. They will ask what happens when the dataset is 500GB and updates daily. A senior candidate must immediately identify the storage duplication trade-off. You should explain that DVC is for artifact versioning, not database versioning. For daily updates at scale, you would advocate for Delta Lake to query the data "as of" a timestamp, or a Feature Store that guarantees point-in-time correctness, rather than hashing flat files.&lt;/p&gt;

&lt;p&gt;The follow-up question that separates practitioners from theorists is: "What happens if someone modifies the data file but forgets to run &lt;code&gt;dvc add&lt;/code&gt; before committing to Git?" The practitioner knows exactly what happens: the Git commit contains the old hash, the new data is left untracked in the workspace, and anyone pulling the repository will get the old data, silently invalidating the training run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your tasks
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Configure a local remote.&lt;/strong&gt; By default, DVC stores the heavy files in &lt;code&gt;.dvc/cache&lt;/code&gt;. Create a new directory outside your Git repository (e.g., &lt;code&gt;mkdir /tmp/dvc-remote&lt;/code&gt;). Use &lt;code&gt;dvc remote add -d myremote /tmp/dvc-remote&lt;/code&gt; to configure it. Run &lt;code&gt;dvc push&lt;/code&gt;. Delete your local &lt;code&gt;.dvc/cache&lt;/code&gt; directory, then run &lt;code&gt;dvc pull&lt;/code&gt;. Verify &lt;code&gt;dataset.csv&lt;/code&gt; is restored.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Break the sync.&lt;/strong&gt; Open &lt;code&gt;dataset.csv&lt;/code&gt; and manually change a value. Do not run &lt;code&gt;dvc add&lt;/code&gt;. Run &lt;code&gt;dvc status&lt;/code&gt;. It will tell you the file is modified. Use the DVC CLI to discard your manual changes and restore the file to match the Git pointer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the pointer.&lt;/strong&gt; You do not strictly need the DVC CLI to know what data a Git commit used. Write a three-line Python script that opens &lt;code&gt;dataset.csv.dvc&lt;/code&gt;, parses it as YAML (or just string splits it), and prints the MD5 hash. This is how CI systems verify data without downloading the whole DVC binary.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Your tasks this week
&lt;/h2&gt;

&lt;p&gt;Do the exercises above before the next chapter. Reading a tutorial and doing&lt;br&gt;
one are different activities and only one of them changes what you can build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stuck on any of them?&lt;/strong&gt; Say so — describe what you tried and what happened:&lt;br&gt;
&lt;a href="https://www.amitchakraborty.dev/learn/mlops-bootcamp#stuck" rel="noopener noreferrer"&gt;tell me where you got stuck&lt;/a&gt;. I read every one, and the questions&lt;br&gt;
that come back more than twice get answered in the next chapter.&lt;/p&gt;

&lt;h2&gt;
  
  
  MLOps: models that survive production
&lt;/h2&gt;

&lt;p&gt;Chapter 2 of 12. New chapter every Thursday night.&lt;br&gt;
Next: &lt;strong&gt;Features: stores, and the simpler thing that usually works&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;· &lt;a href="https://www.amitchakraborty.dev/learn/mlops-bootcamp" rel="noopener noreferrer"&gt;The full syllabus and every chapter so far&lt;/a&gt;&lt;br&gt;
· Subscribers also get the condensed notes for this chapter, the running&lt;br&gt;
  recap of everything the series has covered, and the extended guidance:&lt;br&gt;
  &lt;a href="https://www.amitchakraborty.dev/#newsletter" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by Amit Chakraborty — founding engineer and senior architect: React Native, AI and RAG systems, production architecture. &lt;a href="https://www.amitchakraborty.dev?utm_source=curriculum&amp;amp;utm_medium=content&amp;amp;utm_campaign=mlops-bootcamp" rel="noopener noreferrer"&gt;Portfolio&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Need this built, reviewed or taught to your team? &lt;a href="https://www.amitchakraborty.dev/#contact" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt; or email &lt;a href="mailto:amit@devamit.co.in"&gt;amit@devamit.co.in&lt;/a&gt;. Available for senior and founding engineering roles, consulting and training, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>aiengineering</category>
      <category>mlops</category>
      <category>whygitcannotholdyourdata</category>
    </item>
    <item>
      <title>Defending the 16ms Frame: A Performance Budget for React Native TurboModules</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Thu, 08 Oct 2026 16:30:03 +0000</pubDate>
      <link>https://dev.to/techamit95ch/defending-the-16ms-frame-a-performance-budget-for-react-native-turbomodules-496n</link>
      <guid>https://dev.to/techamit95ch/defending-the-16ms-frame-a-performance-budget-for-react-native-turbomodules-496n</guid>
      <description>&lt;p&gt;I was the first engineering hire at Synapsis Medical Technologies, where I owned the architecture for a HealthTech AI platform from day zero. We were integrating real-time health data from wearables and running HIPAA-aligned RAG pipelines. During a critical phase of our scaling—where I grew the team from zero to 21 engineers in 13 months—we hit a wall. &lt;/p&gt;

&lt;p&gt;We were building a clinical dashboard that needed to stream high-frequency biometric data. On the old Bridge architecture, the serialisation overhead was killing us. Frames were dropping, the UI felt "heavy," and our clinical users—doctors who don't have patience for laggy interfaces—started complaining. We had moved to React Native 0.76 to leverage the New Architecture, but simply "using TurboModules" didn't fix the jank.&lt;/p&gt;

&lt;p&gt;The problem wasn't the technology; it was the lack of a ceiling. We were treating native calls as "free" when they were actually the most expensive part of our render loop. I had to implement a strict performance budget to stop the team from inadvertently shipping 100ms blocks to the main thread. This moved us from a vague "make it faster" mandate to a measurable 16.6ms frame budget that we defended in every PR.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cost of the Invisible Bridge
&lt;/h2&gt;

&lt;p&gt;In the legacy architecture, the Bridge was a bottleneck because of asynchronous JSON serialisation. TurboModules solve this by using JSI (JavaScript Interface), allowing direct C++ communication. However, I’ve found that developers often treat JSI as a license to move massive amounts of data frequently.&lt;/p&gt;

&lt;p&gt;If you pass a 5MB clinical dataset over JSI every 100ms, you won't see a "Bridge Busy" warning, but you will see the JavaScript thread stall while the C++ layer handles the memory allocation. In our case, the cost was a two-day release cycle that I eventually cut to four hours by automating the validation of these boundaries. If the codegen didn't match the performance spec, the build failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why JSI isn't a silver bullet
&lt;/h2&gt;

&lt;p&gt;The documentation tells you how to write a &lt;code&gt;Spec&lt;/code&gt; file and run &lt;code&gt;node-modules/.bin/react-native codegen&lt;/code&gt;. It doesn't tell you that C++ handles memory differently than JavaScript's garbage collector. &lt;/p&gt;

&lt;p&gt;When you define a TurboModule, you are creating a synchronous contract. If your native method takes 20ms to execute, your JavaScript thread is frozen for 20ms. You have just dropped a frame. On a 60Hz display, you have 16.6ms to do everything. If your TurboModule takes 10ms, you have 6.6ms left for React rendering, layout, and event handling. That is a razor-thin margin.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix: Implementing a Codegen Performance Budget
&lt;/h2&gt;

&lt;p&gt;This approach ensures that native modules are not just "fast," but "predictably fast."&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Define the Strict Spec with Primitive Types
&lt;/h3&gt;

&lt;p&gt;Avoid passing generic &lt;code&gt;Object&lt;/code&gt; or &lt;code&gt;Array&lt;/code&gt; types in your TypeScript specs. Every time JSI has to iterate over a generic object to find a key, you lose microseconds. Be explicit.&lt;/p&gt;

&lt;p&gt;In your &lt;code&gt;NativeBiometricScanner.ts&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;TurboModule&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-native&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;TurboModuleRegistry&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-native&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;Spec&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nx"&gt;TurboModule&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// BAD: getHeartRateData(): Object; &lt;/span&gt;
  &lt;span class="c1"&gt;// GOOD: Explicit primitives reduce JSI conversion overhead&lt;/span&gt;
  &lt;span class="nf"&gt;getHeartRateData&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;patientId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;bpm&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nx"&gt;TurboModuleRegistry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;getEnforced&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Spec&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;NativeBiometricScanner&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why:&lt;/strong&gt; Codegen uses these types to generate C++ structs. Primitives map directly to C++ types (&lt;code&gt;double&lt;/code&gt;, &lt;code&gt;bool&lt;/code&gt;, &lt;code&gt;std::string&lt;/code&gt;). Generic objects require a &lt;code&gt;jsi::Object&lt;/code&gt; lookup, which is significantly slower under load.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Instrumented Codegen Execution
&lt;/h3&gt;

&lt;p&gt;Don't just run the codegen; wrap it in a script that validates the generated C++ glue code for "fat" signatures. I used a simple grep-based check in our CI/CD pipeline to flag any TurboModule using &lt;code&gt;jsi::Value&lt;/code&gt; or &lt;code&gt;jsi::Object&lt;/code&gt; where a primitive could have been used.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Run codegen&lt;/span&gt;
node node_modules/react-native/scripts/generate-codegen-artifacts.js &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--path&lt;/span&gt; ./ &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--outputPath&lt;/span&gt; ./generated

&lt;span class="c"&gt;# Confirm it worked: Check for forbidden generic types in the generated Header&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="s2"&gt;"jsi::Object"&lt;/span&gt; ./generated/NativeBiometricScannerSpec.h&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Error: Generic Object detected in TurboModule Spec. Use explicit types."&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. The 16ms Telemetry Wrapper
&lt;/h3&gt;

&lt;p&gt;You cannot manage what you do not measure. We wrapped our TurboModule calls in a telemetry utility. If a call exceeded 8ms (half our frame budget), it triggered a warning in our staging environment.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;performance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;NativeBiometricScanner&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getHeartRateData&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pt-123&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;performance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;end&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`TurboModule Budget Exceeded: getHeartRateData took &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;end&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;ms`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The Number:&lt;/strong&gt; In my experience, a single TurboModule call should ideally take less than 2ms. If you are seeing 8ms+, you are either passing too much data or doing heavy computation on the caller's thread instead of a background thread.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Offloading to Background Threads
&lt;/h3&gt;

&lt;p&gt;If the telemetry shows you are over budget, you must move the work. TurboModules are synchronous by default, but you can return a &lt;code&gt;Promise&lt;/code&gt; which moves the execution to a native background thread.&lt;/p&gt;

&lt;p&gt;In your Objective-C++ or C++ implementation, ensure you aren't blocking &lt;code&gt;RCTQueue&lt;/code&gt;. Use &lt;code&gt;dispatch_async&lt;/code&gt; or &lt;code&gt;std::thread&lt;/code&gt; for anything involving file I/O or complex AI processing (like the RAG pipelines we ran).&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, and when not to do it
&lt;/h2&gt;

&lt;p&gt;This strict approach has a high development cost. Writing explicit specs takes 30% longer than throwing an &lt;code&gt;any&lt;/code&gt; type at the problem. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When to skip this:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you are building a CRUD app with low data frequency (e.g., a simple form entry).&lt;/li&gt;
&lt;li&gt;If your app doesn't target low-end Android devices where JSI overhead is more pronounced.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;When it is mandatory:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;High-frequency data (health, finance, sensors).&lt;/li&gt;
&lt;li&gt;Complex animations that must stay synced with native state.&lt;/li&gt;
&lt;li&gt;When scaling a team. Without these guardrails, junior engineers will treat TurboModules like standard JS functions, eventually leading to a "death by a thousand cuts" performance profile that is nearly impossible to debug later.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  At your level
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting out
&lt;/h3&gt;

&lt;p&gt;Stop using &lt;code&gt;any&lt;/code&gt; in your TypeScript files immediately. TurboModule codegen relies on your types to build the C++ interface; if your types are vague, your native performance will be too. Focus on learning how &lt;code&gt;TurboModuleRegistry&lt;/code&gt; differs from &lt;code&gt;NativeModules&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working engineer
&lt;/h3&gt;

&lt;p&gt;Implement the telemetry wrapper mentioned in Step 3. You likely have modules that are slow, but you don't know which ones. Identify the top three slowest calls and refactor their specs to use primitives instead of objects.&lt;/p&gt;

&lt;h3&gt;
  
  
  Senior or staff
&lt;/h3&gt;

&lt;p&gt;You own the architecture. Your job is to prevent the "jank" before it's written. Set up the CI/CD check to fail builds that use generic &lt;code&gt;jsi::Object&lt;/code&gt; in TurboModule specs. This enforces a performance-first culture without you having to manually review every line of code.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead or director
&lt;/h3&gt;

&lt;p&gt;Performance is a budget, not a feature. If your team is missing release deadlines due to "unexplained lag," it's likely a boundary issue. Allocate the 20% extra time needed for proper codegen specs now to avoid the 200% time cost of a performance refactor six months down the line.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "How do TurboModules improve performance over the legacy Bridge?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Weak Answer:&lt;/strong&gt; "They are faster because they use JSI instead of JSON serialisation, so there's less overhead when calling native code." &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it's weak:&lt;/strong&gt; It's a textbook answer. It doesn't show you've actually dealt with the consequences of JSI, such as thread blocking or memory management.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Strong Answer:&lt;/strong&gt; "TurboModules leverage JSI to provide synchronous access to native methods, eliminating the asynchronous JSON serialisation bottleneck. However, the real advantage is the type-safety provided by Codegen, which generates C++ structs. A strong implementation avoids &lt;code&gt;jsi::Object&lt;/code&gt; in favour of primitives to minimize lookup overhead. The trade-off is that because it's synchronous, a poorly written native method can now block the JavaScript thread and drop frames, which wasn't as direct a risk with the asynchronous Bridge. You have to defend a 16ms frame budget by offloading heavy work to background threads and returning Promises."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Senior/Staff Follow-up:&lt;/strong&gt; "How do you handle large data transfers between JS and Native in the New Architecture?"&lt;br&gt;
&lt;strong&gt;The "I've done this" Answer:&lt;/strong&gt; "I wouldn't pass large blobs over JSI if I can avoid it. Even with TurboModules, moving a 10MB string is expensive. I’d prefer to pass a reference to a memory-mapped file or use &lt;code&gt;ArrayBuffer&lt;/code&gt; / &lt;code&gt;SharedArrayBuffer&lt;/code&gt; to share the memory space between the C++ layer and the JS engine. This avoids the copy-overhead entirely."&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=turbomodules-and-codegen-a-performance-budget-approach" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>turbomodules</category>
      <category>jsi</category>
      <category>performance</category>
    </item>
    <item>
      <title>The hidden cost of TurboModules: Why Codegen is not a free lunch</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Thu, 08 Oct 2026 08:30:04 +0000</pubDate>
      <link>https://dev.to/techamit95ch/the-hidden-cost-of-turbomodules-why-codegen-is-not-a-free-lunch-inc</link>
      <guid>https://dev.to/techamit95ch/the-hidden-cost-of-turbomodules-why-codegen-is-not-a-free-lunch-inc</guid>
      <description>&lt;p&gt;I was leading the architecture for a clinical AI platform at Synapsis Medical Technologies, where we were integrating real-time health data from wearables via Bluetooth Low Energy (BLE). We were hitting a wall with the legacy React Native bridge. High-frequency data packets—heart rate, SpO2, and accelerometer streams—were saturating the bridge, causing noticeable UI stutters and delaying the execution of our RAG-based clinical alerts. &lt;/p&gt;

&lt;p&gt;To solve this, we migrated our native modules to the New Architecture (TurboModules) in React Native 0.7x. On paper, it was the perfect solution: synchronous execution, type safety, and no bridge congestion. Instead, our CI/CD pipeline, which I had previously optimized to a 4-hour release cycle, bloated by 25 minutes per build. We saw cryptic C++ compilation errors that stalled our 21-person engineering team for two days. The "type safety" we were promised required a rigid boilerplate that made simple feature additions a multi-file choreography across TypeScript, C++, and Java.&lt;/p&gt;

&lt;p&gt;If you are moving to TurboModules because the documentation says it is "faster," you are only hearing half the story. The performance gains are real, but the tax on your developer experience and build infrastructure is significant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the bottleneck moves from the Bridge to the Build
&lt;/h2&gt;

&lt;p&gt;In the legacy architecture, the bridge was a JSON-based message bus. It was slow because every call had to be serialized, passed across the boundary, and deserialized. TurboModules eliminate this by using JSI (JavaScript Interface), allowing JavaScript to hold a direct reference to C++ host objects.&lt;/p&gt;

&lt;p&gt;However, JSI requires C++ glue code to bind the JavaScript calls to the native platform code (Objective-C or Java). React Native uses a tool called &lt;code&gt;codegen&lt;/code&gt; to generate this glue code. &lt;/p&gt;

&lt;p&gt;The failure mode nobody warns you about is the &lt;strong&gt;Codegen-Build Loop&lt;/strong&gt;. In the old way, you changed a Java file and re-ran the app. Now, if you change a single property in your TypeScript interface, &lt;code&gt;codegen&lt;/code&gt; must re-run, generating new C++ headers, which then triggers a full recompilation of the native bridge layer. On a standard M2 MacBook Pro, this adds a non-trivial overhead to every native change. If your CI is running on resource-constrained runners, this is where your 4-hour release cycle starts to creep toward 5 hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  The migration: A step-by-step hardening
&lt;/h2&gt;

&lt;p&gt;Do not follow the "Hello World" tutorials that suggest moving your entire library at once. In my experience shipping 18+ production apps, the only way to survive this without breaking the build for the rest of your team is a staggered implementation.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Define the Spec strictly
&lt;/h3&gt;

&lt;p&gt;Create a &lt;code&gt;Native[ModuleName].ts&lt;/code&gt; file. The naming convention is not optional; &lt;code&gt;codegen&lt;/code&gt; looks for the &lt;code&gt;Native&lt;/code&gt; prefix.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// NativeHealthScanner.ts&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;TurboModule&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-native&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;TurboModuleRegistry&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-native&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;Spec&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nx"&gt;TurboModule&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Use exact types. 'any' will fail codegen.&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;getSensorData&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sensorId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;startScan&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;frequency&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nx"&gt;TurboModuleRegistry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;getEnforcing&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Spec&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;HealthScanner&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why:&lt;/strong&gt; &lt;code&gt;codegen&lt;/code&gt; is a parser, not a compiler. If you use complex TypeScript unions or external type imports inside this file, the parser will crash with an unhelpful "Task :app:generateCodegenArtifactsFromSchema Failed" error. Keep this file isolated.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Trigger Codegen manually to verify
&lt;/h3&gt;

&lt;p&gt;Before you try to compile the whole app, run the script to generate the specs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# For iOS&lt;/span&gt;
&lt;span class="nb"&gt;cd &lt;/span&gt;ios &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; bundle &lt;span class="nb"&gt;exec &lt;/span&gt;pod &lt;span class="nb"&gt;install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Confirm it worked:&lt;/strong&gt; Look into &lt;code&gt;ios/build/generated/ios&lt;/code&gt;. You should see &lt;code&gt;HealthScannerSpec.h&lt;/code&gt; and &lt;code&gt;HealthScannerSpec-generated.mm&lt;/code&gt;. If these files aren't there, your &lt;code&gt;package.json&lt;/code&gt; configuration for &lt;code&gt;codegenConfig&lt;/code&gt; is likely pointing to the wrong directory.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Implement the C++ boilerplate (iOS)
&lt;/h3&gt;

&lt;p&gt;This is where the "cost" becomes visible. You cannot just write Objective-C anymore. You must provide a bridge implementation that conforms to the generated C++ protocol.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight objective_c"&gt;&lt;code&gt;&lt;span class="c1"&gt;// HealthScanner.mm&lt;/span&gt;
&lt;span class="cp"&gt;#import "HealthScanner.h"
&lt;/span&gt;
&lt;span class="k"&gt;@implementation&lt;/span&gt; &lt;span class="nc"&gt;HealthScanner&lt;/span&gt;
&lt;span class="n"&gt;RCT_EXPORT_MODULE&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;std&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;shared_ptr&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;facebook&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;react&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;TurboModule&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;getTurboModule&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="n"&gt;facebook&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;react&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;ObjCTurboModule&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;InitParams&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nv"&gt;params&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;std&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;make_shared&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;facebook&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;react&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;NativeHealthScannerSpecJSI&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Your actual logic&lt;/span&gt;
&lt;span class="k"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;void&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;startScan&lt;/span&gt;&lt;span class="p"&gt;:(&lt;/span&gt;&lt;span class="kt"&gt;double&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nv"&gt;frequency&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// Implementation&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;@end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The Trade-off:&lt;/strong&gt; You are now maintaining a C++ shared pointer inside an Objective-C++ file. If your team consists purely of React developers, you have just increased the bus factor of your native infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hidden costs of the New Architecture
&lt;/h2&gt;

&lt;p&gt;Having scaled a team from 0 to 21, I’ve seen how these architectural choices impact velocity.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Binary Size:&lt;/strong&gt; TurboModules require the inclusion of the JSI and often the Hermes engine (though technically separate, they are designed to work together). In my experience, migrating a medium-sized app to full TurboModules and Fabric (the new renderer) can increase the initial APK size by 3MB to 5MB due to the C++ runtime requirements.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Compilation Time:&lt;/strong&gt; As mentioned, our build times increased. Specifically, the &lt;code&gt;ninja&lt;/code&gt; build process for C++ is CPU-intensive. If your developers are on 8GB RAM machines, they will feel the swap pressure during every native rebuild.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;The "Sync" Trap:&lt;/strong&gt; TurboModules allow synchronous calls to native. This is tempting for getting constants or simple state. However, if your native method takes more than 5ms (e.g., a quick disk read), you will drop frames on the UI thread because you’ve blocked the JS thread which is now coupled to the native execution.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  At your level
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting out
&lt;/h3&gt;

&lt;p&gt;Do not start your project with TurboModules unless you have a specific performance requirement (like high-frequency data). Stick to the legacy bridge; it is better documented and has more StackOverflow coverage. Focus on learning how to pass data across the boundary before trying to optimize the boundary itself.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working engineer
&lt;/h3&gt;

&lt;p&gt;When you encounter a "Bridge Busy" warning in your logs, that is your signal to migrate &lt;em&gt;that specific module&lt;/em&gt; only. You can run TurboModules and legacy modules side-by-side. Move your heaviest data-producers first—camera buffers, sensor streams, or large file processing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Senior or staff
&lt;/h3&gt;

&lt;p&gt;You own the build pipeline. Before enabling TurboModules, benchmark your CI. You may need to upgrade your GitHub Actions runners to &lt;code&gt;macos-13-xlarge&lt;/code&gt; to offset the C++ compilation time. Establish a strict linting rule for &lt;code&gt;Native[Module].ts&lt;/code&gt; files to prevent developers from importing complex types that break the &lt;code&gt;codegen&lt;/code&gt; parser.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead or director
&lt;/h3&gt;

&lt;p&gt;Understand that moving to the New Architecture is a hiring decision. You are moving away from "JavaScript developers who do a bit of mobile" toward "Systems engineers." Your team will need to be comfortable debugging C++ stack traces and understanding memory management in a hybrid environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "Why would you choose TurboModules over the standard React Native bridge?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Weak Answer:&lt;/strong&gt; "Because it's faster and it's the new way React Native works." This is weak because it ignores the implementation cost and assumes "new" equals "better."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Strong Answer:&lt;/strong&gt; A strong answer focuses on the &lt;strong&gt;JSI (JavaScript Interface)&lt;/strong&gt;. You should explain that TurboModules allow for synchronous execution and eliminate the JSON serialization overhead. You must mention the trade-off: the increased complexity of the &lt;code&gt;codegen&lt;/code&gt; process and the requirement for C++ glue code. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Senior-level Follow-up:&lt;/strong&gt; "How does the migration affect your CI/CD and developer velocity?" &lt;br&gt;
A candidate who has actually done this will talk about the increase in build times and the brittleness of the &lt;code&gt;codegen&lt;/code&gt; parser. They will mention that while runtime performance improves, the feedback loop for native development slows down due to the re-compilation of generated C++ headers. They might also discuss the memory management implications of holding long-lived C++ objects via JSI versus the fire-and-forget nature of the legacy bridge.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=turbomodules-and-codegen-the-trade-offs-nobody-documents" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>turbomodules</category>
      <category>jsi</category>
      <category>architecture</category>
    </item>
    <item>
      <title>The DevOps bootcamp · 2. The Linux you actually need</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Thu, 08 Oct 2026 04:30:08 +0000</pubDate>
      <link>https://dev.to/techamit95ch/the-devops-bootcamp-2-the-linux-you-actually-need-32bd</link>
      <guid>https://dev.to/techamit95ch/the-devops-bootcamp-2-the-linux-you-actually-need-32bd</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The DevOps bootcamp&lt;/strong&gt; · Chapter 2 of 14 · DevOps · new chapter every Thursday morning&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;By the end of this chapter:&lt;/strong&gt; Work confidently with processes, files, permissions and networking.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Say you are a junior engineer on your first backend project. You need to test a branch on the team's shared staging server. You SSH in, pull your code, and run the start command. The terminal spits back &lt;code&gt;Error: listen EADDRINUSE: address already in use :::8080&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;You try changing the port, but the proxy routing traffic to your application expects 8080. You assume the previous version of the application is still running, but you do not know how to find it. You try running your new application as the superuser to force it through, and it starts, but immediately crashes with &lt;code&gt;Permission denied: /var/log/staging-api/error.log&lt;/code&gt;. &lt;/p&gt;

&lt;p&gt;Because you do not know how to ask Linux what is holding that port, or how to read the permissions on that file, you do the only thing you know: you reboot the staging server. It takes ten minutes to come back. When it does, your application works, but you have just taken down three other services your team was actively testing, costing four engineers an hour of lost work and earning a sharp message from your tech lead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you start
&lt;/h2&gt;

&lt;p&gt;You need a Linux environment. Ubuntu 24.04 LTS is the current standard for server deployments. You can use Windows Subsystem for Linux (WSL2), a local virtual machine, or a cloud instance.&lt;/p&gt;

&lt;p&gt;Verify your environment by checking the operating system release:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /etc/os-release
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output must include &lt;code&gt;PRETTY_NAME="Ubuntu 24.04 LTS"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Verify you have standard utilities installed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output should be &lt;code&gt;Python 3.12.x&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you find what is holding a port?
&lt;/h2&gt;

&lt;p&gt;When an application claims a network port, the Linux kernel prevents any other process from claiming it. If a deployment script fails, if an application crashes poorly, or if a process is sent to the background and forgotten, the process remains alive and the port remains locked. &lt;/p&gt;

&lt;p&gt;You cannot ask the port to release the process. You must ask the kernel which process ID (PID) owns the port, and then terminate that PID.&lt;/p&gt;

&lt;p&gt;The tool for querying network sockets is &lt;code&gt;ss&lt;/code&gt; (socket statistics). By asking it to list listening TCP ports and the processes attached to them, you get the exact PID blocking your deployment. You must run this command as the superuser (&lt;code&gt;sudo&lt;/code&gt;); if you run it as a standard user, the kernel will hide the PIDs of processes owned by other users, leaving the process column blank and you guessing.&lt;/p&gt;

&lt;p&gt;Once you have the PID, you use the &lt;code&gt;kill&lt;/code&gt; command. Despite the name, &lt;code&gt;kill&lt;/code&gt; does not kill processes; it sends signals to them. A standard &lt;code&gt;kill &amp;lt;PID&amp;gt;&lt;/code&gt; sends &lt;code&gt;SIGTERM&lt;/code&gt; (signal 15), which politely asks the application to shut down. A well-written application will catch this, finish writing its logs, release its port, and exit. A hung application will ignore it. When that happens, you use &lt;code&gt;kill -9 &amp;lt;PID&amp;gt;&lt;/code&gt;, which sends &lt;code&gt;SIGKILL&lt;/code&gt;. This signal goes directly to the kernel, which instantly destroys the process without giving the application a chance to react.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does root own the log file?
&lt;/h2&gt;

&lt;p&gt;Linux treats almost everything as a file, and every file has an owner, a group, and a strict set of permissions. Permissions are divided into read (&lt;code&gt;r&lt;/code&gt;), write (&lt;code&gt;w&lt;/code&gt;), and execute (&lt;code&gt;x&lt;/code&gt;). They are applied to three categories: the user who owns the file, the group assigned to the file, and everyone else.&lt;/p&gt;

&lt;p&gt;When you run an application with &lt;code&gt;sudo&lt;/code&gt;, it executes as &lt;code&gt;root&lt;/code&gt; (user ID 0). Any file it creates is owned by &lt;code&gt;root&lt;/code&gt;. If you realise your mistake and later try to run that same application as a standard user, the application will attempt to append to its existing log file, discover it does not have write permissions to a root-owned file, and crash.&lt;/p&gt;

&lt;p&gt;You view these permissions using &lt;code&gt;ls -l&lt;/code&gt;. The output begins with a ten-character string, such as &lt;code&gt;-rw-r--r--&lt;/code&gt;. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The first character is the type (&lt;code&gt;-&lt;/code&gt; for a file, &lt;code&gt;d&lt;/code&gt; for a directory).&lt;/li&gt;
&lt;li&gt;The next three are the owner's permissions (&lt;code&gt;rw-&lt;/code&gt; means read and write, but not execute).&lt;/li&gt;
&lt;li&gt;The next three are the group's permissions (&lt;code&gt;r--&lt;/code&gt; means read only).&lt;/li&gt;
&lt;li&gt;The final three are everyone else's permissions (&lt;code&gt;r--&lt;/code&gt; means read only).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You change who owns a file with &lt;code&gt;chown&lt;/code&gt; (change owner). You change the permissions themselves with &lt;code&gt;chmod&lt;/code&gt; (change mode). While you can use letters to change permissions, engineers use octal numbers because they are exact and overwrite the previous state entirely. Read is worth 4, write is worth 2, and execute is worth 1. You add them together for each category. &lt;/p&gt;

&lt;p&gt;A permission of &lt;code&gt;755&lt;/code&gt; means the owner gets 7 (4+2+1: read, write, execute), the group gets 5 (4+1: read, execute), and everyone else gets 5. A permission of &lt;code&gt;644&lt;/code&gt; means the owner gets 6 (4+2: read, write), and everyone else gets 4 (read only).&lt;/p&gt;

&lt;h2&gt;
  
  
  Build it
&lt;/h2&gt;

&lt;p&gt;We will recreate the staging server failure on your machine, diagnose the port collision, kill the ghost process, and fix the resulting permission trap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Create the ghost process&lt;/strong&gt;&lt;br&gt;
Start a dummy web server in the background, bound to port 8080. The &lt;code&gt;&amp;amp;&lt;/code&gt; at the end tells the shell to run this in the background, returning control of the terminal to you.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; http.server 8080 &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null 2&amp;gt;&amp;amp;1 &amp;amp;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; This simulates the previous version of your application that failed to shut down.&lt;br&gt;
&lt;em&gt;How to confirm:&lt;/em&gt; Run &lt;code&gt;curl http://localhost:8080&lt;/code&gt;. It will output a directory listing of your current folder.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2: Identify the blockage&lt;/strong&gt;&lt;br&gt;
Query the kernel for all listening TCP ports and the processes holding them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;ss &lt;span class="nt"&gt;-lptn&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; :8080
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; &lt;code&gt;l&lt;/code&gt; shows listening sockets, &lt;code&gt;p&lt;/code&gt; shows processes, &lt;code&gt;t&lt;/code&gt; shows TCP only, and &lt;code&gt;n&lt;/code&gt; forces numeric output (preventing Linux from trying to resolve 8080 to a service name, which is slow). We pipe (&lt;code&gt;|&lt;/code&gt;) the output to &lt;code&gt;grep&lt;/code&gt; to filter for our specific port.&lt;br&gt;
&lt;em&gt;How to confirm:&lt;/em&gt; You will see output resembling:&lt;br&gt;
&lt;code&gt;LISTEN 0 5 0.0.0.0:8080 0.0.0.0:* users:(("python3",pid=14235,fd=3))&lt;/code&gt;&lt;br&gt;
The number after &lt;code&gt;pid=&lt;/code&gt; (in this case, 14235) is what you need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3: Terminate the ghost process&lt;/strong&gt;&lt;br&gt;
Send a kill signal to the PID you found in Step 2. Replace &lt;code&gt;14235&lt;/code&gt; with your actual PID.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;kill&lt;/span&gt; &lt;span class="nt"&gt;-9&lt;/span&gt; 14235
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; We use &lt;code&gt;-9&lt;/code&gt; to force the kernel to destroy the process immediately, simulating clearing a hung application.&lt;br&gt;
&lt;em&gt;How to confirm:&lt;/em&gt; Run &lt;code&gt;sudo ss -lptn | grep :8080&lt;/code&gt; again. The output will be completely empty. The port is free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4: Create the permissions trap&lt;/strong&gt;&lt;br&gt;
Create a log directory and file as the root user.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; /var/log/staging-api
&lt;span class="nb"&gt;sudo touch&lt;/span&gt; /var/log/staging-api/error.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; This simulates the moment you tried to force your application to start using &lt;code&gt;sudo&lt;/code&gt;, leaving behind root-owned state.&lt;br&gt;
&lt;em&gt;How to confirm:&lt;/em&gt; Run &lt;code&gt;ls -l /var/log/staging-api/error.log&lt;/code&gt;. The output will show &lt;code&gt;root root&lt;/code&gt; as the owner and group.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5: Trigger the failure&lt;/strong&gt;&lt;br&gt;
Attempt to write to the log file as your standard user.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Starting application..."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; /var/log/staging-api/error.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; This simulates your application trying to start up and write its first log line.&lt;br&gt;
&lt;em&gt;How to confirm:&lt;/em&gt; The terminal will output exactly: &lt;code&gt;bash: /var/log/staging-api/error.log: Permission denied&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 6: Fix the ownership and run successfully&lt;/strong&gt;&lt;br&gt;
Change the ownership of the directory and its contents to your current user. &lt;code&gt;$USER&lt;/code&gt; is an environment variable containing your username.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo chown&lt;/span&gt; &lt;span class="nt"&gt;-R&lt;/span&gt; &lt;span class="nv"&gt;$USER&lt;/span&gt;:&lt;span class="nv"&gt;$USER&lt;/span&gt; /var/log/staging-api
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; The &lt;code&gt;-R&lt;/code&gt; flag applies the change recursively to the directory and everything inside it. The &lt;code&gt;$USER:$USER&lt;/code&gt; syntax sets both the owner and the group to your user.&lt;br&gt;
&lt;em&gt;How to confirm:&lt;/em&gt; Run the application write command again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Starting application..."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; /var/log/staging-api/error.log
&lt;span class="nb"&gt;cat&lt;/span&gt; /var/log/staging-api/error.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output will successfully print &lt;code&gt;Starting application...&lt;/code&gt;. You have cleared the port, fixed the permissions, and started the service without rebooting the machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this breaks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; &lt;code&gt;ss: command not found&lt;/code&gt;&lt;br&gt;
&lt;strong&gt;Cause:&lt;/strong&gt; You are running a minimal Linux distribution, often inside a container, that has stripped out network utilities to save space.&lt;br&gt;
&lt;strong&gt;Fix:&lt;/strong&gt; Install the package that provides &lt;code&gt;ss&lt;/code&gt;. On Ubuntu/Debian, run &lt;code&gt;sudo apt-get update &amp;amp;&amp;amp; sudo apt-get install iproute2&lt;/code&gt;. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; &lt;code&gt;kill -9&lt;/code&gt; returns silently, but the process still appears in &lt;code&gt;ss&lt;/code&gt; or &lt;code&gt;ps&lt;/code&gt;.&lt;br&gt;
&lt;strong&gt;Cause:&lt;/strong&gt; The process is a zombie (listed as &lt;code&gt;Z&lt;/code&gt; state in &lt;code&gt;ps&lt;/code&gt;) or in uninterruptible sleep (listed as &lt;code&gt;D&lt;/code&gt; state). A zombie is already dead, but its parent process has not acknowledged its death. Uninterruptible sleep means the process is waiting on hardware, like a hung network drive, and the kernel will not deliver the kill signal until the hardware responds.&lt;br&gt;
&lt;strong&gt;Fix:&lt;/strong&gt; For a zombie, you must find and kill its parent. Find the parent PID (PPID) with &lt;code&gt;ps -o ppid= -p &amp;lt;PID&amp;gt;&lt;/code&gt;, then kill the parent. For a process in uninterruptible sleep waiting on dead hardware, you cannot kill it. You must reboot the machine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Symptom:&lt;/strong&gt; &lt;code&gt;bash: /var/log/staging-api/error.log: Permission denied&lt;/code&gt; even when you run &lt;code&gt;sudo echo "test" &amp;gt;&amp;gt; /var/log/staging-api/error.log&lt;/code&gt;.&lt;br&gt;
&lt;strong&gt;Cause:&lt;/strong&gt; The shell parses the command line before executing it. It sees the redirect (&lt;code&gt;&amp;gt;&amp;gt;&lt;/code&gt;) and attempts to open the file using your standard user permissions &lt;em&gt;before&lt;/em&gt; it runs &lt;code&gt;sudo echo&lt;/code&gt;. The &lt;code&gt;sudo&lt;/code&gt; only applies to the &lt;code&gt;echo&lt;/code&gt; command, not the file redirection.&lt;br&gt;
&lt;strong&gt;Fix:&lt;/strong&gt; Pass the string to the &lt;code&gt;tee&lt;/code&gt; command, running &lt;code&gt;tee&lt;/code&gt; as root.&lt;br&gt;
&lt;code&gt;echo "test" | sudo tee -a /var/log/staging-api/error.log&lt;/code&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Manual Linux administration is imperative: you are issuing commands that change the state of the machine. The cost of this approach is configuration drift. &lt;/p&gt;

&lt;p&gt;If you fix a permissions error on a staging server by running &lt;code&gt;chown&lt;/code&gt;, you have fixed it once, on one machine. The next time the server is rebuilt, or when you attempt to deploy to production, the error will return. The server's true configuration now exists only in the command history of your terminal, not in source control. &lt;/p&gt;

&lt;p&gt;This commits you to remembering what you typed. It scales poorly, guarantees failures during high-stress deployments, and turns servers into fragile pets that engineers are afraid to replace. We learn these commands so we can diagnose failures and understand what the operating system requires, but we do not use them to deploy software. &lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;A standard infrastructure interview question is: "You are deploying a new version of a backend service to a Linux machine, and it fails to start. Walk me through how you debug it."&lt;/p&gt;

&lt;p&gt;A weak answer is: "I would check the application logs, and if that doesn't work, I would restart the server." This is weak because it assumes the application lived long enough to write logs (which it cannot do if the port is blocked or the log directory is restricted), and it treats restarting a production server as a valid diagnostic step.&lt;/p&gt;

&lt;p&gt;A strong answer names the specific sequence of investigation. You check process status with &lt;code&gt;ps&lt;/code&gt; or &lt;code&gt;systemctl&lt;/code&gt;, check port bindings with &lt;code&gt;ss -lptn&lt;/code&gt;, check file permissions with &lt;code&gt;ls -la&lt;/code&gt;, and read the application's standard error output directly. &lt;/p&gt;

&lt;p&gt;The follow-up questions will change based on the level you are interviewing for. A junior candidate is expected to know the commands (&lt;code&gt;ss&lt;/code&gt;, &lt;code&gt;kill&lt;/code&gt;, &lt;code&gt;chmod&lt;/code&gt;) and what they do. A senior candidate is expected to explain the lifecycle: &lt;em&gt;why&lt;/em&gt; the old version did not release the port (perhaps the deployment script failed to send a &lt;code&gt;SIGTERM&lt;/code&gt;, or the application does not handle signals gracefully) and &lt;em&gt;why&lt;/em&gt; the permissions are wrong. An engineering manager or staff-level loop will probe the systemic failure: why is a human SSHing into a machine to deploy in the first place, and how do we move this process to an automated, immutable pipeline?&lt;/p&gt;

&lt;h2&gt;
  
  
  Your tasks
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Find the SSH daemon.&lt;/strong&gt; Find the PID of the process listening on port 22. 
&lt;em&gt;Done looks like:&lt;/em&gt; You have a number. Running &lt;code&gt;ps -p &amp;lt;number&amp;gt;&lt;/code&gt; outputs a line ending in &lt;code&gt;sshd&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create a read-only secret.&lt;/strong&gt; Create a file named &lt;code&gt;secret.txt&lt;/code&gt;. Change its permissions so that only your user can read it, and absolutely nobody (including you) can write to it or execute it. 
&lt;em&gt;Done looks like:&lt;/em&gt; Running &lt;code&gt;ls -l secret.txt&lt;/code&gt; shows &lt;code&gt;-r--------&lt;/code&gt;. Running &lt;code&gt;echo "test" &amp;gt; secret.txt&lt;/code&gt; fails with &lt;code&gt;Permission denied&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hunt a disconnected process.&lt;/strong&gt; Start a background process by running &lt;code&gt;sleep 3600 &amp;amp;&lt;/code&gt;. Close your terminal completely to log out of the shell. Open a new terminal, log back in, find that specific sleep process, and kill it. 
&lt;em&gt;Done looks like:&lt;/em&gt; Running &lt;code&gt;ps aux | grep sleep&lt;/code&gt; shows nothing after you have killed it.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Your tasks this week
&lt;/h2&gt;

&lt;p&gt;Do the exercises above before the next chapter. Reading a tutorial and doing&lt;br&gt;
one are different activities and only one of them changes what you can build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stuck on any of them?&lt;/strong&gt; Say so — describe what you tried and what happened:&lt;br&gt;
&lt;a href="https://www.amitchakraborty.dev/learn/devops-bootcamp#stuck" rel="noopener noreferrer"&gt;tell me where you got stuck&lt;/a&gt;. I read every one, and the questions&lt;br&gt;
that come back more than twice get answered in the next chapter.&lt;/p&gt;

&lt;h2&gt;
  
  
  The DevOps bootcamp
&lt;/h2&gt;

&lt;p&gt;Chapter 2 of 14. New chapter every Thursday morning.&lt;br&gt;
Next: &lt;strong&gt;The shell: pipes, scripts and automating your own machine&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;· &lt;a href="https://www.amitchakraborty.dev/learn/devops-bootcamp" rel="noopener noreferrer"&gt;The full syllabus and every chapter so far&lt;/a&gt;&lt;br&gt;
· Subscribers also get the condensed notes for this chapter, the running&lt;br&gt;
  recap of everything the series has covered, and the extended guidance:&lt;br&gt;
  &lt;a href="https://www.amitchakraborty.dev/#newsletter" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by Amit Chakraborty — founding engineer and senior architect: React Native, AI and RAG systems, production architecture. &lt;a href="https://www.amitchakraborty.dev?utm_source=curriculum&amp;amp;utm_medium=content&amp;amp;utm_campaign=devops-bootcamp" rel="noopener noreferrer"&gt;Portfolio&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Need this built, reviewed or taught to your team? &lt;a href="https://www.amitchakraborty.dev/#contact" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt; or email &lt;a href="mailto:amit@devamit.co.in"&gt;amit@devamit.co.in&lt;/a&gt;. Available for senior and founding engineering roles, consulting and training, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>devops</category>
      <category>thedevopsbootcamp</category>
      <category>howdoyoufindwhatisholdingaport</category>
    </item>
    <item>
      <title>The Silent Deadlock: Debugging Fabric Renderer Stalls in Production</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Thu, 08 Oct 2026 04:30:03 +0000</pubDate>
      <link>https://dev.to/techamit95ch/the-silent-deadlock-debugging-fabric-renderer-stalls-in-production-3n5</link>
      <guid>https://dev.to/techamit95ch/the-silent-deadlock-debugging-fabric-renderer-stalls-in-production-3n5</guid>
      <description>&lt;p&gt;I was serving as the founding engineer at a HealthTech AI platform, Synapsis Medical Technologies, where we were managing a stack that integrated real-time wearable data with a HIPAA-aligned RAG pipeline. We had just migrated our core mobile interface to React Native 0.76 to leverage the New Architecture. On paper, the Fabric renderer promised synchronous layout and better interop with our native C++ modules. &lt;/p&gt;

&lt;p&gt;Then the clinical feedback started hitting my desk. &lt;/p&gt;

&lt;p&gt;Physicians using the app reported that the UI would "freeze" for three to five seconds when opening a patient’s longitudinal record—a view heavy with complex SVG charts and streaming LLM responses. Our crash reporting (Sentry) showed zero crashes. Our JavaScript thread was idle. But the screen was unresponsive to touch. We were losing clinical trust, and in a health-tech environment, a five-second UI stall isn't just a bug; it's a failure of the tool.&lt;/p&gt;

&lt;p&gt;I spent 72 hours tracing this to the way Fabric handles the "Shadow Tree" during high-frequency state updates. This is the reality of the New Architecture: you trade the asynchronous bridge bottleneck for a synchronous thread-locking problem that is significantly harder to debug.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Fabric Stalls Under Load
&lt;/h2&gt;

&lt;p&gt;In the old architecture, the Yoga layout engine ran on a background "shadow thread." If the layout took too long, the UI thread remained free to scroll or show a loading spinner. In the Fabric renderer, the layout calculation can be triggered synchronously on the UI thread to prevent the "white flash" of unstyled content.&lt;/p&gt;

&lt;p&gt;The failure happens when you combine three factors:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;High-frequency updates:&lt;/strong&gt; A streaming LLM response or a 60Hz wearable data feed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deep Component Trees:&lt;/strong&gt; A complex FHIR-compliant medical record with nested lists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synchronous Triggers:&lt;/strong&gt; Using &lt;code&gt;useLayoutEffect&lt;/code&gt; or certain native gestures that force Fabric to commit a new Shadow Tree immediately.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When these collide, the UI thread enters a state of "Commit Contention." The C++ Shadow Tree is being mutated so rapidly that the UI thread spends its entire budget calculating layout, leaving no time to process the touch events sitting in the queue. Because the JS thread is actually finished with its work, your standard performance monitors will tell you the app is healthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Detecting and Fixing Shadow Tree Contention
&lt;/h2&gt;

&lt;p&gt;You cannot fix this by optimizing JavaScript. You have to change how the renderer commits changes to the native side.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Identify the "Ghost Stall"
&lt;/h3&gt;

&lt;p&gt;First, confirm it is a Fabric stall and not a JS event loop block. Run the app with the &lt;code&gt;Perf Monitor&lt;/code&gt; visible. If the JS FPS is 60 but the UI FPS drops to 0-5 during an interaction, you are looking at a Fabric commit bottleneck.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Batching with &lt;code&gt;useTransition&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;In React Native 0.76+, you must stop treating all state updates as equal. For our clinical AI stream, I moved the text updates into a transition. This tells Fabric that the update is low priority and can be interrupted by a touch event.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;useTransition&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;PatientRecord&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setData&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;isPending&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;startTransition&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useTransition&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="c1"&gt;// WRONG: This forces a synchronous Fabric commit on every chunk&lt;/span&gt;
  &lt;span class="c1"&gt;// onChunk(chunk =&amp;gt; setData(prev =&amp;gt; prev + chunk));&lt;/span&gt;

  &lt;span class="c1"&gt;// RIGHT: This allows Fabric to prioritize UI responsiveness over the text update&lt;/span&gt;
  &lt;span class="nf"&gt;onChunk&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;startTransition&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nf"&gt;setData&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Audit Native Component Synchronicity
&lt;/h3&gt;

&lt;p&gt;If you use &lt;code&gt;react-native-svg&lt;/code&gt; or custom Fabric components, check if they are calling &lt;code&gt;invalidate&lt;/code&gt; on the main thread too frequently. In our case, the wearable heart-rate graph was forcing a layout pass on every data point. &lt;/p&gt;

&lt;p&gt;To confirm this, I used the &lt;strong&gt;ETW (Event Tracing for Windows)&lt;/strong&gt; equivalent on iOS: &lt;strong&gt;Instruments (System Trace)&lt;/strong&gt;. Look for &lt;code&gt;common_mount&lt;/code&gt; and &lt;code&gt;layout&lt;/code&gt; calls. If you see a wall of &lt;code&gt;[Fabric] commit&lt;/code&gt; blocks longer than 16ms, you must wrap the data source in a throttle.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Implement a "Commit Guard"
&lt;/h3&gt;

&lt;p&gt;In my experience shipping 18+ production apps, the most effective way to handle this under load is to decouple the data frequency from the render frequency. Even with Fabric, the human eye cannot perceive updates faster than 30-60fps.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// A simple throttle for high-frequency clinical data&lt;/span&gt;
&lt;span class="nf"&gt;useEffect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;interval&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nf"&gt;setRenderData&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nx"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// Target ~30fps for heavy UI updates&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;interval&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Cost of the New Architecture
&lt;/h2&gt;

&lt;p&gt;Migrating to Fabric is not a free performance upgrade. In our CI/CD overhaul where I cut release cycles from 2 days to 4 hours, we had to add specific automated "jank tests" because Fabric failures are silent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Trade-off:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Memory:&lt;/strong&gt; Fabric keeps the Shadow Tree in C++. While this is faster, I observed a ~15-20% increase in baseline memory usage for complex views compared to the old architecture.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Initialization:&lt;/strong&gt; The "New Architecture" has a heavier startup cost. On low-end Android devices, we saw the &lt;code&gt;onCreate&lt;/code&gt; to &lt;code&gt;onContentAppear&lt;/code&gt; time increase by roughly 200ms because of the C++ TurboModule initialization.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not enable the New Architecture if your app is a simple CRUD tool with flat lists. The complexity of debugging the C++ layer outweighs the benefits unless you are doing high-frequency data visualization or complex animations that require synchronous layout.&lt;/p&gt;

&lt;h2&gt;
  
  
  At Your Level
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting out
&lt;/h3&gt;

&lt;p&gt;Stop using &lt;code&gt;useLayoutEffect&lt;/code&gt; unless you specifically need to measure a DOM/Shadow node before the user sees it. In Fabric, this hook is a common source of UI-blocking stalls. Stick to &lt;code&gt;useEffect&lt;/code&gt; to keep the UI thread free.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working engineer
&lt;/h3&gt;

&lt;p&gt;Learn to use the &lt;strong&gt;React DevTools Profiler&lt;/strong&gt; specifically to look for "Commit" times. If your "Commit" phase is consistently over 10ms, your component tree is too deep for Fabric to handle efficiently. Flatten your hierarchy—use &lt;code&gt;Fragment&lt;/code&gt; instead of &lt;code&gt;View&lt;/code&gt; where possible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Senior or staff
&lt;/h3&gt;

&lt;p&gt;You own the architecture. You must implement a strategy for "Concurrent Metadata." This means deciding which data streams (like our AI RAG pipeline) are allowed to block the UI and which must be deferred. Establish a lint rule or a wrapper around &lt;code&gt;setState&lt;/code&gt; for high-frequency events.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead or director
&lt;/h3&gt;

&lt;p&gt;Understand that moving to the New Architecture increases your "debugging tax." You will need engineers who understand the C++ lifecycle of React Native, not just React. Budget time for a comprehensive regression of your most complex screens, as the performance characteristics will flip—what was fast might become slow due to lock contention.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the Interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "How does the Fabric renderer improve performance over the old architecture, and what are the risks?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Weak Answer:&lt;/strong&gt; "It uses a C++ core and removes the bridge, making everything faster and smoother." This is a marketing answer that ignores the implementation reality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Strong Answer:&lt;/strong&gt; A senior candidate will discuss &lt;strong&gt;synchronous layout capabilities&lt;/strong&gt; and the &lt;strong&gt;removal of the three-tree problem&lt;/strong&gt; (JS, Shadow, and Native trees). They should name the specific trade-off: while it eliminates the "async jump" (where the UI and JS threads are out of sync), it introduces the risk of &lt;strong&gt;UI thread starvation&lt;/strong&gt;. If the JS thread pushes too many synchronous commits, the UI thread cannot process gestures, leading to a "frozen" app that hasn't actually crashed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Follow-up:&lt;/strong&gt; "How do you debug a UI freeze where the JS thread is idle?"&lt;br&gt;
The candidate should mention using &lt;strong&gt;System Trace (Instruments)&lt;/strong&gt; or &lt;strong&gt;systrace&lt;/strong&gt; to look for C++ mutex contention or long &lt;code&gt;Mounting&lt;/code&gt; phases in the Fabric renderer. If they only mention &lt;code&gt;console.log&lt;/code&gt; or the JS profiler, they haven't lived with the New Architecture under load.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=fabric-renderer-internals-what-actually-breaks-in-production" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>fabric</category>
      <category>architecture</category>
      <category>newarchitecture</category>
    </item>
    <item>
      <title>FastAPI in production · 1. FastAPI in one hour</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Wed, 07 Oct 2026 16:30:07 +0000</pubDate>
      <link>https://dev.to/techamit95ch/fastapi-in-production-1-fastapi-in-one-hour-529o</link>
      <guid>https://dev.to/techamit95ch/fastapi-in-production-1-fastapi-in-one-hour-529o</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;FastAPI in production&lt;/strong&gt; · Chapter 1 of 12 · Backend · new chapter every Wednesday night&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;By the end of this chapter:&lt;/strong&gt; Get a typed, documented API running and understand what generated the documentation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Say you are an engineer tasked with standing up a new user-preferences service. The frontend team needs it by Friday. You write a quick Flask app, define a &lt;code&gt;/preferences&lt;/code&gt; endpoint, and write up a wiki page explaining that the endpoint expects a JSON payload with a &lt;code&gt;user_id&lt;/code&gt; (an integer) and a &lt;code&gt;theme&lt;/code&gt; (a string). &lt;/p&gt;

&lt;p&gt;On Thursday afternoon, the frontend team deploys to staging. The service immediately starts throwing 500 Internal Server Errors. You check the logs and see a &lt;code&gt;TypeError&lt;/code&gt;. The frontend sent &lt;code&gt;{"user_id": "12345", "theme": "dark"}&lt;/code&gt;. Because &lt;code&gt;user_id&lt;/code&gt; came in as a string, your database query failed. You patch the code to cast the ID to an integer, deploy the hotfix, and tell the frontend team to try again. Ten minutes later, they complain that the API is returning a 200 OK, but the response shape does not match the wiki page you wrote. The wiki is out of date, the validation is manual, and you are spending your afternoon acting as a human type-checker.&lt;/p&gt;

&lt;p&gt;This is the exact failure mode FastAPI was built to eliminate. It ties the shape of your data, the validation of incoming requests, and the documentation of your API to a single source of truth: Python type hints.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you start
&lt;/h2&gt;

&lt;p&gt;You need Python 3.10 or newer. We will use FastAPI 0.112.2, which introduced the &lt;code&gt;[standard]&lt;/code&gt; extra to bundle the framework with its recommended server, Uvicorn.&lt;/p&gt;

&lt;p&gt;Create a new directory, set up a virtual environment, and install the package:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"fastapi[standard]==0.112.2"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Verify the installation by checking the version in your terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import fastapi; print(fastapi.__version__)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This must output &lt;code&gt;0.112.2&lt;/code&gt; before you continue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why documentation drifts from code
&lt;/h2&gt;

&lt;p&gt;In a traditional Python web framework, the code that handles a request and the documentation that describes it are entirely separate. You write a function that accepts a request object, you parse the JSON body, and you write a docstring or a separate YAML file detailing what that JSON should look like. &lt;/p&gt;

&lt;p&gt;When a requirement changes—say, &lt;code&gt;theme&lt;/code&gt; becomes an optional field—you must update the parsing logic, the validation logic, and the documentation. If you forget one, the system degrades. The documentation lies to the client, or the validation fails to protect the database.&lt;/p&gt;

&lt;p&gt;FastAPI prevents this by generating an OpenAPI schema directly from your function signatures. OpenAPI is a language-agnostic specification for describing REST APIs. Because FastAPI reads the type hints on your Python functions at startup, the generated OpenAPI schema is always an exact reflection of the code that will execute. If you change a type hint from &lt;code&gt;str&lt;/code&gt; to &lt;code&gt;int&lt;/code&gt;, the documentation updates automatically, and the framework immediately begins rejecting requests that cannot be parsed as integers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The role of Pydantic
&lt;/h2&gt;

&lt;p&gt;FastAPI does not do the validation itself. It delegates data parsing and validation to Pydantic, a library that enforces type hints at runtime. &lt;/p&gt;

&lt;p&gt;When a request arrives, FastAPI inspects the signature of your route function. If it sees a Pydantic model in the signature, it takes the incoming JSON payload, passes it to Pydantic, and attempts to construct that model. If the payload is valid, your function receives a fully instantiated Python object. If the payload is invalid, Pydantic raises an error, which FastAPI catches and translates into a &lt;code&gt;422 Unprocessable Entity&lt;/code&gt; HTTP response. Your route function never even executes.&lt;/p&gt;

&lt;p&gt;This means you can write your business logic under the assumption that the data is perfectly formed. You do not need to check if a field is missing, or if a string was passed instead of an integer. If the code inside your route function is running, the contract was met.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build it
&lt;/h2&gt;

&lt;p&gt;We will build a single-file API that accepts user preferences, validates them, and returns a structured response. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Define the data contract&lt;/strong&gt;&lt;br&gt;
Create a file named &lt;code&gt;main.py&lt;/code&gt;. Import FastAPI and Pydantic, and define the shape of the data you expect the client to send.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;UserPreferences&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;
    &lt;span class="n"&gt;theme&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;notifications_enabled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; Inheriting from &lt;code&gt;BaseModel&lt;/code&gt; tells Pydantic to treat this class as a data validator. &lt;code&gt;user_id&lt;/code&gt; and &lt;code&gt;theme&lt;/code&gt; are required. &lt;code&gt;notifications_enabled&lt;/code&gt; is optional and defaults to &lt;code&gt;True&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2: Initialize the application&lt;/strong&gt;&lt;br&gt;
Add the FastAPI application instance to &lt;code&gt;main.py&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Preferences API&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1.0.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; This &lt;code&gt;app&lt;/code&gt; object is the ASGI (Asynchronous Server Gateway Interface) application. The server will look for this exact variable to route incoming network traffic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3: Write the route&lt;/strong&gt;&lt;br&gt;
Create an endpoint that accepts a POST request. Use the Pydantic model as a parameter.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/preferences&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;update_preferences&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prefs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;UserPreferences&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# If we reach this line, 'prefs' is guaranteed to be a valid UserPreferences object.
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Preferences updated successfully&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prefs&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; The &lt;code&gt;@app.post&lt;/code&gt; decorator tells FastAPI to route HTTP POST requests for &lt;code&gt;/preferences&lt;/code&gt; to this function. By typing the &lt;code&gt;prefs&lt;/code&gt; argument as &lt;code&gt;UserPreferences&lt;/code&gt;, you instruct FastAPI to read the request body, validate it against the model, and pass the result to your function.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4: Run the server&lt;/strong&gt;&lt;br&gt;
Start the application using the FastAPI CLI (which wraps Uvicorn). Run this command in your terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;fastapi dev main.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Why:&lt;/em&gt; The &lt;code&gt;dev&lt;/code&gt; command starts a local server on port 8000 and watches your files for changes, reloading the server automatically. &lt;br&gt;
&lt;em&gt;How to confirm:&lt;/em&gt; You should see output ending with &lt;code&gt;Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5: Test the validation&lt;/strong&gt;&lt;br&gt;
Open a second terminal and send a valid request using &lt;code&gt;curl&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST http://127.0.0.1:8000/preferences &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"user_id": 42, "theme": "dark"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Expected output:&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Preferences updated successfully"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"data"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"user_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"theme"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"dark"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"notifications_enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that &lt;code&gt;notifications_enabled&lt;/code&gt; was automatically injected with its default value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 6: Inspect the generated documentation&lt;/strong&gt;&lt;br&gt;
Open your web browser and navigate to &lt;code&gt;http://127.0.0.1:8000/docs&lt;/code&gt;. &lt;br&gt;
&lt;em&gt;Why:&lt;/em&gt; FastAPI has generated an interactive Swagger UI based on your code. You will see the &lt;code&gt;/preferences&lt;/code&gt; endpoint. If you click into it and look at the "Schema" section, you will see the exact JSON structure required, including the types and default values. You did not write this documentation; it was compiled from your Python type hints.&lt;/p&gt;
&lt;h2&gt;
  
  
  When this breaks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The 422 Unprocessable Entity&lt;/strong&gt;&lt;br&gt;
&lt;em&gt;Symptom:&lt;/em&gt; The client sends a request and receives a 422 HTTP status code, and your route function never executes.&lt;br&gt;
&lt;em&gt;Cause:&lt;/em&gt; The incoming data violated the Pydantic contract. For example, sending &lt;code&gt;"theme": ["dark"]&lt;/code&gt; (a list instead of a string).&lt;br&gt;
&lt;em&gt;Fix:&lt;/em&gt; Read the JSON body of the 422 response. FastAPI provides an exact path to the failure.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"detail"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string_type"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"loc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"body"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"theme"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"msg"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Input should be a valid string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"dark"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The client must fix their payload to match the schema.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blocking the event loop&lt;/strong&gt;&lt;br&gt;
&lt;em&gt;Symptom:&lt;/em&gt; Under load, the API suddenly becomes unresponsive. Requests queue up and time out.&lt;br&gt;
&lt;em&gt;Cause:&lt;/em&gt; You defined a route with &lt;code&gt;async def&lt;/code&gt; but performed a synchronous, blocking operation inside it (like a &lt;code&gt;time.sleep()&lt;/code&gt;, a synchronous database call, or a heavy CPU calculation).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/slow&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;slow_route&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# This blocks the entire server
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;done&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;FastAPI runs on a single-threaded event loop. If an &lt;code&gt;async def&lt;/code&gt; function blocks, it prevents the server from handling any other requests.&lt;br&gt;
&lt;em&gt;Fix:&lt;/em&gt; If you are using synchronous libraries, define your route with a standard &lt;code&gt;def&lt;/code&gt; instead of &lt;code&gt;async def&lt;/code&gt;. FastAPI will automatically run standard &lt;code&gt;def&lt;/code&gt; functions in a separate threadpool, keeping the main event loop free to accept new requests. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Address already in use&lt;/strong&gt;&lt;br&gt;
&lt;em&gt;Symptom:&lt;/em&gt; When running &lt;code&gt;fastapi dev main.py&lt;/code&gt;, the terminal outputs &lt;code&gt;[Errno 98] Address already in use&lt;/code&gt; or &lt;code&gt;[Errno 48] Address already in use&lt;/code&gt;.&lt;br&gt;
&lt;em&gt;Cause:&lt;/em&gt; Another process (often a previous instance of your FastAPI app that did not shut down cleanly) is already listening on port 8000.&lt;br&gt;
&lt;em&gt;Fix:&lt;/em&gt; Find the process using the port and kill it. On Linux or macOS: &lt;code&gt;lsof -i :8000&lt;/code&gt;, then &lt;code&gt;kill -9 &amp;lt;PID&amp;gt;&lt;/code&gt;. Alternatively, start FastAPI on a different port: &lt;code&gt;fastapi dev main.py --port 8080&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;The primary cost of FastAPI's approach is strictness. In a Flask app, if you want to accept an arbitrary, deeply nested JSON blob and just store it in a database without inspecting it, you simply call &lt;code&gt;request.get_json()&lt;/code&gt; and pass the resulting dictionary along. &lt;/p&gt;

&lt;p&gt;In FastAPI, to get the documentation and validation benefits, you must define the exact shape of that blob using Pydantic models. If the upstream service sending you data frequently adds random fields, and you use a strict Pydantic model, those requests will either be rejected or the extra fields will be silently stripped out, depending on your configuration. &lt;/p&gt;

&lt;p&gt;Furthermore, Pydantic validation is not free. While Pydantic V2 (written in Rust) is exceptionally fast, parsing and validating a large JSON payload into Python objects takes more CPU cycles than Python's native &lt;code&gt;json.loads()&lt;/code&gt;. For 99% of web applications, this overhead is invisible. But if you are building a high-throughput ingestion pipeline processing tens of thousands of large payloads per second, the validation step will become a bottleneck.&lt;/p&gt;

&lt;p&gt;Finally, you are buying into an ecosystem. Your business logic becomes tightly coupled to Pydantic models and FastAPI's dependency injection system. Migrating a complex FastAPI codebase to another framework later requires significant rewriting, as the framework's concepts bleed deeply into your route handlers.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the interview
&lt;/h2&gt;

&lt;p&gt;When interviewing for a backend role, a common question is: "Why would you choose FastAPI over Flask or Django for a new microservice?"&lt;/p&gt;

&lt;p&gt;A weak answer focuses purely on speed: "FastAPI is faster because it is asynchronous." This is weak because raw framework speed rarely dictates the performance of a real-world application; database queries and network calls do. Furthermore, it shows a surface-level understanding of the tool's actual value proposition.&lt;/p&gt;

&lt;p&gt;A strong answer centers on developer velocity and contract enforcement. "I would choose FastAPI because it eliminates the drift between the API implementation and its documentation. By using Pydantic models as type hints, the framework guarantees that my route handlers only execute if the incoming payload strictly matches the contract. This prevents an entire class of type-casting bugs and offloads manual validation boilerplate."&lt;/p&gt;

&lt;p&gt;If you are a junior candidate, the interviewer expects you to explain how the OpenAPI docs are generated and how a 422 error is produced. If you are a senior candidate, the interviewer will likely probe the trade-offs: they will ask you when you would &lt;em&gt;not&lt;/em&gt; use &lt;code&gt;async def&lt;/code&gt; for a route, or how you handle sharing Pydantic models across different services without creating a distributed monolith. An engineering manager will want to hear about onboarding: how the auto-generated Swagger UI allows frontend teams to unblock themselves without waiting for backend engineers to write documentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your tasks
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Add a GET endpoint with a query parameter.&lt;/strong&gt; &lt;br&gt;
Create an endpoint at &lt;code&gt;@app.get("/users")&lt;/code&gt;. Have it accept an integer query parameter called &lt;code&gt;limit&lt;/code&gt; with a default value of 10. Return a JSON object echoing the limit. Open the &lt;code&gt;/docs&lt;/code&gt; UI and verify that &lt;code&gt;limit&lt;/code&gt; appears as a query parameter, not a request body.&lt;br&gt;
&lt;em&gt;Done when:&lt;/em&gt; &lt;code&gt;curl "http://127.0.0.1:8000/users?limit=5"&lt;/code&gt; returns &lt;code&gt;{"limit": 5}&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Trigger a validation failure.&lt;/strong&gt;&lt;br&gt;
Using the POST &lt;code&gt;/preferences&lt;/code&gt; endpoint from the chapter, send a &lt;code&gt;curl&lt;/code&gt; request where the &lt;code&gt;user_id&lt;/code&gt; is a string that cannot be cast to an integer (e.g., &lt;code&gt;"user_id": "abc"&lt;/code&gt;). &lt;br&gt;
&lt;em&gt;Done when:&lt;/em&gt; You receive a 422 status code and can identify the exact &lt;code&gt;loc&lt;/code&gt; (location) in the JSON response that points to &lt;code&gt;user_id&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Nest a Pydantic model.&lt;/strong&gt;&lt;br&gt;
Create a new Pydantic model called &lt;code&gt;Address&lt;/code&gt; containing a &lt;code&gt;city&lt;/code&gt; (string) and &lt;code&gt;country&lt;/code&gt; (string). Update the &lt;code&gt;UserPreferences&lt;/code&gt; model to include an &lt;code&gt;address&lt;/code&gt; field typed as &lt;code&gt;Address&lt;/code&gt;. &lt;br&gt;
&lt;em&gt;Done when:&lt;/em&gt; You can successfully send a POST request with a nested JSON object for the address, and the &lt;code&gt;/docs&lt;/code&gt; UI correctly displays the nested schema.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Your tasks this week
&lt;/h2&gt;

&lt;p&gt;Do the exercises above before the next chapter. Reading a tutorial and doing&lt;br&gt;
one are different activities and only one of them changes what you can build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stuck on any of them?&lt;/strong&gt; Say so — describe what you tried and what happened:&lt;br&gt;
&lt;a href="https://www.amitchakraborty.dev/learn/fastapi-production#stuck" rel="noopener noreferrer"&gt;tell me where you got stuck&lt;/a&gt;. I read every one, and the questions&lt;br&gt;
that come back more than twice get answered in the next chapter.&lt;/p&gt;

&lt;h2&gt;
  
  
  FastAPI in production
&lt;/h2&gt;

&lt;p&gt;Chapter 1 of 12. New chapter every Wednesday night.&lt;br&gt;
Next: &lt;strong&gt;Pydantic as your API contract&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;· &lt;a href="https://www.amitchakraborty.dev/learn/fastapi-production" rel="noopener noreferrer"&gt;The full syllabus and every chapter so far&lt;/a&gt;&lt;br&gt;
· Subscribers also get the condensed notes for this chapter, the running&lt;br&gt;
  recap of everything the series has covered, and the extended guidance:&lt;br&gt;
  &lt;a href="https://www.amitchakraborty.dev/#newsletter" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written by Amit Chakraborty — founding engineer and senior architect: React Native, AI and RAG systems, production architecture. &lt;a href="https://www.amitchakraborty.dev?utm_source=curriculum&amp;amp;utm_medium=content&amp;amp;utm_campaign=fastapi-production" rel="noopener noreferrer"&gt;Portfolio&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Need this built, reviewed or taught to your team? &lt;a href="https://www.amitchakraborty.dev/#contact" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt; or email &lt;a href="mailto:amit@devamit.co.in"&gt;amit@devamit.co.in&lt;/a&gt;. Available for senior and founding engineering roles, consulting and training, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>backend</category>
      <category>fastapiinproduction</category>
      <category>whydocumentationdriftsfromcode</category>
    </item>
    <item>
      <title>The Hidden Performance Ceiling of React Native Skia in Production</title>
      <dc:creator>Amit chakraborty</dc:creator>
      <pubDate>Wed, 07 Oct 2026 16:30:02 +0000</pubDate>
      <link>https://dev.to/techamit95ch/the-hidden-performance-ceiling-of-react-native-skia-in-production-2nfk</link>
      <guid>https://dev.to/techamit95ch/the-hidden-performance-ceiling-of-react-native-skia-in-production-2nfk</guid>
      <description>&lt;p&gt;I was three weeks out from a major release for a clinical HealthTech platform when we hit a wall that didn't exist in our staging environment. We were using &lt;code&gt;react-native-skia&lt;/code&gt; to render real-time physiological waveforms—ECG and SpO2 data—streaming from wearable devices. On an iPhone 15 Pro, the 60 FPS animations were fluid. On the mid-range Android tablets used in the clinics, the UI thread stayed responsive, but the Skia drawing layer began to lag three seconds behind the actual data stream.&lt;/p&gt;

&lt;p&gt;The symptom wasn't a crash or a Redbox error. It was "frame accumulation." Because Skia runs on the JavaScript thread by default (or the UI thread depending on your configuration), high-frequency updates to a &lt;code&gt;SkiaView&lt;/code&gt; can saturate the bridge or the JSI, leading to a mounting execution queue. We lost three days of development trying to reconcile why the same code that passed automated performance tests failed when subjected to the sustained, 24-hour data streams required for clinical monitoring.&lt;/p&gt;

&lt;p&gt;In my experience shipping 18 production applications, Skia is often touted as the "Flutter-like" saviour for React Native performance. The reality is that it introduces a new category of failure modes—specifically memory leaks in the C++ layer and drawing cache invalidation storms—that you won't see in a "Hello World" or a basic chart demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the "Drawing Loop" Fails Under Load
&lt;/h2&gt;

&lt;p&gt;Most developers treat &lt;code&gt;Canvas&lt;/code&gt; like a standard React component. If the props change, the component re-renders. However, &lt;code&gt;react-native-skia&lt;/code&gt; uses the JSI (JavaScript Interface) to host C++ objects. When you pass a large array of points to a &lt;code&gt;Path&lt;/code&gt; object 60 times a second, you aren't just changing a pointer; you are asking the Skia engine to re-tessellate geometry.&lt;/p&gt;

&lt;p&gt;The failure happens because of the &lt;strong&gt;Asynchronous Mismatch&lt;/strong&gt;. Even with the New Architecture (TurboModules and Fabric) in React Native 0.76, the bridge between the JS logic calculating the path and the GPU executing the draw call has a finite throughput. If your calculation takes 12ms and your draw takes 6ms, you’ve already exceeded the 16.6ms window for 60 FPS. Unlike standard View components, Skia won't just "drop" the frame; it will often attempt to process the backlog, leading to the "rubber-banding" effect where animations speed up unnaturally to catch up to the current state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix: Step-by-Step Production Hardening
&lt;/h2&gt;

&lt;p&gt;To move from a prototype to a production-grade Skia implementation that survives sustained load, follow these steps.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Externalise the Value State
&lt;/h3&gt;

&lt;p&gt;Do not use &lt;code&gt;useState&lt;/code&gt; for high-frequency Skia data. React’s reconciliation is too slow for 60Hz drawing. Use &lt;code&gt;useSharedValue&lt;/code&gt; from &lt;code&gt;react-native-reanimated&lt;/code&gt;, which Skia supports natively.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Action:&lt;/strong&gt; Replace your data arrays with &lt;code&gt;SkiaMutableValue&lt;/code&gt; or Reanimated Shared Values.&lt;br&gt;
&lt;strong&gt;Why:&lt;/strong&gt; This keeps the data updates entirely on the UI/Worklet thread, bypassing the JS thread’s event loop.&lt;br&gt;
&lt;strong&gt;How to confirm:&lt;/strong&gt; Open the Flipper Performance Monitor. The "JS" thread usage should remain flat even as the animation runs.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. Manual Path Management and Point Capping
&lt;/h3&gt;

&lt;p&gt;A common mistake is passing a raw array of 1,000+ points to a Skia &lt;code&gt;Path&lt;/code&gt;. In a production environment, especially on Android devices with Mali GPUs, this causes the GPU driver to hang.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Action:&lt;/strong&gt; Implement a point-decimation algorithm (like Douglas-Peucker) or a simple radial distance filter before the data hits the path.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Example of a simple distance-based decimation&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;simplifiedPoints&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;points&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;points&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Confirm it worked:&lt;/strong&gt; Monitor the &lt;code&gt;GPU&lt;/code&gt; usage in Xcode Instruments. You should see a stable memory footprint rather than a "sawtooth" pattern.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Dispose of Skia Objects Manually
&lt;/h3&gt;

&lt;p&gt;Skia objects like &lt;code&gt;SkImage&lt;/code&gt;, &lt;code&gt;SkTypeface&lt;/code&gt;, and &lt;code&gt;SkRuntimeEffect&lt;/code&gt; are C++ wrappers. While the JS garbage collector (GC) eventually cleans them up, it is not aware of the memory pressure on the C++ side. In the HIPAA-aligned systems I’ve architected, unmanaged Skia objects led to OOM (Out of Memory) crashes after 40 minutes of continuous use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Action:&lt;/strong&gt; Use the &lt;code&gt;useImage&lt;/code&gt; or &lt;code&gt;useFont&lt;/code&gt; hooks provided by the library, but for custom shaders or paths created inside a &lt;code&gt;useMemo&lt;/code&gt;, you must use the &lt;code&gt;.delete()&lt;/code&gt; or &lt;code&gt;.dispose()&lt;/code&gt; methods if they are available in your version, or ensure they are wrapped in &lt;code&gt;useRawValue&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Offscreen Composition
&lt;/h3&gt;

&lt;p&gt;If you have static elements (like a grid or background) and dynamic elements (the waveform), do not redraw the static elements every frame.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Action:&lt;/strong&gt; Use a &lt;code&gt;Picture&lt;/code&gt; or &lt;code&gt;Image&lt;/code&gt; snapshot for the static background.&lt;br&gt;
&lt;strong&gt;Why:&lt;/strong&gt; &lt;code&gt;SkPicture&lt;/code&gt; records drawing commands and plays them back without re-calculating the geometry. I have seen this reduce GPU load by 30-40% in complex dashboard UIs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it Costs and When to Avoid It
&lt;/h2&gt;

&lt;p&gt;The trade-off for this performance is &lt;strong&gt;architectural complexity&lt;/strong&gt;. Once you move to &lt;code&gt;useSharedValue&lt;/code&gt; and manual path management, you lose the "declarative" simplicity of React. You are essentially writing C++ logic in a JavaScript syntax.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do not use Skia if:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Accessibility is a priority:&lt;/strong&gt; Skia draws to a canvas. The elements inside the &lt;code&gt;Canvas&lt;/code&gt; are invisible to Screen Readers (VoiceOver/TalkBack) unless you manually build a parallel "shadow" tree of accessible views.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text Heavy UIs:&lt;/strong&gt; Skia's text rendering (especially with &lt;code&gt;SkParagraph&lt;/code&gt;) requires manual font management and lacks the sophisticated hyphenation and RTL support built into the native iOS/Android &lt;code&gt;Text&lt;/code&gt; components.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Simple Layouts:&lt;/strong&gt; If you can achieve the design with &lt;code&gt;View&lt;/code&gt; and &lt;code&gt;CSS&lt;/code&gt;, do it. The overhead of bringing the Skia engine into your binary (adding ~4-8MB to the APK size) is not worth it for rounded corners or simple shadows.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  At Your Level
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Starting Out
&lt;/h3&gt;

&lt;p&gt;Focus on the difference between the "JS Thread" and the "UI Thread." Your first goal is to get a Skia animation running without the JS thread dropping below 58 FPS. Use &lt;code&gt;react-native-reanimated&lt;/code&gt; for all Skia transforms immediately; do not learn the "wrong" way first.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working Engineer
&lt;/h3&gt;

&lt;p&gt;Profile your app on a low-end Android device (e.g., a Samsung A-series). If you see "jank" or stuttering, look at your &lt;code&gt;Path&lt;/code&gt; objects. Are you recreating the path string or the path object on every render? Move all path construction into a &lt;code&gt;useDerivedValue&lt;/code&gt; to keep it off the React render cycle.&lt;/p&gt;

&lt;h3&gt;
  
  
  Senior or Staff
&lt;/h3&gt;

&lt;p&gt;You own the memory lifecycle. You must establish patterns for "Snapshotting"—capturing complex static vector groups as bitmapped images to save draw calls. You should also be looking at &lt;code&gt;SkiaRuntimeEffects&lt;/code&gt; (shaders) to offload heavy visual processing from the CPU to the GPU.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead or Director
&lt;/h3&gt;

&lt;p&gt;Understand the binary size vs. UX trade-off. Adding Skia is a long-term commitment to a specific rendering engine. Ensure your team has the capacity to handle the "Accessibility Gap" created by Canvas-based rendering, as this often requires 20% more effort in the QA phase to meet compliance standards like WCAG.&lt;/p&gt;

&lt;h2&gt;
  
  
  In the Interview
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Question:&lt;/strong&gt; "We're seeing significant frame drops when rendering a real-time chart with React Native Skia. How do you diagnose and fix this?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Weak Answer:&lt;/strong&gt; "I would use &lt;code&gt;useMemo&lt;/code&gt; to memoize the data and make sure I'm not doing too many re-renders. I might also try to reduce the number of points in the chart." This is weak because it assumes the bottleneck is React's reconciliation, whereas the bottleneck in Skia is usually the JSI bridge or GPU tessellation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Strong Answer:&lt;/strong&gt; A strong candidate will identify the &lt;strong&gt;Thread Boundary&lt;/strong&gt;. They will discuss moving the data stream into a &lt;code&gt;SharedValue&lt;/code&gt; to stay on the UI thread and avoid the bridge. They will mention &lt;strong&gt;GPU Overdraw&lt;/strong&gt;—the cost of drawing pixels that are later covered by other pixels—and how to use &lt;code&gt;SkPicture&lt;/code&gt; to cache static layers. They should also mention the &lt;strong&gt;Garbage Collection lag&lt;/strong&gt; between JS and C++, and how sustained high-frequency updates can lead to memory pressure that the JS GC doesn't respond to quickly enough.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Senior/Staff Follow-up:&lt;/strong&gt; "How does Skia's rendering impact the main thread's ability to handle user interactions like scrolling?"&lt;br&gt;
The answer should touch on the fact that if Skia is configured to run on the UI thread, a heavy draw call can block touch events. A senior engineer will suggest offloading complex path calculations to a Web Worker-like pattern or a background thread via a TurboModule, then passing the final result to the UI thread only for the final draw.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: &lt;a href="https://www.amitchakraborty.dev?utm_source=article&amp;amp;utm_medium=content&amp;amp;utm_campaign=react-native-skia-rendering-what-actually-breaks-in-production" rel="noopener noreferrer"&gt;www.amitchakraborty.dev&lt;/a&gt; · &lt;a href="https://linkedin.com/in/devamitch" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/devamitch" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. Open to senior and founding engineering roles, remote worldwide.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>skia</category>
      <category>performance</category>
      <category>react</category>
    </item>
  </channel>
</rss>
