<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: binadit</title>
    <description>The latest articles on DEV Community by binadit (@binadit).</description>
    <link>https://dev.to/binadit</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3853937%2F7b742322-ef72-44c9-92e2-8a32b6f3aa67.png</url>
      <title>DEV Community: binadit</title>
      <link>https://dev.to/binadit</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/binadit"/>
    <language>en</language>
    <item>
      <title>How to set up website server performance that holds under real traffic</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Fri, 25 Sep 2026 07:53:52 +0000</pubDate>
      <link>https://dev.to/binadit/how-to-set-up-website-server-performance-that-holds-under-real-traffic-31c5</link>
      <guid>https://dev.to/binadit/how-to-set-up-website-server-performance-that-holds-under-real-traffic-31c5</guid>
      <description>&lt;h2&gt;
  
  
  Your server works fine until it doesn't
&lt;/h2&gt;

&lt;p&gt;Every infrastructure engineer knows this moment: the server hums along fine in staging, then a traffic spike hits production and everything falls over. Requests queue, workers pin at max, response times spike from 200ms to 8 seconds. The usual reaction is to throw more CPU at it. That's rarely the actual problem.&lt;/p&gt;

&lt;p&gt;Most servers that collapse under load aren't under-resourced, they're misconfigured. The default settings for PHP-FPM, Nginx, and your database connections were never meant for production traffic. Here's the exact sequence of changes that fixes this, in the order that actually matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 0: measure before you touch anything
&lt;/h2&gt;

&lt;p&gt;You can't prove a fix worked without a baseline. Run this first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ab &lt;span class="nt"&gt;-n&lt;/span&gt; 1000 &lt;span class="nt"&gt;-c&lt;/span&gt; 50 https://yourdomain.com/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Record requests per second, mean response time, and failed request count. Check &lt;code&gt;uptime&lt;/code&gt; and &lt;code&gt;vmstat 1 5&lt;/code&gt; during the test too. Write it all down; you'll compare against it later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 1: PHP-FPM is probably misconfigured
&lt;/h2&gt;

&lt;p&gt;The default pool config is not production-ready. Open it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;nano /etc/php/8.3/fpm/pool.d/www.conf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Size your workers based on actual memory usage, not guesswork. Check real footprint per worker with &lt;code&gt;ps aux | grep php-fpm&lt;/code&gt; (typically 30-60MB for WordPress/Laravel), then divide available RAM by that number.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="py"&gt;pm&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;dynamic&lt;/span&gt;
&lt;span class="py"&gt;pm.max_children&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;40&lt;/span&gt;
&lt;span class="py"&gt;pm.start_servers&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;10&lt;/span&gt;
&lt;span class="py"&gt;pm.min_spare_servers&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;5&lt;/span&gt;
&lt;span class="py"&gt;pm.max_spare_servers&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;15&lt;/span&gt;
&lt;span class="py"&gt;pm.max_requests&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;500&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pm.max_requests&lt;/code&gt; is underrated. It recycles workers periodically, which stops slow memory leaks from degrading your server over hours of uptime.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 2: Nginx is dropping connections it doesn't need to
&lt;/h2&gt;

&lt;p&gt;Reduce TCP handshake overhead with proper keepalive and connection settings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;worker_processes&lt;/span&gt; &lt;span class="s"&gt;auto&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;worker_rlimit_nofile&lt;/span&gt; &lt;span class="mi"&gt;65535&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;events&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;worker_connections&lt;/span&gt; &lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="s"&gt;epoll&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;multi_accept&lt;/span&gt; &lt;span class="no"&gt;on&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;http&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;keepalive_timeout&lt;/span&gt; &lt;span class="mi"&gt;65&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;keepalive_requests&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;sendfile&lt;/span&gt; &lt;span class="no"&gt;on&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;tcp_nopush&lt;/span&gt; &lt;span class="no"&gt;on&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;tcp_nodelay&lt;/span&gt; &lt;span class="no"&gt;on&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you're proxying to PHP-FPM over a Unix socket, bump the backlog to match expected concurrency:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="py"&gt;listen&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;/run/php/php8.3-fpm.sock&lt;/span&gt;
&lt;span class="py"&gt;listen.backlog&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;1024&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Fix 3: your database connections are too expensive
&lt;/h2&gt;

&lt;p&gt;Each unpooled connection costs a handshake, auth round trip, and memory allocation. At scale, this is a silent killer. For Postgres, PgBouncer is the standard fix:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;pgbouncer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[databases]&lt;/span&gt;
&lt;span class="py"&gt;yourapp&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;host=127.0.0.1 port=5432 dbname=yourapp&lt;/span&gt;

&lt;span class="nn"&gt;[pgbouncer]&lt;/span&gt;
&lt;span class="py"&gt;listen_port&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;6432&lt;/span&gt;
&lt;span class="py"&gt;auth_type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;md5&lt;/span&gt;
&lt;span class="py"&gt;pool_mode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;transaction&lt;/span&gt;
&lt;span class="py"&gt;max_client_conn&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;500&lt;/span&gt;
&lt;span class="py"&gt;default_pool_size&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;25&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Point your app at port 6432 instead of 5432. This typically cuts connection overhead by 60-80% under concurrent load. MySQL users: look at ProxySQL or mysqlnd_mux.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 4: cache what shouldn't hit the database every time
&lt;/h2&gt;

&lt;p&gt;Install Redis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;redis-server
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl &lt;span class="nb"&gt;enable &lt;/span&gt;redis-server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set eviction policy so it doesn't just run out of memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="err"&gt;maxmemory&lt;/span&gt; &lt;span class="err"&gt;512mb&lt;/span&gt;
&lt;span class="err"&gt;maxmemory-policy&lt;/span&gt; &lt;span class="err"&gt;allkeys-lru&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For custom apps, wrap expensive queries in a cache-aside pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="nv"&gt;$cacheKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"product:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;$id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="nv"&gt;$product&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$redis&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$cacheKey&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nv"&gt;$product&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nv"&gt;$product&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$db&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"SELECT * FROM products WHERE id = ?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;$id&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
    &lt;span class="nv"&gt;$redis&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;setex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$cacheKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;serialize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$product&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Fix 5: stop static assets from hitting your app server
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="p"&gt;~&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt; &lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="s"&gt;.(jpg|jpeg|png|gif|css|js|woff2)&lt;/span&gt;$ &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;expires&lt;/span&gt; &lt;span class="s"&gt;30d&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;add_header&lt;/span&gt; &lt;span class="s"&gt;Cache-Control&lt;/span&gt; &lt;span class="s"&gt;"public,&lt;/span&gt; &lt;span class="s"&gt;immutable"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If there's a CDN in front, confirm it respects your origin headers instead of overriding with its own default TTL.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying it actually worked
&lt;/h2&gt;

&lt;p&gt;Re-run the same load test, same concurrency:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ab &lt;span class="nt"&gt;-n&lt;/span&gt; 1000 &lt;span class="nt"&gt;-c&lt;/span&gt; 50 https://yourdomain.com/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare against baseline. RPS should climb, often 30-100%. Response time variance between p50 and p95 should tighten. Failed requests should hit zero at your target concurrency.&lt;/p&gt;

&lt;p&gt;Watch PHP-FPM workers scale during load:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;watch &lt;span class="nt"&gt;-n&lt;/span&gt; 1 &lt;span class="s1"&gt;'ps aux | grep php-fpm | wc -l'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Workers should scale up smoothly and settle, not spike immediately to &lt;code&gt;max_children&lt;/code&gt; and stay there. If they pin instantly, either your sizing is wrong or a slow query is holding connections open too long.&lt;/p&gt;

&lt;p&gt;Check PgBouncer is actually in the path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;psql &lt;span class="nt"&gt;-h&lt;/span&gt; 127.0.0.1 &lt;span class="nt"&gt;-p&lt;/span&gt; 6432 &lt;span class="nt"&gt;-U&lt;/span&gt; youruser yourapp &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"SHOW POOLS;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And check Redis hit rate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;redis-cli info stats | &lt;span class="nb"&gt;grep &lt;/span&gt;keyspace
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Healthy cache-aside logic should show above 80% hit ratio within minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistakes that undo all of this
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Oversizing max_children.&lt;/strong&gt; More workers than RAM supports causes swapping, which is worse than queued requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Caching without invalidation.&lt;/strong&gt; Stale cached data is a harder bug to catch than a slow endpoint. Always set a TTL or invalidate on write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pooling without adjusting app-level assumptions.&lt;/strong&gt; Transaction-mode pooling breaks session-level features like advisory locks if your app still expects a persistent connection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skipping the baseline.&lt;/strong&gt; No before/after numbers means no proof anything worked.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't a full re-architecture. It's five configuration changes applied in order, with measurement at each step. That's usually enough to take a server from falling over at the first spike to handling concurrency predictably.&lt;/p&gt;

&lt;p&gt;Full walkthrough with more context: &lt;a href="https://binadit.com/blog/website-server-performance-managed-cloud-provider-europe-setup" rel="noopener noreferrer"&gt;How to set up website server performance that holds under real traffic&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/website-server-performance-managed-cloud-provider-europe-setup" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Your email is GDPR-compliant today. Will it still be next year?</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Sun, 20 Sep 2026 13:07:11 +0000</pubDate>
      <link>https://dev.to/binadit/your-email-is-gdpr-compliant-today-will-it-still-be-next-year-3o5b</link>
      <guid>https://dev.to/binadit/your-email-is-gdpr-compliant-today-will-it-still-be-next-year-3o5b</guid>
      <description>&lt;h2&gt;
  
  
  Your GDPR compliance depends on a court case you're not tracking
&lt;/h2&gt;

&lt;p&gt;Here's an uncomfortable fact: the legal basis for your Google Workspace or Microsoft 365 email is a single framework that's currently being challenged in front of the same court that already killed its two predecessors. If you run infrastructure or make hosting decisions, this is worth ten minutes of your time.&lt;/p&gt;

&lt;p&gt;Ask your IT partner if EU email in a US productivity suite is GDPR-compliant and you'll get a confident yes: both providers are certified under the EU-US Data Privacy Framework. Technically true. Also missing the point.&lt;/p&gt;

&lt;h3&gt;
  
  
  The legal stack, in short
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Sending personal data to a US company is a "transfer" under GDPR Chapter V (Articles 44-50).&lt;/li&gt;
&lt;li&gt;That's only lawful today because of the EU-US Data Privacy Framework (DPF), adopted July 2023.&lt;/li&gt;
&lt;li&gt;Its two predecessors, Safe Harbor and Privacy Shield, were both struck down by the Court of Justice of the EU (CJEU).&lt;/li&gt;
&lt;li&gt;The DPF is now under appeal at that same court.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three attempts. Two collapses. One pending verdict. That's not a stable foundation to build a compliance program on.&lt;/p&gt;

&lt;h3&gt;
  
  
  "We use an EU data center" doesn't fix it
&lt;/h3&gt;

&lt;p&gt;This is the part engineers get wrong most often. Region selection is not jurisdiction. Under the US CLOUD Act, a US company has to hand over data it controls when a US authority demands it, regardless of which data center it's sitting in. A Frankfurt or Amsterdam region owned by a US parent is still in scope.&lt;/p&gt;

&lt;p&gt;Don't take my word for it. In June 2025, Microsoft France's director of legal affairs testified under oath before a French Senate inquiry. Asked directly if he could guarantee French citizens' data would never be handed to the US government without French authorization, his answer was: "No, I cannot guarantee that."&lt;/p&gt;

&lt;p&gt;That's the provider, on the record, under oath.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where things stand right now
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;A challenge to the DPF (Latombe v. Commission) was dismissed by the General Court in September 2025, but only on facts as they existed in 2023. Nothing since then was considered. It's now on appeal to the CJEU.&lt;/li&gt;
&lt;li&gt;Microsoft was granted leave to intervene in that appeal. A company doesn't intervene in EU court proceedings unless it has serious skin in the game.&lt;/li&gt;
&lt;li&gt;The US oversight board that underpins the DPF (PCLOB) has been without quorum since January 2025.&lt;/li&gt;
&lt;li&gt;In June 2026, the US Supreme Court ruled FTC commissioners can be removed at will by the President. The FTC is the enforcement body for the DPF on the US side.&lt;/li&gt;
&lt;li&gt;FISA Section 702, the surveillance law at the center of the original Schrems II ruling, lapsed in June 2026 after failed extensions. Collection continues under existing court certifications regardless.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this means the DPF will definitely fall. It means your compliance posture depends on a decision you don't control, on a timeline you can't predict. Last time this happened (Schrems II, July 2020), companies had zero notice and a three-year gap before a replacement existed.&lt;/p&gt;

&lt;h3&gt;
  
  
  The AI layer adds more surface area
&lt;/h3&gt;

&lt;p&gt;Copilot and Gemini are now embedded directly in mail and file storage. That's more data flows to account for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Questions worth asking your provider contract, literally:&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Does input get used for model training? (check tier, not marketing page)
&lt;span class="p"&gt;-&lt;/span&gt; Can you delete a specific individual's data from a trained model? (usually no)
&lt;span class="p"&gt;-&lt;/span&gt; Do you know which model processed which record, in which region?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The EU AI Act's transparency rules are already active (since August 2026). The heavier documentation requirements for HR, credit, and access-to-services use cases were pushed to December 2027, but that's a delay, not a way out. If you can't answer the questions above today, you won't be able to answer them then either.&lt;/p&gt;

&lt;h3&gt;
  
  
  The actual fix
&lt;/h3&gt;

&lt;p&gt;If email, calendar, and file storage sit with a provider incorporated and operating entirely in the EU, no CLOUD Act exposure exists because there's no US legal entity to compel. That's not a mitigation, it's the absence of the transfer problem entirely.&lt;/p&gt;

&lt;p&gt;As a data point: a Dutch provider offering this setup starts at €1.99 per mailbox per month, often cheaper than what teams already pay for Workspace or M365 licensing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Practical takeaway
&lt;/h3&gt;

&lt;p&gt;Check who actually handles mail for your domain (MX records will tell you the provider, not the jurisdiction, but it's a start). Then ask your provider the same question the French Senate asked Microsoft: can you guarantee this data is never handed over without our government's consent? If the answer is "no, but it hasn't happened," that's not a compliance guarantee, that's a probability bet.&lt;/p&gt;

&lt;p&gt;Full breakdown of the legal timeline and case law: &lt;a href="https://binadit.com/blog/email-gdpr-compliant-google-workspace-microsoft-365" rel="noopener noreferrer"&gt;Your email is GDPR-compliant today. Will it still be next year?&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/email-gdpr-compliant-google-workspace-microsoft-365" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>gdpr</category>
      <category>googleworkspace</category>
      <category>microsoft365</category>
      <category>cloudact</category>
    </item>
    <item>
      <title>Solving the deploy frequency wall: from weekly releases to multiple daily deploys without new incidents</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Fri, 18 Sep 2026 07:59:30 +0000</pubDate>
      <link>https://dev.to/binadit/solving-the-deploy-frequency-wall-from-weekly-releases-to-multiple-daily-deploys-without-new-4i8h</link>
      <guid>https://dev.to/binadit/solving-the-deploy-frequency-wall-from-weekly-releases-to-multiple-daily-deploys-without-new-4i8h</guid>
      <description>&lt;h1&gt;
  
  
  Why your team keeps reverting to weekly deploys (and how to actually fix it)
&lt;/h1&gt;

&lt;p&gt;Here's a pattern that shows up constantly in teams of 15-40 engineers: leadership pushes for faster releases, the team bumps deploy frequency up, and within a few weeks someone quietly walks it back to weekly. No mandate, no meeting. Just two or three rough releases in a row, and the team self-polices back to "safe" cadence.&lt;/p&gt;

&lt;p&gt;It looks like a discipline problem. It isn't. It's an infrastructure problem, and it's fixable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real issue: your pipeline was built for a different era
&lt;/h2&gt;

&lt;p&gt;When you deploy weekly, you can afford sloppiness because a human has time to babysit each release. Watch the error rate for 20 minutes, glance at a dashboard, ship it. Try that five times a day and the manual check either disappears or becomes theater nobody trusts.&lt;/p&gt;

&lt;p&gt;Under the hood, this usually comes down to four coupled problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deploys aren't atomic.&lt;/strong&gt; Code, migrations, and config all ship in one shot. If the migration is slow, the code can't go out without it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rollback means redeploying.&lt;/strong&gt; Reverting takes the same pipeline, the same time, the same risk as deploying forward.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitoring lags deploy speed.&lt;/strong&gt; Alert thresholds and log delays were tuned for hourly checks, not minute-by-minute validation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blast radius is 100%.&lt;/strong&gt; Every deploy hits the whole fleet at once, so any bad release is automatically a full incident.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is about code quality or test coverage. It's about the deploy path, routing layer, and monitoring stack not being built for the frequency you're asking of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: decouple, automate rollback, shrink the blast radius
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Split schema changes from code deploys
&lt;/h3&gt;

&lt;p&gt;Migrations are usually the biggest source of deploy anxiety. Use the expand/contract pattern so schema changes never ship in the same step as code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Step 1: expand (safe, backward compatible)&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;ADD&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;shipping_method_v2&lt;/span&gt; &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Step 2: dual-write in app code, backfill&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;shipping_method_v2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;shipping_method&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;shipping_method_v2&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Step 3: contract (separate deploy, days later)&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;shipping_method&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Code deploys stop waiting on migrations, and migrations stop blocking rollbacks.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Make rollback a routing flip, not a redeploy
&lt;/h3&gt;

&lt;p&gt;If reverting means rerunning the pipeline backward, people will hesitate to ship under pressure. Rollback should be a traffic switch, done in seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Nginx upstream weight shift for canary rollback&lt;/span&gt;
&lt;span class="k"&gt;upstream&lt;/span&gt; &lt;span class="s"&gt;backend&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;app-v124-1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;8080&lt;/span&gt; &lt;span class="s"&gt;weight=0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# new version, was live&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;app-v123-1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;8080&lt;/span&gt; &lt;span class="s"&gt;weight=100&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;# previous stable, restored&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;# reload takes effect in under 1 second&lt;/span&gt;
&lt;span class="k"&gt;nginx&lt;/span&gt; &lt;span class="s"&gt;-s&lt;/span&gt; &lt;span class="s"&gt;reload&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or on Kubernetes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl rollout undo deployment/checkout-api
kubectl rollout status deployment/checkout-api &lt;span class="nt"&gt;--timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;30s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ship to a small slice of instances first. Healthy? Shift more weight over. Not healthy? Shift back to zero. No twelve-minute rebuild required to recover.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Automate the canary check
&lt;/h3&gt;

&lt;p&gt;A human staring at Grafana for 20 minutes works at weekly cadence. At five deploys a day, it's neither sustainable nor reliable, people get worse at spotting anomalies the more times they repeat the check.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;canary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;metrics&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;error_rate&lt;/span&gt;
      &lt;span class="na"&gt;threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;baseline + 0.5%&lt;/span&gt;
      &lt;span class="na"&gt;window&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;p95_latency&lt;/span&gt;
      &lt;span class="na"&gt;threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;baseline + 15%&lt;/span&gt;
      &lt;span class="na"&gt;window&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5xx_count&lt;/span&gt;
      &lt;span class="na"&gt;threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;baseline + &lt;/span&gt;&lt;span class="m"&gt;10&lt;/span&gt;
      &lt;span class="na"&gt;window&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
  &lt;span class="na"&gt;action_on_fail&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;auto_rollback&lt;/span&gt;
  &lt;span class="na"&gt;promotion_steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;5%&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;25%&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;50%&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;100%&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;step_duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is probably the single highest-leverage change for teams moving to daily deploys. It turns a subjective judgment call into a repeatable gate.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Shrink the blast radius with progressive rollout
&lt;/h3&gt;

&lt;p&gt;Deploying to 100% of instances at once makes every release a full-fleet bet. Instead:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ship to 5% of instances or one availability zone&lt;/li&gt;
&lt;li&gt;Hold through a full traffic cycle, including cron jobs and batch work&lt;/li&gt;
&lt;li&gt;Auto-promote on healthy metrics, auto-rollback on bad ones&lt;/li&gt;
&lt;li&gt;Only reach 100% after clearing every stage&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same principle as safe database migrations: keep each irreversible step small.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to know it actually worked
&lt;/h2&gt;

&lt;p&gt;Don't just track "deploy success." Watch these specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Change failure rate&lt;/strong&gt;: percentage of deploys triggering a rollback or hotfix. Aim under 15% (DORA elite benchmark), trending down as frequency goes up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MTTR&lt;/strong&gt;: with automated rollback via traffic weight shift, this should drop under 5 minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rollback execution time&lt;/strong&gt;: measure from "threshold breached" to "traffic fully reverted." Should be seconds to low minutes, not tied to build time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Canary false negative rate&lt;/strong&gt;: how often something bad slips past the gate and gets caught by users instead. If it happens more than once a quarter, retune thresholds, don't remove the gate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deploy frequency vs. incident count&lt;/strong&gt;: these two lines should decouple. If incidents still climb with frequency, your canary metrics are probably missing a business signal like checkout completion or cart abandonment.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Give this at least four weeks at the new cadence before declaring victory. One good week proves nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping it from backsliding
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Alert on deploy frequency itself. A quiet drop without an explicit decision usually means the pipeline is causing avoidance.&lt;/li&gt;
&lt;li&gt;Treat canary thresholds as living config, not a one-time setup. Traffic patterns shift, thresholds need to shift with them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full technical breakdown, including the zero-downtime migration mechanics: &lt;a href="https://binadit.com/blog/infrastructure-performance-optimization-multiple-daily-deploys" rel="noopener noreferrer"&gt;read the original article&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/infrastructure-performance-optimization-multiple-daily-deploys" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to stabilize a system while it is actively failing</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Wed, 16 Sep 2026 07:32:01 +0000</pubDate>
      <link>https://dev.to/binadit/how-to-stabilize-a-system-while-it-is-actively-failing-2ph3</link>
      <guid>https://dev.to/binadit/how-to-stabilize-a-system-while-it-is-actively-failing-2ph3</guid>
      <description>&lt;h1&gt;
  
  
  Stop diagnosing, start stabilizing: a field guide for live incidents
&lt;/h1&gt;

&lt;p&gt;Your error rate just spiked. Someone is already asking "what changed?" in Slack. Here's the uncomfortable truth: the fastest way to make an incident worse is to start debugging the root cause before you've stopped the bleeding. This is the sequence we use to stabilize production systems first and investigate second, whether that's a single Node server or a full multi-node setup behind a load balancer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two mistakes that make incidents worse
&lt;/h2&gt;

&lt;p&gt;Under pressure, teams almost always do one of these:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Change too many things at once, so nobody can tell what fixed (or broke) what&lt;/li&gt;
&lt;li&gt;Start root-causing before the system is even stable&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Everything below exists to prevent those two mistakes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you need before an incident happens
&lt;/h2&gt;

&lt;p&gt;You can't stabilize what you can't see or touch. Minimum bar:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Shell access to affected hosts (SSH or a bastion)&lt;/li&gt;
&lt;li&gt;Basic monitoring: CPU, memory, disk I/O, error rates (Prometheus, Datadog, or even &lt;code&gt;htop&lt;/code&gt; plus logs)&lt;/li&gt;
&lt;li&gt;A load balancer or reverse proxy in front of the app (Nginx, HAProxy, cloud LB)&lt;/li&gt;
&lt;li&gt;A known, tested rollback procedure&lt;/li&gt;
&lt;li&gt;A second person to sanity-check decisions, even just in a Slack thread&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're missing any of this, that's your next project, not something to build mid-incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: freeze all changes
&lt;/h2&gt;

&lt;p&gt;This isn't technical, it's procedural. Post this immediately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INCIDENT: elevated error rate on checkout-api since 14:32 UTC.
No deploys, no config changes, no manual DB edits until stabilized.
Incident channel: #inc-2024-checkout
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This stops the classic "quick fix" from a well-meaning teammate that stacks a second problem on top of the first one you haven't diagnosed yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: check what changed in the last 24 hours
&lt;/h2&gt;

&lt;p&gt;Most live failures trace back to something recent: a deploy, a config edit, a traffic spike, or an upstream update.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Recent deploys&lt;/span&gt;
git log &lt;span class="nt"&gt;--since&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"24 hours ago"&lt;/span&gt; &lt;span class="nt"&gt;--oneline&lt;/span&gt;

&lt;span class="c"&gt;# Recent config changes (if infra is version controlled)&lt;/span&gt;
git &lt;span class="nt"&gt;-C&lt;/span&gt; /etc/nginx log &lt;span class="nt"&gt;--since&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"24 hours ago"&lt;/span&gt;

&lt;span class="c"&gt;# System-level changes&lt;/span&gt;
last &lt;span class="nt"&gt;-x&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-20&lt;/span&gt;
&lt;span class="nb"&gt;cat&lt;/span&gt; /var/log/dpkg.log | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%Y-%m-%d&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If timing lines up, roll back before you dig further:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Tagged release rollback&lt;/span&gt;
git checkout tags/v2.14.1
./deploy.sh production

&lt;span class="c"&gt;# Kubernetes rollback&lt;/span&gt;
kubectl rollout undo deployment/checkout-api
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 3: shed load before you preserve correctness
&lt;/h2&gt;

&lt;p&gt;No obvious recent change, or the rollback didn't help? Reduce load on the failing component. You're trading functionality for stability, temporarily.&lt;/p&gt;

&lt;p&gt;Ordered roughly by disruption:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rate limit at the edge (per IP or per token)&lt;/li&gt;
&lt;li&gt;Kill non-critical endpoints: recommendations, analytics beacons, background reports&lt;/li&gt;
&lt;li&gt;Maintenance page for non-essential routes, keep checkout/login alive&lt;/li&gt;
&lt;li&gt;Scale horizontally if the bottleneck is compute, not a shared resource&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Emergency Nginx rate limit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;limit_req_zone&lt;/span&gt; &lt;span class="nv"&gt;$binary_remote_addr&lt;/span&gt; &lt;span class="s"&gt;zone=emergency:10m&lt;/span&gt; &lt;span class="s"&gt;rate=5r/s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/api/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;limit_req&lt;/span&gt; &lt;span class="s"&gt;zone=emergency&lt;/span&gt; &lt;span class="s"&gt;burst=10&lt;/span&gt; &lt;span class="s"&gt;nodelay&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;limit_req_status&lt;/span&gt; &lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://backend&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nginx &lt;span class="nt"&gt;-t&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; nginx &lt;span class="nt"&gt;-s&lt;/span&gt; reload
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 4: the database goes first
&lt;/h2&gt;

&lt;p&gt;Databases usually fail last and recover slowest. A saturated DB will keep everything else unstable no matter what you fix upstream.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Connections vs max&lt;/span&gt;
&lt;span class="n"&gt;psql&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="nv"&gt;"SELECT count(*) FROM pg_stat_activity;"&lt;/span&gt;
&lt;span class="n"&gt;psql&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="nv"&gt;"SHOW max_connections;"&lt;/span&gt;

&lt;span class="c1"&gt;-- Long-running queries&lt;/span&gt;
&lt;span class="n"&gt;psql&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="k"&gt;c&lt;/span&gt; &lt;span class="nv"&gt;"SELECT pid, now() - query_start AS duration, query
         FROM pg_stat_activity
         WHERE state = 'active'
         ORDER BY duration DESC LIMIT 10;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kill runaway queries holding locks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;pg_terminate_backend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;pg_stat_activity&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;pid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;14832&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If connection exhaustion is the pattern, PgBouncer fixes this without touching app code. No pooler running? That's your action item once the fire's out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: fail over instead of fighting a bad node
&lt;/h2&gt;

&lt;p&gt;Disk pressure, memory leak, bad kernel state on one box: pull it from rotation instead of debugging it live.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# HAProxy&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"disable server backend/web03"&lt;/span&gt; | socat stdio /var/run/haproxy.sock

&lt;span class="c"&gt;# Kubernetes&lt;/span&gt;
kubectl cordon node-3
kubectl drain node-3 &lt;span class="nt"&gt;--ignore-daemonsets&lt;/span&gt; &lt;span class="nt"&gt;--delete-emptydir-data&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once it's out of rotation, it's a forensic artifact you can inspect at your own pace, not a liability sitting in the traffic path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: communicate facts, not guesses
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;STATUS 15:10 UTC: error rate back to baseline (0.2%) after rolling back
deploy v2.14.2 and draining node-3. Root cause under investigation.
Next update in 30 min or on change.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't close the incident the second the dashboard looks calm.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to actually verify stability
&lt;/h2&gt;

&lt;p&gt;One green graph isn't proof. Hold these steady for 30 to 60 minutes at real traffic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Error rate&lt;/strong&gt;: back to baseline, checked in 5-minute buckets, not instant values&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency&lt;/strong&gt;: p50, p95, and p99 all recovered, a good average can hide a bad p99&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resource headroom&lt;/strong&gt;: CPU, memory, connection pools have actual margin&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Queue depth&lt;/strong&gt;: backlogs draining, not just growing slower
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Nginx status code breakdown&lt;/span&gt;
&lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 5000 /var/log/nginx/access.log | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $9}'&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; | &lt;span class="nb"&gt;uniq&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt;

&lt;span class="c"&gt;# Live connection stats&lt;/span&gt;
watch &lt;span class="nt"&gt;-n&lt;/span&gt; 2 &lt;span class="s2"&gt;"ss -s"&lt;/span&gt;

&lt;span class="c"&gt;# DB connection headroom&lt;/span&gt;
psql &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"SELECT count(*), max_conn FROM pg_stat_activity, (SELECT setting::int AS max_conn FROM pg_settings WHERE name='max_connections') s GROUP BY max_conn;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare current traffic-normalized error rate to the same time last week, not to five minutes ago. Post-load-shedding quiet periods lie.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pitfalls to avoid
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Root-causing mid-incident&lt;/strong&gt;: adds risk, rarely speeds recovery&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rolling back and rolling forward in the same window&lt;/strong&gt;: give one change time to prove itself&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Declaring victory too early&lt;/strong&gt;: five calm minutes isn't a stable system&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skipping the change freeze&lt;/strong&gt;: uncoordinated fixes are how you get two incidents instead of one&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stabilize first. Root-cause after. Every time.&lt;/p&gt;

&lt;p&gt;Full original write-up with more context: &lt;a href="https://binadit.com/blog/stabilize-system-active-failure-infrastructure-management-services" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/stabilize-system-active-failure-infrastructure-management-services" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>From 4.2s to 380ms: debugging latency in a high availability infrastructure setup</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:05:45 +0000</pubDate>
      <link>https://dev.to/binadit/from-42s-to-380ms-debugging-latency-in-a-high-availability-infrastructure-setup-5gj</link>
      <guid>https://dev.to/binadit/from-42s-to-380ms-debugging-latency-in-a-high-availability-infrastructure-setup-5gj</guid>
      <description>&lt;h1&gt;
  
  
  Five silent bottlenecks that turned a 400ms API into a 4.2s crawl
&lt;/h1&gt;

&lt;p&gt;No deploys. No schema changes. No new integrations. Just a scheduling and resource-planning SaaS with 40k users watching its p95 response time climb from 400ms to 4.2 seconds over six months. Nobody could point to a cause, because there wasn't one big cause, there were five small ones stacked on top of each other.&lt;/p&gt;

&lt;p&gt;This is the writeup of how we found and fixed them, without a rearchitecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  The symptoms before the numbers existed
&lt;/h2&gt;

&lt;p&gt;Support tickets called the dashboard "laggy" long before anyone had hard metrics. By the time it got measured:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;p95 latency: 400ms to 4.2s over ~6 months&lt;/li&gt;
&lt;li&gt;Trial-to-paid conversion down 11% in the same window&lt;/li&gt;
&lt;li&gt;Sales getting asked about performance during renewals&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The team had already tried the usual moves: bigger instances, a read replica, scheduled restarts. None of it held. When scaling vertically twice does nothing, that's a strong signal the problem is architectural, not capacity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step one: measure before touching anything
&lt;/h2&gt;

&lt;p&gt;We spent the first week purely on instrumentation, tracing requests across the API gateway, app servers, database, and cache layer under normal load. No fixes yet. Changing the system while you're trying to measure it just adds noise.&lt;/p&gt;

&lt;p&gt;Five issues surfaced. None of them alone explained 4.2 seconds, but together they did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five issues
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. N+1 queries in a hot path
&lt;/h3&gt;

&lt;p&gt;A dashboard endpoint fetched a customer's project list, then hit the DB once per project for status. 40 projects meant 41 round trips for one page load.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Connection pool exhaustion
&lt;/h3&gt;

&lt;p&gt;8 app servers x 20 connections each = 160 possible connections, against a Postgres &lt;code&gt;max_connections&lt;/code&gt; of 100. At peak, requests queued for a connection slot before a query even ran.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Cache invalidation nuking whole namespaces
&lt;/h3&gt;

&lt;p&gt;Redis had a 30s TTL on project status, reasonable on paper. But a background job flushed entire cache namespaces on &lt;em&gt;any&lt;/em&gt; write, including unrelated ones. Actual hit ratio: 34%.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Cross-AZ chatter
&lt;/h3&gt;

&lt;p&gt;App tier and Redis cluster weren't pinned to the same AZ. ~40% of Redis calls crossed zones, adding 2 to 4ms each. At volume, that's hundreds of milliseconds per request.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. A synchronous third-party call in the request path
&lt;/h3&gt;

&lt;p&gt;A usage-tracking webhook to an external analytics provider ran synchronously inside the request cycle. On a bad day that provider took 800ms to 1.5s, and every request waited on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How we prioritized fixes
&lt;/h2&gt;

&lt;p&gt;Impact vs. risk, not ease of implementation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Sync third-party call, high impact, low risk&lt;/li&gt;
&lt;li&gt;N+1 queries, high impact, low risk&lt;/li&gt;
&lt;li&gt;Connection pool exhaustion, high impact, medium risk&lt;/li&gt;
&lt;li&gt;Cache invalidation, medium impact, medium risk&lt;/li&gt;
&lt;li&gt;Cross-AZ placement, medium impact, low risk&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We explicitly skipped touching instance sizes again. Two rounds of vertical scaling with zero improvement was itself the data point: the bottleneck was I/O and architecture, not compute.&lt;/p&gt;

&lt;p&gt;We also shipped fixes one at a time, measuring after each. Bundling everything into one release makes it impossible to know what actually helped, or what quietly broke something else.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fixes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Move the analytics call off the request path
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// before: synchronous call blocking the response&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;analyticsClient&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;track&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// 800ms-1.5s on slow days&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// after: fire-and-forget via queue&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;queue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;analytics.track&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// ~2ms&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Used the existing Redis infra as the queue backend, no new dependency. A worker process handles delivery with retries and a dead-letter queue. This alone cut 800ms to 1.5s off the median request on affected endpoints, and when the analytics provider had two outages during the engagement, users never noticed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kill the N+1
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- before&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;projects&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;account_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="c1"&gt;-- then per project:&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;project_status&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;project_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- after&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;updated_at&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;projects&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;project_status&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;project_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;41 round trips became 1. Query time for that endpoint went from ~620ms to ~45ms average.&lt;/p&gt;

&lt;h3&gt;
  
  
  Retune the connection pool
&lt;/h3&gt;

&lt;p&gt;Dropped per-server pool size from 20 to 12 (8 x 12 = 96, under the 100 limit) and added PgBouncer in transaction pooling mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[databases]&lt;/span&gt;
&lt;span class="py"&gt;app_db&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;host=127.0.0.1 port=5432 dbname=app_production&lt;/span&gt;

&lt;span class="nn"&gt;[pgbouncer]&lt;/span&gt;
&lt;span class="py"&gt;pool_mode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;transaction&lt;/span&gt;
&lt;span class="py"&gt;max_client_conn&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;500&lt;/span&gt;
&lt;span class="py"&gt;default_pool_size&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;25&lt;/span&gt;
&lt;span class="py"&gt;reserve_pool_size&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;5&lt;/span&gt;
&lt;span class="py"&gt;reserve_pool_timeout&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;3&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Connection wait time went from ~180ms average at peak to under 5ms.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fix cache invalidation
&lt;/h3&gt;

&lt;p&gt;Switched from namespace-wide flushes on any write to key-level invalidation tied to the specific project changed. Kept the 30s TTL as a safety net, but it stopped being the primary mechanism.&lt;/p&gt;

&lt;p&gt;Cache hit ratio: 34% to 91% within a week. Small diff, outsized effect on database load.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;None of these fixes were exotic. This is what accumulated technical debt looks like in a system that grew for a few years without a dedicated performance pass. If your latency crept up gradually with no single obvious cause, look for a stack of small architectural issues before you reach for bigger hardware.&lt;/p&gt;

&lt;p&gt;Full writeup with more detail: &lt;a href="https://binadit.com/blog/debugging-latency-high-availability-infrastructure-case-study" rel="noopener noreferrer"&gt;Debugging latency in a high availability infrastructure setup&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/debugging-latency-high-availability-infrastructure-case-study" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>backend</category>
      <category>debugging</category>
      <category>infrastructure</category>
      <category>performance</category>
    </item>
    <item>
      <title>How Docker networking broke checkout under load: a container ecommerce infrastructure case study</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Wed, 09 Sep 2026 07:02:54 +0000</pubDate>
      <link>https://dev.to/binadit/how-docker-networking-broke-checkout-under-load-a-container-ecommerce-infrastructure-case-study-4cja</link>
      <guid>https://dev.to/binadit/how-docker-networking-broke-checkout-under-load-a-container-ecommerce-infrastructure-case-study-4cja</guid>
      <description>&lt;h2&gt;
  
  
  Checkout was timing out at 900 req/s, and it had nothing to do with CPU
&lt;/h2&gt;

&lt;p&gt;A marketplace client came to us with a scary but familiar symptom: checkout p95 latency jumping from 280ms to over 2.1 seconds during traffic spikes. Their instinct was to throw more containers at it. That made things &lt;em&gt;worse&lt;/em&gt;. Here's what was actually going on, and how we fixed it without touching a line of application code.&lt;/p&gt;

&lt;p&gt;Link to the full writeup: &lt;a href="https://binadit.com/blog/docker-networking-production-ecommerce-infrastructure-case-study" rel="noopener noreferrer"&gt;Docker networking broke checkout under load&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The setup
&lt;/h3&gt;

&lt;p&gt;PHP monolith, containerized about a year prior, running on a single host via Docker Compose: app containers, Redis, a worker queue, behind a managed load balancer. 40k DAU, peak ~900 req/s. It ran fine for six months. Then flash sales started producing "the site is freezing" tickets.&lt;/p&gt;

&lt;p&gt;The team's first move was scaling app containers, assuming CPU/memory pressure. Query times on Postgres stayed stable at 8-14ms the whole time, so the database wasn't the culprit either. The bottleneck was hiding in the network layer between containers, a place default tooling barely gives you visibility into.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the audit found
&lt;/h3&gt;

&lt;p&gt;Three compounding issues, none fatal alone:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Default bridge network overhead.&lt;/strong&gt; Every inter-container hop (app to Redis, app to Postgres, app to search) was going through userland proxying on the default bridge. We measured ~1.8ms added latency per hop. At 900 req/s with 3-4 internal calls per request, that adds up fast.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Conntrack table exhaustion.&lt;/strong&gt; &lt;code&gt;nf_conntrack_max&lt;/code&gt; was still at the kernel default of 65,536. Short-lived Redis/Postgres connections churned through the table during spikes, filled it, and the kernel silently started dropping packets. Nobody noticed because syslog wasn't shipped anywhere useful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. DNS resolution overhead.&lt;/strong&gt; Docker's embedded DNS (127.0.0.11) was resolving service names on every new connection instead of the app caching results. Under load, with connection churn, lookups started queuing behind each other.&lt;/p&gt;

&lt;p&gt;None of these show up with 10 test users in staging. All three show up hard with 900 real ones under concurrent load.&lt;/p&gt;

&lt;h3&gt;
  
  
  What we didn't do
&lt;/h3&gt;

&lt;p&gt;We ruled out "just move to Kubernetes." It wouldn't have fixed anything here; K8s has its own version of the same problems (CNI choice, kube-proxy mode, CoreDNS caching). Swapping orchestrators without fixing the root cause just relocates it.&lt;/p&gt;

&lt;p&gt;We also ruled out rewriting the app to reduce internal calls. The call pattern was normal. The network layer needed to handle it efficiently, not the other way around.&lt;/p&gt;

&lt;h3&gt;
  
  
  The fix, in four layers
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Kernel tuning for conntrack:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;net&lt;/span&gt;.&lt;span class="n"&gt;netfilter&lt;/span&gt;.&lt;span class="n"&gt;nf_conntrack_max&lt;/span&gt; = &lt;span class="m"&gt;262144&lt;/span&gt;
&lt;span class="n"&gt;net&lt;/span&gt;.&lt;span class="n"&gt;netfilter&lt;/span&gt;.&lt;span class="n"&gt;nf_conntrack_tcp_timeout_established&lt;/span&gt; = &lt;span class="m"&gt;600&lt;/span&gt;
&lt;span class="n"&gt;net&lt;/span&gt;.&lt;span class="n"&gt;ipv4&lt;/span&gt;.&lt;span class="n"&gt;tcp_tw_reuse&lt;/span&gt; = &lt;span class="m"&gt;1&lt;/span&gt;
&lt;span class="n"&gt;net&lt;/span&gt;.&lt;span class="n"&gt;core&lt;/span&gt;.&lt;span class="n"&gt;somaxconn&lt;/span&gt; = &lt;span class="m"&gt;4096&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Applied via &lt;code&gt;/etc/sysctl.d/99-docker-network.conf&lt;/code&gt; and &lt;code&gt;sysctl --system&lt;/code&gt;. We also started monitoring conntrack table utilization going forward, since a full table fails silently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Moved internal traffic off the default bridge:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker network create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--driver&lt;/span&gt; bridge &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--opt&lt;/span&gt; com.docker.network.bridge.enable_icc&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--opt&lt;/span&gt; com.docker.network.driver.mtu&lt;span class="o"&gt;=&lt;/span&gt;9000 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--subnet&lt;/span&gt; 172.28.0.0/16 &lt;span class="se"&gt;\&lt;/span&gt;
  internal-services
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;App containers, Redis, and the search sidecar moved onto this network, keeping east-west traffic off Docker's default NAT path. The public-facing load balancer stayed on a separate network; no change to external attack surface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Local DNS caching with dnsmasq as a sidecar:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="c"&gt;# dnsmasq.conf
&lt;/span&gt;&lt;span class="n"&gt;no&lt;/span&gt;-&lt;span class="n"&gt;resolv&lt;/span&gt;
&lt;span class="n"&gt;server&lt;/span&gt;=&lt;span class="m"&gt;127&lt;/span&gt;.&lt;span class="m"&gt;0&lt;/span&gt;.&lt;span class="m"&gt;0&lt;/span&gt;.&lt;span class="m"&gt;11&lt;/span&gt;
&lt;span class="n"&gt;cache&lt;/span&gt;-&lt;span class="n"&gt;size&lt;/span&gt;=&lt;span class="m"&gt;1000&lt;/span&gt;
&lt;span class="n"&gt;local&lt;/span&gt;-&lt;span class="n"&gt;ttl&lt;/span&gt;=&lt;span class="m"&gt;10&lt;/span&gt;
&lt;span class="n"&gt;neg&lt;/span&gt;-&lt;span class="n"&gt;ttl&lt;/span&gt;=&lt;span class="m"&gt;5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 10-second TTL absorbed connection churn during spikes without causing stale resolution issues when containers got replaced on deploy. We specifically tested that a container restart got picked up within one TTL window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PgBouncer for connection pooling:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[databases]&lt;/span&gt;
&lt;span class="py"&gt;marketplace&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;host=postgres-primary port=5432 dbname=marketplace&lt;/span&gt;

&lt;span class="nn"&gt;[pgbouncer]&lt;/span&gt;
&lt;span class="py"&gt;pool_mode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;transaction&lt;/span&gt;
&lt;span class="py"&gt;max_client_conn&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;2000&lt;/span&gt;
&lt;span class="py"&gt;default_pool_size&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;50&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was sequenced last on purpose. Fewer short-lived connections means fewer conntrack entries and fewer DNS lookups, so it made everything upstream easier once the network layer was already sound.&lt;/p&gt;

&lt;h3&gt;
  
  
  Takeaway
&lt;/h3&gt;

&lt;p&gt;If your containerized app slows down only under load and CPU/memory graphs look fine, stop scaling containers and go look at conntrack, your bridge driver, and DNS resolution behavior. That's usually where it's actually hiding.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/docker-networking-production-ecommerce-infrastructure-case-study" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>docker</category>
      <category>infrastructure</category>
      <category>networking</category>
      <category>performance</category>
    </item>
    <item>
      <title>Solving the real ROI question behind moving off US hyperscalers</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Wed, 02 Sep 2026 07:23:04 +0000</pubDate>
      <link>https://dev.to/binadit/solving-the-real-roi-question-behind-moving-off-us-hyperscalers-2j8k</link>
      <guid>https://dev.to/binadit/solving-the-real-roi-question-behind-moving-off-us-hyperscalers-2j8k</guid>
      <description>&lt;h1&gt;
  
  
  Why your hyperscaler exit math keeps failing finance review
&lt;/h1&gt;

&lt;p&gt;You run the numbers on leaving AWS. Compute is cheaper elsewhere. Storage is cheaper elsewhere. Then finance asks about egress fees, migration hours, and cutover risk, and the whole business case falls apart. Sound familiar? This isn't a bad migration plan, it's a broken cost model.&lt;/p&gt;

&lt;p&gt;Most engineers comparing hyperscaler costs against alternatives only look at list prices for compute and storage. That's maybe 40% of the real picture. Here's the rest of it, and a framework to fix your ROI math.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the comparison breaks down
&lt;/h2&gt;

&lt;p&gt;Hyperscalers don't really compete on compute price. They compete on making it painful to leave. That's baked into the architecture, not just the contract.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Egress fees are the real anchor.&lt;/strong&gt; AWS charges roughly $0.09/GB after the first GB out. For a platform pushing 50TB/month (video, API responses, backups, CDN pulls), that's about $4,500/month just to move data, before touching anything else. This fee isn't about bandwidth cost, bandwidth is cheap. It's a moat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Managed services lock you in harder than compute does.&lt;/strong&gt; If you're on RDS, DynamoDB, SQS, and Lambda, you're not renting VMs, you're depending on proprietary APIs. Migrating off means rewriting your data access layer, not just moving instances. This is the line item most spreadsheets underestimate by 3-5x.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reserved instance pricing hides a bad assumption.&lt;/strong&gt; Finance often compares a competitor's list price to a hyperscaler's 3-year reserved rate and calls it a win for the hyperscaler. But that rate assumes flat usage for 36 months. Most SaaS workloads have seasonal spikes and growth curves. A 3-year commitment is a liability wearing a discount's clothes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Support cost is invisible until you actually need it.&lt;/strong&gt; A named TAM tier starts around $15,000/month, and you're still in a ticket queue for anything real. Compare that to a partner where a senior engineer just picks up.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: five cost categories, not one
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Baseline 12 months of actual spend, not one month
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ce get-cost-and-usage &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--time-period&lt;/span&gt; &lt;span class="nv"&gt;Start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2024-11-01,End&lt;span class="o"&gt;=&lt;/span&gt;2025-11-01 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--granularity&lt;/span&gt; MONTHLY &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--metrics&lt;/span&gt; &lt;span class="s1"&gt;'UnblendedCost'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--group-by&lt;/span&gt; &lt;span class="nv"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;DIMENSION,Key&lt;span class="o"&gt;=&lt;/span&gt;SERVICE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Break it into compute, storage, egress, managed services, support, and RI amortization. Egress alone is usually 8-15% of total spend for content-heavy platforms, and it almost never makes the first draft of a migration business case.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Model the actual migration cost
&lt;/h3&gt;

&lt;p&gt;This means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Engineering hours to replace managed-service dependencies (DynamoDB to Postgres, Lambda to containers, SQS to self-hosted queues)&lt;/li&gt;
&lt;li&gt;One-time data transfer cost to move the dataset out&lt;/li&gt;
&lt;li&gt;Parallel-run cost during validation (typically 4-8 weeks)&lt;/li&gt;
&lt;li&gt;A downtime contingency line, even with a zero-downtime plan&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a mid-sized SaaS app with a 2TB database and 15 microservices, expect 200-450 engineering hours. At 80 euros/hour loaded cost, that's 16,000-36,000 euros, upfront. Put it in the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Recalculate egress under the new architecture
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# AWS us-east-1 to internet: $0.09/GB after first 1GB free&lt;/span&gt;
&lt;span class="c"&gt;# EU provider with peering to major IXPs: $0.01-0.02/GB typical&lt;/span&gt;

&lt;span class="c"&gt;# 50TB/month egress:&lt;/span&gt;
&lt;span class="c"&gt;# AWS: 50,000GB * $0.09 = $4,500/month&lt;/span&gt;
&lt;span class="c"&gt;# EU provider: 50,000GB * $0.015 = $750/month&lt;/span&gt;
&lt;span class="c"&gt;# Annual difference: $45,000&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Colocate with your CDN's regional PoPs and egress typically drops 40-70%.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Swap proprietary APIs for open standards
&lt;/h3&gt;

&lt;p&gt;This is the highest-leverage move in the whole exit. It's what makes the savings durable instead of a one-time discount you slowly erode by re-adopting vendor tooling:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DynamoDB → PostgreSQL or self-managed MongoDB&lt;/li&gt;
&lt;li&gt;Lambda → containers on Kubernetes or Nomad&lt;/li&gt;
&lt;li&gt;SQS/SNS → self-hosted RabbitMQ or Kafka&lt;/li&gt;
&lt;li&gt;CloudWatch → Prometheus + Grafana&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Cut over with zero downtime
&lt;/h3&gt;

&lt;p&gt;Replicate continuously, run both environments in parallel, lower DNS TTL ahead of time, keep the old environment warm for at least one billing cycle.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 48-72 hours before cutover&lt;/span&gt;
example.com.  300  IN  A  203.0.113.10

&lt;span class="c"&gt;# Monitor both origins during cutover&lt;/span&gt;
curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code} %{time_total}s\n'&lt;/span&gt; https://old-origin.example.com/health
curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code} %{time_total}s\n'&lt;/span&gt; https://new-origin.example.com/health
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For database-backed apps, run logical replication for days beforehand, not a single export/import window.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to know it actually worked
&lt;/h2&gt;

&lt;p&gt;Track these for at least two full billing cycles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Total spend vs. your 12-month baseline, normalized for traffic growth&lt;/li&gt;
&lt;li&gt;Egress as % of total spend (expect 8-15% → 2-4% with good peering)&lt;/li&gt;
&lt;li&gt;p95/p99 latency on your top 10 endpoints, before and after&lt;/li&gt;
&lt;li&gt;Error rate and uptime, 30 days before vs. 30 days after&lt;/li&gt;
&lt;li&gt;Engineering hours on infra ops per month&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One sanity check that matters: if spend drops 30% but p99 latency degrades 25%, you haven't won anything. You've just shifted cost into a metric that'll eventually cost you conversions.&lt;/p&gt;

&lt;p&gt;Full framework and migration mechanics in the original piece: &lt;a href="https://binadit.com/blog/cloud-cost-optimization-services-roi-moving-off-us-hyperscalers" rel="noopener noreferrer"&gt;Solving the real ROI question behind moving off US hyperscalers&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/cloud-cost-optimization-services-roi-moving-off-us-hyperscalers" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>cloud</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>How to move ecommerce infrastructure from a single VPS to HA without rewriting the application</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Wed, 26 Aug 2026 07:22:59 +0000</pubDate>
      <link>https://dev.to/binadit/how-to-move-ecommerce-infrastructure-from-a-single-vps-to-ha-without-rewriting-the-application-48h</link>
      <guid>https://dev.to/binadit/how-to-move-ecommerce-infrastructure-from-a-single-vps-to-ha-without-rewriting-the-application-48h</guid>
      <description>&lt;h2&gt;
  
  
  Your VPS will fail at the worst possible time
&lt;/h2&gt;

&lt;p&gt;Here's an uncomfortable truth: that single VPS running your store isn't a matter of if it fails, it's when. A kernel update reboot, a Black Friday memory spike, a disk that fills up overnight. None of these are exotic failure modes, they're Tuesday.&lt;/p&gt;

&lt;p&gt;The good news: fixing this is an infrastructure problem, not a rewrite. If you're running PHP, Node, or Python against a relational database, you can go from one box to a highly available setup without touching your application code. This applies whether you're on WooCommerce, Magento, or something custom.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you touch anything
&lt;/h2&gt;

&lt;p&gt;Check these boxes first:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your app can run statelessly across multiple servers (or can be made to, sessions and uploads are the usual culprits)&lt;/li&gt;
&lt;li&gt;You have root, not just FTP&lt;/li&gt;
&lt;li&gt;You've scheduled a maintenance window for DNS and DB cutover&lt;/li&gt;
&lt;li&gt;You have a tested, recent backup&lt;/li&gt;
&lt;li&gt;You know your traffic pattern: peak RPS, DB connections, payload size&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This walkthrough assumes a LAMP/LEMP-ish stack: Nginx or Apache, PHP-FPM or Node, MySQL or Postgres, Redis. Swap tooling as needed for other stacks, the pattern holds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Get state off local disk
&lt;/h2&gt;

&lt;p&gt;This is the real blocker to scaling horizontally. Local disk state has to go before you add a second server.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sessions&lt;/strong&gt; → move to Redis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="c"&gt;; php.ini
&lt;/span&gt;&lt;span class="py"&gt;session.save_handler&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;redis&lt;/span&gt;
&lt;span class="py"&gt;session.save_path&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"tcp://10.0.0.5:6379"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Uploaded media&lt;/strong&gt; → object storage (S3-compatible) or shared NFS. For WooCommerce, WP Offload Media handles this out of the box. For custom apps, swap filesystem writes for an SDK call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="nv"&gt;$s3&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;putObject&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
    &lt;span class="s1"&gt;'Bucket'&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s1"&gt;'store-uploads'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'Key'&lt;/span&gt;    &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;$filename&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'Body'&lt;/span&gt;   &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;fopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$tmpPath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'r'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Cache&lt;/strong&gt; → Redis or Memcached instead of local disk/OPcache-only.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Put a load balancer in front, even with one backend
&lt;/h2&gt;

&lt;p&gt;Spin up a small VM or a managed LB now, before you have a second server. Point DNS at the LB's IP immediately. This decouples DNS from your server count forever, no more DNS changes for future scaling.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;upstream&lt;/span&gt; &lt;span class="s"&gt;app_servers&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;10.0.0.10&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt; &lt;span class="s"&gt;max_fails=3&lt;/span&gt; &lt;span class="s"&gt;fail_timeout=30s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;10.0.0.11&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt; &lt;span class="s"&gt;max_fails=3&lt;/span&gt; &lt;span class="s"&gt;fail_timeout=30s&lt;/span&gt; &lt;span class="s"&gt;backup&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;443&lt;/span&gt; &lt;span class="s"&gt;ssl&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;shop.example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://app_servers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Host&lt;/span&gt; &lt;span class="nv"&gt;$host&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;X-Real-IP&lt;/span&gt; &lt;span class="nv"&gt;$remote_addr&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 3: Clone the app server
&lt;/h2&gt;

&lt;p&gt;Once sessions and media are externalized, your app server is basically stateless. Build a second node from the same provisioning script, Ansible, Docker, or a shell script, doesn't matter, just make it reproducible.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# minimal provisioning sanity check&lt;/span&gt;
php &lt;span class="nt"&gt;-v&lt;/span&gt;
nginx &lt;span class="nt"&gt;-v&lt;/span&gt;
composer &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;span class="nb"&gt;cat&lt;/span&gt; /etc/php/8.2/fpm/pool.d/www.conf | &lt;span class="nb"&gt;grep &lt;/span&gt;pm.max_children
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add it to the upstream block, drop the &lt;code&gt;backup&lt;/code&gt; flag, confirm both nodes handle real traffic before moving to the database layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Replicate the database
&lt;/h2&gt;

&lt;p&gt;Highest-risk step. Set up primary-replica replication or move to a managed cluster.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="c"&gt;# primary my.cnf
&lt;/span&gt;&lt;span class="py"&gt;server-id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;1&lt;/span&gt;
&lt;span class="py"&gt;log_bin&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;/var/log/mysql/mysql-bin.log&lt;/span&gt;
&lt;span class="py"&gt;binlog_do_db&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;shop_production&lt;/span&gt;

&lt;span class="c"&gt;# replica my.cnf
&lt;/span&gt;&lt;span class="py"&gt;server-id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;2&lt;/span&gt;
&lt;span class="py"&gt;relay-log&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;/var/log/mysql/mysql-relay-bin.log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;CHANGE&lt;/span&gt; &lt;span class="n"&gt;MASTER&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt;
  &lt;span class="n"&gt;MASTER_HOST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'10.0.0.5'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;MASTER_USER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'replicator'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;MASTER_PASSWORD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'***'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;MASTER_LOG_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'mysql-bin.000003'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;MASTER_LOG_POS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;154&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;START&lt;/span&gt; &lt;span class="n"&gt;SLAVE&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check lag before cutover:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SHOW&lt;/span&gt; &lt;span class="n"&gt;SLAVE&lt;/span&gt; &lt;span class="n"&gt;STATUS&lt;/span&gt;&lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="k"&gt;G&lt;/span&gt;
&lt;span class="c1"&gt;-- Seconds_Behind_Master should be 0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once stable, cut over writes using a virtual IP or ProxySQL, not a manual config edit during an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Real health checks, not port pings
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;upstream&lt;/span&gt; &lt;span class="s"&gt;app_servers&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;10.0.0.10&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt; &lt;span class="s"&gt;max_fails=2&lt;/span&gt; &lt;span class="s"&gt;fail_timeout=10s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;10.0.0.11&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt; &lt;span class="s"&gt;max_fails=2&lt;/span&gt; &lt;span class="s"&gt;fail_timeout=10s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# haproxy / nginx plus&lt;/span&gt;
&lt;span class="k"&gt;option&lt;/span&gt; &lt;span class="s"&gt;httpchk&lt;/span&gt; &lt;span class="s"&gt;GET&lt;/span&gt; &lt;span class="n"&gt;/health&lt;/span&gt;
&lt;span class="s"&gt;http-check&lt;/span&gt; &lt;span class="s"&gt;expect&lt;/span&gt; &lt;span class="s"&gt;status&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your &lt;code&gt;/health&lt;/code&gt; endpoint needs to actually check dependencies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nv"&gt;$pdo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;PDO&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$dsn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$pass&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nv"&gt;$redis&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Redis&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="nv"&gt;$redis&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'10.0.0.5'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;6379&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nb"&gt;http_response_code&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'ok'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Exception&lt;/span&gt; &lt;span class="nv"&gt;$e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nb"&gt;http_response_code&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'unhealthy'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Prove it actually works
&lt;/h2&gt;

&lt;p&gt;Don't trust the config, test it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Kill a node during low traffic, confirm the LB routes around it inside your &lt;code&gt;fail_timeout&lt;/code&gt; window and response times stay flat&lt;/li&gt;
&lt;li&gt;Watch &lt;code&gt;Seconds_Behind_Master&lt;/code&gt; during peak checkout load, not just idle&lt;/li&gt;
&lt;li&gt;Load test both nodes: &lt;code&gt;ab -n 5000 -c 50 https://shop.example.com/&lt;/code&gt; and diff the access logs&lt;/li&gt;
&lt;li&gt;Simulate a DB failover in staging, measure reconnect time (should be seconds, not minutes)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Baseline to aim for: zero customer-visible errors when one app node dies, database failover under 30 seconds with retry logic in place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistakes that will bite you
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sticky sessions instead of externalized sessions&lt;/strong&gt;: works fine until a node dies and half your logged-in users get booted&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicate cron jobs&lt;/strong&gt;: order processing and cache warming running on both nodes doubles the work, pin them to one node or use a job queue&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shallow health checks&lt;/strong&gt;: pinging port 80 says nothing about whether the DB connection is alive&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replication as a backup strategy&lt;/strong&gt;: it protects against hardware failure, not against a bad &lt;code&gt;DELETE&lt;/code&gt; that replicates instantly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never testing failover&lt;/strong&gt;: the first failover shouldn't happen during a real outage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full walkthrough with more context here: &lt;a href="https://binadit.com/blog/single-vps-to-ha-ecommerce-infrastructure-without-rewrite" rel="noopener noreferrer"&gt;How to move ecommerce infrastructure from a single VPS to HA without rewriting the application&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/single-vps-to-ha-ecommerce-infrastructure-without-rewrite" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Measuring what FISA 702 reauthorization actually changes for EU SaaS infrastructure</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Fri, 21 Aug 2026 08:24:39 +0000</pubDate>
      <link>https://dev.to/binadit/measuring-what-fisa-702-reauthorization-actually-changes-for-eu-saas-infrastructure-2nlb</link>
      <guid>https://dev.to/binadit/measuring-what-fisa-702-reauthorization-actually-changes-for-eu-saas-infrastructure-2nlb</guid>
      <description>&lt;h2&gt;
  
  
  Your "EU region" toggle probably isn't doing what you think
&lt;/h2&gt;

&lt;p&gt;Here's an uncomfortable fact: setting your AWS region to &lt;code&gt;eu-central-1&lt;/code&gt; does nothing to shield you from FISA 702. None of it. If your provider is a US company, that data can still be compelled, regardless of which data center it physically sits in. We ran the numbers to find out what it actually takes to fix this, and whether the fix costs you performance.&lt;/p&gt;

&lt;p&gt;FISA 702 got reauthorized in April 2024, running through 2026. It'll come up again, and every time it does, compliance teams ask engineering the same question: can we safely keep running on US-owned cloud infra? The answer isn't legal, it's architectural.&lt;/p&gt;

&lt;h3&gt;
  
  
  The setup
&lt;/h3&gt;

&lt;p&gt;We benchmarked the same Laravel API workload (PostgreSQL 16, Redis 7, object storage) across three configs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Config A&lt;/strong&gt;: AWS &lt;code&gt;eu-central-1&lt;/code&gt;, standard hyperscaler, US-headquartered&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config B&lt;/strong&gt;: EU-owned provider (OVHcloud), data residency guarantees&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config C&lt;/strong&gt;: Dedicated private cloud, Rotterdam data center, no shared hyperscaler control plane&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same hardware baseline everywhere: 8 vCPU / 32GB app nodes, NVMe DB nodes at 4 vCPU / 16GB, PgBouncer in front of Postgres, Nginx 1.25 as reverse proxy.&lt;/p&gt;

&lt;p&gt;Load generated with k6, from a Frankfurt runner to kill long-haul network noise as a variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// k6 load profile&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;stages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;10m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="c1"&gt;// ramp 50 -&amp;gt; 2000 VUs&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;20m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="c1"&gt;// sustained peak&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="c1"&gt;// mix: 70% GET, 30% POST/PUT&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We measured two separate things: raw performance, and legal exposure surface (who can be compelled to hand over your data, under what instrument, with or without notifying you).&lt;/p&gt;

&lt;h3&gt;
  
  
  Result 1: performance is basically a wash
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;p50&lt;/th&gt;
&lt;th&gt;p95&lt;/th&gt;
&lt;th&gt;p99&lt;/th&gt;
&lt;th&gt;Max sustained req/s&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A: AWS eu-central-1&lt;/td&gt;
&lt;td&gt;42ms&lt;/td&gt;
&lt;td&gt;118ms&lt;/td&gt;
&lt;td&gt;210ms&lt;/td&gt;
&lt;td&gt;3,150&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B: OVHcloud (EU)&lt;/td&gt;
&lt;td&gt;47ms&lt;/td&gt;
&lt;td&gt;134ms&lt;/td&gt;
&lt;td&gt;245ms&lt;/td&gt;
&lt;td&gt;2,890&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C: Private cloud (Rotterdam)&lt;/td&gt;
&lt;td&gt;39ms&lt;/td&gt;
&lt;td&gt;102ms&lt;/td&gt;
&lt;td&gt;178ms&lt;/td&gt;
&lt;td&gt;3,020&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The private cloud setup wins on tail latency, but only by 15-20%. If you're picking your infra provider based on speed alone, don't bother, the difference is inside normal variance for most apps. This is not a performance decision.&lt;/p&gt;

&lt;h3&gt;
  
  
  Result 2: the legal exposure gap is real and it's structural
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Subject to FISA 702&lt;/th&gt;
&lt;th&gt;Compellable without notifying you&lt;/th&gt;
&lt;th&gt;Legal instrument&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A: AWS eu-central-1&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes, gag orders are standard&lt;/td&gt;
&lt;td&gt;FISA 702 + CLOUD Act&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B: OVHcloud&lt;/td&gt;
&lt;td&gt;No direct exposure, but sub-processor risk&lt;/td&gt;
&lt;td&gt;Depends on your stack&lt;/td&gt;
&lt;td&gt;GDPR only, if fully EU-owned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C: Private cloud&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;GDPR + Dutch law only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here's the part that should worry you: &lt;strong&gt;60% of the "EU cloud" test deployments we checked were still routing through US-owned CDNs, DNS, or email providers.&lt;/strong&gt; You think you're compliant because your DB lives in an EU region, but your DNS resolver, your CDN, and your transactional email are all silently reintroducing FISA 702 exposure.&lt;/p&gt;

&lt;p&gt;The region setting is not the control point. The corporate parent of every vendor in your chain is.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it actually costs to fix
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;US-owned default&lt;/th&gt;
&lt;th&gt;EU-owned alternative&lt;/th&gt;
&lt;th&gt;Effort&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DNS&lt;/td&gt;
&lt;td&gt;Route 53&lt;/td&gt;
&lt;td&gt;deSEC / EU-hosted BIND&lt;/td&gt;
&lt;td&gt;Low, 2-4 hrs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CDN&lt;/td&gt;
&lt;td&gt;CloudFront&lt;/td&gt;
&lt;td&gt;Bunny CDN (EU entity)&lt;/td&gt;
&lt;td&gt;Medium, 1-2 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Email&lt;/td&gt;
&lt;td&gt;SES / SendGrid&lt;/td&gt;
&lt;td&gt;Mailjet (FR entity)&lt;/td&gt;
&lt;td&gt;Medium, 1 day + DNS propagation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Object storage&lt;/td&gt;
&lt;td&gt;S3&lt;/td&gt;
&lt;td&gt;OVHcloud Object Storage / self-hosted MinIO&lt;/td&gt;
&lt;td&gt;High, 3-5 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monitoring/APM&lt;/td&gt;
&lt;td&gt;Datadog&lt;/td&gt;
&lt;td&gt;Self-hosted Grafana + Prometheus&lt;/td&gt;
&lt;td&gt;High, 1-2 weeks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notice the pattern: cheap fixes first, expensive fixes require real migration planning.&lt;/p&gt;

&lt;h3&gt;
  
  
  The pragmatic order of operations
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Swap DNS and CDN first.&lt;/strong&gt; Cheap, fast, immediately reduces exposure. A few hours to a couple days.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move transactional email next.&lt;/strong&gt; Slightly more friction from DNS propagation, still low-risk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat object storage as a proper migration, not a cutover.&lt;/strong&gt; This one involves real data transfer volume. Plan it with a testing window, not a weekend deploy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Self-host monitoring last&lt;/strong&gt;, or accept the risk consciously if it's low priority for your threat model. This is the highest-effort item and often the lowest payoff per hour spent.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Caveats worth knowing
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;We tested from a single European location. If you've got US or APAC traffic, a fully EU-only stack will add real latency you need to benchmark separately.&lt;/li&gt;
&lt;li&gt;Legal classification here is based on published transparency reports and EDPB guidance as of early 2025, not a review of your specific contracts. SCCs, the EU-US Data Privacy Framework, and things like AWS's "Digital Sovereignty Pledge" all interact with this picture in ways a pure infra test can't capture.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Bottom line
&lt;/h3&gt;

&lt;p&gt;FISA 702 reauthorization doesn't change the underlying math, it's been in force since 2008. What's changed is that RFPs from regulated industries now explicitly ask vendors to document their exposure. If you haven't audited your sub-processor chain, that's the actual work here, not swapping your primary cloud region.&lt;/p&gt;

&lt;p&gt;Full methodology and extended results: &lt;a href="https://binadit.com/blog/measuring-fisa-702-reauthorization-infrastructure-management-services" rel="noopener noreferrer"&gt;Measuring what FISA 702 reauthorization actually changes for EU SaaS infrastructure&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/measuring-fisa-702-reauthorization-infrastructure-management-services" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How DNS resolution works under the hood: a step by step guide for high availability infrastructure</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Wed, 19 Aug 2026 07:10:37 +0000</pubDate>
      <link>https://dev.to/binadit/how-dns-resolution-works-under-the-hood-a-step-by-step-guide-for-high-availability-infrastructure-51mk</link>
      <guid>https://dev.to/binadit/how-dns-resolution-works-under-the-hood-a-step-by-step-guide-for-high-availability-infrastructure-51mk</guid>
      <description>&lt;h1&gt;
  
  
  DNS resolution: the seven hops nobody thinks about until failover breaks
&lt;/h1&gt;

&lt;p&gt;You push a config change, wait for propagation, and traffic still hits the dead server. Sound familiar? Most DNS incidents come down to not knowing what actually happens between a browser typing a domain and a TCP handshake starting. Let's fix that.&lt;/p&gt;

&lt;p&gt;This is a practical walkthrough of the resolution chain, plus the config patterns that turn DNS into an active part of your high availability setup instead of a silent single point of failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Assumptions
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;You control a domain through Cloudflare, Route 53, or similar, with dashboard or API access&lt;/li&gt;
&lt;li&gt;You have a Linux box with &lt;code&gt;dig&lt;/code&gt;, &lt;code&gt;nslookup&lt;/code&gt;, and &lt;code&gt;tcpdump&lt;/code&gt;/&lt;code&gt;ngrep&lt;/code&gt; available&lt;/li&gt;
&lt;li&gt;You know the basics of IP, TCP, and UDP&lt;/li&gt;
&lt;li&gt;You're comfortable running commands against a real or disposable test domain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We'll use &lt;code&gt;example.com&lt;/code&gt; everywhere below. Swap it for your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The resolution chain, hop by hop
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Browser cache
&lt;/h3&gt;

&lt;p&gt;Before anything hits the network, Chrome checks its own DNS cache:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chrome://net-internals/#dns
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If there's a live entry, resolution stops right there. This is why a DNS change can look "stuck" in your own browser even after the TTL has expired everywhere else.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. OS stub resolver
&lt;/h3&gt;

&lt;p&gt;Next stop is the OS-level resolver, usually &lt;code&gt;systemd-resolved&lt;/code&gt; or &lt;code&gt;nscd&lt;/code&gt; on Linux:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;resolvectl status
resolvectl statistics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It checks &lt;code&gt;/etc/hosts&lt;/code&gt;, then its own cache, before forwarding anything upstream.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Recursive resolver
&lt;/h3&gt;

&lt;p&gt;No local hit means the query goes to a recursive resolver: your ISP's, or a public one like &lt;code&gt;1.1.1.1&lt;/code&gt; or &lt;code&gt;8.8.8.8&lt;/code&gt;. This is where &lt;code&gt;+trace&lt;/code&gt; becomes your best debugging friend:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig @1.1.1.1 example.com +trace
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. Root servers
&lt;/h3&gt;

&lt;p&gt;The recursive resolver asks a root server who's authoritative for &lt;code&gt;.com&lt;/code&gt;. Roots don't know about &lt;code&gt;example.com&lt;/code&gt;, only about the TLD:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;com. 172800 IN NS a.gtld-servers.net.
com. 172800 IN NS b.gtld-servers.net.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  5. TLD servers
&lt;/h3&gt;

&lt;p&gt;The TLD server hands back your domain's authoritative nameservers, the ones set at your registrar:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;example.com. 172800 IN NS ns1.yourdnsprovider.com.
example.com. 172800 IN NS ns2.yourdnsprovider.com.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  6. Authoritative nameserver
&lt;/h3&gt;

&lt;p&gt;Finally, the actual record comes back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;example.com. 300 IN A 203.0.113.42
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the step that matters for HA: which IP gets returned, how fast it can change, and what happens when that IP goes dark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configuring DNS for actual failover
&lt;/h2&gt;

&lt;p&gt;Understanding the chain is step one. Making it work for you is step two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Drop your TTLs on failover-critical records.&lt;/strong&gt; A 3600s TTL means an hour-long tail of stale traffic during failover. Use 60-300s instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;example.com. 60 IN A 203.0.113.42
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll pay for it in query volume against your authoritative nameservers. Worth it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Attach health checks.&lt;/strong&gt; Route 53 and Cloudflare can pull an unhealthy origin out of the response set automatically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws route53 change-resource-record-sets &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--hosted-zone-id&lt;/span&gt; Z1PA6795UKMFR9 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--change-batch&lt;/span&gt; &lt;span class="s1"&gt;'{
    "Changes": [{
      "Action": "UPSERT",
      "ResourceRecordSet": {
        "Name": "example.com",
        "Type": "A",
        "SetIdentifier": "primary",
        "Failover": "PRIMARY",
        "TTL": 60,
        "ResourceRecords": [{"Value": "203.0.113.42"}],
        "HealthCheckId": "abcd1234-healthcheck-id"
      }
    }]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the same mechanism behind a clean zero-downtime migration: DNS shifts traffic instead of forcing a hard cutover.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Return multiple A records.&lt;/strong&gt; Clients can fall back to a second IP:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;example.com. 300 IN A 203.0.113.42
example.com. 300 IN A 203.0.113.43
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fine for stateless services behind a load balancer, but not a real substitute for health-checked failover; clients cache order and don't always retry smartly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying it actually works
&lt;/h2&gt;

&lt;p&gt;Don't trust the dashboard. Check every layer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig example.com +noall +answer
dig example.com @8.8.8.8 +noall +answer
dig example.com @1.1.1.1 +noall +answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Lower-than-configured TTLs from a resolver mean the change is propagating. Original TTLs mean you're still hitting a cached answer.&lt;/p&gt;

&lt;p&gt;Check resolution latency directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig example.com | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"Query time"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Consistently above 100-150ms? Look at resolver placement or a provider with more edge presence. This happens before the TCP handshake even starts, so it's pure overhead on TTFB.&lt;/p&gt;

&lt;p&gt;Simulate failover instead of waiting for an incident to test it for you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;iptables &lt;span class="nt"&gt;-A&lt;/span&gt; INPUT &lt;span class="nt"&gt;-s&lt;/span&gt; &amp;lt;health-check-ip&amp;gt; &lt;span class="nt"&gt;-j&lt;/span&gt; DROP
watch &lt;span class="nt"&gt;-n&lt;/span&gt; 5 &lt;span class="s1"&gt;'dig example.com +short'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Confirm the response flips to your secondary within the configured TTL, then remove the rule.&lt;/p&gt;

&lt;p&gt;Add DNS resolution time to your monitoring as its own metric, separate from full page load. A spike there with stable backend times points straight at a resolver or nameserver problem, not your app.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pitfalls that keep coming back
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default TTLs.&lt;/strong&gt; 3600 or 86400 seconds quietly kills any failover plan built on top.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single nameserver provider.&lt;/strong&gt; One outage there takes your domain down regardless of server health.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DNSSEC misconfiguration.&lt;/strong&gt; Broken signing chains fail silently for validating resolvers while looking fine in tools that skip validation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Testing only from your machine.&lt;/strong&gt; Your local cache hides what real users see. Always check against multiple public resolvers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deep CNAME chains.&lt;/strong&gt; Each extra hop adds latency and another point of failure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Understand the seven hops, and DNS stops being a mystery box that occasionally ruins your day.&lt;/p&gt;

&lt;p&gt;Full original writeup: &lt;a href="https://binadit.com/blog/how-dns-resolution-works-high-availability-infrastructure" rel="noopener noreferrer"&gt;How DNS resolution works under the hood&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/how-dns-resolution-works-high-availability-infrastructure" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>TTFB vs full page load: choosing the right metric for your ecommerce infrastructure</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Fri, 14 Aug 2026 07:41:20 +0000</pubDate>
      <link>https://dev.to/binadit/ttfb-vs-full-page-load-choosing-the-right-metric-for-your-ecommerce-infrastructure-2e57</link>
      <guid>https://dev.to/binadit/ttfb-vs-full-page-load-choosing-the-right-metric-for-your-ecommerce-infrastructure-2e57</guid>
      <description>&lt;h1&gt;
  
  
  Stop chasing TTFB: it's not the metric that's tanking your conversions
&lt;/h1&gt;

&lt;p&gt;Your TTFB dashboard is green. Your conversion rate is red. If that combination sounds familiar, you're not alone, and you're not chasing the wrong fix by accident, you're chasing it because every tool you use puts TTFB front and center.&lt;/p&gt;

&lt;p&gt;Here's the problem: TTFB and full page load metrics answer completely different questions, and treating them as interchangeable wastes engineering time on ecommerce sites where every millisecond on checkout pages has a dollar value attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  What TTFB actually tells you
&lt;/h2&gt;

&lt;p&gt;TTFB is the clock between the request leaving the browser and the first byte coming back. That window includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DNS resolution (if uncached)&lt;/li&gt;
&lt;li&gt;TCP/TLS handshake&lt;/li&gt;
&lt;li&gt;Server-side work: routing, DB queries, template rendering, cache lookups&lt;/li&gt;
&lt;li&gt;Network transit for that first byte&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On a warm HTTP/2 or HTTP/3 connection, server processing time dominates this number. Which makes TTFB genuinely useful, but only for one job: telling you when your backend is under load.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TTFB baseline: 180ms
TTFB during flash sale: 900ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That jump isn't noise. It's your app server, DB, or cache layer running out of headroom before a single byte of HTML goes out. On WooCommerce stacks especially, a TTFB spike during a traffic surge is usually the first hard evidence that your PHP-FPM pool or database is maxed out, well before cart abandonment shows up in analytics.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where TTFB earns its keep
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Catching server-side regressions.&lt;/strong&gt; A slow query, an N+1 bug, a cold cache: TTFB flags these before users notice anything visually.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load testing and capacity planning.&lt;/strong&gt; Watch TTFB climb under simulated load and you know exactly where your app tier breaks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Splitting backend vs frontend blame.&lt;/strong&gt; Fast TTFB + slow-feeling page = stop looking at the server, start looking at JS and render-blocking assets.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Where it lies to you
&lt;/h3&gt;

&lt;p&gt;A 100ms TTFB tells you nothing about what happens after that byte lands. The page can still take 6 seconds to become usable because of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Render-blocking CSS/JS&lt;/li&gt;
&lt;li&gt;Images without dimensions causing layout shift&lt;/li&gt;
&lt;li&gt;A tag manager script hogging the main thread&lt;/li&gt;
&lt;li&gt;Late-loading fonts (hello, flash of invisible text)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And here's the trap most teams miss: if a response is served from a CDN edge cache, TTFB measures the edge, not your origin. Your dashboard can look perfectly healthy while your actual application is quietly degrading, right up until a cache miss exposes it in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  What full page load metrics tell you instead
&lt;/h2&gt;

&lt;p&gt;LCP, Speed Index, fully-loaded time: these track what the user actually sees and feels. That's why Google leans on LCP and CLS for Core Web Vitals; TTFB alone was a bad predictor of both perceived speed and search ranking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strengths:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Correlates with real user experience, not just server response&lt;/li&gt;
&lt;li&gt;Surfaces conversion killers directly (a bloated hero image, a blocking checkout script)&lt;/li&gt;
&lt;li&gt;Feeds into Core Web Vitals and SEO scoring&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Weaknesses:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Noisy. A regressed LCP could be caused by the backend, a new marketing pixel, an unoptimized CMS upload, or a hydration delay in your frontend framework&lt;/li&gt;
&lt;li&gt;Lagging indicator. You need waterfall analysis and resource timing to find the actual cause after LCP tanks in prod&lt;/li&gt;
&lt;li&gt;Out of infra's hands. You can run flawless high-availability infrastructure and still ship a slow LCP because a third-party payment widget is garbage&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The decision framework
&lt;/h2&gt;

&lt;p&gt;Stop asking "is TTFB good." Ask this instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;TTFB_high&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;LCP_high&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;fixBackendFirst&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// nothing downstream matters until this is solved&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;TTFB_low&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;LCP_high&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;profileBrowserWaterfall&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// frontend/third-party problem&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;TTFB_high&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;LCP_low&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// rare, but check if LCP measurement is misleading you&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;servingMostlyFromCDN&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;trackOriginResponseTimeSeparately&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// edge TTFB will lie to you&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Practical ownership split:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Metric to watch&lt;/th&gt;
&lt;th&gt;Who owns it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Debugging capacity/backend regression&lt;/td&gt;
&lt;td&gt;TTFB&lt;/td&gt;
&lt;td&gt;Backend/infra&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debugging conversion drop or CWV score drop&lt;/td&gt;
&lt;td&gt;LCP, Speed Index&lt;/td&gt;
&lt;td&gt;Frontend + infra&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prepping for a sale or launch&lt;/td&gt;
&lt;td&gt;Both, together&lt;/td&gt;
&lt;td&gt;Whole team&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For traffic events specifically: TTFB tells you when the servers start straining, LCP tells you whether shoppers are actually feeling it. Alert on both, but don't confuse one for the other, and don't let a green TTFB dashboard talk you out of investigating a real conversion problem.&lt;/p&gt;

&lt;p&gt;Full writeup with the complete comparison table and more edge cases: &lt;a href="https://binadit.com/blog/ttfb-vs-full-page-load-metric-ecommerce-infrastructure" rel="noopener noreferrer"&gt;TTFB vs full page load: choosing the right metric for your ecommerce infrastructure&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/ttfb-vs-full-page-load-metric-ecommerce-infrastructure" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Measuring how EuroStack changes the procurement conversation for infrastructure management services</title>
      <dc:creator>binadit</dc:creator>
      <pubDate>Wed, 12 Aug 2026 07:12:58 +0000</pubDate>
      <link>https://dev.to/binadit/measuring-how-eurostack-changes-the-procurement-conversation-for-infrastructure-management-services-43md</link>
      <guid>https://dev.to/binadit/measuring-how-eurostack-changes-the-procurement-conversation-for-infrastructure-management-services-43md</guid>
      <description>&lt;h1&gt;
  
  
  We ran the numbers on 41 procurement deals to see if EuroStack is real or just RFP theater
&lt;/h1&gt;

&lt;p&gt;If you've touched infrastructure procurement in the last year, you've probably seen "EU sovereignty" or "EuroStack" show up somewhere in the RFP. The question every engineering lead asks: is this actual signal that changes vendor selection, or is it a compliance checkbox that nobody enforces?&lt;/p&gt;

&lt;p&gt;We had visibility into 41 procurement processes between January 2023 and September 2025, so instead of guessing, we split the data before and after EuroStack-style criteria showed up explicitly in tender docs and measured what actually shifted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;We defined "EuroStack criteria" narrowly. It had to include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;EU legal jurisdiction over the operating entity&lt;/li&gt;
&lt;li&gt;Data processing location inside the EU/EEA&lt;/li&gt;
&lt;li&gt;Disclosure of subcontractor data flows, including support and monitoring tooling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"GDPR compliant" with no operational detail didn't count. That's boilerplate, not a filter.&lt;/p&gt;

&lt;p&gt;Sample: 41 contracts, €40k to €2.1M ACV, across SaaS (14), public sector (11), fintech/payments (9), and e-commerce/logistics (7). 19 processes ran before explicit sovereignty language appeared, 22 ran after.&lt;/p&gt;

&lt;p&gt;We were not benchmarking server performance here. This is purely procurement mechanics: shortlist size, timelines, contract clauses, final price vs. quote.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually moved
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Metric                              Before    After
Avg. vendors shortlisted            5.2       3.6
% non-EU hyperscaler-reseller-only  47%       12%
Avg. RFP-to-signature time          11.4wk    7.9wk
% with subcontractor disclosure     21%       86%
% with real data residency SLA      32%       79%
Avg. final value vs. initial quote  +14%      +4%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The timeline tail matters more than the average:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Percentile   Before (wk)   After (wk)
p50          9.5           6.0
p95          24.0          15.5
p99          31.0          19.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Post-signature scope disputes also dropped hard: 42% of "before" contracts had a change request or renegotiation in the first 6 months, usually over data location or support access. In the "after" group, that fell to 18%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this happens
&lt;/h2&gt;

&lt;p&gt;The shortlist shrinking from 5.2 to 3.6 vendors isn't a loss, it's a filter working correctly. Reseller setups that white-label a US hyperscaler and can't answer "where does support tooling send logs" get cut early instead of surviving three rounds of vague reassurance.&lt;/p&gt;

&lt;p&gt;That filtering effect is also why timelines compressed. When the RFP already specifies data residency and subcontractor disclosure, vendors don't need three rounds of clarifying questions. We saw this in our own bid pipeline: RFP-to-proposal time dropped from 9 days to 4 when the tender pre-answered the questions we'd otherwise have to ask ourselves.&lt;/p&gt;

&lt;p&gt;The contract value delta is the part people underestimate. Vague sovereignty language at RFP stage doesn't disappear, it just resurfaces post-signature as a change order once someone discovers a gap. Explicit criteria up front push that negotiation into the cheap phase instead of the expensive one.&lt;/p&gt;

&lt;p&gt;One more pattern worth flagging: buyers who specified sovereignty criteria were also far more likely to require named engineer contacts instead of ticket-based support (61% vs. 23%). Teams that care about knowing where their data lives also tend to care about knowing who picks up the phone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats before you cite this
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Selection bias&lt;/strong&gt;: we bid on EU-sovereign contracts, so our "before" sample skews toward processes where we were still invited. Fully non-sovereignty procurement is likely faster and larger in scope than what we captured.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small sample&lt;/strong&gt;: 41 deals over 2.5 years shows a trend, not statistical proof. One outlier public-sector tender inflated the before-p99 from 24 to 31 weeks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No official standard exists yet&lt;/strong&gt;: "EuroStack criteria" is our own classification. There's no certification scheme to point to as of late 2025.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No cost-per-compute comparison&lt;/strong&gt;: this says nothing about whether EU infra is cheaper or pricier per server-hour than hyperscaler equivalents. Different question, different analysis.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaway for anyone writing or responding to RFPs
&lt;/h2&gt;

&lt;p&gt;If you're a buyer: specific sovereignty requirements (jurisdiction, data location, subcontractor disclosure) act as a fast, cheap filter that eliminates non-answers early and shrinks your timeline. If you're a vendor: if you can't answer where support tooling sends logs, you're going to keep losing time to vendors who can.&lt;/p&gt;

&lt;p&gt;Full methodology and data breakdown in the &lt;a href="https://binadit.com/blog/measuring-eurostack-infrastructure-management-services-procurement" rel="noopener noreferrer"&gt;original article&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://binadit.com/blog/measuring-eurostack-infrastructure-management-services-procurement" rel="noopener noreferrer"&gt;binadit.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
