<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shrikant Mhatre (श्री)</title>
    <description>The latest articles on DEV Community by Shrikant Mhatre (श्री) (@shrikant_mhatre_0900).</description>
    <link>https://dev.to/shrikant_mhatre_0900</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172579%2F3e3fbd1b-7bd8-4a73-a9fb-9e4bbbfea012.jpg</url>
      <title>DEV Community: Shrikant Mhatre (श्री)</title>
      <link>https://dev.to/shrikant_mhatre_0900</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shrikant_mhatre_0900"/>
    <language>en</language>
    <item>
      <title>We Broke WordPress for 30 Minutes. Nginx Cache Kept Google From Noticing.</title>
      <dc:creator>Shrikant Mhatre (श्री)</dc:creator>
      <pubDate>Fri, 09 Oct 2026 08:16:45 +0000</pubDate>
      <link>https://dev.to/shrikant_mhatre_0900/we-broke-wordpress-for-30-minutes-nginx-cache-kept-google-from-noticing-o41</link>
      <guid>https://dev.to/shrikant_mhatre_0900/we-broke-wordpress-for-30-minutes-nginx-cache-kept-google-from-noticing-o41</guid>
      <description>&lt;p&gt;A WordPress directory migration, a missed Apache vhost config, a pre-warmed Nginx proxy cache, and a Coraza WAF in between. This is the story of how a cache warmer saved a blog from an SEO disaster when the backend silently broke for 30 minutes.&lt;/p&gt;

&lt;p&gt;Bagful runs a technical blog at bagful.net. The blog had been live for six weeks, indexed by Google, and generating its first organic impressions. The stack was straightforward: Nginx as the reverse proxy, Coraza WAF for request filtering, Apache serving WordPress on the backend. The site was small, around 190 posts, but the team had spent weeks optimizing content for search visibility. Then, on a routine Tuesday afternoon, an engineer moved the WordPress directory to a more secure location, ran a find-and-replace with sed, missed one config file, and silently broke every page on the site. This is what happened, how the proxy cache absorbed the failure, and what the team learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture: Four Layers Between the Browser and WordPress
&lt;/h2&gt;

&lt;p&gt;Before the incident, it helps to understand the proxy chain. Every request to bagful.net passes through four layers before reaching WordPress.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1: Nginx (ports 80/443).&lt;/strong&gt; The public-facing reverse proxy. Handles TLS termination, HTTP/2, static file caching, and proxy caching. Nginx caches full HTML responses from the backend so that repeat visitors get served directly from disk without hitting Apache at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2: Coraza WAF.&lt;/strong&gt; An open-source web application firewall running as a reverse proxy between Nginx and Apache. Coraza inspects every request against OWASP Core Rule Set (CRS) rules, blocks SQL injection, XSS, and other malicious payloads before they reach the application. It listens on an internal port and forwards clean requests to Apache.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3: Apache.&lt;/strong&gt; The application server running mod_php. Apache processes PHP files, executes WordPress code, queries the MySQL database, and returns the rendered HTML response back through the chain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 4: WordPress.&lt;/strong&gt; The CMS itself. Themes, plugins, the block editor, the REST API, and WP-CLI all run here. WordPress lives in a directory on the filesystem, and Apache’s &lt;code&gt;DocumentRoot&lt;/code&gt; points to it.&lt;/p&gt;

&lt;p&gt;The proxy cache sits between Layer 1 and Layer 2. When Nginx has a cached response, it never contacts Coraza or Apache. The request is served entirely from Nginx’s disk cache. This detail is what saved the site.&lt;/p&gt;

&lt;p&gt;Free to use, share it in your presentations, blogs, or learning materials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzk2u77sztd4260kzt3by.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzk2u77sztd4260kzt3by.webp" alt="Four layer proxy architecture showing Browser to Nginx with cache to Coraza WAF to Apache WordPress, and how cache serves 200 while backend returns 404 during the incident" width="800" height="495"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The four-layer proxy chain: Browser to Nginx (with cache) to Coraza WAF to Apache/WordPress. During the incident, Nginx served cached 200s while the backend returned 404.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Why the Directory Move Was Necessary
&lt;/h2&gt;

&lt;p&gt;WordPress had been installed at the default Apache document root: a path inside &lt;code&gt;/var/www/html/&lt;/code&gt;. This is where most tutorials tell you to put it. The problem is that &lt;code&gt;/var/www/html/&lt;/code&gt; is the default document root for every Apache installation. If another virtual host is misconfigured, if a default site is accidentally enabled, or if a new service is added without careful vhost isolation, files under that path could be exposed. Security audits consistently flag this as a risk.&lt;/p&gt;

&lt;p&gt;The plan was simple: move WordPress from &lt;code&gt;/var/www/html/cms/&lt;/code&gt; to &lt;code&gt;/var/www/insights/&lt;/code&gt;, a dedicated directory outside the default root. The new path is explicit, isolated, and does not share a parent directory with any other web service. The &lt;a href="https://csrc.nist.gov/publications/detail/sp/800-123/final" rel="noopener noreferrer"&gt;NIST SP 800-123 Guide to General Server Security&lt;/a&gt; recommends isolating web application files from default server roots to reduce the attack surface from misconfiguration.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Migration Plan
&lt;/h2&gt;

&lt;p&gt;The engineer, Ravi, started by investigating every file on the server that referenced the old path. The goal was to identify every config, script, cron job, and binary that would need updating.&lt;/p&gt;

&lt;p&gt;Searching for hardcoded path references&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; ‘/var/www/html/cms’ /etc/apache2/ /etc/nginx/ /etc/systemd/ /etc/crontab
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; ‘/var/www/html/cms’ /root/.bashrc /home/&lt;span class="k"&gt;*&lt;/span&gt;/.bashrc
&lt;span class="nv"&gt;$ &lt;/span&gt;strings /usr/local/bin/site-backup | &lt;span class="nb"&gt;grep&lt;/span&gt; ‘/var/www’
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The search found references in three places: the Apache virtual host config (&lt;code&gt;DocumentRoot&lt;/code&gt; and &lt;code&gt;Directory&lt;/code&gt; directives), a separate Apache SSL vhost, and a Go binary used for automated backups. Nginx config was clean because it only proxies to Coraza’s internal port and never references the filesystem directly. Crontabs were clean. Shell profiles were clean. WordPress’s own &lt;code&gt;wp-config.php&lt;/code&gt; uses relative paths, so it did not need changes.&lt;/p&gt;

&lt;p&gt;With the investigation complete, Ravi laid out the steps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Copy the entire WordPress directory to the new location with permissions preserved&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;sed -i&lt;/code&gt; to replace the old path in all Apache config files&lt;/li&gt;
&lt;li&gt;Restart Apache&lt;/li&gt;
&lt;li&gt;Clear Nginx proxy cache&lt;/li&gt;
&lt;li&gt;Run the cache warmer to pre-warm all pages&lt;/li&gt;
&lt;li&gt;Verify with curl&lt;/li&gt;
&lt;li&gt;Delete the old directory&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Cache Warmer: 190 Pages in 5 Minutes
&lt;/h2&gt;

&lt;p&gt;Before the migration, the team had already deployed a cache warmer. It was a small Go binary that fetched the WordPress sitemap, extracted every URL, and made an HTTP GET request to each page using two concurrent workers. The purpose was to keep Nginx’s proxy cache warm after cache clears, deployments, or server restarts. Every 30 minutes, a cron job ran the warmer to ensure that no visitor ever hit an uncached page.&lt;/p&gt;

&lt;p&gt;The warmer’s output looked like this after a typical run:&lt;/p&gt;

&lt;p&gt;Cache warmer output after a full run&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;=== GSL Cache Warmer ===

Fetching sitemap: https://bagful.net/post-sitemap.xml

Found 190 URLs
[W1] 200 | MISS | TTFB: 245ms | OK | /zero-trust-architecture-explained/

[W2] 200 | MISS | TTFB: 189ms | OK | /nginx-proxy-cache-wordpress-setup/

[W1] 200 | MISS | TTFB: 312ms | OK | /coraza-waf-reverse-proxy-guide/

…
=== Summary ===

Total URLs: 190

Cache MISS: 190 (freshly cached)

Cache HIT: 0

Errors: 0

Time elapsed: 5m12s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every page showed &lt;code&gt;MISS&lt;/code&gt; (freshly cached from the backend). This meant Nginx now had a complete, warm cache of every page on the site. For the next 2 hours (the cache TTL), any visitor requesting any page would get a response directly from Nginx’s disk cache without touching Coraza, Apache, or WordPress.&lt;/p&gt;

&lt;p&gt;The cache warmer ran successfully at 15:30. The migration started at 15:45.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Migration: Copy, Sed, Restart
&lt;/h2&gt;

&lt;p&gt;Step one: copy the files.&lt;/p&gt;

&lt;p&gt;Copying WordPress to the new secure directory&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo cp&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /var/www/html/cms /var/www/insights
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo chown&lt;/span&gt; &lt;span class="nt"&gt;-R&lt;/span&gt; www-data:www-data /var/www/insights
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-la&lt;/span&gt; /var/www/insights/wp-config.php
&lt;span class="nt"&gt;-rw-r&lt;/span&gt;—– 1 www-data www-data 3858 Apr 6 13:54 /var/www/insights/wp-config.php

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;-a&lt;/code&gt; flag preserves ownership, permissions, and timestamps. The copy took about 90 seconds for the 190-post site with its media uploads.&lt;/p&gt;

&lt;p&gt;Step two: update Apache configs with sed.&lt;/p&gt;

&lt;p&gt;Replacing the old path in Apache configs&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo sed&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; ‘s|/var/www/html/cms|/var/www/insights|g’ /etc/apache2/sites-available/blog-ssl.conf
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo sed&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; ‘s|/var/www/html/cms|/var/www/insights|g’ /etc/apache2/sites-available/blog.conf
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep &lt;/span&gt;DocumentRoot /etc/apache2/sites-available/blog-ssl.conf /etc/apache2/sites-available/blog.conf
/etc/apache2/sites-available/blog-ssl.conf: DocumentRoot /var/www/insights

/etc/apache2/sites-available/blog.conf: DocumentRoot /var/www/insights

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ravi ran sed against the two config files he had found during the investigation. Apache config test passed. Apache restarted cleanly.&lt;/p&gt;

&lt;p&gt;Config test and restart&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apache2ctl configtest
Syntax OK

&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl restart apache2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Step three: clear the Nginx cache and test.&lt;/p&gt;

&lt;p&gt;Clear Nginx cache and quick test&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; /var/cache/nginx/&lt;span class="k"&gt;*&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl reload nginx
&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-sk&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; ‘%&lt;span class="o"&gt;{&lt;/span&gt;http_code&lt;span class="o"&gt;}&lt;/span&gt;’ https://bagful.net/
200

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The homepage returned 200. Ravi tested two more pages. Both returned 200. Everything looked good. He ran the cache warmer again to pre-warm the entire site from the new backend location. All 190 pages came back as 200 MISS (freshly cached).&lt;/p&gt;

&lt;p&gt;Step four: check for open files and delete the old directory.&lt;/p&gt;

&lt;p&gt;Check for open files and delete old directory&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;lsof +D /var/www/html/cms/ | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
0

&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; /var/www/html/cms
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; /var/www/html/cms 2&amp;gt;&amp;amp;1
&lt;span class="nb"&gt;ls&lt;/span&gt;: cannot access ‘/var/www/html/cms’: No such file or directory

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero open files. Old directory deleted. Migration complete. Or so Ravi thought.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Config That sed Missed
&lt;/h2&gt;

&lt;p&gt;Twenty minutes later, Ravi ran a routine health check from his laptop.&lt;/p&gt;

&lt;p&gt;Health check from external machine&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-sk&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; ‘%&lt;span class="o"&gt;{&lt;/span&gt;http_code&lt;span class="o"&gt;}&lt;/span&gt;’ https://bagful.net/
200
&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-sk&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; ‘%&lt;span class="o"&gt;{&lt;/span&gt;http_code&lt;span class="o"&gt;}&lt;/span&gt;’ https://bagful.net/wp-admin/
404

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The homepage returned 200. But wp-admin returned 404. That made no sense. WordPress was running. The files existed at the new path. Apache had restarted without errors.&lt;/p&gt;

&lt;p&gt;The homepage was 200 because Nginx was serving it from cache. The wp-admin path was 404 because wp-admin is never cached (the Nginx config skips caching for authenticated paths). So wp-admin hit the full proxy chain: Nginx &amp;gt; Coraza &amp;gt; Apache. And Apache was returning 404.&lt;/p&gt;

&lt;p&gt;Ravi tested each layer individually.&lt;/p&gt;

&lt;p&gt;Testing each layer of the proxy chain&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Test Apache directly&lt;/span&gt;

&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://127.0.0.1:8080/ &lt;span class="nt"&gt;-H&lt;/span&gt; ‘Host: bagful.net’ &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; ‘%&lt;span class="o"&gt;{&lt;/span&gt;http_code&lt;span class="o"&gt;}&lt;/span&gt;’
404
&lt;span class="c"&gt;# Test Coraza (which forwards to Apache)&lt;/span&gt;

&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://127.0.0.1:8082/ &lt;span class="nt"&gt;-H&lt;/span&gt; ‘Host: bagful.net’ &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; ‘%&lt;span class="o"&gt;{&lt;/span&gt;http_code&lt;span class="o"&gt;}&lt;/span&gt;’
404
&lt;span class="c"&gt;# Test Nginx (which serves from cache for public pages)&lt;/span&gt;

&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-sk&lt;/span&gt; https://bagful.net/ &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; ‘%&lt;span class="o"&gt;{&lt;/span&gt;http_code&lt;span class="o"&gt;}&lt;/span&gt;’
200 &lt;span class="o"&gt;(&lt;/span&gt;cached&lt;span class="o"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apache was returning 404 on every request. The cache was masking the problem for public pages. Ravi checked the Apache vhost configs he had updated with sed.&lt;/p&gt;

&lt;p&gt;Checking the Apache configs&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep &lt;/span&gt;DocumentRoot /etc/apache2/sites-available/blog-ssl.conf
DocumentRoot /var/www/insights ← correct
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep &lt;/span&gt;DocumentRoot /etc/apache2/sites-available/blog.conf
DocumentRoot /var/www/insights ← correct
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep &lt;/span&gt;DocumentRoot /etc/apache2/sites-enabled/blog-main.conf
DocumentRoot /var/www/html/cms ← OLD PATH!

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There it was. The &lt;code&gt;sites-enabled&lt;/code&gt; directory had a third config file: &lt;code&gt;blog-main.conf&lt;/code&gt;. This was the vhost listening on the internal port that Coraza forwards to. Ravi’s sed command had updated &lt;code&gt;blog-ssl.conf&lt;/code&gt; and &lt;code&gt;blog.conf&lt;/code&gt;, but missed &lt;code&gt;blog-main.conf&lt;/code&gt; because his investigation had found only two files in &lt;code&gt;sites-available&lt;/code&gt;. The third file existed only in &lt;code&gt;sites-enabled&lt;/code&gt; as a standalone file, not a symlink.&lt;/p&gt;

&lt;p&gt;The proxy chain was: Nginx (443) &amp;gt; Coraza (internal) &amp;gt; Apache (internal port using blog-main.conf). The vhost that actually served requests from the WAF was the one that sed missed. The two vhosts that sed updated were for direct Apache access and legacy SSL, neither of which was in the active request path.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix: One sed Command
&lt;/h2&gt;

&lt;p&gt;Fixing the missed config&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo sed&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; ‘s|/var/www/html/cms|/var/www/insights|g’ /etc/apache2/sites-enabled/blog-main.conf
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl restart apache2&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://127.0.0.1:8080/ &lt;span class="nt"&gt;-H&lt;/span&gt; ‘Host: bagful.net’ &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; ‘%&lt;span class="o"&gt;{&lt;/span&gt;http_code&lt;span class="o"&gt;}&lt;/span&gt;’
200

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One command. Apache now served from the correct directory. The entire proxy chain returned 200.&lt;/p&gt;

&lt;p&gt;Free to use, share it in your presentations, blogs, or learning materials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5uh02ozydjmxknr2o5k7.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5uh02ozydjmxknr2o5k7.webp" alt="Six step investigation flowchart tracing the 404 error through Nginx cache, wp-admin bypass, Apache direct test, Coraza WAF test, finding the missed config, and the one sed command fix" width="800" height="522"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The investigation traced through six steps: testing public URL (cached 200), testing wp-admin (uncached 404), testing Apache directly (404), testing Coraza (pass-through 404), finding the missed config file, and fixing it with one sed command.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  How the Cache Saved SEO: Minute by Minute
&lt;/h2&gt;

&lt;p&gt;Here is the timeline of what happened during the 30-minute window when Apache was broken:&lt;/p&gt;

&lt;p&gt;Free to use, share it in your presentations, blogs, or learning materials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq3o6g03tv5fwvqhqfvfr.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq3o6g03tv5fwvqhqfvfr.webp" alt="Incident timeline from cache warmer run at 15:30 through migration at 15:45, directory deletion at 15:53 causing silent 404, cache absorbing the failure, discovery at 16:20, and fix at 16:22" width="799" height="441"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Minute-by-minute timeline: cache warmer at 15:30, migration at 15:45, directory deleted at 15:53 (backend breaks silently), cache absorbs the failure for 27 minutes, 404 discovered at 16:20, fixed at 16:22.&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;15:30&lt;/strong&gt; Cache warmer runs. All 190 pages cached with 200 responses. Cache TTL: 120 minutes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;15:45&lt;/strong&gt; Migration starts. Files copied. sed runs against two config files. Apache restarts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;15:48&lt;/strong&gt; Nginx cache cleared. Cache warmer runs again. All 190 pages return 200 from the new backend. Cache is fresh&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;15:53&lt;/strong&gt; Old directory deleted. Apache’s internal vhost (blog-main.conf) still points to the deleted path. Apache starts returning 404 for every uncached request&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;15:53 to 16:20&lt;/strong&gt; The broken window. Every request that hits the backend returns 404. But Nginx serves all public pages from cache (200 HIT). Only wp-admin, wp-cron, and logged-in user requests see the 404&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;16:20&lt;/strong&gt; Ravi discovers the 404 on wp-admin. Traces through the proxy chain. Finds the missed config file&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;16:22&lt;/strong&gt; One sed command fixes the config. Apache restarts. Site fully operational&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ravi immediately checked the Nginx access logs for Google crawler activity during the broken window.&lt;/p&gt;

&lt;p&gt;Checking for Googlebot hits during the 404 window&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; ‘googlebot’ /var/log/nginx/access.log | &lt;span class="nb"&gt;grep&lt;/span&gt; ’15:5[3-9]&lt;span class="se"&gt;\|&lt;/span&gt;16:[01]’
&lt;span class="c"&gt;# (empty: zero Googlebot hits during the window)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero Google crawler hits during the entire 30-minute window. Google did not request a single page while the backend was broken. Even if it had, Nginx would have served the cached 200 response. The crawler would have seen a perfectly healthy site.&lt;/p&gt;

&lt;p&gt;The Nginx cache had a 120-minute TTL. The broken window was 30 minutes. The cache would have continued serving valid responses for another 90 minutes before expiring. Even without the fix, the site would have appeared healthy to every external visitor until the cache TTL expired at 17:48.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Cache Warmer Was the Real Hero
&lt;/h2&gt;

&lt;p&gt;The proxy cache alone would not have saved the site. Here is why. Nginx only caches a page after someone requests it. If the cache had been cold (empty) when the backend broke, the first visitor to each page would have received a 404. That 404 would have been cached (Nginx was configured with &lt;code&gt;proxy_cache_valid 404 1m&lt;/code&gt;) and served to every subsequent visitor for one minute. With 190 pages, it would have taken many visitors hitting 404s before anyone noticed.&lt;/p&gt;

&lt;p&gt;The cache warmer eliminated this scenario. By pre-fetching every page before the migration, it ensured that Nginx had a warm, complete cache of every URL on the site. When the backend broke, there was no “first visitor” problem. Every page was already cached with a valid 200 response.&lt;/p&gt;

&lt;p&gt;The warmer ran on a 30-minute cron schedule from a separate server:&lt;/p&gt;

&lt;p&gt;Cache warmer cron job&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Run cache warmer every 30 minutes&lt;/span&gt;

&lt;span class="k"&gt;*&lt;/span&gt;/30 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; /opt/cache-warmer/bin/warmer &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; /opt/cache-warmer/log/warmer.log 2&amp;gt;&amp;amp;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The combination of proxy cache (Nginx) and cache warmer (external cron) created an unintentional safety net. The cache absorbed the backend failure. The warmer ensured the cache was complete. Together, they bought the team 120 minutes of runway to find and fix the problem without any visitor or crawler seeing a broken page.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Nginx Proxy Cache Configuration
&lt;/h2&gt;

&lt;p&gt;For context, here is the Nginx proxy cache configuration that made this possible. The cache zone is defined in the &lt;code&gt;http&lt;/code&gt; block and used in the &lt;code&gt;server&lt;/code&gt; block.&lt;/p&gt;

&lt;p&gt;PHPNginx proxy cache configuration&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt; &lt;span class="k"&gt;1&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;
 &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="c1"&gt;# In http block: define the cache zone&amp;lt;br /&amp;gt;&lt;/span&gt;
 &lt;span class="s"&gt;3proxy_cache_path&lt;/span&gt; &lt;span class="n"&gt;/var/cache/nginx/blog&lt;/span&gt; &lt;span class="s"&gt;levels=1:2&lt;/span&gt; &lt;span class="s"&gt;keys_zone=blog_cache:10m&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;
 &lt;span class="mi"&gt;4&lt;/span&gt; &lt;span class="s"&gt;max_size=500m&lt;/span&gt; &lt;span class="s"&gt;inactive=120m&lt;/span&gt; &lt;span class="s"&gt;use_temp_path=off&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;&amp;lt;/p&amp;gt;&lt;/span&gt;
 &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;p&amp;gt;&lt;/span&gt;&lt;span class="c1"&gt;# In server block: use the cache&amp;lt;br /&amp;gt;&lt;/span&gt;
 &lt;span class="s"&gt;6location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="kn"&gt;&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;
 &lt;span class="mi"&gt;7&lt;/span&gt; &lt;span class="s"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://127.0.0.1:8082&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;# Coraza WAF&amp;lt;/p&amp;gt;&lt;/span&gt;
 &lt;span class="kn"&gt;8&amp;lt;p&amp;gt;&lt;/span&gt; &lt;span class="s"&gt;proxy_cache&lt;/span&gt; &lt;span class="s"&gt;blog_cache&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="kn"&gt;&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;
 &lt;span class="mi"&gt;9&lt;/span&gt; &lt;span class="s"&gt;proxy_cache_valid&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt; &lt;span class="mi"&gt;120m&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;# Cache 200 responses for 2 hours&amp;lt;br /&amp;gt;&lt;/span&gt;
&lt;span class="kn"&gt;10&lt;/span&gt; &lt;span class="s"&gt;proxy_cache_valid&lt;/span&gt; &lt;span class="mi"&gt;301&lt;/span&gt; &lt;span class="mi"&gt;302&lt;/span&gt; &lt;span class="mi"&gt;10m&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;# Cache redirects for 10 minutes&amp;lt;br /&amp;gt;&lt;/span&gt;
&lt;span class="kn"&gt;11&lt;/span&gt; &lt;span class="s"&gt;proxy_cache_valid&lt;/span&gt; &lt;span class="mi"&gt;404&lt;/span&gt; &lt;span class="mi"&gt;1m&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;# Cache 404s for only 1 minute&amp;lt;br /&amp;gt;&lt;/span&gt;
&lt;span class="kn"&gt;12&lt;/span&gt; &lt;span class="s"&gt;proxy_cache_bypass&lt;/span&gt; &lt;span class="nv"&gt;$skip_cache&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="kn"&gt;&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="mi"&gt;13&lt;/span&gt; &lt;span class="s"&gt;proxy_no_cache&lt;/span&gt; &lt;span class="nv"&gt;$skip_cache&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="kn"&gt;&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="mi"&gt;14&lt;/span&gt; &lt;span class="s"&gt;proxy_cache_use_stale&lt;/span&gt; &lt;span class="s"&gt;error&lt;/span&gt; &lt;span class="s"&gt;timeout&lt;/span&gt; &lt;span class="s"&gt;updating&lt;/span&gt; &lt;span class="s"&gt;http_500&lt;/span&gt; &lt;span class="s"&gt;http_502&lt;/span&gt; &lt;span class="s"&gt;http_503&lt;/span&gt; &lt;span class="s"&gt;http_504&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="kn"&gt;&amp;lt;/p&amp;gt;&lt;/span&gt;
&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;p&amp;gt;&lt;/span&gt; &lt;span class="s"&gt;add_header&lt;/span&gt; &lt;span class="s"&gt;X-Cache-Status&lt;/span&gt; &lt;span class="nv"&gt;$upstream_cache_status&lt;/span&gt; &lt;span class="s"&gt;always&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="kn"&gt;&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="err"&gt;}&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;br&lt;/span&gt; &lt;span class="n"&gt;/&amp;gt;&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two settings were critical during the incident:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;proxy_cache_valid 200 120m&lt;/code&gt;&lt;/strong&gt; meant every cached 200 response stayed valid for 2 hours. The backend could be completely down for 2 hours and visitors would still see the site&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;proxy_cache_use_stale error timeout http_500 http_502 http_503 http_504&lt;/code&gt;&lt;/strong&gt; tells Nginx to serve stale (expired) cache if the backend returns an error. This did not activate during the incident because the cache was still fresh, but it would have provided an additional safety net if the outage had lasted longer than the TTL&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;$skip_cache&lt;/code&gt; variable bypasses caching for logged-in users, wp-admin, wp-cron, and POST requests. This is why Ravi saw the 404 on wp-admin but not on public pages. wp-admin always hits the backend.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Coraza WAF in the Middle
&lt;/h2&gt;

&lt;p&gt;The team chose &lt;a href="https://coraza.io/" rel="noopener noreferrer"&gt;Coraza WAF&lt;/a&gt; over ModSecurity for a specific reason: memory footprint. The blog server runs on a modest VPS with limited RAM. ModSecurity 3.x (libmodsecurity) loads as an Apache or Nginx module, sharing the web server’s memory space and adding significant overhead per worker process. Coraza is written in Go and runs as a standalone reverse proxy with its own memory management. In testing, the team measured Coraza using 50MB of RAM with the full &lt;a href="https://coreruleset.org/" rel="noopener noreferrer"&gt;OWASP Core Rule Set&lt;/a&gt; loaded, compared to 120-180MB per Apache worker process with ModSecurity embedded. On a memory-constrained server running WordPress with multiple Apache prefork workers, that difference is the gap between stable operation and OOM kills.&lt;/p&gt;

&lt;p&gt;Coraza is also a modern, actively maintained project. ModSecurity’s development slowed significantly after Trustwave transferred it to OWASP in 2024. Coraza is fully compatible with the OWASP CRS rules, supports the SecLang rule language, and can be deployed without compiling into Nginx or Apache. If you are looking to implement Coraza in your own stack, we have articles that walk through the full setup: &lt;a href="https://blogs.getsetlive.com/building-the-coraza-nginx-waf-connector-on-ubuntu-24-part-1-architecture-and-prerequisites/" rel="noopener noreferrer"&gt;Part 1: Architecture and Prerequisites&lt;/a&gt;, &lt;a href="https://blogs.getsetlive.com/building-the-coraza-nginx-waf-connector-on-ubuntu-24-part-2-compiling-testing-and-findings/" rel="noopener noreferrer"&gt;Part 2: Compiling, Testing, and Findings&lt;/a&gt;, and &lt;a href="https://blogs.getsetlive.com/writing-custom-coraza-waf-rules-for-php-and-wordpress-protection/" rel="noopener noreferrer"&gt;Writing Custom Coraza WAF Rules for PHP and WordPress&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The presence of Coraza in the proxy chain added a layer of complexity that contributed to the missed config. In a simple Nginx &amp;gt; Apache setup, there are only two config files: the Nginx server block and the Apache vhost. With Coraza in between, Apache needs two vhosts: one for the WAF’s internal forwarding port and one for direct/legacy access. Coraza acted as a transparent proxy, forwarding clean requests from Nginx to Apache’s internal vhost.&lt;/p&gt;

&lt;p&gt;Ravi’s investigation found the Apache configs in &lt;code&gt;sites-available&lt;/code&gt; but missed the one in &lt;code&gt;sites-enabled&lt;/code&gt; that was not a symlink. This is a common Apache pattern: most files in &lt;code&gt;sites-enabled&lt;/code&gt; are symlinks to &lt;code&gt;sites-available&lt;/code&gt;, but standalone files can exist there too. A more thorough investigation would have searched both directories independently.&lt;/p&gt;

&lt;p&gt;The investigation command that would have caught it&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Search sites-available AND sites-enabled independently&lt;/span&gt;

&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; ‘/var/www/html/cms’ /etc/apache2/sites-available/ /etc/apache2/sites-enabled/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That single grep against both directories would have revealed the third config file. The original investigation only searched &lt;code&gt;/etc/apache2/&lt;/code&gt; recursively, which should have found it. But the file was found, counted, and then sed was run against only the two files in &lt;code&gt;sites-available&lt;/code&gt;, not the standalone file in &lt;code&gt;sites-enabled&lt;/code&gt;. A manual oversight in a scripted process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Google Search Console Check
&lt;/h2&gt;

&lt;p&gt;The team checked Google Search Console the next morning. The Coverage report showed zero new 404 errors. The crawl stats showed no crawl activity during the 30-minute window. The site’s impressions continued their upward trend without a dip.&lt;/p&gt;

&lt;p&gt;This was not just luck. Google’s crawler does not visit small sites continuously. For a blog with 190 posts and moderate authority, Googlebot typically crawls a few pages per hour, not per minute. The probability of a crawler hit during any specific 30-minute window is low. Combined with the proxy cache serving valid 200 responses, the risk of a Google-visible 404 was near zero.&lt;/p&gt;

&lt;p&gt;However, if the outage had lasted longer (past the 2-hour cache TTL), the risk profile changes dramatically. Once the cache expires, Nginx would request a fresh copy from the backend, receive the 404, and cache that 404 for 1 minute. Googlebot could then see a 404. Google does not immediately deindex a page that returns 404. It retries over several days. But if the 404 persists, the page is dropped from the index within 1-2 weeks. For a new blog still building authority, losing indexed pages would have been devastating.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;p&gt;The team documented five takeaways from the incident.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Always test the full proxy chain, not just the endpoint.&lt;/strong&gt; Ravi tested &lt;code&gt;curl https://bagful.net/&lt;/code&gt; and got 200. That tested Nginx (which served from cache). He should have also tested Apache directly: &lt;code&gt;curl http://127.0.0.1:8080/ -H 'Host: bagful.net'&lt;/code&gt;. Testing every hop in the proxy chain would have revealed the 404 immediately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Search all config directories, not just the expected ones.&lt;/strong&gt; Apache’s &lt;code&gt;sites-enabled&lt;/code&gt; can contain standalone files that are not symlinks to &lt;code&gt;sites-available&lt;/code&gt;. Always search both. Better yet, search the entire &lt;code&gt;/etc/apache2/&lt;/code&gt; tree and process every match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. A cache warmer is not just a performance tool. It is a resilience layer.&lt;/strong&gt; The warmer ensured that every page was cached before the migration. Without it, the first visitor to each page during the outage would have seen a 404.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Set &lt;code&gt;proxy_cache_use_stale&lt;/code&gt; aggressively.&lt;/strong&gt; The &lt;code&gt;error timeout http_500 http_502 http_503 http_504&lt;/code&gt; directive tells Nginx to serve stale cache when the backend fails. Consider adding &lt;code&gt;http_404&lt;/code&gt; to this list for content sites where the URLs are stable. A stale cached page is always better than a 404 for a page that existed 5 minutes ago.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Do not delete the old directory until you have verified the backend independently.&lt;/strong&gt; Ravi deleted the old directory after seeing 200 responses from the public URL. Those responses came from cache, not from the backend. The old directory should have stayed until a direct backend test confirmed the new path was serving correctly.&lt;/p&gt;




&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Nginx, &lt;a href="https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_cache" rel="noopener noreferrer"&gt;ngx_http_proxy_module: proxy_cache documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Nginx, &lt;a href="https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_cache_use_stale" rel="noopener noreferrer"&gt;proxy_cache_use_stale: serving stale content during errors&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Coraza WAF, &lt;a href="https://coraza.io/docs/tutorials/introduction/" rel="noopener noreferrer"&gt;Introduction and Setup Guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NIST, &lt;a href="https://csrc.nist.gov/publications/detail/sp/800-123/final" rel="noopener noreferrer"&gt;SP 800-123: Guide to General Server Security&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google Search Central, &lt;a href="https://developers.google.com/search/docs/crawling-indexing/http-network-errors" rel="noopener noreferrer"&gt;HTTP status codes and network errors for Google Search&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Apache, &lt;a href="https://httpd.apache.org/docs/2.4/vhosts/" rel="noopener noreferrer"&gt;Apache Virtual Host documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OWASP, &lt;a href="https://coreruleset.org/" rel="noopener noreferrer"&gt;Core Rule Set (CRS) for ModSecurity/Coraza&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can Nginx proxy cache prevent SEO damage during downtime?&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;Yes. Nginx proxy cache serves stored HTML responses without contacting the backend. If the backend is down or returning errors, Nginx continues serving cached 200 responses to visitors and crawlers. With proxy_cache_use_stale configured, Nginx can even serve expired cache during errors. Combined with a cache warmer that pre-fetches all pages, the cache acts as a complete safety net. Google crawlers see 200 responses and the site’s search rankings are unaffected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is a cache warmer and why does WordPress need one?&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;A cache warmer is a script or binary that pre-fetches every page on your site so the proxy cache is fully populated. Without it, the first visitor to each page after a cache clear hits the backend directly (cache MISS). For WordPress sites behind Nginx proxy cache, a warmer ensures every page is served from cache instantly. It also creates a resilience layer: if the backend breaks, all pages are already cached with valid responses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is it safe to move WordPress to a different directory?&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;Yes, but you must update every config that references the old path: Apache DocumentRoot and Directory directives, any standalone vhost files in sites-enabled, backup scripts, cron jobs, and compiled binaries with hardcoded paths. Search both sites-available and sites-enabled independently. Test the backend directly (not through a proxy cache) after making changes. Do not delete the old directory until you confirm the backend serves correctly from the new location.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does proxy_cache_use_stale do in Nginx?&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;The proxy_cache_use_stale directive tells Nginx to serve expired (stale) cached content when the backend returns specific errors. For example, proxy_cache_use_stale error timeout http_500 http_502 http_503 http_504 serves stale cache when the backend is unreachable or returns server errors. This prevents visitors from seeing error pages during backend outages. For content sites, consider adding http_404 to serve stale cache instead of 404 errors for previously valid pages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long before Google deindexes a page that returns 404?&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;Google does not immediately deindex a page that returns 404. It retries over several days, typically 3-5 crawl attempts spread over 1-2 weeks. If the 404 persists across all retries, the page is dropped from the index. A brief 404 (minutes to hours) followed by recovery has virtually zero impact on indexing. However, for new sites still building authority, even a temporary loss of indexed pages can slow organic growth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is Coraza WAF and why put it between Nginx and Apache?&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;Coraza is an open-source web application firewall compatible with the OWASP Core Rule Set. Placing it between Nginx and Apache creates a defense-in-depth architecture: Nginx handles TLS and caching, Coraza inspects request payloads for attacks (SQL injection, XSS, command injection), and Apache runs WordPress. This separation means each layer can be updated, scaled, and secured independently. Coraza adds 2-5ms of latency per request for rule evaluation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why should WordPress not be in the default Apache document root?&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;The default Apache document root (/var/www/html/) is shared by default configurations and potentially other virtual hosts. If another service or a default Apache config is accidentally enabled, files in that directory could be exposed. Moving WordPress to a dedicated directory (like /var/www/insights/) isolates it from other services. NIST SP 800-123 recommends isolating web application files from default server roots to reduce misconfiguration attack surface.&lt;/p&gt;

&lt;p&gt;The post &lt;a href="https://blogs.getsetlive.com/we-broke-wordpress-for-30-minutes-nginx-cache-kept-google-from-noticing/" rel="noopener noreferrer"&gt;We Broke WordPress for 30 Minutes. Nginx Cache Kept Google From Noticing.&lt;/a&gt; appeared first on &lt;a href="https://blogs.getsetlive.com" rel="noopener noreferrer"&gt;GetSetLive Blogs&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>wordpress</category>
      <category>nginx</category>
      <category>devops</category>
      <category>webdev</category>
    </item>
    <item>
      <title>What AI Actually Does in Bank Fraud Detection (Part 1): The Architecture</title>
      <dc:creator>Shrikant Mhatre (श्री)</dc:creator>
      <pubDate>Fri, 09 Oct 2026 07:57:21 +0000</pubDate>
      <link>https://dev.to/shrikant_mhatre_0900/what-ai-actually-does-in-bank-fraud-detection-part-1-the-architecture-kpg</link>
      <guid>https://dev.to/shrikant_mhatre_0900/what-ai-actually-does-in-bank-fraud-detection-part-1-the-architecture-kpg</guid>
      <description>&lt;p&gt;A walk through the real-time fraud detection stack we rebuilt with a US mid-size bank: Kafka ingest, Flink features, a three-model ensemble, the analyst’s alert queue, which model does which job, and why the card decline you hate at the grocery store still happens.&lt;/p&gt;

&lt;p&gt;Every fraud vendor who walks into a bank says the same thing: AI will solve your fraud problem. I spent 16 months on a fraud-platform rebuild with a mid-size US bank (4.2 million retail accounts, about 480,000 small-business accounts) whose head of fraud operations had heard that pitch a dozen times and did not believe a word of it. The brief was simple. Build the thing, then write down what AI does and what two humans and a phone still do. This is the first half, the architecture. How a card swipe becomes a fraud score and a decision in under 80 milliseconds, which models sit where, and what holds them together. No magic. Part 2 covers where it broke.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe802dxg6niffz949htrb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe802dxg6niffz949htrb.png" alt="Shri" width="96" height="96"&gt;&lt;/a&gt;&lt;br&gt;
By &lt;a href="https://dev.to/shri/"&gt;Shri&lt;/a&gt; Director of Technology at GetSetLive and Bagful Cloud Hosting, with 25 years across infrastructure, datacenters, cloud platforms, and modern AI/ML. Day-to-day work focuses on consulting clients to make their existing systems AI-aware and AI-ready, and architecting AI/ML solutions tailored to each customer environment. Every article here comes from direct field exposure, written to help the next batch of professionals design and deploy more sophisticated solutions in their own careers. &lt;a href="https://blogs.getsetlive.com/shri/" rel="noopener noreferrer"&gt;Read all posts by Shri →&lt;/a&gt;&lt;br&gt;
Editor, GetSetLive Technical Blogs17 min read&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;AI does not stop fraud. It cut the bank’s alert haystack from 18,000 a day to about 400, and rules plus analysts decide what happens to those 400.&lt;/li&gt;
&lt;li&gt;The stack is an ensemble of three narrow models (XGBoost, an autoencoder, GraphSAGE) behind a rules engine, not one big model; median swipe-to-decision is 72 ms.&lt;/li&gt;
&lt;li&gt;Step-up authentication, the “confirm this purchase” push, only exists because the model outputs a probability. It is the single biggest customer-experience win.&lt;/li&gt;
&lt;li&gt;The feedback loop from analyst labels back into monthly retraining is the part vendors never demo, and the part that keeps the model alive.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Is Inside
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The sentence in every vendor deck&lt;/li&gt;
&lt;li&gt;What the bank ran before the rebuild&lt;/li&gt;
&lt;li&gt;What the model watches that rules could not&lt;/li&gt;
&lt;li&gt;Anatomy of the real-time pipeline&lt;/li&gt;
&lt;li&gt;Which model does which job&lt;/li&gt;
&lt;li&gt;One card swipe, end to end&lt;/li&gt;
&lt;li&gt;A word on SARs and AML&lt;/li&gt;
&lt;li&gt;What comes next&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Sentence in Every Vendor Deck
&lt;/h2&gt;

&lt;p&gt;“AI stops fraud.” I have seen that line on roughly 70% of the fraud vendor slides I have sat through, and it is wrong in a specific, interesting way. AI does not stop fraud. What it did at the bank was shrink the alert haystack from something a human team cannot review (18,000 alerts a day before the rebuild) to something they can (around 400 a day now). Three things together stop the fraud itself: the model’s score, a rule that decides what to do with that score, and, in a small but critical slice of cases, a phone call from an analyst to the customer.&lt;/p&gt;

&lt;p&gt;On my first day the head of fraud operations put it to me in a way I have quoted many times since. A model is a filter. It does not decide anything. The decision is a rule: if the score is above 0.83 and the merchant category is one of 14 high-risk ones, block the transaction and send a push notification. The model knows none of that. It returns a number. Everything around the number is policy, and policy is a person figuring out the bank’s risk appetite for the quarter. I have heard a version of that speech from every fraud operations lead I have worked with, and it is always the truest thing they say all day.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Bank Ran Before the Rebuild
&lt;/h2&gt;

&lt;p&gt;Until late 2023 the bank ran fraud detection the way most regional banks still do. A rules engine, written in a mix of old SAS code and a newer Drools layer, scored every transaction against roughly 240 rules. The rules were the classics. A transaction above 500 dollars in a country the customer has never transacted in. Three card-not-present transactions inside four minutes. A velocity breach on a specific merchant category code. Each rule produced a yes or no. If any rule fired, the transaction was flagged; if enough fired, it was blocked. The fraud team wrote those rules over a decade, almost always after an incident. Someone got defrauded, a rule went in. Someone complained, a rule got tuned.&lt;/p&gt;

&lt;p&gt;The system caught fraud. It also produced around 18,000 alerts per business day, and about 12,000 of those needed an analyst. Fourteen analysts cannot review 12,000 alerts a day, so in practice they reviewed none of them properly. The floor had 14 analysts on two shifts, so each analyst was expected to dispose of roughly 850 alerts per shift. That is 35 an hour, under 90 seconds each. Nobody investigates fraud in 90 seconds. What happened instead, and I watched it during our discovery weeks, is that analysts built reflexes. If it looks like the 300 alerts already closed today that turned out to be nothing, close it. The false positive rate on blocked transactions sat around 68%. Real customers got declined at the grocery store while a script somewhere burned through stolen card numbers unseen.&lt;/p&gt;

&lt;p&gt;Below is the pre-AI stack as we mapped it in the first month. The red stops are where volume, accuracy, or plain blindness was costing the bank money and customers.&lt;/p&gt;

&lt;p&gt;Free to use, share it in your presentations, blogs, or learning materials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lfsq2xenmk3kuzjib7p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lfsq2xenmk3kuzjib7p.png" alt="AI fraud detection banking legacy rule-based system block diagram with five failure points" width="800" height="435"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The pre-AI fraud stack at the bank. Red stops mark where volume, accuracy, or blind spots cost money and customers.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The failures were structural. Better rules would not have fixed them. A rule cannot see beyond a single transaction. There was no memory of “this account has been quiet for 14 months and just lit up”, because that needs a windowed feature across time and the rules engine had no compute for it. There was no view of mule accounts, because mule detection needs a graph and rules do not do graphs. Encrypted mobile banking sessions were a black box. First-party fraud, where the customer is the fraudster, is about 11% of losses and was almost invisible, because every rule assumed the customer was the victim.&lt;/p&gt;

&lt;p&gt;And then there were the SARs. A Suspicious Activity Report is a regulatory filing a bank must make when it suspects money laundering. Each one took an analyst two to three hours to write. A senior analyst might finish two in a day. That number came up in every planning meeting we had, and it turned out to be the easiest thing in the whole programme to fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Model Watches That Rules Could Not
&lt;/h2&gt;

&lt;p&gt;The rebuild took the bank’s data engineering team and us 16 months, four months longer than the plan. Twelve of those months went into data plumbing, not models. On day two I asked the data engineering lead to name the single thing a model does that a rule cannot. No hesitation: context. A rule sees one transaction on its own. A model sees it against the last 10 minutes, the last 30 days, and the links that account has to other accounts. Same data, joined differently, and the joining is the whole game. Rules are a lookup table. Models are a function that takes a vector.&lt;/p&gt;

&lt;p&gt;Here is what the model now watches, roughly in order of how much each one moved our numbers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Windowed transaction features.&lt;/strong&gt; Count in the last 10 minutes, average amount over 30 days, distinct merchant categories in 24 hours, ratio of the current amount to the 90-day average. Flink computes all of these continuously from the Kafka stream.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Device fingerprint and behavioural biometrics.&lt;/strong&gt; Typing cadence on mobile logins, swipe pressure, scroll pattern, and the JA4 TLS fingerprint of the app session, through a BioCatch integration. These tell the real account holder apart from someone who merely has the password.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Graph proximity.&lt;/strong&gt; Is this account within two hops of a known mule? Shared device, shared phone number, shared beneficiary? The graph lives in Neo4j and a PyTorch Geometric GraphSAGE model queries it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Merchant embedding similarity.&lt;/strong&gt; Every merchant the customer has used gets a 64-dimensional embedding. A transaction at a merchant cluster unlike anything this customer has touched is a weak signal on its own, and a useful one in combination.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-channel signals.&lt;/strong&gt; A physical ATM withdrawal in Chicago at 14:32 while the mobile app is active in Atlanta at 14:30 is a red flag even when each event alone looks fine.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are rules. They are features. A rule says “block if X”. A feature is an input to a model that says “the probability this is fraud is 0.73, and here are the six inputs that drove it”. The difference matters because a feature-based system combines weak signals into a strong one. Three weak anomalies on a rules engine are three rules that did not fire. Three weak anomalies on a gradient-boosted tree are a 0.91 fraud score. That, more than any single model choice, is what the 16 months bought.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anatomy of the Real-Time Pipeline
&lt;/h2&gt;

&lt;p&gt;The architecture drawing on the data engineering team’s wall covers most of the wall. The diagram below is the simplified version. It reads left to right: events in on the far left, a decision out on the far right. Median time from swipe to decision is 72 milliseconds, with a 99th percentile of 118. We cannot afford to be slower, because the card networks give an issuer roughly two seconds for the whole authorisation and Visa and Mastercard use a good chunk of that themselves.&lt;/p&gt;

&lt;p&gt;Free to use, share it in your presentations, blogs, or learning materials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7907ar1e8okx0ewlrhaf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7907ar1e8okx0ewlrhaf.png" alt="AI fraud detection banking architecture Kafka Flink Tecton XGBoost GraphSAGE inference pipeline" width="800" height="609"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Six lanes, thirteen components. Median card-swipe-to-decision time is 72 milliseconds.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Lane 1: Event Ingest
&lt;/h3&gt;

&lt;p&gt;Everything the model will ever see enters through Apache Kafka. The bank runs a Confluent Cloud cluster with five topics: card-auth (Visa and Mastercard authorisation messages), ach-events (ACH pushes and pulls), wire-events (Fedwire and SWIFT), mobile-sessions (app login and in-app behaviour), and atm-events. Peak is about 14,000 events per second, and Black Friday 2024 pushed it to 19,000 for six hours, which is the day we learned our partition count was two too low. Every event has a schema in Confluent Schema Registry. Each downstream consumer knows exactly which fields to expect, and a producer cannot quietly change one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lane 2: Stream Processing and Feature Engineering
&lt;/h3&gt;

&lt;p&gt;Apache Flink is the workhorse. A Flink job enriches every event on Kafka with windowed aggregates computed on the fly: events in the last 10 minutes, card-not-present amount in the last 24 hours, distinct IP addresses in the last 7 days, and about 80 other features held in per-account keyed state. The state lives in RocksDB on the Flink task managers, so lookups are local and fast. When a transaction arrives, Flink computes the features and attaches them to the event. It then writes the enriched record to a Kafka topic for model serving and to the online feature store.&lt;/p&gt;

&lt;p&gt;This lane is where our worst early incident lived. In month seven a checkpoint failure during a deploy replayed 40 minutes of the card-auth topic, and for about 25 minutes every “count in last 10 minutes” feature was roughly double. Scores drifted up. The step-up rate tripled. The care queue noticed before our dashboards did. Our fix was idempotent state updates keyed on the authorisation ID, plus an alert on feature distribution shift, which Part 2 comes back to. Nothing about it was a modelling problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lane 3: Feature Store
&lt;/h3&gt;

&lt;p&gt;The bank uses Tecton, the commercial feature store built by people from the Uber Michelangelo team. Its job is the one every feature store has: the features the model learned on must be the exact same features available at serving time. Online lookups hit Redis in under a millisecond. Historical features go to Snowflake for training. Every feature definition is versioned, so a feature can change or retire without breaking the model in silence. This sounds boring. It is boring. It is also the thing that separates a working ML system from a flaky demo, and I would pick it over any model upgrade.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lane 4: Model Inference
&lt;/h3&gt;

&lt;p&gt;This is where the models live, and there are three of them in the transaction path. The primary is an XGBoost classifier that takes around 240 features and outputs a fraud probability between 0 and 1. Beside it runs an autoencoder for unsupervised anomaly detection, and a GraphSAGE graph neural network that scores how close the transaction sits to known fraud in the account graph. A small logistic meta-model combines the three scores into one number. The ensemble runs inside a FastAPI container per region, with TorchServe handling the GraphSAGE inference and XGBoost serving itself. Median inference latency across all three is 9 milliseconds. Twelve Kubernetes pods sit behind an internal load balancer and scale on request rate.&lt;/p&gt;

&lt;p&gt;We argued for a month about whether the autoencoder earned its place. It adds latency and it is hard to explain to a regulator. It stayed because in the shadow period it flagged a card-testing pattern (hundreds of one-dollar authorisations across fresh merchant IDs) two weeks before the supervised model had any labelled examples of it. That is the whole case for an unsupervised model in the ensemble: it sees shapes before anyone has named them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lane 5: Decision Engine
&lt;/h3&gt;

&lt;p&gt;The ensemble score, plus a dozen hard rules that still exist (sanctions list hit, card reported stolen, customer on fraud hold), flow into a Drools decision engine. Drools produces one of five outcomes: allow, allow with customer notification, step-up authentication (an OTP challenge or a biometric check in the app), block and notify, or block and lock the card. The fraud operations team sets the thresholds between those outcomes and reviews them monthly; the model has no say in them. When the retail fraud loss ratio drifts above 0.04% of swipe volume, thresholds tighten. When complaints about declines spike, they loosen. It is a permanent trade between fraud loss and customer experience. Nobody solves it; they manage it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lane 6: Case Management and Feedback
&lt;/h3&gt;

&lt;p&gt;Alerts that need a human land in Actimize, the case management system. Each case carries the transaction, the features, the top 10 SHAP values from XGBoost, the graph neighbourhood, and a one-paragraph summary written by an LLM. Analysts close the case with one of six labels, including “confirmed fraud” and “confirmed legitimate”. Those labels flow back through Kafka into the training set. The monthly retrain on the last 90 days of labelled data then learns from last month’s analyst decisions. This feedback loop is the most important piece of the whole system. Without it the model goes stale within weeks, because the people on the other side adapt faster than any vendor’s release cycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Model Does Which Job
&lt;/h2&gt;

&lt;p&gt;Seven distinct models run across the bank’s fraud and AML stack, three of them in the real-time path. Every choice was deliberate. The data engineering lead had to justify each one to the model risk committee in writing, and that exercise killed two models we liked. The diagram shows the mapping; the table underneath is the short version of those justifications.&lt;/p&gt;

&lt;p&gt;Free to use, share it in your presentations, blogs, or learning materials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ji8xialb6ofjzj91gfd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ji8xialb6ofjzj91gfd.png" alt="AI fraud detection banking model map XGBoost GraphSAGE autoencoder LSTM LLM random forest isolation forest" width="800" height="479"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Seven jobs, seven model families. Each matched to the shape of the problem, not picked for novelty.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Function&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Why this one&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary fraud score at transaction time&lt;/td&gt;
&lt;td&gt;XGBoost&lt;/td&gt;
&lt;td&gt;Fast tabular inference, handles missing values, SHAP-explainable for regulator review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unsupervised anomaly (zero-day fraud)&lt;/td&gt;
&lt;td&gt;Autoencoder&lt;/td&gt;
&lt;td&gt;No labels needed; catches patterns the supervised model has not seen yet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mule ring and collusion detection&lt;/td&gt;
&lt;td&gt;GraphSAGE (GNN)&lt;/td&gt;
&lt;td&gt;Tabular models cannot see account-to-account relationships; the GNN traverses the neighbourhood&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Behavioural biometrics on mobile&lt;/td&gt;
&lt;td&gt;LSTM on keystroke and swipe sequences&lt;/td&gt;
&lt;td&gt;Biometrics are time series; the LSTM captures cadence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Device fingerprinting&lt;/td&gt;
&lt;td&gt;Random Forest on JA4/TLS features&lt;/td&gt;
&lt;td&gt;Hard to evade, interpretable, cheap at scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Synthetic identity detection&lt;/td&gt;
&lt;td&gt;Entity resolution + graph embeddings&lt;/td&gt;
&lt;td&gt;Synthetic identities are stitched across sources; the graph exposes the stitching&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SAR narrative drafting&lt;/td&gt;
&lt;td&gt;Fine-tuned LLM (human-reviewed)&lt;/td&gt;
&lt;td&gt;Only model that writes regulator-grade prose; kept out of every automatic decision&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nothing in that table is a foundation model for fraud. There is no GPT-fraud. What there is, is a set of small specialist models, each picked because its assumptions match the shape of one problem. I have now seen the same pattern at a lender, an insurer, and a bank: narrow models do narrow jobs well, and one big model does nothing well.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Card Swipe, End to End
&lt;/h2&gt;

&lt;p&gt;The best way to feel how the pipeline holds up is to follow one transaction. The example below comes from the bank’s staging replay logs, altered to protect the customer but otherwise real. A 41-year-old debit card holder who lives near the bank’s home city is at an electronics store in Charlotte on a Saturday afternoon, buying a 1,200 dollar laptop. That is unusual for this customer on three counts: never shopped in Charlotte, never at that retailer, and the amount is three standard deviations above his 90-day average.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;T+0 ms:&lt;/strong&gt; The store terminal swipes the card. The authorisation request goes to Mastercard, which forwards it to the bank over ISO 8583.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+4 ms:&lt;/strong&gt; The bank’s payment gateway normalises the authorisation into a JSON event and writes it to the card-auth topic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+9 ms:&lt;/strong&gt; Flink picks up the event and pulls the customer’s 80 windowed features from Tecton. Transactions in the last 10 minutes: 0. Thirty-day average amount: 47 dollars. Distinct cities in 30 days: 1. Features enriched and attached.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+14 ms:&lt;/strong&gt; The enriched event lands on the inference topic and a FastAPI worker reads it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+23 ms:&lt;/strong&gt; XGBoost returns 0.71 (somewhat suspicious). The autoencoder returns 0.68 (high reconstruction error). GraphSAGE returns 0.12 (no fraud-graph proximity). The meta-model combines them to 0.69.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+28 ms:&lt;/strong&gt; A decision rule fires: score between 0.60 and 0.80, plus high amount, plus new merchant city, equals step-up authentication. Response: challenge with a mobile push.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+32 ms:&lt;/strong&gt; The bank returns a conditional response to Mastercard: approve with 3-D Secure step-up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+600 ms:&lt;/strong&gt; The customer’s phone buzzes: “Confirm $1,200 at an electronics store in Charlotte?” He taps yes; Face ID confirms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T+1,840 ms:&lt;/strong&gt; The bank sends final approval to Mastercard. The terminal beeps. He walks out with the laptop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The inference took 9 milliseconds and the decision was out 28 milliseconds after the swipe hit the gateway. The remaining 1.8 seconds was a person reacting to a push notification. Under the old rules engine this purchase would have been either fully approved (no rule caught “first time in Charlotte”) or fully blocked (a velocity rule might have fired on the amount). The middle path, where the customer is asked rather than refused, did not exist in the rules world. The model’s probability is what unlocks it.&lt;/p&gt;

&lt;p&gt;Step-up authentication is the unsung hero of modern fraud detection. It is invisible to the customer who is real and fatal to the attacker who is not. On the bank’s numbers it now resolves about 62% of ambiguous transactions without an analyst and without a decline, which is the single line on the board deck that made the whole programme worth it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Word on SARs and AML
&lt;/h2&gt;

&lt;p&gt;The card pipeline is the fast, real-time side. The anti-money-laundering side is slower and more regulated. In the US a Suspicious Activity Report has to be filed with FinCEN within 30 days of detection. Before the rebuild the bank filed around 3,800 SARs a year. Each one took a senior analyst two to three hours to draft. We added a fine-tuned LLM that reads the case file, the relevant transaction history, and the analyst’s preliminary notes, and produces a first draft of the narrative. An analyst reads it, corrects it, and signs it. That draft saves about 90 minutes per SAR. Today the bank files around 4,600 a year (detection got better, not worse) and the per-SAR effort is down to roughly 45 minutes of senior analyst time.&lt;/p&gt;

&lt;p&gt;The LLM is explicitly not in the decision path. It does not decide whether a SAR gets filed. It writes a draft; a human decides. Every SAR is signed by a named compliance officer who carries personal regulatory exposure if it is wrong. That is the pattern FinCEN and the OCC have said they expect, and it is the only pattern our compliance head would sign off on. For the first month the compliance head read every single draft against the source case before trusting a paragraph of it, and I think that was the right call.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Comes Next
&lt;/h2&gt;

&lt;p&gt;The architecture is the mechanical half. The more interesting half, honestly, is what happens when this stack fails in production, and it does. The bank had three meaningful incidents in the year after go-live, each of which taught the fraud team something about how AI detection breaks at scale. Part 2 covers the named case studies that put real numbers on the scale of the problem (DBS, JPMorgan, HSBC, Danske Bank, Standard Chartered), the brutal arithmetic of false-positive economics on a 0.04% base rate, what SR 11-7 and the EU AI Act mean for a fraud team, and what the senior analyst does on the cases the model refuses to handle. &lt;a href="https://blogs.getsetlive.com/ai-fraud-detection-banking-failures-regulation/" rel="noopener noreferrer"&gt;Part 2 is here&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://blogs.getsetlive.com/ai-loan-underwriting-architecture/" rel="noopener noreferrer"&gt;What AI Actually Does in Loan Underwriting (Part 1): The Architecture&lt;/a&gt;: the same lens on an Indian NBFC’s underwriting stack, where the models are similar and the plumbing is different.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blogs.getsetlive.com/ai-loan-underwriting-failures-human-in-loop/" rel="noopener noreferrer"&gt;What AI Actually Does in Loan Underwriting (Part 2): Where It Breaks and Who Still Signs Off&lt;/a&gt;: thin files, festival drift, and proxy discrimination, the lending versions of the failures Part 2 covers for banking.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blogs.getsetlive.com/7-times-ai-gave-the-wrong-answer-with-proof/" rel="noopener noreferrer"&gt;7 Times AI Gave the Wrong Answer (with Proof)&lt;/a&gt;: concrete examples of AI confidently getting it wrong, which matters when the model is touching money.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Kai Waehner, &lt;a href="https://www.kai-waehner.de/blog/2022/10/25/fraud-detection-with-apache-kafka-ksql-and-apache-flink/" rel="noopener noreferrer"&gt;Fraud Detection with Apache Kafka, KSQL and Apache Flink&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NVIDIA, &lt;a href="https://developer.nvidia.com/blog/supercharging-fraud-detection-in-financial-services-with-graph-neural-networks/" rel="noopener noreferrer"&gt;Supercharging Fraud Detection with Graph Neural Networks&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Fintech Global, &lt;a href="https://fintech.global/2026/02/03/how-ai-is-reshaping-aml-compliance-in-modern-banking/" rel="noopener noreferrer"&gt;How AI Is Reshaping AML Compliance in Modern Banking&lt;/a&gt;, 2026&lt;/li&gt;
&lt;li&gt;Conduktor, &lt;a href="https://www.conduktor.io/glossary/real-time-fraud-detection-with-streaming/" rel="noopener noreferrer"&gt;Real-Time Fraud Detection with Streaming&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Ampcome, &lt;a href="https://www.ampcome.com/post/agentic-ai-examples-banking-compliance-risk" rel="noopener noreferrer"&gt;Agentic AI Examples in Banking: Compliance &amp;amp; Risk&lt;/a&gt;, 2026&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does AI actually stop bank fraud?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not on its own. The models produce a fraud probability for each transaction. Rules set by the fraud operations team decide what to do with that score: allow, challenge with step-up authentication, or block. The model narrows the field; rules and humans make the call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which machine learning model is used for real-time fraud detection?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most banks run an ensemble. XGBoost is the primary supervised classifier on tabular features, an autoencoder catches unsupervised anomalies, and a graph neural network such as GraphSAGE catches mule rings. A small meta-model combines the three scores, and the whole ensemble answers in under 15 milliseconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is step-up authentication in fraud detection?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The middle-ground response between approve and decline. When the fraud score is ambiguous (roughly 0.60 to 0.80), the bank asks the customer to confirm the transaction through a mobile push, OTP, or biometric check. It catches fraud without declining legitimate purchases, and it only works because the model produces a probability rather than a yes or no.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long does an AI fraud decision take?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Model inference runs in 5 to 15 milliseconds. End to end, from card swipe to bank response, the median is 70 to 90 milliseconds. Card networks allow about two seconds for the whole authorisation, so there is headroom, but the pipeline is tuned hard because latency compounds across hops.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is a feature store and why does a fraud system need one?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A feature store guarantees that the features a model was trained on are computed the same way at serving time. Without one, training and production drift apart silently and the model scores wrong with full confidence. The bank in this article uses Tecton with Redis for online lookups and Snowflake for training data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why keep an autoencoder if XGBoost is more accurate?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because the autoencoder needs no labels. It flags transaction shapes the supervised model has never been trained on, such as a new card-testing pattern, weeks before enough labelled examples exist to retrain XGBoost. It costs a few milliseconds and some explainability, and it earns its place on zero-day fraud.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can AI replace fraud analysts?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. AI cut the daily alert volume from a number no team can review to a number that fits in a shift. The analysts who remain handle the hard cases, phone customers, confirm fraud rings, and draft Suspicious Activity Reports. AI does the bulk filtering; humans do the judgement and carry the regulatory exposure.&lt;/p&gt;

&lt;p&gt;The post &lt;a href="https://blogs.getsetlive.com/ai-fraud-detection-banking-architecture/" rel="noopener noreferrer"&gt;What AI Actually Does in Bank Fraud Detection (Part 1): The Architecture&lt;/a&gt; appeared first on &lt;a href="https://blogs.getsetlive.com" rel="noopener noreferrer"&gt;GetSetLive Blogs&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>fintech</category>
      <category>security</category>
    </item>
  </channel>
</rss>
