<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Eduardo Pittol</title>
    <description>The latest articles on DEV Community by Eduardo Pittol (@edpittol).</description>
    <link>https://dev.to/edpittol</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3881058%2F17f670b3-e8eb-4e4e-b7ad-128c71b6e970.jpeg</url>
      <title>DEV Community: Eduardo Pittol</title>
      <link>https://dev.to/edpittol</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/edpittol"/>
    <language>en</language>
    <item>
      <title>WP-Cron and Action Scheduler are a headache in WordPress infrastructure management. On this post I bring good practices to deal and fix gaps to improve the usage and monitoring of them.</title>
      <dc:creator>Eduardo Pittol</dc:creator>
      <pubDate>Thu, 06 Aug 2026 19:22:06 +0000</pubDate>
      <link>https://dev.to/edpittol/wp-cron-and-action-scheduler-are-a-headache-in-wordpress-infrastructure-management-on-this-post-i-564f</link>
      <guid>https://dev.to/edpittol/wp-cron-and-action-scheduler-are-a-headache-in-wordpress-infrastructure-management-on-this-post-i-564f</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/edpittol/disablewpcron-doesnt-close-the-door-its-named-after-34ih" class="crayons-story__hidden-navigation-link"&gt;DISABLE_WP_CRON Doesn't Close the Door It's Named After&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/edpittol" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3881058%2F17f670b3-e8eb-4e4e-b7ad-128c71b6e970.jpeg" alt="edpittol profile" class="crayons-avatar__image" width="460" height="460"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/edpittol" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Eduardo Pittol
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Eduardo Pittol
                
              
              &lt;div id="story-author-preview-content-4334688" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/edpittol" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3881058%2F17f670b3-e8eb-4e4e-b7ad-128c71b6e970.jpeg" class="crayons-avatar__image" alt="" width="460" height="460"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Eduardo Pittol&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/edpittol/disablewpcron-doesnt-close-the-door-its-named-after-34ih" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 6&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/edpittol/disablewpcron-doesnt-close-the-door-its-named-after-34ih" id="article-link-4334688"&gt;
          DISABLE_WP_CRON Doesn't Close the Door It's Named After
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/wordpress"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;wordpress&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/php"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;php&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devops"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devops&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/security"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;security&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/edpittol/disablewpcron-doesnt-close-the-door-its-named-after-34ih#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            23 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>devops</category>
      <category>monitoring</category>
      <category>php</category>
      <category>wordpress</category>
    </item>
    <item>
      <title>DISABLE_WP_CRON Doesn't Close the Door It's Named After</title>
      <dc:creator>Eduardo Pittol</dc:creator>
      <pubDate>Thu, 06 Aug 2026 19:07:18 +0000</pubDate>
      <link>https://dev.to/edpittol/disablewpcron-doesnt-close-the-door-its-named-after-34ih</link>
      <guid>https://dev.to/edpittol/disablewpcron-doesnt-close-the-door-its-named-after-34ih</guid>
      <description>&lt;p&gt;WordPress runs your scheduled work inside a web request unless you stop it. A visitor loads a page, WordPress fires a loopback request to &lt;code&gt;wp-cron.php&lt;/code&gt;, and every due event runs in there — core's own housekeeping, whatever your plugins scheduled, and, if you have Action Scheduler installed, a queue drain along with them. Action Scheduler then goes one step further: it has its own loopback dispatcher, so it can keep draining long after the visitor who triggered it has gone.&lt;/p&gt;

&lt;p&gt;Most of the time that is fine. When it isn't — when scheduled work competes with real traffic, or when you want one place where it runs and one log that records it — the fix is to move cron onto a system scheduler and close the web paths behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two layers, one entry point.&lt;/strong&gt; WP-Cron is WordPress's scheduler: a stored list of due events, and a web request that fires them. Action Scheduler is a plugin that registers &lt;em&gt;one&lt;/em&gt; of those events — &lt;code&gt;action_scheduler_run_queue&lt;/code&gt; — and drains its own queue when that event fires. It is not a system parallel to WP-Cron; it sits on top of it. So everything below about &lt;code&gt;wp-cron.php&lt;/code&gt;, the loopback, and the constant that is supposed to disable it applies to core's events exactly as it applies to the queue. Where Action Scheduler needs a step of its own, it gets one, and it is named as such.&lt;/p&gt;

&lt;p&gt;This post does that in five steps, then explains what each step is actually doing and why it takes the form it does. Every step is a file you can paste or a command you can run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What this post is not about.&lt;/strong&gt; How long a run takes, how many PHP workers it occupies, or how much front-end latency it costs is out of scope here. Those numbers depend entirely on your workload and your pool, and they deserve their own measurements rather than borrowed ones. The question here is narrower and more portable: &lt;strong&gt;which doors can reach your scheduled work, and what happens when you close each one.&lt;/strong&gt; Where §1 gets into what an open door costs you, it names the mechanism and stops there — the numbers stay yours to measure.&lt;/p&gt;

&lt;p&gt;Everything below was measured on WordPress 7.0.2 with Action Scheduler 4.0.0, served by nginx and PHP-FPM, against a queue of 50 pending actions on a single probe hook.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you'll end up with
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;system cron ──&amp;gt; wp cron event run --due-now ──┬─&amp;gt; WordPress core's own events
                                              └─&amp;gt; the Action Scheduler queue

nginx        ──&amp;gt; /wp-cron.php ................ 403, PHP never starts
WordPress    ──&amp;gt; /wp-cron.php reaching PHP ... logged, then 403
WordPress    ──&amp;gt; async loopback dispatch ..... suppressed
WordPress    ──&amp;gt; action_scheduler_run_queue .. no callback attached, except under WP-CLI
WordPress    ──&amp;gt; any non-CLI queue run ....... logged
admin UI     ──&amp;gt; the "Run" row action ........ still works, on purpose
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One scheduled command, three closed doors, one left open deliberately, two tripwires. &lt;code&gt;wp cron event run --due-now&lt;/code&gt; fires the same hook &lt;code&gt;wp-cron.php&lt;/code&gt; would have fired, from a process that step 4 deliberately exempts — so core's own events and the Action Scheduler queue both drain from that single entry, and there is one place to look when something doesn't run.&lt;/p&gt;

&lt;p&gt;The two tripwires watch different things, and both are needed. The &lt;strong&gt;entry-point tripwire&lt;/strong&gt; in step 2b records that someone reached &lt;code&gt;wp-cron.php&lt;/code&gt; at all, whatever they were after. The &lt;strong&gt;queue tripwire&lt;/strong&gt; in step 4 records that Action Scheduler's queue ran outside WP-CLI. A hit on the endpoint that runs only core's own events — a version check, a transient cleanup — trips the first and not the second.&lt;/p&gt;

&lt;p&gt;The order of the steps matters: step 1 is what makes step 2b safe, and step 4's WP-CLI exemption is what keeps step 5 working at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1 — Stop WordPress from dispatching cron itself
&lt;/h2&gt;

&lt;p&gt;In &lt;code&gt;wp-config.php&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="nb"&gt;define&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'DISABLE_WP_CRON'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This stops WordPress from firing its own loopback request to &lt;code&gt;wp-cron.php&lt;/code&gt; at the end of a page load. It does &lt;strong&gt;not&lt;/strong&gt; make &lt;code&gt;wp-cron.php&lt;/code&gt; unreachable — that's step 2, and the reason why is §1.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2a — Close the web entry point in front of PHP
&lt;/h2&gt;

&lt;p&gt;In your nginx server block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Must come before any `location ~ \.php$` block. nginx matches an exact&lt;/span&gt;
&lt;span class="c1"&gt;# `location =` ahead of every regex location, so this is what serves this&lt;/span&gt;
&lt;span class="c1"&gt;# one URI.&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;/wp-cron.php&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;403&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reload nginx. Then check it — this is the one step in this post you can verify by reading a status code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-is&lt;/span&gt; https://example.com/wp-cron.php | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt; &lt;span class="m"&gt;403&lt;/span&gt; &lt;span class="ne"&gt;Forbidden&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PHP never starts for that request. Nothing bootstraps, no worker is consumed, and nothing runs — no core event, no queue drain.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If you are deploying one config across many sites, read §3 before you ship this.&lt;/strong&gt; nginx cannot read a PHP constant, so this block is unconditional: it also blocks sites that still depend on the web loopback, and those sites lose cron entirely. That is measured, not hypothetical.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 2b — Log every hit, then close it again inside WordPress
&lt;/h2&gt;

&lt;p&gt;Create &lt;code&gt;wp-content/mu-plugins/cron-web-entry-point.php&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?php&lt;/span&gt;
&lt;span class="cd"&gt;/**
 * Log — and then refuse — HTTP requests to the cron entry point.
 *
 * Second layer behind the web-server block: this one catches a request
 * that reaches PHP by some route the web server config doesn't cover —
 * another vhost, a rewrite, an environment where that config wasn't
 * deployed.
 *
 * The log line comes FIRST, deliberately. The refusal below ends the
 * request, and §2 explains why the status it returns is invisible: on
 * PHP-FPM the response was already flushed before this file loaded. A
 * log line is the only trace this request will ever leave. Blocking
 * without logging would close the door and destroy the evidence.
 *
 * wp-cron.php defines DOING_CRON before mu-plugins load, and it die()s
 * on any request with a body, so everything reaching this point is a GET
 * aimed at the cron endpoint. WP-CLI is excluded: that is the path we
 * are keeping.
 */&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nb"&gt;defined&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'DOING_CRON'&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="no"&gt;DOING_CRON&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s1"&gt;'cli'&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;PHP_SAPI&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nv"&gt;$clean&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$value&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nb"&gt;preg_replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'/[\x00-\x1F\x7F]+/'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;' '&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nv"&gt;$value&lt;/span&gt; &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="nv"&gt;$due&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;array&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nf"&gt;wp_get_ready_cron_jobs&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nv"&gt;$hooks&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nv"&gt;$due&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;array_merge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$due&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;array_keys&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$hooks&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nb"&gt;error_log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nb"&gt;sprintf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s1"&gt;'[cron-guard] wp-cron.php reached: ip=%s query=%s lock=%s due=%s'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;isset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'REMOTE_ADDR'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="nv"&gt;$clean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'REMOTE_ADDR'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'n/a'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="k"&gt;isset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'QUERY_STRING'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="nv"&gt;$clean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'QUERY_STRING'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;''&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nf"&gt;get_transient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'doing_cron'&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="s1"&gt;'held'&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'free'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nv"&gt;$due&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="nb"&gt;implode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;','&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;array_unique&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$due&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'nothing'&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nb"&gt;defined&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'DISABLE_WP_CRON'&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="no"&gt;DISABLE_WP_CRON&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nb"&gt;headers_sent&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nb"&gt;http_response_code&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="mi"&gt;403&lt;/span&gt; &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two real lines from that file, one from core's own loopback on a site still using it, one from a bare external request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[05-Aug-2026 20:11:19 UTC] [cron-guard] wp-cron.php reached: ip=172.18.0.2 query=doing_wp_cron=1785960679.6381549835205078125000 lock=held due=recovery_mode_clean_expired_keys,wp_privacy_delete_old_export_files,wp_version_check,wp_update_plugins,wp_update_themes,action_scheduler_run_queue,wp_site_health_scheduled_check
[05-Aug-2026 20:10:06 UTC] [cron-guard] wp-cron.php reached: ip=172.18.0.2 query= lock=free due=recovery_mode_clean_expired_keys,wp_privacy_delete_old_export_files,wp_version_check,wp_update_plugins,wp_update_themes,action_scheduler_run_queue,wp_site_health_scheduled_check
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second line is the one worth an alert: no &lt;code&gt;doing_wp_cron&lt;/code&gt; key, so nothing internal sent it, and the lock was free, so core was about to run all seven of those hooks. Both requests here came from inside the same container network, which is why &lt;code&gt;ip=&lt;/code&gt; doesn't separate them — on a real site an outside hit carries an outside address. &lt;code&gt;query=&lt;/code&gt; and &lt;code&gt;lock=&lt;/code&gt; are what tell the two apart regardless.&lt;/p&gt;

&lt;p&gt;Note what &lt;code&gt;due=&lt;/code&gt; contains. Six of those seven hooks are core's own — version checks, a privacy cleanup, a site-health check — and only one belongs to Action Scheduler. A request that runs just the other six drains no queue at all, so nothing downstream of Action Scheduler would ever notice it happened.&lt;/p&gt;

&lt;p&gt;Both layers can be armed at once. When nothing is wrong, the whole file costs one &lt;code&gt;defined()&lt;/code&gt; check.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why the log and the block live in one file.&lt;/strong&gt; Splitting them looks tidier and quietly breaks: mu-plugins load in alphabetical order by filename — &lt;code&gt;wp_get_mu_plugins()&lt;/code&gt; ends with a plain &lt;code&gt;sort()&lt;/code&gt; — so a logger in a separate file only runs first if its name happens to sort earlier. Keeping both in one file, log before block, makes the order a property of the code rather than of a filename someone may rename later.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 3 — Suppress Action Scheduler's async dispatch
&lt;/h2&gt;

&lt;p&gt;Create &lt;code&gt;wp-content/mu-plugins/cron-suppress-async-dispatch.php&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?php&lt;/span&gt;
&lt;span class="cd"&gt;/**
 * Stop Action Scheduler from opening its own loopback door.
 *
 * On 'shutdown', Action Scheduler decides whether to POST to
 * admin-ajax.php?action=as_async_request_queue_runner to keep draining
 * without waiting for the next cron tick. This filter answers no.
 *
 * This stops WordPress from *sending* that request. It does not gate the
 * endpoint against someone sending an equivalent request directly — see §4.
 */&lt;/span&gt;

&lt;span class="nf"&gt;add_filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'action_scheduler_allow_async_request_runner'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'__return_false'&lt;/span&gt; &lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 4 — Unhook the queue runner outside WP-CLI
&lt;/h2&gt;

&lt;p&gt;Create &lt;code&gt;wp-content/mu-plugins/cron-unhook-queue-runner.php&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?php&lt;/span&gt;
&lt;span class="cd"&gt;/**
 * Detach Action Scheduler's queue runner from the WP-Cron hook.
 *
 * Action Scheduler attaches its own run() method to
 * 'action_scheduler_run_queue' from its 'init' priority-1 callback.
 * Removing it turns WP-Cron, the async loopback, and any direct call to
 * that hook into no-ops: there is no callback left to run.
 *
 * Priority 100 is not decoration. remove_action() only works if it runs
 * *after* the matching add_action(), at the same priority — otherwise it
 * silently removes nothing. See §5 for the measurement.
 *
 * WP-CLI is exempt: that is the path we are keeping.
 */&lt;/span&gt;

&lt;span class="nf"&gt;add_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'init'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nb"&gt;defined&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'WP_CLI'&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="no"&gt;WP_CLI&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nb"&gt;class_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'ActionScheduler_QueueRunner'&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="nf"&gt;remove_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;ActionScheduler_QueueRunner&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="no"&gt;WP_CRON_HOOK&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="n"&gt;ActionScheduler_QueueRunner&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;instance&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="s1"&gt;'run'&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="mi"&gt;100&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Optional but recommended — the second tripwire, &lt;code&gt;wp-content/mu-plugins/cron-log-non-cli-queue-runs.php&lt;/code&gt;. Step 2b's watches the cron endpoint; this one watches the queue itself, and the two catch different things:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?php&lt;/span&gt;
&lt;span class="cd"&gt;/**
 * Log any queue run that isn't WP-CLI. Blocks nothing.
 *
 * Action Scheduler fires this hook at the start of every run() call,
 * whatever triggered it. If the queue is ever processed outside CLI, a
 * line lands in the error log naming the SAPI, the URI and the caller.
 *
 * Request URI and remote address are attacker-controlled — this exists to
 * log hostile requests — so control characters are stripped before they
 * are written, or a crafted request could forge extra log lines.
 *
 * Note the blind spot in §4: one entry point never fires this hook.
 */&lt;/span&gt;

&lt;span class="nf"&gt;add_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'action_scheduler_before_process_queue'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'cli'&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;PHP_SAPI&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="nv"&gt;$clean&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$value&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nb"&gt;preg_replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'/[\x00-\x1F\x7F]+/'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;' '&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nv"&gt;$value&lt;/span&gt; &lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;};&lt;/span&gt;

        &lt;span class="nb"&gt;error_log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="nb"&gt;sprintf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s1"&gt;'[cron-guard] queue run outside CLI: sapi=%s uri=%s ip=%s'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="kc"&gt;PHP_SAPI&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="k"&gt;isset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'REQUEST_URI'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="nv"&gt;$clean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'REQUEST_URI'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'n/a'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="k"&gt;isset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'REMOTE_ADDR'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="nv"&gt;$clean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$_SERVER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'REMOTE_ADDR'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'n/a'&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 5 — Schedule the runner
&lt;/h2&gt;

&lt;p&gt;One entry, standing in for the loopback you just closed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Core's own scheduled events and the Action Scheduler queue, in one pass.
* * * * * cd /var/www/html &amp;amp;&amp;amp; flock -n /tmp/wp-cron.lock wp cron event run --due-now --quiet
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Action Scheduler also ships its own command, &lt;code&gt;wp action-scheduler run&lt;/code&gt;&lt;/strong&gt;, which drains the same queue without going through the WP-Cron hook. It is not a substitute for the entry above: it drains the queue and nothing else, so a site scheduling only that command has closed the web loopback and left core's own events — version checks, privacy cleanups, site health — with nothing to fire them. Schedule it &lt;em&gt;alongside&lt;/em&gt; &lt;code&gt;wp cron event run --due-now&lt;/code&gt; if you want it, never in place of it. If you schedule both, each action's log records which command claimed it — some read &lt;code&gt;WP Cron&lt;/code&gt;, some read &lt;code&gt;WP CLI&lt;/code&gt;, and which one wins a given action is a race. Weighing that split is out of scope for this post, which keeps a single entry point so there is a single place to look. §6 covers what the labels mean.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. &lt;code&gt;DISABLE_WP_CRON&lt;/code&gt; is a scheduling switch, not a lock on the door
&lt;/h2&gt;

&lt;p&gt;The constant's name invites a reasonable assumption: turn it on, and WP-Cron stops running on this site. What it actually does is narrower. It stops WordPress from spawning its own loopback request at the end of a page load — and that is all. In core's whole cron path the constant is read in exactly one place, inside the function that decides whether to spawn. Nothing in &lt;code&gt;wp-cron.php&lt;/code&gt; consults it.&lt;/p&gt;

&lt;p&gt;Core says so itself, in that file's own header: defining &lt;code&gt;DISABLE_WP_CRON&lt;/code&gt; and calling the file directly &lt;em&gt;"are mutually exclusive and the latter does not rely on the former to work."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;So here is the whole of it. Set the constant to &lt;code&gt;true&lt;/code&gt;, change nothing else, and run one command from anywhere on the internet:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://example.com/wp-cron.php
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That request drained &lt;strong&gt;50 of 50&lt;/strong&gt; pending actions — and it ran core's own due events on the way, because that is the endpoint's entire job: it fires everything the schedule says is ready, and the queue drain is one entry on that list. No cookie, no nonce, no capability check, no authentication of any kind — the file is a public endpoint by design, because it has to be callable by whatever is meant to be calling it. The scheduling switch was on the whole time. Cron ran anyway, because reachability was never what the constant governed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Here is the one duration in this post, and it is here for a reason.&lt;/strong&gt; That request returned in &lt;strong&gt;0.008 seconds&lt;/strong&gt; — eight milliseconds — and the fifty actions drained &lt;em&gt;afterwards&lt;/em&gt;, in the same PHP worker, with the connection already closed. This is not a figure about how fast the queue is. It is the cheapest available proof that the response and the work are two separate things: your monitoring saw an 8 ms &lt;code&gt;200&lt;/code&gt;, and fifty actions ran after the client hung up.&lt;/p&gt;

&lt;h3&gt;
  
  
  The same property, read from the other side
&lt;/h3&gt;

&lt;p&gt;Everything above is also a description of a request an attacker can send, and it is worth being explicit about why this endpoint is a favourite.&lt;/p&gt;

&lt;p&gt;Reading core's own order of operations in that file: the response is flushed and the connection closed &lt;strong&gt;first&lt;/strong&gt;; then WordPress boots in full via &lt;code&gt;wp-load.php&lt;/code&gt;; then &lt;code&gt;wp_raise_memory_limit( 'cron' )&lt;/code&gt; lifts that worker's memory ceiling; then, and only then, does core read the &lt;code&gt;doing_cron&lt;/code&gt; transient to find out whether another cron run is already in progress and it should stand down.&lt;/p&gt;

&lt;p&gt;Three consequences follow from that ordering, and none of them is a bug — each is a deliberate choice with a good reason behind it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The attacker pays for a socket that is closed almost immediately.&lt;/strong&gt; Flooding a slow endpoint normally costs the attacker too: they have to hold connections open while the server works. Here the server hangs up first and keeps working. One cheap request buys a full WordPress bootstrap in a worker that the client is no longer waiting on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The lock is checked after the bootstrap, not before it.&lt;/strong&gt; &lt;code&gt;WP_CRON_LOCK_TIMEOUT&lt;/code&gt; defaults to sixty seconds, so the queue itself can only &lt;em&gt;drain&lt;/em&gt; about once a minute no matter how hard the endpoint is hit. But that lock lives in the database and is read by WordPress — which means booting WordPress is what it costs to find out you were supposed to stand down. Rate-limited work, unlimited bootstraps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing in the request needs to be valid.&lt;/strong&gt; There is no signature to forge and no state to guess. The one thing core does reject is a request body: &lt;code&gt;wp-cron.php&lt;/code&gt; calls &lt;code&gt;die()&lt;/code&gt; if &lt;code&gt;$_POST&lt;/code&gt; is non-empty, so this is a &lt;code&gt;GET&lt;/code&gt;-only endpoint. That is the extent of the input validation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is exotic knowledge; it is the reason "block &lt;code&gt;wp-cron.php&lt;/code&gt; and drive cron from the system scheduler" is standard advice at every managed WordPress host, and the reason serious hosts rate-limit or refuse this path at their edge rather than leaving it to each site. What the measurement adds is the part people skip: &lt;strong&gt;&lt;code&gt;DISABLE_WP_CRON&lt;/code&gt; does not participate in any of it.&lt;/strong&gt; A site that has moved cron to the system scheduler, and believes it has therefore turned the web path off, is running exactly the endpoint described above — with the additional property that the operator is not watching it, because they think it is disabled.&lt;/p&gt;

&lt;p&gt;That is the argument for step 2, and the whole reason the tutorial has a step 2 at all. Why closing it inside WordPress is not enough on its own is the next section.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. &lt;code&gt;wp-cron.php&lt;/code&gt; answers 200 no matter what your guard decides
&lt;/h2&gt;

&lt;p&gt;Core's &lt;code&gt;wp-cron.php&lt;/code&gt; calls &lt;code&gt;fastcgi_finish_request()&lt;/code&gt; &lt;strong&gt;before it loads WordPress&lt;/strong&gt;. The order is not subtle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;line 19   ignore_user_abort( true )
line 21   if ( ! headers_sent() ) {
line 22-23    two cache headers — no status set, so PHP's implicit 200 is queued
line 26   // Don't run cron until the request finishes, if possible.
line 28   fastcgi_finish_request()   ←── response flushed and closed HERE
line 33   bail if this is a POST, AJAX, or already DOING_CRON
line 42   define( 'DOING_CRON', true )
line 46   require wp-load.php  ──&amp;gt; wp-config.php ──&amp;gt; mu-plugins load HERE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At the moment of the flush, WordPress does not exist in that process. &lt;code&gt;wp-config.php&lt;/code&gt; hasn't been read, so &lt;code&gt;DISABLE_WP_CRON&lt;/code&gt; isn't defined yet. No mu-plugin has loaded. &lt;strong&gt;The response is gone before any code of yours can have an opinion about it&lt;/strong&gt; — which is why nothing you configure can prevent this.&lt;/p&gt;

&lt;p&gt;By the time step 2b's &lt;code&gt;exit&lt;/code&gt; runs, the client has been answered. And its &lt;code&gt;http_response_code( 403 )&lt;/code&gt; doesn't merely fail to reach the client: measured against a live request, &lt;strong&gt;PHP refuses the call outright.&lt;/strong&gt; It returns &lt;code&gt;false&lt;/code&gt;, and reading the status back afterwards still reports &lt;strong&gt;200&lt;/strong&gt;. The 403 does not exist anywhere.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not fail quietly, though, and this is where the two halves of step 2b meet.&lt;/strong&gt; Called after the flush, &lt;code&gt;http_response_code()&lt;/code&gt; writes a line into PHP's error log:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PHP Warning:  http_response_code(): Cannot set response code - headers already sent in .../cron-web-entry-point.php on line 143
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One per blocked request, in the same file the guard's own log line goes to. That is the only trace the refusal leaves anywhere — and it is the wrong trace: a complaint about a failed function call, sitting next to the record of the request that caused it, saying nothing about who sent it or what was due. It looks like a bug in your mu-plugin, because in a sense it is one.&lt;/p&gt;

&lt;p&gt;Which is why step 2b tests &lt;code&gt;headers_sent()&lt;/code&gt; before making the call at all. The test isn't defensive habit; it keeps the log you are actually going to read free of a warning that fires on every single hit. Suppressing it with &lt;code&gt;@&lt;/code&gt; would work too, and is worse: it hides a real fact about the platform instead of encoding it. Ask whether the headers are gone, and the answer tells you which server you're on — on Apache with mod_php they aren't, so the call proceeds and the client genuinely receives a 403.&lt;/p&gt;

&lt;p&gt;Core does the same thing three lines into the file above. Look again at the timeline: line 21 is &lt;code&gt;if ( ! headers_sent() )&lt;/code&gt;, wrapped around core's own two &lt;code&gt;header()&lt;/code&gt; calls. &lt;code&gt;wp-cron.php&lt;/code&gt; is written by people who knew this file might be reached with the response already gone. Your guard is in the same position, one bootstrap later.&lt;/p&gt;

&lt;p&gt;Here is the same block, armed, at both layers, with four independently recorded statuses for one request:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Where the block lives&lt;/th&gt;
&lt;th&gt;Client read&lt;/th&gt;
&lt;th&gt;Web server logged&lt;/th&gt;
&lt;th&gt;PHP-FPM logged&lt;/th&gt;
&lt;th&gt;PHP's own status at shutdown&lt;/th&gt;
&lt;th&gt;Drained&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;In WordPress&lt;/strong&gt; (mu-plugin)&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 of 50&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;In front of PHP&lt;/strong&gt; (nginx)&lt;/td&gt;
&lt;td&gt;403&lt;/td&gt;
&lt;td&gt;403&lt;/td&gt;
&lt;td&gt;never reached PHP&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 of 50&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;em&gt;No block at all&lt;/em&gt;, for comparison&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;50 of 50&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read the first and third rows together. &lt;strong&gt;Every observable status is identical between a block that stopped everything and no block at all.&lt;/strong&gt; The guard worked — zero drained, zero probe executions logged — and nothing anywhere recorded a refusal. Not the client, not nginx, not PHP-FPM, not PHP itself.&lt;/p&gt;

&lt;p&gt;This is a feature, not a bug, which is why it will not be fixed. Core's own comment, immediately above that block at line 26, reads &lt;em&gt;"Don't run cron until the request finishes, if possible."&lt;/em&gt; The point of flushing early is that the unlucky visitor who triggered a loopback isn't kept waiting for someone else's scheduled work. Response masking is the side effect of that kindness.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Why the block goes in front of PHP — and what it costs
&lt;/h2&gt;

&lt;p&gt;§2 is the whole argument for step 2a. A block inside WordPress is real but unverifiable from outside; a block in the web server is real &lt;em&gt;and&lt;/em&gt; legible. It also never starts PHP, which means an endpoint you have decided is closed stops costing you a worker and a full WordPress bootstrap on every hit.&lt;/p&gt;

&lt;p&gt;The cost is a genuine one, and it is measured rather than warned about. nginx cannot read a PHP constant, so it cannot make the block conditional. Run both blocks against a site where &lt;code&gt;DISABLE_WP_CRON&lt;/code&gt; is &lt;code&gt;false&lt;/code&gt; — a site that still depends on the loopback for cron to run at all:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Block&lt;/th&gt;
&lt;th&gt;&lt;code&gt;DISABLE_WP_CRON&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Client read&lt;/th&gt;
&lt;th&gt;Drained&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;In WordPress (mu-plugin)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;50 of 50&lt;/strong&gt; — goes inert, cron keeps working&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;In front of PHP (nginx)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;403&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0 of 50&lt;/strong&gt; — cron is gone&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The mu-plugin checks the constant and stands down. nginx blocks regardless, and that site's scheduled work stops silently — no error, no log line, just a queue that never moves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So: an observable status, or a conditional block that is safe to deploy everywhere. At one layer you can have either, not both.&lt;/strong&gt; For a single site you are configuring deliberately — which is what the tutorial above assumes — the nginx block is the better default, because you set the constant yourself in step 1 and the condition is already satisfied by hand. For a config rolled out across a fleet, the conditional check in step 2b is the one that won't take down the site that hadn't finished migrating. Installing both, as the tutorial does, is what makes that choice recoverable instead of load-bearing.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Action Scheduler has four entry points; closing one closes nothing
&lt;/h2&gt;

&lt;p&gt;This is the one part of the arrangement that is specifically about Action Scheduler, and it is the reason steps 3 and 4 exist at all. Core's due events have effectively one door: the stored schedule, walked either by &lt;code&gt;wp-cron.php&lt;/code&gt; or by WP-CLI. Close the endpoint and they stop. Action Scheduler adds three more doors of its own, and a block on the cron endpoint cannot see any of them.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;action_scheduler_run_queue&lt;/code&gt; can be reached at least four independent ways. Each was exercised on its own, against the same 50-action queue:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;wp-cron.php&lt;/code&gt; over HTTP.&lt;/strong&gt; 50 of 50, as in §1.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A direct POST to &lt;code&gt;admin-ajax.php?action=as_async_request_queue_runner&lt;/code&gt;.&lt;/strong&gt; Unauthenticated, 50 of 50 drained. With the web-entry-point block, the async suppression, &lt;em&gt;and&lt;/em&gt; the tripwire all armed: &lt;strong&gt;still 50 of 50.&lt;/strong&gt; Suppressing dispatch stops WordPress from &lt;em&gt;sending&lt;/em&gt; that request. It does nothing to gate the endpoint against someone sending an equivalent one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An authenticated admin page load.&lt;/strong&gt; Action Scheduler's own dispatcher reaches its decision on &lt;code&gt;shutdown&lt;/code&gt; and fires the loopback: 50 of 50. With the suppression armed, &lt;strong&gt;0 of 50&lt;/strong&gt; — and the record shows the decision point was genuinely reached and answered no, rather than never consulted at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The "Run" row action&lt;/strong&gt; on the Scheduled Actions admin screen. With &lt;strong&gt;all four&lt;/strong&gt; guards armed, including the unhook: &lt;strong&gt;1 of 50&lt;/strong&gt; — one action ran and completed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Four doors. Each guard closes some of them. Only the unhook closes the hook itself — and the row action still works, because it never goes near the hook.&lt;/p&gt;

&lt;h3&gt;
  
  
  The fourth door is meant to stay open
&lt;/h3&gt;

&lt;p&gt;That last result is the one most likely to be misread, so it is worth being exact about what it is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One is the ceiling, not a leak.&lt;/strong&gt; The row-action handler takes a single action ID from the request and processes that one action. One click, one action. So &lt;code&gt;1 of 50&lt;/code&gt; is the whole of what that path can do per click — not a guard that let one action slip past while catching forty-nine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is nothing like the door in §1.&lt;/strong&gt; The handler requires &lt;code&gt;row_action&lt;/code&gt;, &lt;code&gt;row_id&lt;/code&gt; and &lt;code&gt;nonce&lt;/code&gt; to all be present, verifies the nonce against that specific action ID, and the screen it lives on is registered under &lt;code&gt;manage_options&lt;/code&gt;. That is an authenticated administrator, on a nonce-checked request, running one action they picked by hand. Compare the unauthenticated &lt;code&gt;GET&lt;/code&gt; in §1 that drained the entire queue: same queue, entirely different threat model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nothing in this arrangement is trying to close it&lt;/strong&gt;, and that is deliberate. When a job has failed and you want to retry exactly one action while watching what happens, this is the tool. Keeping it is a feature, not a gap left behind.&lt;/p&gt;

&lt;p&gt;The reason it survives all four guards is structural rather than lucky. The row action calls the queue runner's &lt;code&gt;process_action()&lt;/code&gt; directly, so: the web-entry-point block never fires, because an admin page load doesn't define &lt;code&gt;DOING_CRON&lt;/code&gt;; the async suppression has no dispatch to suppress; and the unhook removes a callback from a hook this path doesn't use. Short of adding a fifth guard aimed specifically at it, it stays reachable — which is the intent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The one thing it does cost you is a detection gap, and a dashboard can hide it.&lt;/strong&gt; On that same measurement the queue tripwire from step 4 reports no line at all, while an action genuinely ran. &lt;code&gt;action_scheduler_before_process_queue&lt;/code&gt; is fired in exactly two places in Action Scheduler: the queue runner's &lt;code&gt;run()&lt;/code&gt; and the WP-CLI runner. &lt;code&gt;process_action()&lt;/code&gt; fires neither. The entry-point tripwire from step 2b says nothing either, and correctly so — this door isn't the cron endpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An empty tripwire log is not evidence that nothing ran.&lt;/strong&gt; The pending count is — and so is Action Scheduler's own execution context, which stamped this run &lt;code&gt;Admin List Table&lt;/code&gt; and recorded it as started and completed. The blind spot is in the tripwire's wiring, not in Action Scheduler's records, which is why §6 argues for reading the context label instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The silent removal-timing trap
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;remove_action()&lt;/code&gt; removes a callback only if it runs &lt;em&gt;after&lt;/em&gt; the matching &lt;code&gt;add_action()&lt;/code&gt; has already executed, at the same priority. Run it earlier, or at a different priority, and it removes nothing — with no warning, no error, and no visible difference until the thing you thought you'd disabled fires anyway.&lt;/p&gt;

&lt;p&gt;Checking &lt;code&gt;has_action()&lt;/code&gt; at six points inside a single real HTTP request shows exactly where the margin is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;&lt;code&gt;has_action( 'action_scheduler_run_queue' )&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;plugins_loaded:10&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;init:1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;action_scheduler_init:10&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;attached, priority 10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;init:99&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;attached, priority 10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;init:101&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;false&lt;/strong&gt; (step 4 armed) / attached, priority 10 (not armed)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;wp_loaded:10&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;false&lt;/strong&gt; (step 4 armed) / attached, priority 10 (not armed)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Action Scheduler hasn't attached the callback yet at &lt;code&gt;plugins_loaded:10&lt;/code&gt; or &lt;code&gt;init:1&lt;/code&gt; — an unhook attempt at either point is a guaranteed silent no-op, and both are plausible places to put one. It has attached by the time its own &lt;code&gt;action_scheduler_init&lt;/code&gt; fires, still inside &lt;code&gt;init&lt;/code&gt; priority 1. Step 4 runs at &lt;code&gt;init:100&lt;/code&gt;; by &lt;code&gt;init:101&lt;/code&gt; the callback reads &lt;code&gt;false&lt;/code&gt;, genuinely removed rather than merely untested.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Step 4 is what keeps step 5 working — and it is one line away from not
&lt;/h2&gt;

&lt;p&gt;This is the load-bearing joint of the whole arrangement, and it is easy to miss because it looks like a detail.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;wp cron event run --due-now&lt;/code&gt; does not know Action Scheduler exists. It walks core's list of due events, and one of those events is &lt;code&gt;action_scheduler_run_queue&lt;/code&gt;. Firing it runs whatever is attached to that hook — which, before step 4, is Action Scheduler's queue runner. That is the entire mechanism: the command reaches the queue through the same hook &lt;code&gt;wp-cron.php&lt;/code&gt; would have used, and it drained &lt;strong&gt;50 of 50&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Step 4 then removes that callback. For every non-CLI request, &lt;code&gt;action_scheduler_run_queue&lt;/code&gt; becomes a hook with nothing on it. &lt;strong&gt;The &lt;code&gt;if ( defined( 'WP_CLI' ) &amp;amp;&amp;amp; WP_CLI ) { return; }&lt;/code&gt; at the top of that guard is the only reason step 5's command still drains anything.&lt;/strong&gt; Delete those two lines, or tighten the exemption to a narrower condition than you meant, and &lt;code&gt;wp cron event run --due-now&lt;/code&gt; keeps working perfectly: it still runs core's scheduled events, still exits zero, still prints success for every event it fired — while running not one Action Scheduler action. Nothing warns you. The pending count simply stops falling.&lt;/p&gt;

&lt;p&gt;That failure mode is why the thing to verify is the pending count, not the command's output. A cron entry that reports success and drains nothing is the same shape of problem as a guard that returns 200 and blocks everything: the signal you'd naturally trust has come apart from the thing you care about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What you get for accepting that fragility&lt;/strong&gt; is one hook with one purpose. &lt;code&gt;action_scheduler_run_queue&lt;/code&gt; stops being a shared entry point that four callers can reach and becomes something exactly one caller uses. A guard aimed at "the WP-Cron hook" now has a single, well-understood meaning — and Action Scheduler's own logs will tell you when that stops being true. Open any action on the Scheduled Actions screen and its log names the trigger, one of four labels:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Label&lt;/th&gt;
&lt;th&gt;What it means after this arrangement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;WP Cron&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Expected. Your scheduled &lt;code&gt;wp cron event run --due-now&lt;/code&gt;, which fires the WP-Cron hook.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;Admin List Table&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Expected, when someone clicked "Run" on a row. A person, not a schedule — see §4.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;Async Request&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Investigate.&lt;/strong&gt; The async dispatch is suppressed, so this label should not appear.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;WP CLI&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Investigate, unless you also scheduled &lt;code&gt;wp action-scheduler run&lt;/code&gt; on purpose.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The point of the arrangement is that this column stops being noise. Before it, &lt;code&gt;WP Cron&lt;/code&gt; could mean your scheduler, a stranger's &lt;code&gt;curl&lt;/code&gt;, or a visitor who happened to trigger a loopback, and there was no way to tell them apart. After it, each label maps to one identifiable cause, and two of the four are things you should go and look at.&lt;/p&gt;

&lt;p&gt;Note that &lt;code&gt;Admin List Table&lt;/code&gt; is not an alarm. That path is deliberately left reachable, it is capability- and nonce-gated, and it runs exactly one action per click — §4 has the detail. What the label buys you is knowing a human did it, rather than wondering why the count moved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For queue activity specifically, this is a better monitor than the tripwire from step 4&lt;/strong&gt;, for two reasons: Action Scheduler records it for you, with none of your code involved, and it has no blind spot at the manual-run path. The "Run" button that never fires the tripwire's hook still gets stamped &lt;code&gt;Admin List Table&lt;/code&gt; and recorded as started and completed.&lt;/p&gt;

&lt;p&gt;It does not replace step 2b's entry-point log, though, and the difference is worth keeping straight. The context label answers &lt;em&gt;how did this action get run&lt;/em&gt;. The entry-point log answers &lt;em&gt;is anyone still knocking on the cron endpoint&lt;/em&gt; — including all the times they knocked and nothing was due, which the context label can never show you because no action ran to be labelled.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;On Action Scheduler's own command.&lt;/strong&gt; &lt;code&gt;wp action-scheduler run&lt;/code&gt; bypasses the WP-Cron hook entirely, so it is immune to the fragility above — and it takes &lt;code&gt;--batch-size&lt;/code&gt;, &lt;code&gt;--batches&lt;/code&gt;, &lt;code&gt;--group&lt;/code&gt; and &lt;code&gt;--hooks&lt;/code&gt;, which lets a slow group get its own schedule and its own lock. Its actions log as &lt;strong&gt;&lt;code&gt;WP CLI&lt;/code&gt;&lt;/strong&gt;. Bypassing the hook is also its limitation: core's own due events live on that hook list and this command never touches them, so it is an addition to step 5's entry rather than a replacement for it. Running both is legitimate and the queue's claim mechanism keeps them from processing the same action twice, but you then read two labels for work you consider identical. When to prefer it is a separate question from this post's, which is about which doors reach your scheduled work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  7. Six ways to fool yourself while testing this
&lt;/h2&gt;

&lt;p&gt;A result of "0 of 50 drained" looks identical whether your block worked or your test was broken. Before believing any measurement here, rule these out — each one produces a convincing false pass on its own:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Empty queue.&lt;/strong&gt; Nothing was pending. Check the count &lt;em&gt;before&lt;/em&gt; the request, not only after.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No callback attached to the action.&lt;/strong&gt; Action Scheduler completes actions as instant no-ops when nothing is hooked to them. The pending count still drops and nothing logs — which reads exactly like a successful drain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A stale claim row&lt;/strong&gt; in &lt;code&gt;wp_actionscheduler_claims&lt;/code&gt;, left by a killed process. It blocks the runner from claiming anything, queue or no queue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A leftover &lt;code&gt;doing_cron&lt;/code&gt; transient.&lt;/strong&gt; WordPress core's own cron lock. A stale copy makes both &lt;code&gt;wp-cron.php&lt;/code&gt; and &lt;code&gt;wp cron event run --due-now&lt;/code&gt; into silent no-ops. &lt;code&gt;wp transient delete doing_cron&lt;/code&gt; clears it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code that isn't the code you think you're testing.&lt;/strong&gt; Opcache holding a previous version of a mu-plugin, a container built from a stale image, a deploy that didn't reach the box you're curling. Read the versions back from the running site rather than from your lockfile.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An empty entry-point log, read as a broken tripwire.&lt;/strong&gt; With step 2a armed, nginx refuses the request and PHP never starts — so step 2b's log line is &lt;em&gt;supposed&lt;/em&gt; to be absent. That is the arrangement working exactly as designed, and it is indistinguishable from a logger you wired up wrong. To test the log itself, take the web-server block out of the way first; to test the whole arrangement, put it back. Confusing the two costs an afternoon.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And the seventh, which is the reason this post exists: &lt;strong&gt;judging a block by the HTTP status it returns.&lt;/strong&gt; On the armed cron entry point, all four independently recorded statuses read 200 — client, web server, process manager, and PHP's own status at shutdown — the same four values the unblocked run produced while draining all fifty actions. Status and outcome are two different facts. Only one of them tells you whether your queue is safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. The whole arrangement, and what it costs
&lt;/h2&gt;

&lt;p&gt;Put together: trigger cron from a system scheduler through WP-CLI's cron command, which carries core's own events and the Action Scheduler queue in one pass; close &lt;code&gt;wp-cron.php&lt;/code&gt; in the web server, where the refusal is observable and PHP never starts; log every hit that reaches PHP anyway and &lt;em&gt;then&lt;/em&gt; close it again inside WordPress, conditionally, so the file is safe to deploy anywhere and no refusal goes unrecorded; suppress Action Scheduler's outbound async dispatch, knowing it closes the door WordPress opens and not the endpoint itself; unhook the queue runner outside WP-CLI at a priority verified to run after Action Scheduler attaches; keep a second tripwire on non-CLI queue runs, with its manual-run blind spot known rather than assumed away; and leave the admin "Run" button alone, because a person retrying one action by hand is not the problem you set out to solve.&lt;/p&gt;

&lt;p&gt;The result is the same hook doing the same work as before, reached from a process you scheduled instead of a request someone else triggered.&lt;/p&gt;

&lt;p&gt;What it costs you, stated plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No web fallback.&lt;/strong&gt; If the system scheduler stops, the queue stops. Monitor the scheduler, or monitor the pending count — something has to watch it now that nothing will accidentally cover for you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One exemption holding up the whole path.&lt;/strong&gt; Step 5 drains the queue only because step 4 stands down under WP-CLI. Anyone tightening that guard later will get a cron entry that reports success and runs zero actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One door deliberately left open.&lt;/strong&gt; The manual "Run" button still works — one action per click, behind a capability check and a nonce — and this arrangement makes no attempt to close it. The cost isn't the door, it's that neither tripwire can see through it, which is why the pending count and the execution-context label, not a log line of your own, are the things to watch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A log that grows in proportion to the noise.&lt;/strong&gt; The entry-point tripwire writes a line for every hit on &lt;code&gt;wp-cron.php&lt;/code&gt; that reaches PHP. Behind the nginx block that should be nothing at all; without it, a busy endpoint can produce real volume. That is information rather than a nuisance — but rotate the log, and if the volume is high, the answer is step 2a rather than deleting the line.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fleet decision you cannot dodge.&lt;/strong&gt; The nginx block is unconditional. Either deploy it per-site, or rely on the conditional in-WordPress block for the sites that haven't migrated yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One thing this post deliberately does not tell you: whether any of it is worth doing for your site. That depends on what a drain costs you, which is a measurement about your workload and your worker pool — and, as promised at the top, not a number anyone should borrow from someone else's stack.&lt;/p&gt;

</description>
      <category>wordpress</category>
      <category>php</category>
      <category>devops</category>
      <category>security</category>
    </item>
    <item>
      <title>Your Stopped GCP VM Is Not Guaranteed to Start Again</title>
      <dc:creator>Eduardo Pittol</dc:creator>
      <pubDate>Wed, 22 Jul 2026 19:46:37 +0000</pubDate>
      <link>https://dev.to/edpittol/your-stopped-gcp-vm-is-not-guaranteed-to-start-again-3159</link>
      <guid>https://dev.to/edpittol/your-stopped-gcp-vm-is-not-guaranteed-to-start-again-3159</guid>
      <description>&lt;p&gt;The setup: an &lt;strong&gt;ephemeral staging box&lt;/strong&gt;. Nobody reviews releases at 2 AM, so a scheduler auto-stops it every evening after business hours — you don't pay for a machine that's asleep. It's a deliberately disposable environment, which is exactly why the next part stung.&lt;/p&gt;

&lt;p&gt;One morning you press start and get this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The zone &lt;code&gt;us-central1-f&lt;/code&gt; does not have enough resources available to fulfill the request.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Under the hood that's &lt;code&gt;ZONE_RESOURCE_POOL_EXHAUSTED&lt;/code&gt;. Not quota. Not billing. Not IAM. The instance is right there in the console, stopped, exactly as you left it — and it won't turn on. Google Cloud has simply &lt;strong&gt;run out of your machine type in that zone&lt;/strong&gt;. It's called a stockout, and "just start it in another zone" turns out to be impossible. This post is the recovery script I keep on hand, the one concept that makes it work, and what I learned about how GCP hands out capacity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a stopped VM can even fail to start
&lt;/h2&gt;

&lt;p&gt;It ran fine yesterday. Why can't it start today?&lt;/p&gt;

&lt;p&gt;Because &lt;strong&gt;capacity errors apply only to &lt;em&gt;new&lt;/em&gt; requests&lt;/strong&gt;, and a stopped VM has already handed its compute back. GCP tracks capacity &lt;strong&gt;per machine type, per zone&lt;/strong&gt; — there is no single generic pool of servers. When your VM stopped, that capacity returned to the &lt;code&gt;e2-medium&lt;/code&gt; pool in that zone. Pressing start is a brand-new request against whatever's free &lt;em&gt;right now&lt;/em&gt;. If the zone filled up with other people's &lt;code&gt;e2-medium&lt;/code&gt; instances overnight, your start loses the race. Stopping to save money quietly means giving up your seat and hoping one's free when you come back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix in one idea: make the disk portable
&lt;/h2&gt;

&lt;p&gt;A stopped VM's boot disk is a &lt;strong&gt;zonal&lt;/strong&gt; resource. It lives in exactly one zone and it cannot move. When you start the VM, GCP has to find capacity for that machine type &lt;em&gt;in that specific zone&lt;/em&gt;, because that's where the disk is nailed down. If the zone is stocked out for your shape, you're holding a disk you can't boot and can't relocate. Retrying &lt;code&gt;start&lt;/code&gt; just hammers the one zone that already said no.&lt;/p&gt;

&lt;p&gt;The escape hatch: &lt;strong&gt;a zonal disk can't cross zones, but an image can.&lt;/strong&gt; So you snapshot the stuck disk into an image, delete the instance, and recreate it from that image in whichever zone (and machine type) actually has room.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp9868lo4pwfylv9fkulw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp9868lo4pwfylv9fkulw.png" alt="A zonal boot disk is pinned to one zone; snapshotting it into a global recovery image lets the VM be recreated in any zone that has capacity." width="800" height="267"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The script
&lt;/h2&gt;

&lt;p&gt;Set the four variables at the top and run it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

&lt;span class="nv"&gt;PROJECT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"my-project"&lt;/span&gt;
&lt;span class="nv"&gt;INSTANCE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"stg-my-app"&lt;/span&gt;
&lt;span class="nv"&gt;IMAGE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;INSTANCE&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;-recovery-image"&lt;/span&gt;

&lt;span class="c"&gt;# Standard-tier fallbacks, cheapest first. All x86 + pd-balanced compatible,&lt;/span&gt;
&lt;span class="c"&gt;# so no Arm families (T2A/C4A/N4A) and no Hyperdisk-only ones (C4/C3D) —&lt;/span&gt;
&lt;span class="c"&gt;# the recovery image simply can't boot on those.&lt;/span&gt;
&lt;span class="nv"&gt;MACHINE_TYPES&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;e2-medium e2-standard-2 n2d-standard-2 t2d-standard-2 n1-standard-2&lt;span class="o"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# 1. Find where the stopped instance currently lives.&lt;/span&gt;
&lt;span class="nv"&gt;ZONE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;gcloud compute instances list &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"name=&lt;/span&gt;&lt;span class="nv"&gt;$INSTANCE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"value(zone.basename())"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ZONE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Couldn't find &lt;/span&gt;&lt;span class="nv"&gt;$INSTANCE&lt;/span&gt;&lt;span class="s2"&gt;."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;# 2. Discover every UP zone in the instance's region — no hardcoded zone&lt;/span&gt;
&lt;span class="c"&gt;#    list to maintain as GCP adds or retires zones.&lt;/span&gt;
&lt;span class="nv"&gt;REGION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;ZONE&lt;/span&gt;&lt;span class="p"&gt;%-*&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;                       &lt;span class="c"&gt;# us-central1-f -&amp;gt; us-central1&lt;/span&gt;
&lt;span class="nv"&gt;ZONES&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;gcloud compute zones list &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"region:&lt;/span&gt;&lt;span class="nv"&gt;$REGION&lt;/span&gt;&lt;span class="s2"&gt; AND status=UP"&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"value(name)"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="k"&gt;${#&lt;/span&gt;&lt;span class="nv"&gt;ZONES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt; &lt;span class="nt"&gt;-gt&lt;/span&gt; 0 &lt;span class="o"&gt;]]&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"No UP zones in &lt;/span&gt;&lt;span class="nv"&gt;$REGION&lt;/span&gt;&lt;span class="s2"&gt;."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;# 3. Read the config we need to recreate the instance faithfully.&lt;/span&gt;
field&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; gcloud compute instances describe &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$INSTANCE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--zone&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ZONE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"value(&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;)"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="nv"&gt;NETWORK&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;field &lt;span class="s2"&gt;"networkInterfaces[0].network.basename()"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;SUBNET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;field &lt;span class="s2"&gt;"networkInterfaces[0].subnetwork.basename()"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;DISK_SIZE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;field &lt;span class="s2"&gt;"disks[0].diskSizeGb"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# 4. Freeze the boot disk into a GLOBAL image. This is the key move:&lt;/span&gt;
&lt;span class="c"&gt;#    a zonal disk can't change zones, but an image can.&lt;/span&gt;
gcloud compute images create &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--source-disk&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$INSTANCE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--source-disk-zone&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ZONE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;span class="c"&gt;# 5. Delete the stuck instance so its name is free to reuse.&lt;/span&gt;
gcloud compute instances delete &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$INSTANCE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--zone&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ZONE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--quiet&lt;/span&gt;

&lt;span class="c"&gt;# 6. Cheapest type first: hold each machine type and sweep every zone before&lt;/span&gt;
&lt;span class="c"&gt;#    escalating to a costlier one. Lands on the cheapest shape that has&lt;/span&gt;
&lt;span class="c"&gt;#    capacity anywhere, and only pays more when a type is stocked out&lt;/span&gt;
&lt;span class="c"&gt;#    region-wide. (MACHINE_TYPES is already ordered cheapest -&amp;gt; priciest.)&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;mtype &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;MACHINE_TYPES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  for &lt;/span&gt;zone &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;ZONES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Trying &lt;/span&gt;&lt;span class="nv"&gt;$mtype&lt;/span&gt;&lt;span class="s2"&gt; in &lt;/span&gt;&lt;span class="nv"&gt;$zone&lt;/span&gt;&lt;span class="s2"&gt;..."&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;gcloud compute instances create &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$INSTANCE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
      &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--zone&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$zone&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--machine-type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$mtype&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
      &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--network&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$NETWORK&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--subnet&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SUBNET&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
      &lt;span class="nt"&gt;--boot-disk-size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;DISK_SIZE&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;GB"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
      &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Back up in &lt;/span&gt;&lt;span class="nv"&gt;$zone&lt;/span&gt;&lt;span class="s2"&gt; as &lt;/span&gt;&lt;span class="nv"&gt;$mtype&lt;/span&gt;&lt;span class="s2"&gt;."&lt;/span&gt;
      gcloud compute images delete &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--quiet&lt;/span&gt;
      &lt;span class="nb"&gt;exit &lt;/span&gt;0
    &lt;span class="k"&gt;fi
  done
done

&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"No capacity anywhere. Recovery image '&lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="s2"&gt;' kept — recreate by hand."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
&lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Step 6 is the heart of it: a nested loop that holds the &lt;strong&gt;cheapest machine type&lt;/strong&gt; and sweeps every zone before moving up to a costlier type — so you land on the cheapest shape that has capacity anywhere, and only pay more when the cheap one is stocked out region-wide. The zone list isn't hardcoded — it's pulled from the instance's own region at runtime, so the script keeps working as Google adds or retires zones.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk11aq8r3q657kh139fzd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk11aq8r3q657kh139fzd.png" alt="A grid of machine types (rows, cheapest on top) by zones (columns); the cheapest type is out of capacity in every zone, and the run succeeds on the next type up in zone b." width="800" height="337"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Cheapest type first: sweep it across every zone before paying for a bigger one, and stop at the first hit.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas that cost me retries
&lt;/h2&gt;

&lt;p&gt;The machine-type list isn't arbitrary. Two constraints are baked in, both learned the hard way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No Arm.&lt;/strong&gt; The image is built from an x86 box. T2A, C4A, and N4A are Arm — the image won't boot on them. That's not a capacity failure, it's a hard incompatibility, and each one wastes a retry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No Hyperdisk-only families.&lt;/strong&gt; C4/C4D/C3D want Hyperdisk boot disks, not the &lt;code&gt;pd-balanced&lt;/code&gt; this instance uses. Same story: they fail for the wrong reason.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the rule I'd tattoo on my hand: &lt;strong&gt;the recovery image is only deleted after a new instance is confirmed up.&lt;/strong&gt; You are deliberately deleting your only instance in step 5. If the script also deleted the image on failure, a total stockout would leave you with &lt;em&gt;nothing&lt;/em&gt;. On failure it keeps the image — that's your recovery point.&lt;/p&gt;

&lt;h2&gt;
  
  
  How GCP hands out capacity (so you can bias the odds)
&lt;/h2&gt;

&lt;p&gt;The fallback list looks cheapest-first, but it's really &lt;em&gt;availability&lt;/em&gt;-first — it climbs GCP's capacity gradient:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You can't check first.&lt;/strong&gt; There's no API that tells you whether a type has free capacity in a zone. On-demand capacity is opaque; the only signal is attempting the create and catching the stockout. Every strategy is fundamentally attempt-and-fall-back. (I did try &lt;code&gt;gcloud alpha compute advice capacity&lt;/code&gt; to get ahead of it — it's gated behind an Alpha allowlist my account isn't on, so: &lt;code&gt;403&lt;/code&gt;.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commodity beats cutting-edge.&lt;/strong&gt; The newest series (C3, C4, GPUs) are the scarcest; E2 and N2 are the safe fallbacks. A ladder like C3 → N2 → E2 usually finds room.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Smaller beats bigger.&lt;/strong&gt; Large shapes need contiguous capacity on a single host, so they stock out more. A smaller shape in the same series is likelier to fit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zone beats type.&lt;/strong&gt; A stockout is per type &lt;em&gt;and&lt;/em&gt; per zone, so sweeping zones is often a bigger lever than swapping types. The loop leans on this: it holds the cheapest type and tries &lt;em&gt;every&lt;/em&gt; zone before escalating — diversifying zones first, and paying more only when a type is out region-wide.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spot and on-demand are separate pools.&lt;/strong&gt; Spot has its own capacity, but it can be preempted — no good when you need the box reliably up for a test session.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Lock down who can run it
&lt;/h2&gt;

&lt;p&gt;"Let a teammate restart the staging box" sounds harmless, but this script &lt;strong&gt;deletes an instance&lt;/strong&gt;. You do not want to hand out &lt;code&gt;compute.admin&lt;/code&gt; for that.&lt;/p&gt;

&lt;p&gt;I scoped a custom role to exactly the verbs the script needs — create/delete instance, create/delete image, read the config, and (for the dynamic zone lookup) &lt;code&gt;compute.zones.list&lt;/code&gt; — and pinned it with an IAM condition so it only applies to &lt;em&gt;this one instance, its disk, and its recovery image&lt;/em&gt;. The blast radius of the restart button should be one VM.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this happens — and why it's not just you
&lt;/h2&gt;

&lt;p&gt;It's tempting to assume you misconfigured something. You didn't. A zone is a finite pile of physical machines, and sometimes the shape you want isn't in the pile right now. Google's own docs are blunt about it: resource errors are unrelated to your quota and apply only to the exact resource you asked for, at the moment you asked — &lt;a href="https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-resource-availability" rel="noopener noreferrer"&gt;try a different zone, a different machine type, or again later&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;And it's a well-worn rake. Google's &lt;a href="https://groups.google.com/g/gce-discussion/c/Lfyk38giqK8" rel="noopener noreferrer"&gt;gce-discussion group&lt;/a&gt; has threads going back years — one stockout stretched close to 24 hours across multiple zones, with a user flatly saying it "affects us to the point we cant use GCP."&lt;/p&gt;

&lt;p&gt;If uptime genuinely matters, this script is the wrong tool, and it's worth knowing what the right ones are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reservations&lt;/strong&gt; hold capacity for a specific machine type in a specific zone. You pay whether you use it or not, but the capacity is &lt;em&gt;yours&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Committed use discounts (CUDs)&lt;/strong&gt; are easy to confuse with reservations, but they're a &lt;em&gt;billing&lt;/em&gt; discount across a series in a region — not a capacity guarantee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MIG instance flexibility&lt;/strong&gt; is GCP's native "try types until one is available" primitive, with a ranked list of machine types. It's built for stateless, interchangeable VMs, though — wrapping a single stateful staging box in a managed instance group is far more machinery than the problem deserves. The &lt;a href="https://medium.com/google-cloud/beyond-try-again-later-a-strategic-guide-to-obtaining-high-demand-resources-in-google-cloud-0f7a2f152b7f" rel="noopener noreferrer"&gt;&lt;em&gt;Beyond "Try Again Later"&lt;/em&gt;&lt;/a&gt; writeup is a good map of these options.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The cheapest fix is prevention
&lt;/h2&gt;

&lt;p&gt;The highest-leverage, lowest-effort move isn't the script at all: &lt;strong&gt;don't run staging on scarce hardware.&lt;/strong&gt; If the box is on C3/C4/GPU, that's very likely why it won't start. Staging rarely needs the newest silicon — move it to E2 or N2 and the stockouts mostly disappear. The script is the seatbelt for the day it happens anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A stopped VM is not a reservation.&lt;/strong&gt; Stopping it releases capacity back to the pool; starting it is a fresh request that can fail.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity is per type, per zone, and opaque.&lt;/strong&gt; There's no API to check it in advance — you attempt and fall back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zonal disks pin you to a zone.&lt;/strong&gt; Crossing zones means routing through a global artifact — an image or a snapshot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commodity types stock out far less.&lt;/strong&gt; The cheapest fix is not being on scarce hardware in the first place.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When a script deletes your only copy, protect the recovery point above everything else.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cloud markets itself as infinite. It isn't — it's a very large, very finite pile of other people's computers, and once in a while the pile you want is empty. Plan for the empty pile.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sources: &lt;a href="https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-resource-availability" rel="noopener noreferrer"&gt;GCP resource-availability troubleshooting&lt;/a&gt; · &lt;a href="https://groups.google.com/g/gce-discussion/c/Lfyk38giqK8" rel="noopener noreferrer"&gt;gce-discussion stockout thread&lt;/a&gt; · &lt;a href="https://medium.com/google-cloud/beyond-try-again-later-a-strategic-guide-to-obtaining-high-demand-resources-in-google-cloud-0f7a2f152b7f" rel="noopener noreferrer"&gt;Beyond "Try Again Later" (Google Cloud Community)&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>gcp</category>
      <category>googlecloud</category>
      <category>devops</category>
      <category>bash</category>
    </item>
    <item>
      <title>I had an application taking 10 seconds to load. Turned out file ops functions on aren't real filesystem calls due S3 stream wrapper. This post brings the results of a POC that I built to demonstrate the impact of these invisible of these network calls.</title>
      <dc:creator>Eduardo Pittol</dc:creator>
      <pubDate>Thu, 16 Jul 2026 18:26:55 +0000</pubDate>
      <link>https://dev.to/edpittol/i-had-an-application-taking-10-seconds-to-load-turned-out-file-ops-functions-on-arent-real-38j5</link>
      <guid>https://dev.to/edpittol/i-had-an-application-taking-10-seconds-to-load-turned-out-file-ops-functions-on-arent-real-38j5</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/edpittol/your-fileexists-is-secretly-a-network-call-4236" class="crayons-story__hidden-navigation-link"&gt;Your file_exists() Is Secretly a Network Call&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/edpittol" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3881058%2F17f670b3-e8eb-4e4e-b7ad-128c71b6e970.jpeg" alt="edpittol profile" class="crayons-avatar__image" width="460" height="460"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/edpittol" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Eduardo Pittol
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Eduardo Pittol
                
              
              &lt;div id="story-author-preview-content-4160257" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/edpittol" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3881058%2F17f670b3-e8eb-4e4e-b7ad-128c71b6e970.jpeg" class="crayons-avatar__image" alt="" width="460" height="460"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Eduardo Pittol&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/edpittol/your-fileexists-is-secretly-a-network-call-4236" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Jul 16&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/edpittol/your-fileexists-is-secretly-a-network-call-4236" id="article-link-4160257"&gt;
          Your file_exists() Is Secretly a Network Call
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/php"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;php&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/performance"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;performance&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/s3"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;s3&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/aws"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;aws&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/edpittol/your-fileexists-is-secretly-a-network-call-4236#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            6 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>aws</category>
      <category>backend</category>
      <category>performance</category>
      <category>php</category>
    </item>
    <item>
      <title>Your file_exists() Is Secretly a Network Call</title>
      <dc:creator>Eduardo Pittol</dc:creator>
      <pubDate>Thu, 16 Jul 2026 18:14:10 +0000</pubDate>
      <link>https://dev.to/edpittol/your-fileexists-is-secretly-a-network-call-4236</link>
      <guid>https://dev.to/edpittol/your-fileexists-is-secretly-a-network-call-4236</guid>
      <description>&lt;h2&gt;
  
  
  The one-line horror
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nb"&gt;file_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="nv"&gt;$path&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// ... use the cached file&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing about that line looks dangerous. It's the kind of guard clause you've written a thousand times. But suppose &lt;code&gt;$path&lt;/code&gt; is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;s3://my-bucket/cache/style-4f2a.css
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now that single &lt;code&gt;file_exists()&lt;/code&gt; is not a syscall. It's a &lt;code&gt;HeadObject&lt;/code&gt; request to Amazon S3 — a network round trip that leaves your server, waits for S3, and comes back. Against distant S3, that round trip costs tens of milliseconds.&lt;/p&gt;

&lt;p&gt;Here's the twist that makes it dangerous: the AWS SDK caches that result in memory, so a &lt;em&gt;second&lt;/em&gt; check of the same path is instant. The cost hides behind the cache. But every web request is a fresh PHP process with an empty cache, so the first touch of every path pays full price — every request. And writes never get even that reprieve: nothing caches a &lt;code&gt;PutObject&lt;/code&gt;, so every &lt;code&gt;file_put_contents()&lt;/code&gt; over S3 is a full round trip, every time. Reads can ride the cache; writes always cross the wire.&lt;/p&gt;

&lt;p&gt;Local &lt;code&gt;stat()&lt;/code&gt; costs microseconds. The code doesn't change — only the string in &lt;code&gt;$path&lt;/code&gt; does. PHP stream wrappers make remote storage look like a local disk, and code written for local-disk economics keeps compiling, keeps passing tests, and silently falls off a performance cliff the moment the path points at a bucket.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why an innocent function makes a network call
&lt;/h2&gt;

&lt;p&gt;A stream wrapper is a class registered against a URL scheme:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="nb"&gt;stream_wrapper_register&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="s1"&gt;'s3'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;S3\StreamWrapper&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;class&lt;/span&gt; &lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once &lt;code&gt;s3://&lt;/code&gt; is registered, PHP routes every filesystem call for that scheme through the wrapper: &lt;code&gt;fopen()&lt;/code&gt;, &lt;code&gt;fread()&lt;/code&gt;, &lt;code&gt;file_get_contents()&lt;/code&gt;, &lt;code&gt;stat()&lt;/code&gt;, &lt;code&gt;is_dir()&lt;/code&gt;, &lt;code&gt;unlink()&lt;/code&gt;. The whole point is that your code doesn't have to know or care whether it's talking to a disk or a bucket — the API surface is identical.&lt;/p&gt;

&lt;p&gt;That's exactly the trap. Each operation quietly maps to an S3 API request:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PHP call&lt;/th&gt;
&lt;th&gt;S3 request&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;file_exists()&lt;/code&gt; / &lt;code&gt;stat()&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;HeadObject&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;is_dir()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ListObjects&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;unlink()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;DeleteObject&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;file_put_contents()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;PutObject&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you've ever hunted down an &lt;strong&gt;N+1 query&lt;/strong&gt; in an ORM, you already understand the failure mode. N+1 is a loop that fires one database round trip per item instead of batching them. This is the same shape — one network round trip per filesystem call — except the round trip hides behind a function whose name says "filesystem." Nobody profiles &lt;code&gt;file_exists()&lt;/code&gt;. Static analysis won't flag it. And in local development, where the wrapper points at a fast disk or a nearby S3-compatible service, it's instant. The cost only shows up as latency against real, distant S3 in production — the worst possible place to discover it.&lt;/p&gt;

&lt;p&gt;The numbers below come from a controlled rig: a local MinIO standing in for S3, with &lt;a href="https://github.com/Shopify/toxiproxy" rel="noopener noreferrer"&gt;Toxiproxy&lt;/a&gt; injecting a fixed round-trip time (RTT) of 0, 10, 20, or 40 ms — roughly the spread from same-AZ to cross-region.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark: cost scales with the network
&lt;/h2&gt;

&lt;p&gt;Median latency per call, S3 backend, as RTT climbs (local disk stays ≤ 0.01 ms for all operations at every RTT):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;S3 operation&lt;/th&gt;
&lt;th&gt;RTT 0 ms&lt;/th&gt;
&lt;th&gt;10 ms&lt;/th&gt;
&lt;th&gt;20 ms&lt;/th&gt;
&lt;th&gt;40 ms&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;file_exists&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.002 ms&lt;/td&gt;
&lt;td&gt;0.022 ms&lt;/td&gt;
&lt;td&gt;0.025 ms&lt;/td&gt;
&lt;td&gt;0.026 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;stat&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.002 ms&lt;/td&gt;
&lt;td&gt;0.022 ms&lt;/td&gt;
&lt;td&gt;0.029 ms&lt;/td&gt;
&lt;td&gt;0.032 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;file_put_contents&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.460 ms&lt;/td&gt;
&lt;td&gt;17.961 ms&lt;/td&gt;
&lt;td&gt;29.626 ms&lt;/td&gt;
&lt;td&gt;49.551 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things jump out. &lt;strong&gt;Reads barely move with latency.&lt;/strong&gt; Checking the same path in a loop, &lt;code&gt;file_exists()&lt;/code&gt; stays around 0.02 ms no matter the RTT — the SDK's in-memory stat cache absorbs the repeats, so they never cross the wire twice. &lt;strong&gt;Writes track RTT almost linearly&lt;/strong&gt;, because nothing caches a &lt;code&gt;PutObject&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Now scale to a page. A single real page in the incident below fired &lt;strong&gt;88&lt;/strong&gt; filesystem ops, so reconstruct the page cost as median × 88:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;S3 op × 88&lt;/th&gt;
&lt;th&gt;RTT 0 ms&lt;/th&gt;
&lt;th&gt;10 ms&lt;/th&gt;
&lt;th&gt;20 ms&lt;/th&gt;
&lt;th&gt;40 ms&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;file_exists&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.14 ms&lt;/td&gt;
&lt;td&gt;1.95 ms&lt;/td&gt;
&lt;td&gt;2.19 ms&lt;/td&gt;
&lt;td&gt;2.25 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;stat&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.17 ms&lt;/td&gt;
&lt;td&gt;1.94 ms&lt;/td&gt;
&lt;td&gt;2.55 ms&lt;/td&gt;
&lt;td&gt;2.81 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;file_put_contents&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;128.5 ms&lt;/td&gt;
&lt;td&gt;1,580.6 ms&lt;/td&gt;
&lt;td&gt;2,607.1 ms&lt;/td&gt;
&lt;td&gt;4,360.5 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 40 ms RTT — a cross-region hop — 88 writes cost &lt;strong&gt;~4.4 seconds&lt;/strong&gt;. That's the cliff, and every extra millisecond between you and S3 makes it steeper.&lt;/p&gt;

&lt;p&gt;But look at the read rows: 2 ms for 88 checks. That looks harmless — and it's the most misleading number in the table. It's low only because the benchmark checks the same paths repeatedly inside one long-lived process, so the SDK's in-memory stat cache absorbs the repeats. Nothing crosses the wire twice.&lt;/p&gt;

&lt;p&gt;PHP in production doesn't work that way. Every web request is served by a fresh process with an empty cache, and that cache is gone the moment the request ends — the &lt;code&gt;LruArrayCache&lt;/code&gt; lives in process memory, not in Redis or on disk. So the first touch of each &lt;em&gt;distinct&lt;/em&gt; path is always a full round trip, and nothing carries over to the next request. If a page checks 73 distinct files, that's 73 cold round trips — every request. Which is exactly what happened next.&lt;/p&gt;

&lt;h2&gt;
  
  
  The war story: a 10-second homepage
&lt;/h2&gt;

&lt;p&gt;This isn't hypothetical. A high-traffic WordPress site, with its media library backed by S3, had a homepage that took &lt;strong&gt;10 seconds&lt;/strong&gt; to return.&lt;/p&gt;

&lt;p&gt;New Relic told the story immediately. &lt;code&gt;GET /&lt;/code&gt; returned HTTP 200 in 10.07 s — with an &lt;strong&gt;empty database-queries tab&lt;/strong&gt;. No slow SQL. Roughly 90% of the time was synchronous S3 traffic, all inside a single page view:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Segment&lt;/th&gt;
&lt;th&gt;Calls&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;s3.amazonaws.com&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;73&lt;/td&gt;
&lt;td&gt;3,281 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Guzzle &lt;code&gt;CurlMultiHandler::tick&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;59&lt;/td&gt;
&lt;td&gt;2,623 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stream wrapper closure&lt;/td&gt;
&lt;td&gt;88&lt;/td&gt;
&lt;td&gt;1,702 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AwsClient::execute&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;124&lt;/td&gt;
&lt;td&gt;1,582 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Guzzle &lt;code&gt;CurlMultiHandler::execute&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;td&gt;873 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;73 real S3 calls to render one homepage.&lt;/strong&gt; The theme's CSS layer was the culprit: for every style handle, on every breakpoint, on every request, it called &lt;code&gt;file_exists()&lt;/code&gt; on the generated stylesheet to decide whether to regenerate — dozens of &lt;code&gt;HeadObject&lt;/code&gt;s per page — &lt;em&gt;even though it already held a persisted flag saying the cache was valid.&lt;/em&gt; Those were 73 &lt;strong&gt;distinct&lt;/strong&gt; paths, so the in-process cache never helped; and because the SDK's stat cache was request-scoped — an in-memory cache that dies with the PHP process — nothing survived to the next request either. Every check was a cold round trip.&lt;/p&gt;

&lt;p&gt;The fix was a tour of &lt;strong&gt;where a cache can live&lt;/strong&gt; — each layer trading something different:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Persistent stat cache.&lt;/strong&gt; Replace the request-scoped &lt;code&gt;LruArrayCache&lt;/code&gt; with an adapter over the WordPress object cache (Redis). It implements the SDK's &lt;code&gt;CacheInterface&lt;/code&gt; and is injected when the wrapper is registered — no plugin patching required. Now a stat survives across requests. The cost: a Redis hop instead of memory, still far cheaper than S3.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local-disk-first.&lt;/strong&gt; For generated CSS, write to local disk, mirror to S3 asynchronously after the write, and restore from S3 only if the local copy goes missing. Reads become genuine microsecond &lt;code&gt;stat()&lt;/code&gt;s again, and — crucially — the writes stop hitting S3 on the request path. The cost: you trade strong read-after-write consistency across nodes for latency, a deliberate, documented trade-off.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The page dropped from ~10 seconds to well under a second. Same features, same S3 bucket — the only thing that changed was refusing to let a network call keep masquerading as a syscall.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to spot this in your own code
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Audit filesystem calls near wrapped paths.&lt;/strong&gt; Grep for &lt;code&gt;file_exists&lt;/code&gt;, &lt;code&gt;stat&lt;/code&gt;, &lt;code&gt;is_dir&lt;/code&gt;, &lt;code&gt;unlink&lt;/code&gt;, and &lt;code&gt;file_put_contents&lt;/code&gt;, and ask, for each, whether the path could ever be &lt;code&gt;s3://&lt;/code&gt; (or &lt;code&gt;gs://&lt;/code&gt;, or any wrapper). Pay special attention to anything inside a loop and to &lt;strong&gt;distinct&lt;/strong&gt; paths — those never hit the cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read APM segment breakdowns, not just the total.&lt;/strong&gt; A slow request with an &lt;em&gt;empty&lt;/em&gt; SQL tab is the tell. Look for time in the storage SDK and the HTTP handler.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distrust request-scoped caches.&lt;/strong&gt; A cache that isn't shared across processes does nothing for a per-request PHP model. Confirm where your cache actually lives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remember that a cache can't save a write.&lt;/strong&gt; Reads can be cached; every &lt;code&gt;PutObject&lt;/code&gt; is a real round trip. Batch writes, defer them off the request path, or keep them local and mirror asynchronously.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the principle underneath all of it: &lt;strong&gt;when an abstraction changes the cost model by orders of magnitude, it has to leak that cost somewhere.&lt;/strong&gt; An abstraction that hides a 50 ms network call behind a microsecond-shaped function isn't a convenience — it's a latency bug waiting for production traffic. Put the cost back where you can see it: a persistent cache, a batched call, or a local-first layer. Don't let &lt;code&gt;file_exists()&lt;/code&gt; keep wearing a syscall's clothes.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All the measurements in this post come from a small, self-contained benchmark rig — MinIO standing in for S3, Toxiproxy injecting the round-trip latency, and the PHP scripts that produced every table above. It's on GitHub: &lt;a href="https://github.com/edpittol/s3-stream-multiple-operations" rel="noopener noreferrer"&gt;edpittol/s3-stream-multiple-operations&lt;/a&gt;. Clone it and reproduce the numbers.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>php</category>
      <category>performance</category>
      <category>s3</category>
      <category>aws</category>
    </item>
    <item>
      <title>Stop your team from rebuilding the same AI skills: a shared catalog that maintains itself</title>
      <dc:creator>Eduardo Pittol</dc:creator>
      <pubDate>Sun, 14 Jun 2026 18:39:06 +0000</pubDate>
      <link>https://dev.to/edpittol/stop-your-team-from-rebuilding-the-same-ai-skills-a-shared-catalog-that-maintains-itself-35l5</link>
      <guid>https://dev.to/edpittol/stop-your-team-from-rebuilding-the-same-ai-skills-a-shared-catalog-that-maintains-itself-35l5</guid>
      <description>&lt;p&gt;If your team has started using skills for AI coding agents, you've probably hit this already: someone uses an existing skill or builds one to run the test suite a certain way. Two weeks later, someone else does it &lt;em&gt;their own&lt;/em&gt; way for the same job — different steps, different conventions, a different result — because they had no idea the first one existed. Now the agent does the "same" task one way on your machine and another way on theirs.&lt;/p&gt;

&lt;p&gt;That's the part that actually hurts. It isn't just the wasted hour rebuilding something that already existed (though that adds up). It's that the team quietly drifts into a dozen slightly different ways of solving the &lt;em&gt;same&lt;/em&gt; problem — inconsistent results, no shared standard, and no single place to fix anything when it breaks. Multiply it across a few people and a few months, and "skills" stops being leverage and starts being entropy.&lt;/p&gt;

&lt;p&gt;I ran into exactly this. The fix wasn't more discipline — it was a shared catalog that's so easy to contribute to that nobody has an excuse not to — and the whole team reaches for &lt;em&gt;one&lt;/em&gt; skill instead of each rolling their own. The trick that made it stick: a skill whose only job is to add other skills to the catalog. This post is how that works, and how to build your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea: one catalog, one source of truth
&lt;/h2&gt;

&lt;p&gt;The catalog is a single repo. Skills live under &lt;code&gt;skills/&amp;lt;category&amp;gt;/&lt;/code&gt;, where the categories are deliberately boring:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;engineering&lt;/strong&gt; — development, testing, code review, automation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;productivity&lt;/strong&gt; — organization, writing, planning, day-to-day workflows&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;misc&lt;/strong&gt; — everything else&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At the root there's a master table — one row per skill, with its category and a link back to the source it came from. That table is the single source of truth for what the team has. New teammate? Point them at the table. Wondering if a skill already exists? Check the table before you build.&lt;/p&gt;

&lt;p&gt;And here's the payoff that's easy to miss: when everyone reaches for the &lt;em&gt;same&lt;/em&gt; skill, the catalog becomes a feedback loop. Someone hits a rough edge — a flaky step, a missing case, a prompt the agent keeps misreading — and they fix it in that one canonical skill. The fix lands for the whole team at once. Troubleshooting stops being something each person quietly redoes in private and starts &lt;em&gt;accumulating&lt;/em&gt;: a shared skill gets sharper every time anyone uses it. Twelve private copies just stay broken in twelve different ways.&lt;/p&gt;

&lt;p&gt;It's a nice idea. It also rots the instant adding a skill becomes a chore.&lt;/p&gt;

&lt;h2&gt;
  
  
  The catch: cataloging by hand is exactly the friction that kills it
&lt;/h2&gt;

&lt;p&gt;Look at what "just add it to the catalog" actually means by hand:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pick the right category and copy the folder into &lt;code&gt;skills/&amp;lt;category&amp;gt;/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Write an entry in that category's README — in the right alphabetical spot.&lt;/li&gt;
&lt;li&gt;Resolve a permalink back to the source: the upstream repo URL, the &lt;code&gt;org/repo&lt;/code&gt;, and the exact commit hash.&lt;/li&gt;
&lt;li&gt;Regenerate the master table at the root, re-sorted alphabetically, without breaking the existing rows.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nobody wants to do that after every skill. So they don't — and the catalog drifts out of date until it's useless. The friction is the failure mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: a skill that catalogs skills
&lt;/h2&gt;

&lt;p&gt;So I made the bookkeeping itself a skill: &lt;code&gt;catalog-a-skill&lt;/code&gt;. You point it at a skill folder and it does all four steps. Here's the abridged &lt;code&gt;SKILL.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;catalog-a-skill&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Catalogs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;an&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;existing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;skill&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;into&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;this&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;repo's&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;structure.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;when&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;skill&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;folder&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;valid&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;SKILL.md&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;already&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;exists&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;needs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;be&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;added&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;catalog."&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gh"&gt;# Catalog a Skill&lt;/span&gt;

&lt;span class="gu"&gt;## Workflow&lt;/span&gt;

&lt;span class="gu"&gt;### 1. Read the skill&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Read &lt;span class="sb"&gt;`&amp;lt;skill-dir&amp;gt;/SKILL.md`&lt;/span&gt; and extract &lt;span class="sb"&gt;`name`&lt;/span&gt; and &lt;span class="sb"&gt;`description`&lt;/span&gt;.
&lt;span class="p"&gt;-&lt;/span&gt; Gate check 1 — missing frontmatter.
&lt;span class="p"&gt;-&lt;/span&gt; Gate check 2 — folder name ≠ frontmatter &lt;span class="sb"&gt;`name`&lt;/span&gt;.

&lt;span class="gu"&gt;### 2. Choose the category&lt;/span&gt;
Pick engineering / productivity / misc from the description.

&lt;span class="gu"&gt;### 3. Position the skill&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Copy to &lt;span class="sb"&gt;`skills/&amp;lt;category&amp;gt;/&amp;lt;skill-name&amp;gt;/`&lt;/span&gt;.
&lt;span class="p"&gt;-&lt;/span&gt; Resolve the upstream reference from the SOURCE repo:
  commit hash (&lt;span class="sb"&gt;`git rev-parse HEAD`&lt;/span&gt;), remote URL, and &lt;span class="sb"&gt;`org/repo`&lt;/span&gt;.

&lt;span class="gu"&gt;### 4. Update the category README&lt;/span&gt;
Insert a section at its alphabetical position.

&lt;span class="gu"&gt;### 5. Regenerate the master table&lt;/span&gt;
One row per skill, sorted by name, with a permalink to the source commit.

&lt;span class="gu"&gt;### 6. Verify&lt;/span&gt;
Run &lt;span class="sb"&gt;`npx skills list`&lt;/span&gt; and confirm the skill appears.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The whole point: the contributor does the &lt;em&gt;fun&lt;/em&gt; part (writing a useful skill) and hands the &lt;em&gt;tedious, rule-bound&lt;/em&gt; part to the agent. That's the kind of work agents are genuinely good at — deterministic, fiddly, easy to get subtly wrong by hand.&lt;/p&gt;

&lt;h3&gt;
  
  
  The part that makes it trustworthy: it refuses to do the wrong thing
&lt;/h3&gt;

&lt;p&gt;A naive version would happily corrupt the catalog. The real value is in the gates — the cases where it &lt;em&gt;stops&lt;/em&gt; instead of cataloging:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Missing frontmatter.&lt;/strong&gt; If &lt;code&gt;name&lt;/code&gt; or &lt;code&gt;description&lt;/code&gt; is absent, it copies the skill into &lt;code&gt;misc&lt;/code&gt; but does &lt;strong&gt;not&lt;/strong&gt; touch the README or the master table. Then it tells you exactly which field is missing. A broken skill never silently pollutes the source of truth.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name divergence.&lt;/strong&gt; If the folder is named &lt;code&gt;foo&lt;/code&gt; but the frontmatter says &lt;code&gt;name: bar&lt;/code&gt;, that's a red flag — it copies under the original folder name, refuses to catalog, and reports the mismatch.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That "refuse and report" behavior is what turns a convenient script into something a team can trust to run on its catalog.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to build your own
&lt;/h2&gt;

&lt;p&gt;You don't need my repo — you need the pattern. To replicate it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Make the catalog repo.&lt;/strong&gt; &lt;code&gt;skills/&amp;lt;category&amp;gt;/&lt;/code&gt; directories, a root README with a master table, one README per category. Keep the categories few and obvious.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide your source of truth.&lt;/strong&gt; I use the root table; the rule is "if it's not in the table, it doesn't exist." Pick yours and make it explicit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the meta-skill.&lt;/strong&gt; Encode the four steps above as a &lt;code&gt;SKILL.md&lt;/code&gt;. The non-obvious work is the bookkeeping rules: alphabetical insertion, permalink resolution from the &lt;em&gt;source&lt;/em&gt; repo's git remote and commit, and consistent table regeneration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add the gates.&lt;/strong&gt; Decide what "invalid" means for you (missing fields, name mismatch, wrong structure) and make the skill &lt;em&gt;refuse and report&lt;/em&gt; instead of pushing bad entries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lower the barrier to zero.&lt;/strong&gt; The whole bet is that a catalog survives only when contributing costs nothing. The meta-skill is what buys that.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's it. The catalog stays current because keeping it current is now a single command, not a chore — and the team stops rebuilding skills it already has.&lt;/p&gt;

&lt;p&gt;The full working version, gate cases and all, is here: &lt;strong&gt;&lt;a href="https://github.com/edpittol/skills" rel="noopener noreferrer"&gt;github.com/edpittol/skills&lt;/a&gt;&lt;/strong&gt;. Take it apart, adapt the categories and rules to your team, and let your agent keep the books.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>claude</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
