<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Libme</title>
    <description>The latest articles on DEV Community by Libme (@libme).</description>
    <link>https://dev.to/libme</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4062668%2F1762b3d3-a3e8-4856-a46b-270e29821fed.png</url>
      <title>DEV Community: Libme</title>
      <link>https://dev.to/libme</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/libme"/>
    <language>en</language>
    <item>
      <title>Your LLM Bill Jumped After You Added Context: Find the Cache Miss Before You Downgrade the Model</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Sun, 04 Oct 2026 00:44:13 +0000</pubDate>
      <link>https://dev.to/libme/your-llm-bill-jumped-after-you-added-context-find-the-cache-miss-before-you-downgrade-the-model-31io</link>
      <guid>https://dev.to/libme/your-llm-bill-jumped-after-you-added-context-find-the-cache-miss-before-you-downgrade-the-model-31io</guid>
      <description>&lt;p&gt;If your LLM API spend climbed after you added retrieval, a longer system prompt, or tool definitions, the cause is almost always input tokens being reprocessed at full price on every request — not the model you picked. Check the cache fields in the response &lt;code&gt;usage&lt;/code&gt; object first: if cached reads are zero across repeated requests that share a prefix, you have a bug, not a pricing problem. Fixing prefix stability and breakpoint placement is free; downgrading the model costs you quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is the bill rising when output length didn't change?
&lt;/h2&gt;

&lt;p&gt;The first thing to do is stop reasoning about the bill and start logging per-request token accounting. On the Anthropic API, every response carries a &lt;code&gt;usage&lt;/code&gt; object that splits input tokens three ways:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;STABLE_PREAMBLE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# instructions, schema, few-shot examples
&lt;/span&gt;            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;u&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;uncached=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;input_tokens&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;written=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cache_creation_input_tokens&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cache_read_input_tokens&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;output_tokens&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The field that trips people up is &lt;code&gt;input_tokens&lt;/code&gt;: it is the &lt;em&gt;uncached remainder&lt;/em&gt;, not the prompt size. Total prompt size is &lt;code&gt;input_tokens + cache_creation_input_tokens + cache_read_input_tokens&lt;/code&gt;. I have watched a team conclude their prompt was small because &lt;code&gt;input_tokens&lt;/code&gt; read 4K, while the other two fields accounted for ten times that. OpenAI's API exposes the analogous number as &lt;code&gt;usage.prompt_tokens_details.cached_tokens&lt;/code&gt;; the arithmetic lesson is the same.&lt;/p&gt;

&lt;p&gt;Run the same request twice in a row and compare. A healthy second request reads the shared prefix and writes only the delta. If the second request shows a cache write roughly equal to the full prompt and a read of zero, your prefix is not stable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: before you argue about model pricing, prove whether your repeated prefix is being read from cache or reprocessed from scratch.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What silently breaks prompt caching?
&lt;/h2&gt;

&lt;p&gt;Prompt caching is a prefix match over the exact rendered bytes. The rendering order is &lt;code&gt;tools&lt;/code&gt; → &lt;code&gt;system&lt;/code&gt; → &lt;code&gt;messages&lt;/code&gt;, and any byte that changes early invalidates everything after it. That makes the failure mode sneaky: the request still succeeds, the output is still correct, and only the bill notices.&lt;/p&gt;

&lt;p&gt;The classic offender is a system prompt that looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Every request produces a different prefix — nothing after this ever caches
&lt;/span&gt;&lt;span class="n"&gt;system&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are a support assistant.
Current time: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;isoformat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
User: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (plan: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;plan&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)
...500 lines of instructions...
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three invalidators in four lines. The fix is to freeze the prefix and push volatile values past the last breakpoint — into a later message, where a change at turn five invalidates nothing before turn five:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;system&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;STATIC_INSTRUCTIONS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
           &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}}]&lt;/span&gt;

&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Current time: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;. Plan: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;plan&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;]},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Other invalidators worth grepping for in any code that builds a prompt: &lt;code&gt;json.dumps()&lt;/code&gt; without &lt;code&gt;sort_keys=True&lt;/code&gt;, iteration over a &lt;code&gt;set&lt;/code&gt;, tool lists assembled per user (tools render at position zero, so a per-user tool set means no cross-user reuse), and conditional system sections where every feature-flag combination becomes its own distinct prefix. Switching models invalidates too — caches are model-scoped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: anything that varies per request belongs after the last cache breakpoint, and "per request" includes the clock.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where should the cache breakpoint actually go?
&lt;/h2&gt;

&lt;p&gt;Marking the end of the prompt feels right and is usually wrong. If the last block is the user's unique question or freshly retrieved rows, the breakpoint lands after bytes that will never repeat — so every request pays the write premium and no request ever reads it back. The signature in your logs is a cache write on &lt;em&gt;every&lt;/em&gt; request while reads never cover the shared prefix.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Prompt shape&lt;/th&gt;
&lt;th&gt;Where the breakpoint goes&lt;/th&gt;
&lt;th&gt;Failure if you get it wrong&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Big fixed preamble, varying question&lt;/td&gt;
&lt;td&gt;End of the shared preamble&lt;/td&gt;
&lt;td&gt;Write premium on every request, reads near zero&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Growing multi-turn conversation&lt;/td&gt;
&lt;td&gt;Last block of the newest turn&lt;/td&gt;
&lt;td&gt;Whole history reprocessed each turn&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefix differs from byte one per request&lt;/td&gt;
&lt;td&gt;Nowhere — don't cache&lt;/td&gt;
&lt;td&gt;Pure surcharge, no reads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sections with different change rates&lt;/td&gt;
&lt;td&gt;One breakpoint per stability boundary&lt;/td&gt;
&lt;td&gt;Daily-changing context invalidates never-changing tools&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two more constraints that bite in practice, as of late 2026: there is a minimum cacheable prefix (model-dependent, and &lt;em&gt;not&lt;/em&gt; monotonic across generations — newer models cache shorter prefixes than some older ones), and below it caching silently no-ops with no error. And parallel requests over an identical prefix all miss, because an entry only becomes readable once the first response starts streaming. For fan-out, send one request, wait for its first token, then fire the rest.&lt;/p&gt;

&lt;p&gt;The economics are a ratio, not a mystery: a cache write costs more than a plain input token and a cache read costs a fraction of one, so a prefix read back even twice is already ahead. Longer-lived cache entries cost more to write, which means they only pay off across gaps that the default short-lived entry can't bridge. If you want the version of this that requires no breakpoint bookkeeping at all, the Anthropic API's automatic caching places the breakpoint for you and is the right default for ordinary multi-turn chat — it places exactly one breakpoint, so prompts with several stability boundaries still need explicit markers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: put the breakpoint at the end of what repeats, not at the end of the prompt.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When is the batch endpoint worth it?
&lt;/h2&gt;

&lt;p&gt;For any workload where a human isn't waiting on the response, the batch endpoints from the major providers run asynchronously at roughly half the standard per-token rate (as of late 2026 — confirm the current discount in your provider's pricing docs). Nightly summarization, backfills, evaluation runs, bulk classification, and embedding generation all qualify.&lt;/p&gt;

&lt;p&gt;Three things to plan for. Results arrive in arbitrary order, so key them by your own &lt;code&gt;custom_id&lt;/code&gt; and never by position. Completion is a window, not a latency target — build the job so a delayed batch degrades a dashboard rather than blocking a user request. And batch does not compose with everything: pairing it with caching tricks or streaming-dependent code paths usually means reshaping the request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: if no human is blocked on the answer, half-price asynchronous processing is the highest-leverage change you can make without touching prompt quality.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you route to a smaller model?
&lt;/h2&gt;

&lt;p&gt;Eventually, maybe — but it belongs last in the order, after the free wins, because it is the only lever that trades quality. Two things make cascades less attractive than the napkin math suggests. First, caches are model-scoped, so splitting traffic across two models forfeits cache reuse between them; a cascade that saves 40% on paper can lose most of it to cold prefixes. Second, the honest unit is cost per &lt;em&gt;completed task&lt;/em&gt;, not cost per request — a cheaper model that needs a retry, a longer chain, or human cleanup is not cheaper.&lt;/p&gt;

&lt;p&gt;Measure the simpler alternative first: the same capable model at a lower effort or reasoning setting, on the same traffic. On current-generation models, a reduced-effort run often lands where the previous generation's maximum effort did, and you keep one cache namespace. Tune that per route — classification and extraction routes usually hold quality at the low end, while coding and long-horizon agent loops do not.&lt;/p&gt;

&lt;p&gt;For attributing spend to routes before you cut anything, an LLM observability layer such as Langfuse gives you per-trace token and cost breakdowns without writing your own aggregation, at the cost of running one more service (or paying for the hosted tier) and adding a logging dependency to your request path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: prove the cheap model wins on cost per completed task — including retries and lost cache reuse — before you ship the cascade.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why is &lt;code&gt;cache_read_input_tokens&lt;/code&gt; always 0?&lt;/strong&gt;&lt;br&gt;
Something in the prefix changes between requests. Log two consecutive full request bodies, strip the &lt;code&gt;cache_control&lt;/code&gt; markers, and diff them — the first difference inside the overlapping region is your invalidator. Timestamps, UUIDs, unsorted JSON, and per-user tool lists cause most of these.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does prompt caching reduce latency or just cost?&lt;/strong&gt;&lt;br&gt;
Both, in practice: cached prefix tokens are not reprocessed, so time to first token drops on long prompts. Cost is the more reliable win; latency improvement depends on how much of the prompt was cacheable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is the LLM batch API worth it for a small app?&lt;/strong&gt;&lt;br&gt;
Yes, for any job a user isn't waiting on — roughly half price for the same model and prompt, as of late 2026. It is the rare optimization with no quality tradeoff, only a latency one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Start with measurement: log &lt;code&gt;input_tokens&lt;/code&gt;, cache writes, and cache reads per request, and treat a zero read rate on a repeated prefix as a bug with a ticket. Fix prefix stability and breakpoint placement before anything else — they cost nothing in quality and often account for most of the surprise. Move every job with no human in the loop to the batch endpoint. Only then consider lower effort settings, and only after that a smaller model, judged on cost per completed task rather than cost per call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/your-postgres-backups-are-untested-until-you-restore-one-a-drill-for-small-teams-3paj"&gt;Your Postgres Backups Are Untested Until You Restore One: A Drill for Small Teams&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/your-postgres-migration-runner-needs-a-retry-contract-not-just-a-lock-timeout-1oii"&gt;Your Postgres Migration Runner Needs a Retry Contract, Not Just a Lock Timeout&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/add-full-text-search-to-your-app-before-reaching-for-elasticsearch-lmc"&gt;Add Full-Text Search to Your App Before Reaching for Elasticsearch&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>performance</category>
    </item>
    <item>
      <title>Postgres Multi-Tenancy: Row-Level Security, tenant_id Filters, or a Schema per Tenant?</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Fri, 02 Oct 2026 15:28:02 +0000</pubDate>
      <link>https://dev.to/libme/postgres-multi-tenancy-row-level-security-tenantid-filters-or-a-schema-per-tenant-11j9</link>
      <guid>https://dev.to/libme/postgres-multi-tenancy-row-level-security-tenantid-filters-or-a-schema-per-tenant-11j9</guid>
      <description>&lt;p&gt;For most B2B SaaS apps, a &lt;code&gt;tenant_id&lt;/code&gt; column plus row-level security (RLS) on every table is the right default: you keep one schema and one migration path, and the database enforces isolation even when someone forgets a &lt;code&gt;WHERE&lt;/code&gt; clause. Schema-per-tenant only pays off when you have a small number of large tenants with genuinely different needs (per-tenant restores, per-tenant data residency). Database-per-tenant is an operations decision, not a data-modeling one — pick it when tenants must be backed up, moved, or deleted independently.&lt;/p&gt;

&lt;p&gt;The part nobody warns you about: RLS fails &lt;em&gt;silently&lt;/em&gt; in both directions. Misconfigured one way, it returns zero rows and your app looks broken. Misconfigured the other way, it returns everyone's rows and nothing looks wrong at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does "multi-tenancy" actually mean at the storage layer?
&lt;/h2&gt;

&lt;p&gt;Three models, in increasing order of isolation and operational cost:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Isolation enforced by&lt;/th&gt;
&lt;th&gt;Migration cost&lt;/th&gt;
&lt;th&gt;Per-tenant restore&lt;/th&gt;
&lt;th&gt;Realistic tenant count&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;tenant_id&lt;/code&gt; column, app filters&lt;/td&gt;
&lt;td&gt;Your application code&lt;/td&gt;
&lt;td&gt;One migration&lt;/td&gt;
&lt;td&gt;Hard (row-level surgery)&lt;/td&gt;
&lt;td&gt;Unlimited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;tenant_id&lt;/code&gt; column + RLS&lt;/td&gt;
&lt;td&gt;Postgres, per query&lt;/td&gt;
&lt;td&gt;One migration&lt;/td&gt;
&lt;td&gt;Hard&lt;/td&gt;
&lt;td&gt;Unlimited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Schema per tenant&lt;/td&gt;
&lt;td&gt;Postgres, per &lt;code&gt;search_path&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;One migration × N schemas&lt;/td&gt;
&lt;td&gt;Easy (&lt;code&gt;pg_dump -n&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Tens to low hundreds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database per tenant&lt;/td&gt;
&lt;td&gt;Postgres, per connection&lt;/td&gt;
&lt;td&gt;One migration × N databases&lt;/td&gt;
&lt;td&gt;Easy (&lt;code&gt;pg_dump&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Tens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first row is where most teams start and where most tenant-data leaks come from. A single endpoint that builds a query without the tenant predicate — a new report, an admin tool, a background job reusing a helper — is enough. Code review catches this most of the time, which is exactly the problem: "most of the time" is not an isolation guarantee.&lt;/p&gt;

&lt;p&gt;Takeaway: if isolation depends on every future query being written correctly, you do not have isolation, you have a convention.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I set up RLS so it actually enforces anything?
&lt;/h2&gt;

&lt;p&gt;Two statements per table, plus a policy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;invoices&lt;/span&gt; &lt;span class="n"&gt;ENABLE&lt;/span&gt; &lt;span class="k"&gt;ROW&lt;/span&gt; &lt;span class="k"&gt;LEVEL&lt;/span&gt; &lt;span class="k"&gt;SECURITY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;invoices&lt;/span&gt; &lt;span class="k"&gt;FORCE&lt;/span&gt; &lt;span class="k"&gt;ROW&lt;/span&gt; &lt;span class="k"&gt;LEVEL&lt;/span&gt; &lt;span class="k"&gt;SECURITY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="n"&gt;POLICY&lt;/span&gt; &lt;span class="n"&gt;tenant_isolation&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;invoices&lt;/span&gt;
  &lt;span class="k"&gt;USING&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tenant_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;current_setting&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'app.current_tenant'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="k"&gt;CHECK&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tenant_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;current_setting&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'app.current_tenant'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;ENABLE&lt;/code&gt; turns policies on. &lt;code&gt;FORCE&lt;/code&gt; is the one people skip, and it matters: without it, the table's &lt;strong&gt;owner&lt;/strong&gt; bypasses every policy. If your application connects with the same role that created the tables — the default in a lot of small setups — then &lt;code&gt;ENABLE ROW LEVEL SECURITY&lt;/code&gt; alone does nothing for your app's queries. It looks configured. It enforces nothing.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;USING&lt;/code&gt; filters reads (and which rows &lt;code&gt;UPDATE&lt;/code&gt;/&lt;code&gt;DELETE&lt;/code&gt; can see). &lt;code&gt;WITH CHECK&lt;/code&gt; validates rows being written. Omit &lt;code&gt;WITH CHECK&lt;/code&gt; and a tenant can insert rows labeled with someone else's &lt;code&gt;tenant_id&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On the application side, set the tenant inside the transaction, never as a session-level &lt;code&gt;SET&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// pg (node-postgres). set_config(..., true) == SET LOCAL, scoped to this transaction.&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;withTenant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;BEGIN&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;SELECT set_config($1, $2, true)&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;app.current_tenant&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;COMMIT&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="k"&gt;catch&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ROLLBACK&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;SET LOCAL&lt;/code&gt; cannot take a bind parameter, which is why &lt;code&gt;set_config($1, $2, true)&lt;/code&gt; is the version you want — string-concatenating a tenant id into a &lt;code&gt;SET&lt;/code&gt; statement is an injection point in the one place you least want one.&lt;/p&gt;

&lt;p&gt;Takeaway: &lt;code&gt;ENABLE&lt;/code&gt; without &lt;code&gt;FORCE&lt;/code&gt;, or &lt;code&gt;USING&lt;/code&gt; without &lt;code&gt;WITH CHECK&lt;/code&gt;, is a policy that reads as secure in a migration diff and is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does RLS return zero rows (or everyone's rows) in production?
&lt;/h2&gt;

&lt;p&gt;These are the failure modes I keep running into, in the order they bite.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero rows, no error.&lt;/strong&gt; &lt;code&gt;current_setting('app.current_tenant', true)&lt;/code&gt; with the second argument &lt;code&gt;true&lt;/code&gt; returns &lt;code&gt;NULL&lt;/code&gt; when the setting is missing instead of raising &lt;code&gt;unrecognized configuration parameter&lt;/code&gt;. &lt;code&gt;tenant_id = NULL&lt;/code&gt; is never true, so every query returns an empty set. The symptom is a page that renders with no data and no stack trace, usually in a code path that got a connection outside your &lt;code&gt;withTenant&lt;/code&gt; wrapper — a health check, a migration script, a queue worker. Drop the &lt;code&gt;true&lt;/code&gt; during development so you get a loud error instead of an empty list.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everyone's rows, no error.&lt;/strong&gt; Three causes: the app role owns the tables and you did not &lt;code&gt;FORCE&lt;/code&gt;; the app role has &lt;code&gt;BYPASSRLS&lt;/code&gt; (superuser always does — this is why your app should never connect as &lt;code&gt;postgres&lt;/code&gt;); or the table is new and nobody enabled RLS on it. The third is the common one. Put it in CI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Fails the build if any table in the schema is missing RLS.&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relname&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;pg_class&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
&lt;span class="k"&gt;JOIN&lt;/span&gt; &lt;span class="n"&gt;pg_namespace&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;oid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relnamespace&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nspname&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'public'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relkind&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'r'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relrowsecurity&lt;/span&gt; &lt;span class="k"&gt;OR&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relforcerowsecurity&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Assert that query returns no rows, with an explicit allowlist for genuinely shared tables (plans, feature flags, country codes).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cross-tenant bleed under a connection pooler.&lt;/strong&gt; In transaction pooling mode, PgBouncer or Supavisor hand your next transaction whatever server connection is free. A session-level &lt;code&gt;SET app.current_tenant&lt;/code&gt; survives on that server connection and is then inherited by a different tenant's transaction. &lt;code&gt;SET LOCAL&lt;/code&gt; / &lt;code&gt;set_config(..., true)&lt;/code&gt; resets at commit, which is what makes RLS and transaction pooling compatible at all.&lt;/p&gt;

&lt;p&gt;Takeaway: test isolation with a real query as a real tenant role in CI — a passing migration proves the policy exists, not that it applies.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is schema-per-tenant worth the operational cost?
&lt;/h2&gt;

&lt;p&gt;When a tenant is a unit of &lt;em&gt;operations&lt;/em&gt;, not just a unit of data. Concretely: you need &lt;code&gt;pg_dump -n tenant_42&lt;/code&gt; to restore one customer without touching the others, you have contractual data-residency requirements, or a few large tenants need different indexes or retention.&lt;/p&gt;

&lt;p&gt;The cost is real. Every migration runs N times, so a schema change is now a job with partial-failure semantics rather than a single statement. Thousands of schemas bloat the system catalogs and make &lt;code&gt;pg_dump&lt;/code&gt;, autovacuum scheduling, and query planning noticeably slower. Under a pooler, &lt;code&gt;search_path&lt;/code&gt; has the same leak problem as session-level &lt;code&gt;SET&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If you want per-tenant isolation without hand-rolling the migration fan-out, Neon's database branching gives you cheap copy-on-write databases you can create and drop per tenant or per environment, at the cost of tying that piece of your infrastructure to one vendor's control plane. If you are already on Postgres-as-your-backend and want RLS wired to your auth layer, Supabase enforces policies against JWT claims by default, which is genuinely good hygiene but means an incorrect policy is directly internet-facing rather than behind your API. For tenants that outgrow one machine, Citus distributes tables by &lt;code&gt;tenant_id&lt;/code&gt; so a tenant's rows stay colocated on one node, with the tradeoff that cross-tenant analytical queries and some schema changes get more expensive.&lt;/p&gt;

&lt;p&gt;Takeaway: choose schema- or database-per-tenant when you need per-tenant &lt;em&gt;restore and delete&lt;/em&gt;, not because it feels more secure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does RLS cost performance?
&lt;/h2&gt;

&lt;p&gt;It adds a predicate to every query on the table, so the honest answer is: only if your indexes ignore the tenant. Make &lt;code&gt;tenant_id&lt;/code&gt; the leading column of the indexes that back your hot queries — &lt;code&gt;(tenant_id, created_at DESC)&lt;/code&gt; rather than &lt;code&gt;(created_at DESC)&lt;/code&gt; — and the RLS predicate rides along for free.&lt;/p&gt;

&lt;p&gt;The subtler cost: Postgres evaluates RLS quals before user-supplied quals unless the operators involved are marked &lt;code&gt;LEAKPROOF&lt;/code&gt;, since a non-leakproof function could otherwise expose values from rows you should never see. That restriction occasionally blocks a plan the planner would have chosen otherwise. Check &lt;code&gt;EXPLAIN&lt;/code&gt; on your slowest tenant-scoped query with RLS on and off before assuming it is free.&lt;/p&gt;

&lt;p&gt;Takeaway: RLS is cheap when &lt;code&gt;tenant_id&lt;/code&gt; leads your indexes and worth measuring when it doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can I use row-level security with PgBouncer in transaction mode?&lt;/strong&gt;&lt;br&gt;
Yes, as long as you set the tenant with &lt;code&gt;SET LOCAL&lt;/code&gt; or &lt;code&gt;set_config('app.current_tenant', $1, true)&lt;/code&gt; inside the transaction. A plain session-level &lt;code&gt;SET&lt;/code&gt; persists on the pooled server connection and will be inherited by another tenant's transaction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does my RLS policy return no rows even though the data exists?&lt;/strong&gt;&lt;br&gt;
The configuration parameter the policy reads is unset on that connection, so &lt;code&gt;current_setting('app.current_tenant', true)&lt;/code&gt; returns NULL and the comparison is never true. It almost always means a code path acquired a connection without going through your per-request transaction wrapper.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need one database per tenant for compliance?&lt;/strong&gt;&lt;br&gt;
Usually no — RLS satisfies most logical-isolation requirements. Separate databases matter when you need per-tenant backup, restore, deletion, or data residency as an operational guarantee rather than a query-level one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Start with a &lt;code&gt;tenant_id&lt;/code&gt; column and RLS with both &lt;code&gt;ENABLE&lt;/code&gt; and &lt;code&gt;FORCE&lt;/code&gt;, an app role that is not the table owner and has no &lt;code&gt;BYPASSRLS&lt;/code&gt;, and &lt;code&gt;SET LOCAL&lt;/code&gt; inside every request transaction. Add a CI check that fails when a new table ships without a policy, because that is how the leak actually happens. Move to schema- or database-per-tenant only when per-tenant restore, deletion, or residency becomes a requirement — and accept that your migration pipeline becomes a fan-out job that day. Whatever you pick, write one test that logs in as tenant A and asserts it cannot read tenant B's row; it is the only evidence that any of this works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/cutting-your-side-projects-cloud-bill-a-checklist-that-doesnt-sacrifice-uptime-17kf"&gt;Cutting Your Side Project's Cloud Bill: A Checklist That Doesn't Sacrifice Uptime&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/websocket-closes-every-60-seconds-with-code-1006-finding-the-proxy-idle-timeout-and-fixing-it-with-278e"&gt;WebSocket Closes Every 60 Seconds With Code 1006: Finding the Proxy Idle Timeout and Fixing It With Heartbeats&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/chatgpt-free-vs-go-vs-plus-vs-pro-which-tier-should-a-coding-beginner-actually-pay-for-jke"&gt;ChatGPT Free vs Go vs Plus vs Pro: Which Tier Should a Coding Beginner Actually Pay For?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>security</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Users Randomly Logged Out? Your Redis Eviction Policy Is Deleting Sessions</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Thu, 01 Oct 2026 23:54:44 +0000</pubDate>
      <link>https://dev.to/libme/users-randomly-logged-out-your-redis-eviction-policy-is-deleting-sessions-4722</link>
      <guid>https://dev.to/libme/users-randomly-logged-out-your-redis-eviction-policy-is-deleting-sessions-4722</guid>
      <description>&lt;p&gt;If a subset of your users get logged out at unpredictable times — no deploy, no session-secret change, worse during traffic peaks — check &lt;code&gt;maxmemory-policy&lt;/code&gt; on your Redis instance before you touch cookie code. When Redis hits its memory ceiling with an &lt;code&gt;allkeys-*&lt;/code&gt; policy, it deletes whatever keys are least recently used to make room, and it does not care that some of those keys are login sessions or queued jobs. The fix is not a bigger instance; it's separating keys you can afford to lose from keys you can't.&lt;/p&gt;

&lt;p&gt;I lost most of a day to this on a side project where sessions, a page cache, and a BullMQ queue all shared one small managed Redis. The cache was doing its job — filling memory — and Redis was doing its job — throwing things out. Together they produced a third behavior nobody asked for: random logouts and jobs that silently never ran.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does Redis delete keys that still have time left on their TTL?
&lt;/h2&gt;

&lt;p&gt;Redis with &lt;code&gt;maxmemory&lt;/code&gt; set is a bounded store, and &lt;code&gt;maxmemory-policy&lt;/code&gt; decides what happens when a write would cross that bound. The policies split into three families:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;noeviction&lt;/code&gt; — refuse the write. The client gets &lt;code&gt;OOM command not allowed when used memory &amp;gt; 'maxmemory'&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;volatile-lru&lt;/code&gt;, &lt;code&gt;volatile-lfu&lt;/code&gt;, &lt;code&gt;volatile-random&lt;/code&gt;, &lt;code&gt;volatile-ttl&lt;/code&gt; — evict only keys that have a TTL set. If there are no such keys, writes fail like &lt;code&gt;noeviction&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;allkeys-lru&lt;/code&gt;, &lt;code&gt;allkeys-lfu&lt;/code&gt;, &lt;code&gt;allkeys-random&lt;/code&gt; — evict anything, TTL or not.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trap is that eviction is invisible from the application's side. A deleted session key is indistinguishable from an expired one: your middleware looks up the session, gets &lt;code&gt;nil&lt;/code&gt;, and correctly concludes "not logged in." There's no error to catch, no log line, nothing that says &lt;em&gt;this key was taken from you&lt;/em&gt;. Same for a queue: BullMQ or Sidekiq asks for a job hash that isn't there anymore and the job just doesn't exist.&lt;/p&gt;

&lt;p&gt;Two details make this harder to reason about than it should be. First, &lt;code&gt;volatile-*&lt;/code&gt; policies feel safe ("only expiring keys get evicted") right up until you realize sessions almost always have a TTL — they're prime eviction candidates, and the ones with the longest remaining life under &lt;code&gt;volatile-ttl&lt;/code&gt; survive while your active-but-recently-renewed ones may not. Second, &lt;code&gt;SELECT&lt;/code&gt;-ing a different logical database (&lt;code&gt;db 0&lt;/code&gt; vs &lt;code&gt;db 1&lt;/code&gt;) does &lt;strong&gt;not&lt;/strong&gt; isolate anything: all databases on an instance share one memory budget and one eviction policy, and Redis Cluster drops multiple databases entirely. Prefixing keys and switching DB numbers buys you tidiness, not safety.&lt;/p&gt;

&lt;p&gt;Takeaway: eviction is a silent &lt;code&gt;DEL&lt;/code&gt; issued by the server, so any key whose absence changes correctness must live somewhere eviction cannot reach.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I confirm eviction is what's happening?
&lt;/h2&gt;

&lt;p&gt;Three commands, in this order. Don't trust the default documented in the docs — managed providers ship their own parameter defaults, and someone may have changed yours years ago.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. What policy is actually in force, and how close to the ceiling are we?&lt;/span&gt;
redis-cli INFO memory | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'used_memory_human|maxmemory_human|maxmemory_policy'&lt;/span&gt;

&lt;span class="c"&gt;# 2. The smoking gun: a non-zero, *increasing* counter.&lt;/span&gt;
redis-cli INFO stats | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'evicted_keys|expired_keys'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;evicted_keys&lt;/code&gt; is a monotonic counter since the last restart. Any non-zero value on an instance holding sessions or jobs is a bug report. Sample it twice a minute apart — a rising number during the window your users complain about is as close to proof as you'll get after the fact.&lt;/p&gt;

&lt;p&gt;If you need to catch it live, keyspace notifications will tell you exactly which keys are going:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 'E' = keyevent channels, 'e' = evicted events&lt;/span&gt;
redis-cli CONFIG SET notify-keyspace-events Ee
redis-cli PSUBSCRIBE &lt;span class="s1"&gt;'__keyevent@*__:evicted'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each message carries the evicted key name, so you'll see &lt;code&gt;sess:...&lt;/code&gt; or &lt;code&gt;bull:mail:...&lt;/code&gt; scroll past and the argument ends there. Two honest caveats: notifications are fire-and-forget (a disconnected subscriber misses events entirely, so this is a debugging tool, not an audit log), and &lt;code&gt;CONFIG SET&lt;/code&gt; doesn't survive a restart unless you &lt;code&gt;CONFIG REWRITE&lt;/code&gt; or change it in your provider's parameter group.&lt;/p&gt;

&lt;p&gt;The dead ends I burned time on first, so you can skip them: rotating the session secret, &lt;code&gt;SameSite&lt;/code&gt;/&lt;code&gt;Secure&lt;/code&gt; cookie flags, sticky-session config on the load balancer, and clock skew on token expiry. All plausible causes of logouts, none of which produce a rising &lt;code&gt;evicted_keys&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Takeaway: &lt;code&gt;evicted_keys &amp;gt; 0&lt;/code&gt; on an instance storing sessions or jobs is not a tuning opportunity, it's a correctness defect.&lt;/p&gt;

&lt;h2&gt;
  
  
  What policy should each workload get?
&lt;/h2&gt;

&lt;p&gt;Pick per role, not per cluster. The point of the table is that "which policy is best" is the wrong question — the right one is "what does this instance hold."&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Instance role&lt;/th&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;th&gt;Failure mode you're accepting&lt;/th&gt;
&lt;th&gt;Alert on&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Page/query cache&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;allkeys-lru&lt;/code&gt; (or &lt;code&gt;allkeys-lfu&lt;/code&gt; for skewed access)&lt;/td&gt;
&lt;td&gt;Cache misses, extra DB load&lt;/td&gt;
&lt;td&gt;Hit ratio drop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sessions / auth&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;noeviction&lt;/code&gt; + TTL on every key&lt;/td&gt;
&lt;td&gt;Writes fail loudly at the ceiling&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;evicted_keys &amp;gt; 0&lt;/code&gt;, OOM errors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Job queue / broker&lt;/td&gt;
&lt;td&gt;&lt;code&gt;noeviction&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;New enqueues rejected; existing jobs safe&lt;/td&gt;
&lt;td&gt;Any OOM error, memory &amp;gt; 70%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rate limit counters&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;volatile-lru&lt;/code&gt;, short TTLs&lt;/td&gt;
&lt;td&gt;Some limits reset early (fails open)&lt;/td&gt;
&lt;td&gt;Memory trend only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Locks / idempotency keys&lt;/td&gt;
&lt;td&gt;&lt;code&gt;noeviction&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Lock acquisition fails instead of double-granting&lt;/td&gt;
&lt;td&gt;&lt;code&gt;evicted_keys &amp;gt; 0&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;noeviction&lt;/code&gt; feels scary because it turns a memory problem into visible request failures. That's the point: an OOM error is a page you can respond to, while an evicted lock key is a duplicate charge a customer tells you about. Sidekiq's documentation has recommended &lt;code&gt;noeviction&lt;/code&gt; for its Redis for years for exactly this reason.&lt;/p&gt;

&lt;p&gt;Enforce it at boot rather than trusting a runbook. This has caught a mis-restored parameter group for me more than once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;Redis&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ioredis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cache&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Redis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;REDIS_CACHE_URL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;// evictable&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;durable&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Redis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;REDIS_DURABLE_URL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;// sessions, queue, locks&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;assertRedisPolicies&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[,&lt;/span&gt; &lt;span class="nx"&gt;policy&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;durable&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;GET&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;maxmemory-policy&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;policy&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;noeviction&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="s2"&gt;`durable Redis has maxmemory-policy=&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;policy&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;; sessions and jobs can be silently evicted`&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fail the deploy on that, and the class of bug disappears instead of getting rediscovered next year.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is a second instance worth the cost?
&lt;/h2&gt;

&lt;p&gt;For most small teams, yes, and it's often cheaper than the upsize you were about to buy. Cache is the part that wants headroom; sessions, locks, and queue metadata for a modest app are megabytes, so the durable node can stay tiny.&lt;/p&gt;

&lt;p&gt;If you want a Redis-compatible store where eviction isn't a correctness risk at all, Amazon MemoryDB is the one that treats memory as a cache over a durable multi-AZ transaction log rather than as the only copy of your data — it costs meaningfully more than ElastiCache per node, and it inherits the same single-threaded hot-key ceiling, so it solves durability, not throughput. If you'd rather not pay for a second always-on node, Upstash bills per request instead of per instance, which makes creating one database per role effectively free — the tradeoff is that a chatty cache client becomes a line item, so you watch command volume there instead of memory. And if you moved to Valkey after the Redis licensing change, note that it kept the same eviction semantics and &lt;code&gt;INFO&lt;/code&gt; fields, so both the bug and every diagnostic above migrate with you unchanged.&lt;/p&gt;

&lt;p&gt;Takeaway: the second instance isn't about capacity, it's about giving eviction a blast radius you chose on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why are my Redis keys disappearing before the TTL expires?&lt;/strong&gt;&lt;br&gt;
Almost always eviction under &lt;code&gt;maxmemory&lt;/code&gt; pressure. Run &lt;code&gt;redis-cli INFO stats | grep evicted_keys&lt;/code&gt; — if that number is climbing, Redis is deleting keys to stay under its memory limit, and an &lt;code&gt;allkeys-*&lt;/code&gt; policy lets it take keys that have plenty of TTL left.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is it safe to use the same Redis for caching and sessions?&lt;/strong&gt;&lt;br&gt;
Only if the policy is &lt;code&gt;noeviction&lt;/code&gt;, and then your cache stops absorbing growth gracefully and starts failing writes instead. Separate instances are the safe answer; separate logical databases (&lt;code&gt;SELECT 1&lt;/code&gt;) are not, because all databases on one instance share the same memory budget and eviction policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the best maxmemory-policy for Redis?&lt;/strong&gt;&lt;br&gt;
There is no single best one — it depends on whether losing a key is acceptable. Use &lt;code&gt;allkeys-lru&lt;/code&gt; for pure caches, and &lt;code&gt;noeviction&lt;/code&gt; for sessions, job queues, locks, and idempotency keys where a missing key changes program behavior rather than just costing you a recompute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;If users report random logouts or jobs that vanish without a trace, check &lt;code&gt;evicted_keys&lt;/code&gt; and &lt;code&gt;maxmemory-policy&lt;/code&gt; before you audit a single line of auth code. Caches should be evictable and everything whose absence changes correctness should not be, which in practice means two Redis instances with two policies and a boot-time assertion that nobody has quietly changed them. Keep &lt;code&gt;noeviction&lt;/code&gt; on the durable side even though it converts memory pressure into loud request failures — loud is the feature. Bigger instances only move the date this happens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/your-password-reset-emails-are-going-to-spam-choosing-between-resend-postmark-and-amazon-ses-3jak"&gt;Your Password Reset Emails Are Going to Spam: Choosing Between Resend, Postmark, and Amazon SES&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/plausible-vs-umami-vs-posthog-when-self-hosting-analytics-stops-being-cheaper-5952"&gt;Plausible vs Umami vs PostHog: When Self-Hosting Analytics Stops Being Cheaper&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/postgres-full-text-search-in-production-how-to-load-test-the-index-and-pin-down-relevance-282b"&gt;Postgres Full-Text Search in Production: How to Load-Test the Index and Pin Down Relevance&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>database</category>
      <category>backend</category>
      <category>devops</category>
      <category>performance</category>
    </item>
    <item>
      <title>Plausible vs Umami vs PostHog: When Self-Hosting Analytics Stops Being Cheaper</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Wed, 30 Sep 2026 15:27:08 +0000</pubDate>
      <link>https://dev.to/libme/plausible-vs-umami-vs-posthog-when-self-hosting-analytics-stops-being-cheaper-5952</link>
      <guid>https://dev.to/libme/plausible-vs-umami-vs-posthog-when-self-hosting-analytics-stops-being-cheaper-5952</guid>
      <description>&lt;p&gt;If all you need is pageviews, referrers, and top pages, self-hosting Umami on a small VPS is genuinely cheap and stays cheap. If you need session replay, funnels, and feature flags, self-hosting PostHog is the most expensive "free" decision on this list. Plausible sits in between: pleasant to run, but it drags a ClickHouse instance along with it, and that's the part people underestimate.&lt;/p&gt;

&lt;p&gt;I've run all three — Umami and Plausible self-hosted, PostHog on their cloud — on side projects and a small internal product. What follows is the decision I'd hand a colleague, including the failure mode that costs people a week before they notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are you actually buying when you pay for analytics?
&lt;/h2&gt;

&lt;p&gt;Not the dashboard. The dashboard is the easy part, which is why there are a dozen good open-source ones.&lt;/p&gt;

&lt;p&gt;You're paying for storage and retention of an append-only event stream, an ingestion endpoint that survives your traffic spikes, and — the underrated one — somebody else owning the upgrade path of a columnar database you don't want to learn. Analytics data is write-heavy, immutable, and queried with wide aggregations, which is exactly the workload that makes people reach for ClickHouse and exactly the workload that makes a single Postgres box feel fine right up until it doesn't.&lt;/p&gt;

&lt;p&gt;The honest framing: hosted analytics is a bet that your time is worth more than the subscription, and self-hosting is a bet that it isn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plausible, Umami, PostHog: what is each one actually for?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Umami&lt;/th&gt;
&lt;th&gt;Plausible&lt;/th&gt;
&lt;th&gt;PostHog&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core job&lt;/td&gt;
&lt;td&gt;Lightweight web analytics&lt;/td&gt;
&lt;td&gt;Web analytics + goals/funnels&lt;/td&gt;
&lt;td&gt;Product analytics suite&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data store (self-host)&lt;/td&gt;
&lt;td&gt;Postgres or MySQL&lt;/td&gt;
&lt;td&gt;Postgres &lt;strong&gt;and&lt;/strong&gt; ClickHouse&lt;/td&gt;
&lt;td&gt;ClickHouse, Kafka, Redis, Postgres&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;License (self-host)&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;AGPL (Community Edition)&lt;/td&gt;
&lt;td&gt;MIT core, some features source-available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session replay&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feature flags / experiments&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Realistic self-host floor&lt;/td&gt;
&lt;td&gt;One small VPS&lt;/td&gt;
&lt;td&gt;One mid-size VPS&lt;/td&gt;
&lt;td&gt;Not a small-VPS workload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud pricing model (as of Sep 2026)&lt;/td&gt;
&lt;td&gt;Event/site tiers, free tier&lt;/td&gt;
&lt;td&gt;Pageview tiers&lt;/td&gt;
&lt;td&gt;Usage-based per product, generous free tier&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The shape of that table is the whole decision. Umami answers "what pages do people read." Plausible answers that plus "did they do the thing I care about." PostHog answers "why did this cohort drop off at step three, and what happens if I flip this flag for 10% of them."&lt;/p&gt;

&lt;p&gt;If you want the lightest possible self-hosted option that runs next to your app on the database you already operate, Umami is the one that doesn't add a second data store to your stack.&lt;/p&gt;

&lt;p&gt;Drawbacks, because every tool gets one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Umami&lt;/strong&gt;: the reporting is deliberately shallow. No funnel debugging worth the name, and once your events table gets large on Postgres you will be the one adding indexes and thinking about partitioning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plausible&lt;/strong&gt;: the Community Edition is intentionally behind the hosted product on some features, and self-hosting means you now operate ClickHouse — including its memory appetite and its own backup story.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PostHog&lt;/strong&gt;: enormous surface area, and as of mid-2026 their self-host guidance steers small teams toward the hobby Docker deployment rather than a supported production cluster. Session replay also makes your event volume — and your bill — grow in a way pageview counting never does.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Takeaway: pick the tool by which question you're going to ask at 11pm, not by which dashboard looks nicest in the screenshots.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does my analytics show fewer visits than my server logs?
&lt;/h2&gt;

&lt;p&gt;This is the first real bug everyone hits, and it isn't a bug.&lt;/p&gt;

&lt;p&gt;Client-side analytics scripts are blocked by content blockers, and the block rate skews hard by audience — a developer-tools site loses far more than a recipe blog. Your access logs count requests; your analytics counts successfully executed JavaScript. Those numbers will never match, and if they do match exactly, something is double-counting.&lt;/p&gt;

&lt;p&gt;The partial fix is to serve the script and the event endpoint from your own domain so they aren't third-party requests. For Plausible on Next.js:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// next.config.js&lt;/span&gt;
&lt;span class="nx"&gt;module&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;exports&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;rewrites&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/stats/js/script.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;destination&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://plausible.io/js/script.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/stats/api/event&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;destination&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://plausible.io/api/event&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;];&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;script
  &lt;/span&gt;&lt;span class="na"&gt;defer&lt;/span&gt;
  &lt;span class="na"&gt;data-domain=&lt;/span&gt;&lt;span class="s"&gt;"example.com"&lt;/span&gt;
  &lt;span class="na"&gt;data-api=&lt;/span&gt;&lt;span class="s"&gt;"/stats/api/event"&lt;/span&gt;
  &lt;span class="na"&gt;src=&lt;/span&gt;&lt;span class="s"&gt;"/stats/js/script.js"&lt;/span&gt;
&lt;span class="nt"&gt;&amp;gt;&amp;lt;/script&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;data-api&lt;/code&gt; attribute is the part people forget. Proxy the script but leave the event endpoint pointing at the vendor and you've solved nothing — the beacon is still the third-party request that gets blocked.&lt;/p&gt;

&lt;p&gt;If you terminate TLS yourself, the Caddy equivalent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;example.com {
    handle /stats/js/script.js {
        rewrite * /js/script.js
        reverse_proxy https://plausible.io {
            header_up Host plausible.io
        }
    }

    handle /stats/api/event {
        rewrite * /api/event
        reverse_proxy https://plausible.io {
            header_up Host plausible.io
            header_up X-Forwarded-For {remote_host}
        }
    }

    reverse_proxy localhost:3000
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here's the failure mode that wastes the week. If you proxy the event endpoint without forwarding the real client IP, every visitor arrives from your proxy's address. Analytics tools that derive country and the daily visitor hash from the client IP will then report your entire audience as one country and collapse distinct visitors together. The dashboard doesn't error — it just quietly goes wrong, and it looks like a traffic change rather than a config change. If your geography chart flattens to a single row on the day you shipped a proxy, that's your bug, not your users.&lt;/p&gt;

&lt;p&gt;The same rule applies when self-hosting behind nginx or a CDN: the analytics container must see a trustworthy &lt;code&gt;X-Forwarded-For&lt;/code&gt;, and it must be configured to trust that header from your proxy only.&lt;/p&gt;

&lt;p&gt;Takeaway: proxying analytics recovers blocked traffic, but it moves client IP handling into your infrastructure — and every downstream metric that depends on IP silently degrades if you get it wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does self-hosting actually cost?
&lt;/h2&gt;

&lt;p&gt;Not the VPS. The VPS is the cheap part.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Memory, not CPU.&lt;/strong&gt; ClickHouse is happy to use whatever RAM you give it, and a Plausible or PostHog stack on an undersized box gets OOM-killed during ingestion bursts, not during queries. Budget for headroom you aren't using yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backups you've actually restored.&lt;/strong&gt; An analytics database nobody backs up is a decision, just an unstated one. Restoring ClickHouse is not the same procedure as restoring Postgres, and the time to learn it is not the morning after.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upgrades.&lt;/strong&gt; Minor version bumps are fine. The painful ones are the migrations that change the event schema; those arrive on the maintainer's schedule, not yours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retention.&lt;/strong&gt; Hosted plans often trim raw event retention on lower tiers. Self-hosting gives you unlimited retention and the corresponding unlimited disk growth — set a TTL on day one instead of discovering it at 90% full.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rough rule from running these: a self-hosted analytics stack costs a few hours to stand up, then a few hours a year, until the day it breaks — and then it costs whatever your afternoon is worth, at the least convenient moment.&lt;/p&gt;

&lt;p&gt;Takeaway: self-hosting analytics isn't free, it's prepaid in hours, and you pay in a currency you can't budget in advance.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is the paid cloud tier worth it?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;What I'd do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Side project, pageviews only&lt;/td&gt;
&lt;td&gt;Self-host Umami, one VPS, done&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Client sites you bill for&lt;/td&gt;
&lt;td&gt;Hosted tier — you're selling reliability, not operating it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Privacy/compliance is the reason you're here&lt;/td&gt;
&lt;td&gt;Self-host Plausible CE, budget for ClickHouse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;You need funnels, replay, and flags&lt;/td&gt;
&lt;td&gt;PostHog Cloud, and cap your event volume early&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Traffic is spiky and unpredictable&lt;/td&gt;
&lt;td&gt;Hosted — burst ingestion is the hardest part to self-run&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The tipping point isn't traffic volume, it's whether losing a week of data would matter. Hobby projects can absorb a gap; anything you report to someone else can't.&lt;/p&gt;

&lt;p&gt;If you want product analytics without operating a multi-service data pipeline, PostHog Cloud is the option that gives you funnels, replay, and feature flags behind one SDK — with the tradeoff that event volume, not user count, drives what you pay.&lt;/p&gt;

&lt;p&gt;Takeaway: pay when the data becomes evidence someone else relies on.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is self-hosted Plausible free?&lt;/strong&gt;&lt;br&gt;
The Community Edition is free software under the AGPL, but running it requires both Postgres and ClickHouse, plus backups and upgrades you perform yourself. It's free of licensing cost, not free of operating cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does Google Analytics show different numbers than Plausible or Umami?&lt;/strong&gt;&lt;br&gt;
They count different things. GA4 models sessions and applies its own filtering and modeling; privacy-focused tools typically count a visitor via a daily rotating hash with no cross-day identity. Neither number is wrong — they answer different questions, so never compare them directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I self-host PostHog on a single small server?&lt;/strong&gt;&lt;br&gt;
You can run the hobby Docker deployment for evaluation, but as of mid-2026 PostHog directs production users to its cloud rather than supporting large self-hosted clusters. If you're choosing PostHog for its full feature set, plan on the hosted version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Choose Umami if you want pageviews on infrastructure you already run, and accept shallow reporting in exchange. Choose Plausible when goals and conversions matter and you want a clean, privacy-respecting dashboard — self-host it if compliance demands it, otherwise let them operate the ClickHouse. Choose PostHog when your real question is about product behavior rather than traffic, and pay for the cloud tier instead of assembling the pipeline yourself. Whichever you pick, proxy the script through your own domain on day one and verify your geography chart afterward — that single check catches the most common silent misconfiguration in this entire category.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/webhook-reliability-101-retries-idempotency-and-the-bugs-that-bite-at-2-am-3e5e"&gt;Webhook Reliability 101: Retries, Idempotency, and the Bugs That Bite at 2 A.M.&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/how-to-test-search-relevance-before-you-ship-a-ranking-change-29o"&gt;How to Test Search Relevance Before You Ship a Ranking Change&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/automate-your-code-reviews-with-an-llm-without-annoying-your-team-5h2n"&gt;Automate Your Code Reviews with an LLM Without Annoying Your Team&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>webdev</category>
      <category>privacy</category>
      <category>opensource</category>
      <category>saas</category>
    </item>
    <item>
      <title>Scrubbing Test Data Without Nulling It Out: The Reserved Ranges That Are Safe to Seed</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Wed, 30 Sep 2026 06:42:58 +0000</pubDate>
      <link>https://dev.to/libme/scrubbing-test-data-without-nulling-it-out-the-reserved-ranges-that-are-safe-to-seed-2nc5</link>
      <guid>https://dev.to/libme/scrubbing-test-data-without-nulling-it-out-the-reserved-ranges-that-are-safe-to-seed-2nc5</guid>
      <description>&lt;p&gt;If your scrub script sets &lt;code&gt;phone = NULL&lt;/code&gt; and &lt;code&gt;email = NULL&lt;/code&gt;, every feature that depends on those fields silently stops existing in your test environment — SMS opt-in, required-field validation, notification fan-out. The fix is not random generation, which can dial a stranger. Standards bodies and telecom regulators have set aside ranges that are syntactically valid but guaranteed unroutable, and writing those into your scrubbed data keeps the code paths alive while making delivery impossible.&lt;/p&gt;

&lt;p&gt;A commenter on an earlier post made this point about phone numbers, and it generalizes to almost every PII column you touch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually breaks when you NULL a column?
&lt;/h2&gt;

&lt;p&gt;Three distinct failures, and only the first one is obvious.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validation paths never execute.&lt;/strong&gt; Your checkout form requires a phone number. In the scrubbed environment nobody has one, so the edit form renders a blank required field, the E.164 normalizer is never called, and the branch that handles "user has a number but it failed verification" is unreachable. You ship the bug that only fires for users who &lt;em&gt;do&lt;/em&gt; have data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Downstream code gets a type it did not plan for.&lt;/strong&gt; A NULL phone reaches your SMS client and you get &lt;code&gt;TypeError: Cannot read properties of null (reading 'replace')&lt;/code&gt; in your formatter — or, if the null makes it to the provider, a rejection like Twilio's &lt;code&gt;21211 Invalid 'To' Phone Number&lt;/code&gt;. Neither error tells you anything about the feature you were testing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Randomly generated data is worse than NULL.&lt;/strong&gt; The moment someone points a test environment at live credentials — a misconfigured secret, a provider client that defaults to production when an env var is missing — randomly generated digits are routable digits. Random data fails open: it reaches a real person. Reserved data fails closed.&lt;/p&gt;

&lt;p&gt;The takeaway: NULL removes the code path, random data removes the safety, and reserved ranges remove neither.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Shape is valid&lt;/th&gt;
&lt;th&gt;Exercises the code path&lt;/th&gt;
&lt;th&gt;Safe if creds leak&lt;/th&gt;
&lt;th&gt;Deterministic&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;NULL&lt;/code&gt; everything&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Random digits/strings&lt;/td&gt;
&lt;td&gt;Usually&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Only if seeded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scrubbed production values&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reserved ranges&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes, if hashed from the PK&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Which ranges are actually reserved?
&lt;/h2&gt;

&lt;p&gt;These are set aside by the relevant authority specifically so that fiction and documentation cannot collide with a real subscriber or host.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Data&lt;/th&gt;
&lt;th&gt;Reserved range&lt;/th&gt;
&lt;th&gt;Set aside by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Phone, North America&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;555-0100&lt;/code&gt; to &lt;code&gt;555-0199&lt;/code&gt; in any area code&lt;/td&gt;
&lt;td&gt;NANPA, for fictional use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Phone, UK mobile&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;07700 900000&lt;/code&gt;–&lt;code&gt;900999&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Ofcom drama range&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Phone, UK landline&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;020 7946 0000&lt;/code&gt;–&lt;code&gt;0999&lt;/code&gt;, &lt;code&gt;01632 960000&lt;/code&gt;–&lt;code&gt;960999&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Ofcom drama range&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Email domain&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;example.com&lt;/code&gt;, &lt;code&gt;example.net&lt;/code&gt;, &lt;code&gt;example.org&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;RFC 2606&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hostname / TLD&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.test&lt;/code&gt;, &lt;code&gt;.example&lt;/code&gt;, &lt;code&gt;.invalid&lt;/code&gt;, &lt;code&gt;.localhost&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;RFC 2606&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IPv4&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;192.0.2.0/24&lt;/code&gt;, &lt;code&gt;198.51.100.0/24&lt;/code&gt;, &lt;code&gt;203.0.113.0/24&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;RFC 5737 (documentation)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IPv6&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2001:db8::/32&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;RFC 3849 (documentation)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Card numbers&lt;/td&gt;
&lt;td&gt;Your gateway's published test numbers, sandbox only&lt;/td&gt;
&lt;td&gt;Payment provider docs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two details that bite people. The whole &lt;code&gt;555&lt;/code&gt; prefix is &lt;em&gt;not&lt;/em&gt; reserved — only &lt;code&gt;555-0100&lt;/code&gt; through &lt;code&gt;555-0199&lt;/code&gt;; &lt;code&gt;555-1212&lt;/code&gt; is real directory assistance in much of the NANP. And other regulators publish their own drama ranges, but the specific blocks differ by country, so look them up at the regulator rather than pattern-matching from the US ones.&lt;/p&gt;

&lt;p&gt;For email, &lt;code&gt;example.com&lt;/code&gt; is safe to &lt;em&gt;write&lt;/em&gt; but mail to it goes nowhere, which means you cannot open a signup confirmation. If your tests need to read the message, keep the reserved domain in the database and point your provider at a mail-capture inbox in the preview environment's config — separate concerns, separate places.&lt;/p&gt;

&lt;p&gt;The takeaway: reserved ranges are valid enough to pass your validators and dead enough that no carrier or DNS resolver will deliver anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I generate them deterministically?
&lt;/h2&gt;

&lt;p&gt;Hash the primary key. Same input, same fake value, every run — so a re-scrub does not churn your test fixtures, and the same user looks identical in &lt;code&gt;users&lt;/code&gt;, &lt;code&gt;invoices&lt;/code&gt;, and your search index.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Run only on a machine allowed to see production data.&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;OR&lt;/span&gt; &lt;span class="k"&gt;REPLACE&lt;/span&gt; &lt;span class="k"&gt;FUNCTION&lt;/span&gt; &lt;span class="n"&gt;fake_phone_e164&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seed&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;RETURNS&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;
  &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="s1"&gt;'+1'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;npa&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s1"&gt;'555'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;lpad&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;))::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'0'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
           &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ARRAY&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'202'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'212'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'312'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'415'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'503'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'617'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'713'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'808'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
             &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;npa&lt;/span&gt;
    &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'x'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;substr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;))::&lt;/span&gt;&lt;span class="nb"&gt;bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;
  &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="err"&gt;$$&lt;/span&gt; &lt;span class="k"&gt;LANGUAGE&lt;/span&gt; &lt;span class="k"&gt;sql&lt;/span&gt; &lt;span class="k"&gt;IMMUTABLE&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt;
  &lt;span class="n"&gt;phone&lt;/span&gt;      &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;fake_phone_e164&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;email&lt;/span&gt;      &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'user-'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;substr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s1"&gt;'@example.com'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;full_name&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'Test User '&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;substr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;last_ip&lt;/span&gt;    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'203.0.113.'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="s1"&gt;'x'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;substr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;))::&lt;/span&gt;&lt;span class="nb"&gt;bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;250&lt;/span&gt;&lt;span class="p"&gt;)))::&lt;/span&gt;&lt;span class="n"&gt;inet&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;bit(24)::int&lt;/code&gt; is always positive, which saves you from the &lt;code&gt;abs()&lt;/code&gt; overflow footgun on the more common &lt;code&gt;bit(32)::int&lt;/code&gt; version of this trick.&lt;/p&gt;

&lt;p&gt;Keep the shape of the data, not just the type. If 20% of your users have no phone in production, preserve that: &lt;code&gt;CASE WHEN phone IS NULL THEN NULL ELSE fake_phone_e164(id::text) END&lt;/code&gt;. A column where every row is populated tests a world you do not ship to.&lt;/p&gt;

&lt;p&gt;The takeaway: derive fakes from the primary key so the scrub is idempotent and referentially consistent across tables.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I prove the scrub actually worked?
&lt;/h2&gt;

&lt;p&gt;Assert it in the same transaction, and let the assertion fail the job. A scrub that half-ran is more dangerous than no scrub, because everyone downstream assumes it completed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DO&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;
&lt;span class="k"&gt;DECLARE&lt;/span&gt; &lt;span class="n"&gt;leaked&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;BEGIN&lt;/span&gt;
  &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;leaked&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;
   &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;phone&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;phone&lt;/span&gt; &lt;span class="o"&gt;!~&lt;/span&gt; &lt;span class="s1"&gt;'^&lt;/span&gt;&lt;span class="se"&gt;\+&lt;/span&gt;&lt;span class="s1"&gt;1&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="s1"&gt;{3}55501&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="s1"&gt;{2}$'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="n"&gt;leaked&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;THEN&lt;/span&gt;
    &lt;span class="n"&gt;RAISE&lt;/span&gt; &lt;span class="n"&gt;EXCEPTION&lt;/span&gt; &lt;span class="s1"&gt;'scrub incomplete: % rows outside reserved phone range'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;leaked&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;leaked&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;
   &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="o"&gt;!~*&lt;/span&gt; &lt;span class="s1"&gt;'@example&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="s1"&gt;(com|net|org)$'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="n"&gt;leaked&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;THEN&lt;/span&gt;
    &lt;span class="n"&gt;RAISE&lt;/span&gt; &lt;span class="n"&gt;EXCEPTION&lt;/span&gt; &lt;span class="s1"&gt;'scrub incomplete: % rows with non-reserved email domain'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;leaked&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run this as the last step of the export job, and again as the first step of the import job on the untrusted side. The second run is the one that catches a snapshot that was copied from the wrong bucket.&lt;/p&gt;

&lt;p&gt;If you want the managed version of this, PostgreSQL Anonymizer is the open-source extension that lets you declare masking rules as security labels on the columns themselves, so the rule lives next to the schema instead of in a script that drifts. Its cost is real: it is a C extension, so it has to be available on your managed Postgres provider, and the rules need database-level privileges to manage. Neosync covers similar ground as a service with referential-integrity-aware transformers if you would rather not run an extension at all, with the usual tradeoff of sending your schema to a third party.&lt;/p&gt;

&lt;p&gt;The takeaway: a scrub without a machine-checked assertion is a hope, and hopes do not fail CI.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about unique constraints and volume?
&lt;/h2&gt;

&lt;p&gt;This is where the approach has a hard edge. The North American reserved block gives you exactly 100 numbers per area code. Eight area codes is 800 distinct values; every valid NPA is on the order of tens of thousands. If &lt;code&gt;users.phone&lt;/code&gt; carries a unique index and you have more rows than that, the scrub will fail on a duplicate key.&lt;/p&gt;

&lt;p&gt;Do not widen the range to fix that. The right response is to shrink the dataset: a preview or test database needs hundreds of representative rows, not your full user table. Subset first, scrub second, and the ceiling stops mattering. If you genuinely need a full-size dataset for load testing, drop the unique constraint in that environment explicitly and write down why — as of late 2026 there is no reserved range large enough to satisfy a million-row unique phone column, and pretending otherwise means routable numbers.&lt;/p&gt;

&lt;p&gt;The takeaway: running out of reserved values is a signal that your test dataset is too big, not that the range is too small.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What phone numbers are safe to use for test data?&lt;/strong&gt;&lt;br&gt;
In North America, &lt;code&gt;555-0100&lt;/code&gt; through &lt;code&gt;555-0199&lt;/code&gt; in any area code are reserved by NANPA for fictional use and will not connect to a subscriber. In the UK, Ofcom reserves &lt;code&gt;07700 900000&lt;/code&gt;–&lt;code&gt;900999&lt;/code&gt; for mobile and &lt;code&gt;020 7946 0000&lt;/code&gt;–&lt;code&gt;0999&lt;/code&gt; for London landlines. Other prefixes, including the rest of &lt;code&gt;555&lt;/code&gt;, can be real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I send test emails to example.com?&lt;/strong&gt;&lt;br&gt;
You can safely store &lt;code&gt;@example.com&lt;/code&gt; addresses — RFC 2606 reserves the domain so nobody can register it or receive mail there. But nothing will be delivered, so if your test needs to open the email, route the preview environment's SMTP to a mail-capture inbox instead of relying on the address.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I keep fake data consistent across tables?&lt;/strong&gt;&lt;br&gt;
Generate every fake value from a hash of the row's primary key rather than from a random source. The same user ID produces the same phone number and email on every run and in every table that references it, which makes the scrub idempotent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Null-based scrubbing is the default because it is one line of SQL, and it quietly deletes the test coverage you thought you had. Write reserved-range values instead: they pass validation, exercise the same code paths as real data, and cannot reach a human even if someone wires the environment to live credentials by mistake. Derive them deterministically from the primary key, assert the result with a query that can fail the job, and subset your data before you scrub it so unique constraints never force you outside the reserved block. If you would rather declare the rules than maintain a script, PostgreSQL Anonymizer is the sane starting point on self-managed Postgres.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/preview-environments-per-pull-request-how-to-seed-them-without-cloning-production-1f08"&gt;Preview Environments Per Pull Request: How to Seed Them Without Cloning Production&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/curl-returns-200-but-your-http-client-gets-a-404-debugging-vary-and-cdn-cache-variants-38je"&gt;curl Returns 200 but Your HTTP Client Gets a 404: Debugging Vary and CDN Cache Variants&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/netlify-pros-and-cons-when-its-the-right-host-and-when-youll-outgrow-it-2ka1"&gt;Netlify Pros and Cons: When It's the Right Host, and When You'll Outgrow It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>testing</category>
      <category>privacy</category>
    </item>
    <item>
      <title>CORS Errors Only in Production: Preflight, Credentials, and the Cache That Hides Your Fix</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:26:02 +0000</pubDate>
      <link>https://dev.to/libme/cors-errors-only-in-production-preflight-credentials-and-the-cache-that-hides-your-fix-4mdk</link>
      <guid>https://dev.to/libme/cors-errors-only-in-production-preflight-credentials-and-the-cache-that-hides-your-fix-4mdk</guid>
      <description>&lt;p&gt;If your frontend talks to your API fine on localhost and dies with a CORS error in production, the cause is almost never "CORS is broken." It is that locally the browser saw one origin (a dev-server proxy) and in production it sees two, and the first cross-origin request your app makes is now a preflight &lt;code&gt;OPTIONS&lt;/code&gt; that your backend answers wrong — often with a &lt;code&gt;401&lt;/code&gt;, a redirect, or a &lt;code&gt;500&lt;/code&gt; that carries no CORS headers at all. Fix it in this order: read the actual preflight response, decide where CORS belongs (proxy, app, or gateway), then make sure your CDN is not caching the answer for the wrong origin.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does it work locally and fail in production?
&lt;/h2&gt;

&lt;p&gt;In development, Vite, Next.js, and Create React App all proxy &lt;code&gt;/api&lt;/code&gt; to your backend. The browser sees &lt;code&gt;http://localhost:5173/api/items&lt;/code&gt; — same origin, no CORS at all. In production you deploy the frontend to &lt;code&gt;https://app.example.com&lt;/code&gt; and the API to &lt;code&gt;https://api.example.com&lt;/code&gt;, and every request is suddenly cross-origin.&lt;/p&gt;

&lt;p&gt;Cross-origin does not automatically mean preflight. The browser sends a preflight only when the request is not "simple": a method other than &lt;code&gt;GET&lt;/code&gt;, &lt;code&gt;HEAD&lt;/code&gt;, or &lt;code&gt;POST&lt;/code&gt;, a &lt;code&gt;Content-Type&lt;/code&gt; outside &lt;code&gt;application/x-www-form-urlencoded&lt;/code&gt;, &lt;code&gt;multipart/form-data&lt;/code&gt;, and &lt;code&gt;text/plain&lt;/code&gt;, or any header you set yourself. That is why the failure often looks arbitrary: your &lt;code&gt;GET /health&lt;/code&gt; works, and &lt;code&gt;POST /v1/items&lt;/code&gt; with &lt;code&gt;Content-Type: application/json&lt;/code&gt; and an &lt;code&gt;Authorization&lt;/code&gt; header fails. The JSON content type alone is enough to trigger the preflight.&lt;/p&gt;

&lt;p&gt;The first thing to establish is not "is CORS configured" but "did this request preflight at all" — those are two different bugs with two different fixes.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I read the real error instead of guessing?
&lt;/h2&gt;

&lt;p&gt;Chrome's console messages are specific, and each one points at a different header:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Console message (abridged)&lt;/th&gt;
&lt;th&gt;What it actually means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;No 'Access-Control-Allow-Origin' header is present&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The response never went through your CORS layer — wrong route, error path, or middleware order&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Response to preflight request doesn't pass access control check&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The &lt;code&gt;OPTIONS&lt;/code&gt; response is the problem, not your &lt;code&gt;POST&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Request header field authorization is not allowed by Access-Control-Allow-Headers&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Preflight answered, but the allow-list is missing a header you send&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Method PATCH is not allowed by Access-Control-Allow-Methods&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Same, for the method&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;...must not be the wildcard '*' when the request's credentials mode is 'include'&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;You are sending cookies; &lt;code&gt;*&lt;/code&gt; is illegal in that mode&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Redirect is not allowed for a preflight request&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Your &lt;code&gt;OPTIONS&lt;/code&gt; got a &lt;code&gt;301&lt;/code&gt;/&lt;code&gt;302&lt;/code&gt; (usually HTTP→HTTPS or a trailing-slash rule)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Then reproduce the preflight by hand. This is the single most useful command in the whole investigation, because it shows you exactly what the browser sees without the browser in the way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; OPTIONS https://api.example.com/v1/items &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Origin: https://app.example.com'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Access-Control-Request-Method: POST'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Access-Control-Request-Headers: content-type,authorization'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You want &lt;code&gt;204&lt;/code&gt; (or &lt;code&gt;200&lt;/code&gt;) plus &lt;code&gt;Access-Control-Allow-Origin&lt;/code&gt;, &lt;code&gt;Access-Control-Allow-Methods&lt;/code&gt;, and &lt;code&gt;Access-Control-Allow-Headers&lt;/code&gt; covering what you asked for. What I usually find instead is a &lt;code&gt;401&lt;/code&gt;: the preflight carries no cookies and no &lt;code&gt;Authorization&lt;/code&gt; header by design, so any auth middleware mounted before the CORS layer rejects it. The browser then reports a generic CORS failure and never sends the real request.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;curl -i -X OPTIONS&lt;/code&gt; that returns 401 or 302 is your answer — stop reading CORS docs and fix middleware order or redirect rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where should CORS live: proxy, app, or gateway?
&lt;/h2&gt;

&lt;p&gt;This is the decision that actually matters, and most teams make it by accident.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Good when&lt;/th&gt;
&lt;th&gt;Real drawback&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Same-origin path routing (&lt;code&gt;/api/*&lt;/code&gt; → backend at the edge)&lt;/td&gt;
&lt;td&gt;You control DNS and want CORS to disappear entirely&lt;/td&gt;
&lt;td&gt;Another hop to operate; cookie &lt;code&gt;Domain&lt;/code&gt;/&lt;code&gt;Path&lt;/code&gt; and cache rules need review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CORS in the app (framework middleware)&lt;/td&gt;
&lt;td&gt;Small number of services, origins known at deploy time&lt;/td&gt;
&lt;td&gt;Every service re-implements it; error responses easily bypass it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CORS at the gateway&lt;/td&gt;
&lt;td&gt;Many services behind one entry point&lt;/td&gt;
&lt;td&gt;Two places can emit the headers, and duplicated &lt;code&gt;Access-Control-Allow-Origin&lt;/code&gt; is itself a hard failure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same-origin routing is underrated. If your frontend already sits behind a CDN, forwarding &lt;code&gt;app.example.com/api/*&lt;/code&gt; to the API removes preflights, third-party cookie problems, and the entire &lt;code&gt;Allow-Headers&lt;/code&gt; allow-list in one move. If you want that without running a reverse proxy yourself, Cloudflare Workers can sit in front of the existing hostname and forward &lt;code&gt;/api/*&lt;/code&gt; to the backend, which keeps the browser on one origin. The cost is real: you now own a routing layer, and you have to be deliberate about what it caches.&lt;/p&gt;

&lt;p&gt;If you keep CORS in the app, mount it first and make sure it also wraps error handlers. In Express:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;express&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;express&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cors&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;cors&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;express&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;allowed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://app.example.com&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://staging.example.com&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;use&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;cors&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;origin&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;origin&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;cb&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// no Origin header: curl, server-to-server, same-origin&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;origin&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;allowed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;origin&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;cb&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;cb&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`origin not allowed: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;origin&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;credentials&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;methods&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;GET&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PATCH&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;DELETE&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;allowedHeaders&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Authorization&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;maxAge&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}))&lt;/span&gt;

&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;use&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;express&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/v1/items&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;

&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two notes from getting this wrong. The &lt;code&gt;cors&lt;/code&gt; middleware already answers &lt;code&gt;OPTIONS&lt;/code&gt; when mounted with &lt;code&gt;app.use&lt;/code&gt;, so a separate catch-all &lt;code&gt;OPTIONS&lt;/code&gt; route is unnecessary — and on Express 5 a bare &lt;code&gt;app.options('*', ...)&lt;/code&gt; no longer behaves as it did on 4, because the path-matching rules changed. Second, if auth middleware comes before this block, you are back to the &lt;code&gt;401&lt;/code&gt; preflight.&lt;/p&gt;

&lt;p&gt;In FastAPI the equivalent is &lt;code&gt;CORSMiddleware&lt;/code&gt;, and the trap is different: passing &lt;code&gt;allow_origins=["*"]&lt;/code&gt; together with &lt;code&gt;allow_credentials=True&lt;/code&gt; is invalid per spec, and frameworks resolve that contradiction inconsistently rather than erroring.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi.middleware.cors&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;CORSMiddleware&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_middleware&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;CORSMiddleware&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;allow_origins&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://app.example.com&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;allow_credentials&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;allow_methods&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GET&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;POST&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PATCH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DELETE&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;allow_headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;max_age&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Never trust the config object — trust the &lt;code&gt;curl -i&lt;/code&gt; output, because the header your framework actually emits is the only thing the browser evaluates.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes the moment cookies are involved?
&lt;/h2&gt;

&lt;p&gt;Once the frontend sends &lt;code&gt;credentials: 'include'&lt;/code&gt;, three rules bind at once: &lt;code&gt;Access-Control-Allow-Origin&lt;/code&gt; must be an exact origin (no &lt;code&gt;*&lt;/code&gt;), &lt;code&gt;Access-Control-Allow-Credentials: true&lt;/code&gt; must be present, and wildcards in &lt;code&gt;Allow-Headers&lt;/code&gt; or &lt;code&gt;Allow-Methods&lt;/code&gt; stop counting. A cookie crossing sites also needs &lt;code&gt;SameSite=None; Secure&lt;/code&gt; — so the "CORS error" may be a cookie that was never sent.&lt;/p&gt;

&lt;p&gt;The fastest way to tell them apart: if the preflight passes and the real request returns &lt;code&gt;401&lt;/code&gt;, CORS is fine and the cookie is the problem. Check the Network tab's request headers for &lt;code&gt;Cookie&lt;/code&gt; before touching any CORS config.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does the fix work for one user and not another?
&lt;/h2&gt;

&lt;p&gt;Because something cached it. If you echo the request's &lt;code&gt;Origin&lt;/code&gt; into &lt;code&gt;Access-Control-Allow-Origin&lt;/code&gt; — which you must do for credentialed requests — the response varies by a request header, and any shared cache in front of you needs to know that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;Access-Control-Allow-Origin: https://app.example.com
Access-Control-Allow-Credentials: true
Vary: Origin
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without &lt;code&gt;Vary: Origin&lt;/code&gt;, a CDN can serve &lt;code&gt;app.example.com&lt;/code&gt;'s allow-header to &lt;code&gt;staging.example.com&lt;/code&gt; and back, and the symptom is the worst kind: intermittent, user-specific, unreproducible on your machine. I have also watched a correct fix look like a non-fix because a stale preflight was still cached — browsers cache preflight results per &lt;code&gt;Access-Control-Max-Age&lt;/code&gt;, and Chromium caps that at two hours (as of mid-2026). Test fixes in a fresh incognito window, or keep &lt;code&gt;maxAge&lt;/code&gt; low until things are stable.&lt;/p&gt;

&lt;p&gt;Any response whose CORS headers depend on the request's &lt;code&gt;Origin&lt;/code&gt; must send &lt;code&gt;Vary: Origin&lt;/code&gt;, or your cache will eventually hand the wrong answer to the wrong site.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does my API work in Postman but not in the browser?&lt;/strong&gt;&lt;br&gt;
Postman and curl are not browsers and do not enforce CORS — they never send a preflight and never check response headers. A request succeeding in Postman tells you the endpoint works; it tells you nothing about CORS.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I fix a CORS error from the frontend?&lt;/strong&gt;&lt;br&gt;
No. CORS headers come from the server that owns the resource, so the only frontend-side "fixes" are avoiding the cross-origin request entirely (proxy the call through your own origin) or, for third-party APIs you do not control, calling them from your backend.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does my preflight return 401?&lt;/strong&gt;&lt;br&gt;
Preflight &lt;code&gt;OPTIONS&lt;/code&gt; requests deliberately carry no cookies and no &lt;code&gt;Authorization&lt;/code&gt; header, so authentication middleware rejects them unless CORS handling runs first. Mount your CORS layer before auth, or exempt &lt;code&gt;OPTIONS&lt;/code&gt; from authentication.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Debug in this order: confirm whether a preflight is involved, reproduce it with &lt;code&gt;curl -i -X OPTIONS&lt;/code&gt;, and only then edit configuration. If you control both hostnames, same-origin path routing at the edge is the most durable fix, because it deletes the problem class instead of configuring it. If you keep CORS in the app, mount it before auth, list your origins explicitly, and send &lt;code&gt;Vary: Origin&lt;/code&gt;. And if the preflight passes but the real request 401s, you have a cookie problem wearing a CORS costume.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/before-you-set-plancachemode-write-the-regression-test-that-proves-it-worked-23l6"&gt;Before You Set plan_cache_mode, Write the Regression Test That Proves It Worked&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/your-secrets-manager-ends-at-processenv-where-secrets-actually-leak-at-runtime-3b8b"&gt;Your Secrets Manager Ends at process.env: Where Secrets Actually Leak at Runtime&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/your-postgres-migration-runner-needs-a-retry-contract-not-just-a-lock-timeout-1oii"&gt;Your Postgres Migration Runner Needs a Retry Contract, Not Just a Lock Timeout&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>webdev</category>
      <category>api</category>
      <category>backend</category>
      <category>security</category>
    </item>
    <item>
      <title>Docker Image Works on Your Mac but Fails on the Server: Fixing exec format error for Good</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Mon, 28 Sep 2026 15:30:22 +0000</pubDate>
      <link>https://dev.to/libme/docker-image-works-on-your-mac-but-fails-on-the-server-fixing-exec-format-error-for-good-dag</link>
      <guid>https://dev.to/libme/docker-image-works-on-your-mac-but-fails-on-the-server-fixing-exec-format-error-for-good-dag</guid>
      <description>&lt;p&gt;If your container runs fine locally on an Apple Silicon Mac and dies on the server with &lt;code&gt;exec /usr/local/bin/app: exec format error&lt;/code&gt;, you built an arm64 image and deployed it to an amd64 host. The fix is not a flag on &lt;code&gt;docker run&lt;/code&gt; — you need a multi-architecture image, built either by cross-compiling inside the Dockerfile or by building each architecture on its own native machine and merging the results into one manifest. Emulation via QEMU will also work, and it is the slowest option by a wide margin on compile-heavy images.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does &lt;code&gt;exec format error&lt;/code&gt; only show up on the server?
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;docker build&lt;/code&gt; produces an image for the architecture of whatever machine ran the build, unless you tell it otherwise. On an M-series Mac that is &lt;code&gt;linux/arm64&lt;/code&gt;. Your EC2 box, your Hetzner VPS, most CI runners, and most managed container platforms are still &lt;code&gt;linux/amd64&lt;/code&gt;. The kernel on that host loads your ELF binary, reads a machine type it cannot execute, and returns &lt;code&gt;ENOEXEC&lt;/code&gt;. Docker surfaces that as &lt;code&gt;exec format error&lt;/code&gt;, which reads like a corrupt binary and is actually a passport problem.&lt;/p&gt;

&lt;p&gt;You can confirm the mismatch in about ten seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# What did I actually build?&lt;/span&gt;
docker image inspect myapp:latest &lt;span class="nt"&gt;--format&lt;/span&gt; &lt;span class="s1"&gt;'{{.Os}}/{{.Architecture}}'&lt;/span&gt;

&lt;span class="c"&gt;# What architectures does the pushed tag serve?&lt;/span&gt;
docker buildx imagetools inspect ghcr.io/me/myapp:latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the second command lists only one platform, every host that isn't that platform is one deploy away from the same error. The related failure you'll hit on the way is &lt;code&gt;no matching manifest for linux/amd64 in the manifest list entries&lt;/code&gt; — that is the same problem caught earlier, at pull time instead of exec time, which is strictly better.&lt;/p&gt;

&lt;p&gt;There's also a quieter variant. Run an amd64 image on an Apple Silicon Mac and Docker prints &lt;code&gt;WARNING: The requested image's platform (linux/amd64) does not match the detected host platform&lt;/code&gt; and then runs it anyway under emulation. That warning is the one people learn to scroll past, and it is exactly the signal that your local and production architectures have drifted apart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: &lt;code&gt;exec format error&lt;/code&gt; is never a corrupted binary — it means the image's architecture and the host's architecture disagree, and one &lt;code&gt;imagetools inspect&lt;/code&gt; proves which side is wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What are the three ways to build for both architectures?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;arm64 + amd64 build speed&lt;/th&gt;
&lt;th&gt;Setup cost&lt;/th&gt;
&lt;th&gt;Works with CGO / native deps&lt;/th&gt;
&lt;th&gt;Best when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;QEMU emulation on one runner&lt;/td&gt;
&lt;td&gt;Slowest; the emulated half dominates the build&lt;/td&gt;
&lt;td&gt;One extra action, no Dockerfile changes&lt;/td&gt;
&lt;td&gt;Yes, but painfully slow&lt;/td&gt;
&lt;td&gt;Interpreted apps, images you rebuild rarely&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cross-compile in the Dockerfile&lt;/td&gt;
&lt;td&gt;Fast — no emulation at all&lt;/td&gt;
&lt;td&gt;Dockerfile rework&lt;/td&gt;
&lt;td&gt;No (needs a cross toolchain)&lt;/td&gt;
&lt;td&gt;Go and Rust services&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Native runner per architecture, then merge manifests&lt;/td&gt;
&lt;td&gt;Fast, and both halves run in parallel&lt;/td&gt;
&lt;td&gt;Two jobs plus a merge job&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Anything, including Python wheels and node-gyp&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A fourth option is to rent someone else's native builders. If you want a managed version of the two-runner setup without maintaining the matrix yourself, Depot runs native amd64 and arm64 builders behind a single &lt;code&gt;docker build&lt;/code&gt; call and keeps the layer cache on persistent volumes between runs; the trade-off is that your build context now leaves your CI provider for a third party, which some compliance reviews will care about. Docker Build Cloud is the equivalent from Docker itself and has the advantage of reusing your existing Docker Hub identity, with the caveat that its pricing is metered in build minutes, so a repo with a noisy &lt;code&gt;push&lt;/code&gt; trigger can spend more than you expect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: emulation is the cheapest to configure and the most expensive to run, so treat it as a stopgap rather than a resting place.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I set up buildx and QEMU correctly?
&lt;/h2&gt;

&lt;p&gt;The single-runner path needs the QEMU binfmt handlers registered and a buildx builder that uses the &lt;code&gt;docker-container&lt;/code&gt; driver — the default &lt;code&gt;docker&lt;/code&gt; driver cannot build more than one platform at a time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;--privileged&lt;/span&gt; &lt;span class="nt"&gt;--rm&lt;/span&gt; tonistiigi/binfmt &lt;span class="nt"&gt;--install&lt;/span&gt; all
docker buildx create &lt;span class="nt"&gt;--name&lt;/span&gt; multi &lt;span class="nt"&gt;--driver&lt;/span&gt; docker-container &lt;span class="nt"&gt;--use&lt;/span&gt;
docker buildx build &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--platform&lt;/span&gt; linux/amd64,linux/arm64 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-t&lt;/span&gt; ghcr.io/me/myapp:1.4.0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--push&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The thing that trips people up here: you cannot &lt;code&gt;--load&lt;/code&gt; a multi-platform build into your local image store. Docker's classic image store holds one manifest per tag, so &lt;code&gt;--load&lt;/code&gt; with two platforms fails outright. Either push to a registry (as above), or build a single platform locally for testing with &lt;code&gt;--platform linux/amd64 --load&lt;/code&gt;. As of mid-2026 the containerd image store in Docker Desktop removes this limitation, but it's an opt-in setting, so don't write a team runbook that assumes it.&lt;/p&gt;

&lt;p&gt;In GitHub Actions the equivalent is two setup steps before the build:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/setup-qemu-action@v3&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/setup-buildx-action@v3&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/build-push-action@v6&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;platforms&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;linux/amd64,linux/arm64&lt;/span&gt;
    &lt;span class="na"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io/${{ github.repository }}:latest&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway: &lt;code&gt;--platform&lt;/code&gt; with two values and &lt;code&gt;--load&lt;/code&gt; are mutually exclusive on the classic image store — push to a registry, or build one platform at a time locally.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I stop the emulated half from dominating build time?
&lt;/h2&gt;

&lt;p&gt;For compiled languages, skip emulation entirely. BuildKit hands you &lt;code&gt;BUILDPLATFORM&lt;/code&gt; and &lt;code&gt;TARGETARCH&lt;/code&gt;; pin the build stage to the native platform and let the compiler do the cross-targeting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="c"&gt;# syntax=docker/dockerfile:1&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;--platform=$BUILDPLATFORM golang:1.24-alpine&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;AS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;build&lt;/span&gt;
&lt;span class="k"&gt;ARG&lt;/span&gt;&lt;span class="s"&gt; TARGETOS TARGETARCH&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /src&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; go.mod go.sum ./&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;go mod download
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; . .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;&lt;span class="nv"&gt;CGO_ENABLED&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0 &lt;span class="nv"&gt;GOOS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$TARGETOS&lt;/span&gt; &lt;span class="nv"&gt;GOARCH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$TARGETARCH&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;    go build &lt;span class="nt"&gt;-o&lt;/span&gt; /out/app ./cmd/app

&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; alpine:3.20&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=build /out/app /usr/local/bin/app&lt;/span&gt;
&lt;span class="k"&gt;ENTRYPOINT&lt;/span&gt;&lt;span class="s"&gt; ["/usr/local/bin/app"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every stage now runs natively; only the tiny final &lt;code&gt;COPY&lt;/code&gt; differs per architecture. The honest limitation is &lt;code&gt;CGO_ENABLED=0&lt;/code&gt;: the moment you need cgo, sqlite bindings, or any C library, you're back to installing a cross toolchain, and that is usually more work than the third option below.&lt;/p&gt;

&lt;p&gt;For everything else — Python with compiled wheels, Node with native addons, anything with a &lt;code&gt;make install&lt;/code&gt; that assumes a real compiler — build each architecture on a machine that actually is that architecture, then stitch the digests together. As of mid-2026 GitHub offers hosted Linux arm64 runners under labels like &lt;code&gt;ubuntu-24.04-arm&lt;/code&gt;, which makes this pattern available without self-hosting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;matrix&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;include&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;platform&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;linux/amd64&lt;/span&gt;
            &lt;span class="na"&gt;runner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-24.04&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;platform&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;linux/arm64&lt;/span&gt;
            &lt;span class="na"&gt;runner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-24.04-arm&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ matrix.runner }}&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/setup-buildx-action@v3&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/login-action@v3&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;registry&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io&lt;/span&gt;
          &lt;span class="na"&gt;username&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.actor }}&lt;/span&gt;
          &lt;span class="na"&gt;password&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.GITHUB_TOKEN }}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;build&lt;/span&gt;
        &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/build-push-action@v6&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;platforms&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ matrix.platform }}&lt;/span&gt;
          &lt;span class="na"&gt;outputs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;type=image,name=ghcr.io/${{ github.repository }},push-by-digest=true,name-canonical=true,push=true&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;mkdir -p /tmp/digests&lt;/span&gt;
          &lt;span class="s"&gt;touch "/tmp/digests/${DIGEST#sha256:}"&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;DIGEST&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ steps.build.outputs.digest }}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/upload-artifact@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;digests-${{ strategy.job-index }}&lt;/span&gt;
          &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/tmp/digests/*&lt;/span&gt;

  &lt;span class="na"&gt;merge&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;needs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;build&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-24.04&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/download-artifact@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/tmp/digests&lt;/span&gt;
          &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;digests-*&lt;/span&gt;
          &lt;span class="na"&gt;merge-multiple&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/setup-buildx-action@v3&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker/login-action@v3&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;registry&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io&lt;/span&gt;
          &lt;span class="na"&gt;username&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.actor }}&lt;/span&gt;
          &lt;span class="na"&gt;password&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.GITHUB_TOKEN }}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;working-directory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/tmp/digests&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;docker buildx imagetools create \&lt;/span&gt;
            &lt;span class="s"&gt;-t ghcr.io/${{ github.repository }}:latest \&lt;/span&gt;
            &lt;span class="s"&gt;$(printf 'ghcr.io/${{ github.repository }}@sha256:%s ' *)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two build jobs run in parallel, so wall-clock time is roughly the slower of the two native builds rather than the sum of a native build and an emulated one. What tripped me up the first time: the merge job needs its own registry login, because &lt;code&gt;imagetools create&lt;/code&gt; reads the source digests from the registry rather than from any local state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway: parallel native builds plus a manifest merge is the only approach that is both fast and language-agnostic, and it costs you one extra CI job.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Which one should you actually pick?
&lt;/h2&gt;

&lt;p&gt;Start from what breaks if you choose wrong. If your image is a Go or Rust binary, cross-compiling in the Dockerfile is the smallest change with the biggest win, and you can do it in one commit. If your build installs native dependencies, go straight to the matrix-plus-merge workflow — you will fight the cross toolchain for longer than it takes to write the merge job. Reach for emulation when the image is a thin layer over an interpreter and you rebuild it a few times a week; the build being slow simply won't matter at that frequency.&lt;/p&gt;

&lt;p&gt;And regardless of approach, make the mismatch impossible to ship. Adding a pull-time architecture assertion to your deploy script catches the problem before a container ever restarts in a loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail
&lt;span class="nv"&gt;want&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"linux/amd64"&lt;/span&gt;
docker buildx imagetools inspect &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$want&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"image &lt;/span&gt;&lt;span class="nv"&gt;$IMAGE&lt;/span&gt;&lt;span class="s2"&gt; has no &lt;/span&gt;&lt;span class="nv"&gt;$want&lt;/span&gt;&lt;span class="s2"&gt; variant"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway: pick the approach by whether your build needs a C compiler, then verify the published manifest in CI so architecture drift fails the pipeline instead of the deploy.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What does &lt;code&gt;exec /usr/local/bin/app: exec format error&lt;/code&gt; mean in Docker?&lt;/strong&gt;&lt;br&gt;
It means the binary inside the image was compiled for a different CPU architecture than the host running the container — almost always an arm64 image built on an Apple Silicon Mac being run on an amd64 server. Rebuild the image with &lt;code&gt;docker buildx build --platform linux/amd64&lt;/code&gt; or publish a multi-architecture manifest covering both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I build a multi-platform Docker image and load it into my local Docker?&lt;/strong&gt;&lt;br&gt;
Not with the classic image store: &lt;code&gt;docker buildx build --platform linux/amd64,linux/arm64 --load&lt;/code&gt; fails because a local tag can only point at one manifest. Push to a registry instead, or build a single platform with &lt;code&gt;--load&lt;/code&gt; for local testing. Docker Desktop's containerd image store lifts this restriction, but it is opt-in as of mid-2026.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is &lt;code&gt;--platform linux/amd64&lt;/code&gt; on an M1 Mac enough for production?&lt;/strong&gt;&lt;br&gt;
It produces a correct amd64 image, but it runs your whole build under QEMU emulation, which is dramatically slower for anything that compiles code and occasionally exposes emulator bugs in native toolchains. It's fine as a one-off; for CI, build each architecture on a native runner and merge the manifests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/how-to-grade-search-relevance-before-you-touch-the-ranking-weights-3554"&gt;How to Grade Search Relevance Before You Touch the Ranking Weights&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/hotstandbyfeedback-for-a-reporting-replica-query-cancellations-or-primary-bloat-mek"&gt;hot_standby_feedback for a Reporting Replica: Query Cancellations or Primary Bloat?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/preview-environments-per-pull-request-how-to-seed-them-without-cloning-production-1f08"&gt;Preview Environments Per Pull Request: How to Seed Them Without Cloning Production&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>docker</category>
      <category>devops</category>
      <category>cicd</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Preview Environments Per Pull Request: How to Seed Them Without Cloning Production</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Sun, 27 Sep 2026 15:29:30 +0000</pubDate>
      <link>https://dev.to/libme/preview-environments-per-pull-request-how-to-seed-them-without-cloning-production-1f08</link>
      <guid>https://dev.to/libme/preview-environments-per-pull-request-how-to-seed-them-without-cloning-production-1f08</guid>
      <description>&lt;p&gt;Spinning up a preview deploy per pull request is the easy half — every modern host does it with a config flag. The hard half is the database: a shared staging database turns previews into a queue of people breaking each other's data, and a restored production snapshot puts real customer records in an environment that anyone with the PR link can reach. The pattern that survives is an ephemeral database per PR, created from schema plus a small deterministic seed, with production data used only after it has been scrubbed on a machine that is allowed to see it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually breaks in a preview environment?
&lt;/h2&gt;

&lt;p&gt;The deploy almost never fails. What fails looks like this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everything points at one staging database.&lt;/strong&gt; Two PRs land migrations in different orders and the third developer gets &lt;code&gt;ERROR: relation "order_line_items" does not exist&lt;/code&gt; from a branch that never ran the migration adding it — or someone's test loop truncates a table another person was demoing from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Connection exhaustion.&lt;/strong&gt; Each preview app opens its own pool. Ten open PRs at ten connections each, against a small instance with a 100-connection ceiling, gives you &lt;code&gt;FATAL: sorry, too many clients already&lt;/code&gt; in whichever environment connected last — usually the one the reviewer just opened. Previews multiply &lt;em&gt;idle&lt;/em&gt; connections rather than traffic, which is why they surface pooling problems first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Migrations that only run forward.&lt;/strong&gt; Preview infrastructure will create a database for you; it will not un-apply your migration when you force-push a rewritten one. If your runner tracks applied versions in the database, recreate the environment rather than reusing it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Seed scripts that assume an empty database.&lt;/strong&gt; Run them twice on a reused environment and you get &lt;code&gt;duplicate key value violates unique constraint "users_email_key"&lt;/code&gt;. Write every seed as an upsert (&lt;code&gt;INSERT ... ON CONFLICT DO NOTHING&lt;/code&gt;) and you stop caring which state you started from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Secrets.&lt;/strong&gt; Preview builds inherit project-level environment variables by default on most hosts — that is how a preview ends up holding a live payment key. Give the preview scope its own secret set, test-mode credentials only.&lt;/p&gt;

&lt;p&gt;The takeaway: a preview environment is only as isolated as its database and its secrets, and both default to shared.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where should the preview database come from?
&lt;/h2&gt;

&lt;p&gt;Four options, and they differ mostly in how fast they are and how much real data they expose.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Setup cost&lt;/th&gt;
&lt;th&gt;Time to fresh env&lt;/th&gt;
&lt;th&gt;Data realism&lt;/th&gt;
&lt;th&gt;Main risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Shared staging DB&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Instant&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Cross-PR interference, migration drift&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ephemeral Postgres container + seed&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Seconds&lt;/td&gt;
&lt;td&gt;Low (what you seed)&lt;/td&gt;
&lt;td&gt;Seed rot: diverges from real schema use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Branch on managed Postgres (copy-on-write)&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Seconds&lt;/td&gt;
&lt;td&gt;High (mirrors parent)&lt;/td&gt;
&lt;td&gt;Real data in a low-trust environment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Restore of a scrubbed prod snapshot&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Scrub pipeline must be airtight&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For most small teams the ephemeral container plus a seed file is the correct default, and it is free. You reach for branching when the bugs you keep shipping are data-shaped — queries that are fast on 200 seeded rows and hopeless on the real distribution.&lt;/p&gt;

&lt;p&gt;If you want copy-on-write branches without building them, Neon is the managed Postgres whose branching is the product itself: a branch is a cheap pointer at the parent's storage, so a per-PR database costs you the diff rather than a full copy. Supabase offers git-linked preview branches if your app already lives on its stack, with the tradeoff that the branch is a whole project and carries that startup latency. As of mid-2026, treat either as a &lt;em&gt;pre-production&lt;/em&gt; surface no matter how convenient it is — a branch of production is production data.&lt;/p&gt;

&lt;p&gt;The takeaway: pick branching for data realism, containers for isolation and cost, and never a shared database for either.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you seed data without copying production?
&lt;/h2&gt;

&lt;p&gt;Build the seed from three layers, in this order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Schema from your migration runner, never from a dump.&lt;/strong&gt; The whole point is to exercise the migrations the PR contains.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A committed deterministic seed.&lt;/strong&gt; Fixed UUIDs and fixed timestamps, so a failing test is reproducible and reviewers can bookmark &lt;code&gt;/orders/0000...0001&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- seed/01_core.sql — idempotent, safe to re-run&lt;/span&gt;
&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;display_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'00000000-0000-0000-0000-000000000001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'owner@example.test'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="s1"&gt;'Test Owner'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="s1"&gt;'2026-01-02T00:00:00Z'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'00000000-0000-0000-0000-000000000002'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'viewer@example.test'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Test Viewer'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-02T00:00:00Z'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;CONFLICT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;DO&lt;/span&gt; &lt;span class="k"&gt;NOTHING&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;total_cents&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'00000000-0000-0000-0000-0000000100'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;lpad&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'0'&lt;/span&gt;&lt;span class="p"&gt;))::&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'00000000-0000-0000-0000-000000000001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ARRAY&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'pending'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'paid'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'refunded'&lt;/span&gt;&lt;span class="p"&gt;])[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt;
  &lt;span class="mi"&gt;1000&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;37&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;timestamptz&lt;/span&gt; &lt;span class="s1"&gt;'2026-01-03T00:00:00Z'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s1"&gt;' hours'&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="n"&gt;interval&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;generate_series&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;
&lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;CONFLICT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;DO&lt;/span&gt; &lt;span class="k"&gt;NOTHING&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use a reserved test domain like &lt;code&gt;example.test&lt;/code&gt; for every address. Seeds leak into outbound email eventually, and a domain that cannot resolve is the cheapest guardrail there is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Volume, generated rather than copied.&lt;/strong&gt; If the PR touches a query, 50 rows will lie to you. &lt;code&gt;generate_series&lt;/code&gt; into the hot table buys realistic row counts and index behavior without a single real record:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt;
  &lt;span class="s1"&gt;'00000000-0000-0000-0000-000000000001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'page_view'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;jsonb_build_object&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'path'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'/p/'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
  &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s1"&gt;' minutes'&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="n"&gt;interval&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;generate_series&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;200000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ANALYZE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;ANALYZE&lt;/code&gt; matters. Freshly bulk-loaded tables have no statistics, and the planner will pick a plan that has nothing to do with what production does.&lt;/p&gt;

&lt;p&gt;When you genuinely need production shapes, scrub inside the trusted environment and ship only the output — never &lt;code&gt;pg_dump&lt;/code&gt; production to a laptop first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

pg_dump &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROD_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--no-owner&lt;/span&gt; &lt;span class="nt"&gt;--no-privileges&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--exclude-table-data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'audit_log'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--exclude-table-data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sessions'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-Fc&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; /tmp/raw.dump

pg_restore &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SCRUB_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--no-owner&lt;/span&gt; &lt;span class="nt"&gt;--clean&lt;/span&gt; &lt;span class="nt"&gt;--if-exists&lt;/span&gt; /tmp/raw.dump
psql &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SCRUB_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nv"&gt;ON_ERROR_STOP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;-f&lt;/span&gt; scrub.sql
pg_dump &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SCRUB_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--no-owner&lt;/span&gt; &lt;span class="nt"&gt;-Fc&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; /tmp/preview-seed.dump
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- scrub.sql: irreversible, and verified by a test&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt;
  &lt;span class="n"&gt;email&lt;/span&gt;        &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'user'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s1"&gt;'@example.test'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;display_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'User '&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="k"&gt;left&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;phone&lt;/span&gt;        &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;DELETE&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;payment_methods&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scrub needs a test in the same job that fails loudly — a query asserting zero rows still match your real email domain. A scrub nobody verifies is a scrub that stopped matching the schema three migrations ago. If you would rather not hand-roll the rules, PostgreSQL Anonymizer is the extension that lets you declare masking rules on columns, so they live with the schema instead of in a script that drifts.&lt;/p&gt;

&lt;p&gt;The takeaway: generate volume, scrub only where you must, and assert the scrub worked in CI rather than trusting it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wiring it into CI without leaving databases behind
&lt;/h2&gt;

&lt;p&gt;The failure mode here is cost, not correctness: ephemeral environments are only ephemeral if something deletes them. Tie creation and teardown to the same PR lifecycle events.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;preview-db&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;types&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;opened&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;synchronize&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;reopened&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;closed&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;concurrency&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;group&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;preview-db-${{ github.event.pull_request.number }}&lt;/span&gt;
  &lt;span class="na"&gt;cancel-in-progress&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;up&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;github.event.action != 'closed'&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Create or reset the PR database&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;ADMIN_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.PREVIEW_ADMIN_URL }}&lt;/span&gt;
          &lt;span class="na"&gt;DB&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pr_${{ github.event.pull_request.number }}&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;psql "$ADMIN_URL" -v ON_ERROR_STOP=1 \&lt;/span&gt;
            &lt;span class="s"&gt;-c "DROP DATABASE IF EXISTS \"$DB\" WITH (FORCE)" \&lt;/span&gt;
            &lt;span class="s"&gt;-c "CREATE DATABASE \"$DB\""&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./scripts/migrate.sh&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.PREVIEW_BASE_URL }}/pr_${{ github.event.pull_request.number }}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;psql "$DATABASE_URL" -v ON_ERROR_STOP=1 -f seed/01_core.sql&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.PREVIEW_BASE_URL }}/pr_${{ github.event.pull_request.number }}&lt;/span&gt;

  &lt;span class="na"&gt;down&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;github.event.action == 'closed'&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Drop the PR database&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;ADMIN_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.PREVIEW_ADMIN_URL }}&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;psql "$ADMIN_URL" -v ON_ERROR_STOP=1 \&lt;/span&gt;
            &lt;span class="s"&gt;-c "DROP DATABASE IF EXISTS \"pr_${{ github.event.pull_request.number }}\" WITH (FORCE)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;WITH (FORCE)&lt;/code&gt; (Postgres 13+) is what makes the drop reliable — without it, one leftover connection leaves you with &lt;code&gt;database "pr_412" is being accessed by other users&lt;/code&gt; and a database that lives forever. Add a scheduled job that drops any &lt;code&gt;pr_*&lt;/code&gt; database whose PR is closed, because the &lt;code&gt;closed&lt;/code&gt; event does get missed.&lt;/p&gt;

&lt;p&gt;On the app side, Vercel's preview deployments are the least-effort way to get a URL per PR for a frontend, while Render's preview environments declared in &lt;code&gt;render.yaml&lt;/code&gt; cover the case where the PR needs a long-running backend process. Both still expect you to supply the database strategy above; neither isolates your data for you.&lt;/p&gt;

&lt;p&gt;The takeaway: every create path needs a matching delete path plus a sweeper, because webhook-driven teardown is not reliable enough to be the only one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about webhooks and OAuth callbacks that can't reach a preview URL?
&lt;/h2&gt;

&lt;p&gt;Third-party services cannot deliver to an unpredictable per-PR hostname, and OAuth providers reject unregistered redirect URIs. Best first: register one stable proxy URL that routes to the right preview by header or path segment; use the provider's CLI forwarding where it exists; or stub the integration and exercise the real thing only in staging. Do not register thirty redirect URIs — that list becomes permanent.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is paid branching worth it?
&lt;/h2&gt;

&lt;p&gt;The honest baseline is a small always-on Postgres instance plus the CI job above: a few dollars a month and an afternoon of setup. Branching starts paying when your data has enough real-world skew that seeds stop predicting production, when the seed job lands on the critical path of every review, or when reviewers are non-engineers who need plausible data to judge a UI.&lt;/p&gt;

&lt;p&gt;The takeaway: pay for branching to buy data realism, not to skip the seed file — you need the seed file either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can I just use a copy of production for preview environments?&lt;/strong&gt;&lt;br&gt;
Only after scrubbing, and only if the scrub runs inside the environment that is already allowed to hold production data. A preview URL is typically reachable by anyone with the link and protected by far weaker controls than production, so unscrubbed copies turn every open PR into an additional place a breach can start.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I give each pull request its own database on Postgres?&lt;/strong&gt;&lt;br&gt;
Create a database named after the PR number in a CI job triggered by &lt;code&gt;pull_request&lt;/code&gt;, run your migrations against it, apply an idempotent seed, and drop it on the &lt;code&gt;closed&lt;/code&gt; event with &lt;code&gt;DROP DATABASE ... WITH (FORCE)&lt;/code&gt;. Add a scheduled sweeper for the PRs whose close event got lost, otherwise abandoned databases accumulate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does my preview environment run out of database connections?&lt;/strong&gt;&lt;br&gt;
Each preview app holds its own connection pool, so idle previews consume connections even with no traffic. Either put a pooler in front of the instance and let previews connect through it in transaction mode, or cap preview pools to one or two connections each — previews serve one reviewer, not production traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;With a handful of open PRs, run an ephemeral database per PR with a committed idempotent seed plus a &lt;code&gt;generate_series&lt;/code&gt; block for volume: free, isolated, and it exercises your migrations on every push. If seeded data keeps failing to predict production, move to copy-on-write branching on a managed Postgres and apply production-level access controls to those branches. If you need production shapes without production risk, build the scrub pipeline once, run it inside the trusted environment, and test the scrub in CI. Whichever you pick, write the teardown job before the create job — the environments that cost real money are the ones nobody remembers to delete.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/postgres-says-too-many-clients-already-diagnose-it-before-you-add-a-pooler-5dbb"&gt;Postgres Says "Too Many Clients Already": Diagnose It Before You Add a Pooler&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/token-streaming-works-locally-but-arrives-all-at-once-in-production-finding-the-buffer-1o9g"&gt;Token Streaming Works Locally but Arrives All at Once in Production: Finding the Buffer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/postgres-as-a-job-queue-vs-redis-vs-sqs-when-does-just-use-your-database-stop-working-1f24"&gt;Postgres as a Job Queue vs Redis vs SQS: When Does "Just Use Your Database" Stop Working?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devops</category>
      <category>cicd</category>
      <category>postgres</category>
      <category>testing</category>
    </item>
    <item>
      <title>Why Page 500 Is Slow and Shows Duplicate Rows: Offset vs Keyset Pagination</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Sat, 26 Sep 2026 15:10:23 +0000</pubDate>
      <link>https://dev.to/libme/why-page-500-is-slow-and-shows-duplicate-rows-offset-vs-keyset-pagination-32jl</link>
      <guid>https://dev.to/libme/why-page-500-is-slow-and-shows-duplicate-rows-offset-vs-keyset-pagination-32jl</guid>
      <description>&lt;p&gt;&lt;code&gt;LIMIT 20 OFFSET 10000&lt;/code&gt; asks the database to produce 10,020 rows in sort order and throw 10,000 of them away, so the deeper a user scrolls the more work each request does. Worse, offsets describe a &lt;em&gt;position&lt;/em&gt; in a result set that keeps changing — if rows are inserted or deleted between two requests, the reader sees the same item twice or never sees it at all. Keyset pagination (also called cursor or seek pagination) filters on the last row you sent instead of counting rows, which makes every page cost the same and makes the sequence stable.&lt;/p&gt;

&lt;p&gt;This is the kind of bug that never shows up in development, because your seed data has 200 rows and nobody inserts anything while you click.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually happens when you say OFFSET 10000?
&lt;/h2&gt;

&lt;p&gt;There is no "skip ahead" primitive in a B-tree index scan. Postgres (and MySQL, and SQLite) implements OFFSET by fetching rows in order and discarding them until the count is satisfied. You can see it in the plan — the &lt;code&gt;Limit&lt;/code&gt; node reports the rows it emitted, but the node underneath it has already produced offset + limit rows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;EXPLAIN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;ANALYZE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;BUFFERS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;feed_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="k"&gt;OFFSET&lt;/span&gt; &lt;span class="mi"&gt;10000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at &lt;code&gt;actual rows&lt;/code&gt; on the child node and at the &lt;code&gt;Buffers: shared hit/read&lt;/code&gt; counts. Both scale with the offset, not with the page size. That is the whole problem: page 1 and page 500 are not the same query in disguise, they are a cheap query and an expensive one that happen to share syntax.&lt;/p&gt;

&lt;p&gt;If the sort column isn't indexed it's worse — the planner sorts the entire filtered set before it can discard anything, and deep pages can tip into a disk-based sort (&lt;code&gt;Sort Method: external merge&lt;/code&gt;). But even with a perfect index, offset cost grows linearly with depth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway: an offset is not a bookmark, it's an instruction to re-walk the list from the beginning every time.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug that isn't slowness: duplicated and skipped rows
&lt;/h2&gt;

&lt;p&gt;Slowness gets noticed. This one gets filed as "the API is flaky."&lt;/p&gt;

&lt;p&gt;A client reads page 1 (&lt;code&gt;OFFSET 0 LIMIT 20&lt;/code&gt;) of a feed sorted newest-first. Before it requests page 2, three new rows are inserted. Now the row that was at index 19 has shifted to index 22 — so &lt;code&gt;OFFSET 20&lt;/code&gt; returns rows the client already displayed. Deletions produce the mirror image: rows shift up, and items slide past the window unseen.&lt;/p&gt;

&lt;p&gt;For an infinite-scroll feed this shows up as duplicate cards in the list. For a background job that pages through a table to export or reprocess it, the same mechanism silently skips records, which is much more expensive to discover. I have chased "missing rows" in a nightly export that turned out to be nothing but &lt;code&gt;OFFSET&lt;/code&gt; racing against concurrent inserts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway: if anything writes to the table while a client is paging through it, offset pagination is not just slow — it is incorrect.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I write a keyset pagination query in Postgres?
&lt;/h2&gt;

&lt;p&gt;Keyset pagination replaces "skip N rows" with "give me the rows after this one." You need a sort key that is &lt;strong&gt;totally ordered&lt;/strong&gt; — that means adding a unique tiebreaker to whatever the user actually sorts by, because &lt;code&gt;created_at&lt;/code&gt; alone will have ties and rows will be dropped or repeated at page boundaries.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- page 1&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;feed_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- page N+1: pass the last row's (created_at, id) back in&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;feed_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The row-constructor comparison &lt;code&gt;(created_at, id) &amp;lt; ($1, $2)&lt;/code&gt; is the important part. It is not the same as &lt;code&gt;created_at &amp;lt; $1 AND id &amp;lt; $2&lt;/code&gt; (that drops rows), and unlike the hand-expanded &lt;code&gt;created_at &amp;lt; $1 OR (created_at = $1 AND id &amp;lt; $2)&lt;/code&gt;, Postgres can turn the row comparison into a single index seek on a composite index:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;posts_feed_created_id_idx&lt;/span&gt;
  &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;feed_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now every page is the same cost: seek to a position in the index, read 20 entries, stop. Page 500 costs what page 1 costs.&lt;/p&gt;

&lt;p&gt;Two constraints to check before you ship this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;All sort columns must point the same direction.&lt;/strong&gt; Row comparison assumes a single ordering. &lt;code&gt;ORDER BY created_at DESC, id ASC&lt;/code&gt; can't be expressed as &lt;code&gt;(created_at, id) &amp;lt; (...)&lt;/code&gt;; either make the directions agree or write the expanded OR form and accept a worse plan.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NULLs break it.&lt;/strong&gt; A nullable sort column with &lt;code&gt;NULLS LAST&lt;/code&gt; won't compare the way you expect. Sort on a &lt;code&gt;NOT NULL&lt;/code&gt; column, or on a &lt;code&gt;COALESCE(...)&lt;/code&gt; expression that you also index.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The takeaway: keyset pagination is a WHERE clause that mirrors your ORDER BY exactly — the moment they diverge, rows go missing at page boundaries.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How should the cursor be encoded in the API?
&lt;/h2&gt;

&lt;p&gt;Don't expose &lt;code&gt;?created_at=...&amp;amp;last_id=...&lt;/code&gt; as separate query parameters. Clients will start constructing them by hand, and then you can never change the sort key. Encode the position into one opaque token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// cursor.js&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;encode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
  &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;t&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toISOString&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="na"&gt;i&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;base64url&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;base64url&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isInteger&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;bad cursor&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;listPosts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;feedId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;limit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;cursor&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt;
    &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s2"&gt;`SELECT id, title, created_at FROM posts
         WHERE feed_id = $1 AND (created_at, id) &amp;lt; ($2, $3)
         ORDER BY created_at DESC, id DESC LIMIT $4`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;feedId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;limit&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])).&lt;/span&gt;&lt;span class="nx"&gt;rows&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s2"&gt;`SELECT id, title, created_at FROM posts
         WHERE feed_id = $1
         ORDER BY created_at DESC, id DESC LIMIT $2`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;feedId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;limit&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])).&lt;/span&gt;&lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hasMore&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;nextCursor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;hasMore&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;at&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details worth copying: fetching &lt;code&gt;limit + 1&lt;/code&gt; rows is how you answer "is there a next page" without a second &lt;code&gt;COUNT(*)&lt;/code&gt;, and validating the decoded shape matters because a cursor is user-controlled input that goes into a WHERE clause. Base64 is encoding, not protection — if the sort key is something you'd rather not leak, sign the cursor or store it server-side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway: an opaque cursor is a version boundary — it lets you change the sort key later without breaking every client that saved a URL.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When is OFFSET still the right call?
&lt;/h2&gt;

&lt;p&gt;Keyset pagination cannot jump to an arbitrary page, because page 47 has no meaning without reading pages 1 through 46. If your UI has numbered page buttons over an admin table of a few thousand rows that nobody is writing to, OFFSET is fine and simpler.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Offset/limit&lt;/th&gt;
&lt;th&gt;Keyset/cursor&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cost of deep pages&lt;/td&gt;
&lt;td&gt;Grows with depth&lt;/td&gt;
&lt;td&gt;Flat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stable under concurrent writes&lt;/td&gt;
&lt;td&gt;No — duplicates and skips&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jump to page N&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total page count&lt;/td&gt;
&lt;td&gt;Easy (&lt;code&gt;COUNT(*)&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Needs a separate estimate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arbitrary user-chosen sort&lt;/td&gt;
&lt;td&gt;Any column&lt;/td&gt;
&lt;td&gt;Needs an index per sort order&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Implementation cost&lt;/td&gt;
&lt;td&gt;Trivial&lt;/td&gt;
&lt;td&gt;Cursor encode/decode + composite index&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Total counts are the real trade. &lt;code&gt;COUNT(*)&lt;/code&gt; on a large filtered set is its own performance problem, so most feeds that switch to cursors also drop the exact total and show "load more" instead. If you need a number, an approximate count from planner statistics is usually enough for a UI hint — just never present an estimate as an exact figure.&lt;/p&gt;

&lt;p&gt;Framework support, as of mid-2026: Django REST Framework ships &lt;code&gt;CursorPagination&lt;/code&gt; out of the box, and it is the one that gets you a correct opaque cursor without writing SQL — at the price of requiring a stable ordering field and giving you no total count. In the Node/TypeScript world, Prisma's &lt;code&gt;cursor&lt;/code&gt; + &lt;code&gt;take&lt;/code&gt; arguments implement the same idea, and the awkward part is that the cursor must be a unique field, so a compound "sort by date, break ties by id" ordering still needs care. If your list is served from a search engine rather than a relational database, Elasticsearch's &lt;code&gt;search_after&lt;/code&gt; is the equivalent primitive, and it needs a tiebreaker field plus a point-in-time to stay consistent across pages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway: choose offset for bounded, browsable, mostly-static tables; choose keyset for anything that grows or that a machine pages through.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does OFFSET get slower on higher page numbers?&lt;/strong&gt;&lt;br&gt;
Because the database has to generate and discard every skipped row in sort order before it can return your page. &lt;code&gt;OFFSET 10000 LIMIT 20&lt;/code&gt; reads 10,020 rows internally. The cost grows linearly with the offset, so the last pages of a long list are the most expensive queries in your system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does keyset pagination require a unique column?&lt;/strong&gt;&lt;br&gt;
It requires a sort key that is unique &lt;em&gt;as a whole&lt;/em&gt;. Sorting by a non-unique column like &lt;code&gt;created_at&lt;/code&gt; is fine as long as you append a unique tiebreaker (&lt;code&gt;id&lt;/code&gt;) to both the ORDER BY and the cursor comparison. Without the tiebreaker, rows sharing a timestamp at a page boundary get duplicated or skipped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I still show a total page count with cursor pagination?&lt;/strong&gt;&lt;br&gt;
Not cheaply. You can run a separate &lt;code&gt;COUNT(*)&lt;/code&gt; with the same filters, but that query scans the matching rows and often costs more than the page itself. Most cursor-paginated APIs replace the page count with a "has more" boolean derived from fetching one extra row.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;If your endpoint backs an infinite scroll, a public feed, or any job that walks a table while other processes write to it, switch to keyset pagination — the correctness argument is stronger than the performance one, because offset-based paging genuinely loses rows under concurrent inserts. Keep offset for admin tables with page-number UI, bounded row counts, and low write traffic. The migration is usually one composite index, one row-constructor WHERE clause, and an opaque cursor in the response; the thing you give up is jumping to an arbitrary page, so confirm the UI can live with "load more" before you start.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/why-your-p99-looks-fine-while-users-complain-averaged-percentiles-and-histogram-buckets-eej"&gt;Why Your p99 Looks Fine While Users Complain: Averaged Percentiles and Histogram Buckets&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/getting-429-too-many-requests-fix-the-client-before-you-ask-for-a-higher-quota-7c1"&gt;Getting 429 Too Many Requests? Fix the Client Before You Ask for a Higher Quota&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/from-cron-jobs-to-event-driven-migrating-scheduled-tasks-to-serverless-functions-4i0a"&gt;From Cron Jobs to Event-Driven: Migrating Scheduled Tasks to Serverless Functions&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>api</category>
      <category>performance</category>
    </item>
    <item>
      <title>PgBouncer vs Supavisor vs RDS Proxy: Which Postgres Pooler Survives Transaction Mode?</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Fri, 25 Sep 2026 23:38:35 +0000</pubDate>
      <link>https://dev.to/libme/pgbouncer-vs-supavisor-vs-rds-proxy-which-postgres-pooler-survives-transaction-mode-1l52</link>
      <guid>https://dev.to/libme/pgbouncer-vs-supavisor-vs-rds-proxy-which-postgres-pooler-survives-transaction-mode-1l52</guid>
      <description>&lt;p&gt;Picking a Postgres connection pooler is really picking which session features you are willing to lose. All three of these poolers give you the same win — hundreds of app connections multiplexed onto a few dozen backend connections — and all three do it by taking your connection away between transactions, which is exactly what breaks prepared statements, session &lt;code&gt;SET&lt;/code&gt;s, &lt;code&gt;LISTEN/NOTIFY&lt;/code&gt; and cross-transaction advisory locks. PgBouncer gives you the most control and the most operational work, Supavisor is the reasonable pick if you're already on Supabase or want a clustered pooler you don't hand-tune, and RDS Proxy is the one to choose when IAM auth and failover handling matter more than pooling efficiency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does adding a pooler break my app?
&lt;/h2&gt;

&lt;p&gt;The failure that sends you looking for a pooler is familiar:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FATAL: sorry, too many clients already
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Postgres allocates a backend process per connection, so &lt;code&gt;max_connections&lt;/code&gt; is a real ceiling, and serverless functions or a scaled-out app blow through it. A pooler fixes that by keeping a small set of server connections and lending them out.&lt;/p&gt;

&lt;p&gt;The catch is &lt;em&gt;when&lt;/em&gt; it takes the connection back. In &lt;strong&gt;session pooling&lt;/strong&gt;, a client holds one server connection until it disconnects — safe, but it barely helps, because your idle app connections still pin backends. In &lt;strong&gt;transaction pooling&lt;/strong&gt;, the server connection returns to the pool at &lt;code&gt;COMMIT&lt;/code&gt;, so 500 clients can share 20 backends. That's the mode everyone actually wants, and it's the mode that changes the semantics of your database connection.&lt;/p&gt;

&lt;p&gt;Your driver notices first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;prepared&lt;/span&gt; &lt;span class="k"&gt;statement&lt;/span&gt; &lt;span class="nv"&gt;"S_1"&lt;/span&gt; &lt;span class="n"&gt;already&lt;/span&gt; &lt;span class="k"&gt;exists&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or, with asyncpg:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;prepared&lt;/span&gt; &lt;span class="k"&gt;statement&lt;/span&gt; &lt;span class="nv"&gt;"__asyncpg_stmt_3__"&lt;/span&gt; &lt;span class="n"&gt;does&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="n"&gt;exist&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two different clients landed on the same server connection, and the second one's prepared statement cache is now describing a different backend than it thinks. The dangerous version of this bug is the one that throws no error at all: &lt;code&gt;pg_advisory_lock()&lt;/code&gt; taken outside a transaction, released to the pool mid-flight, and silently held by a connection your next request never gets back. &lt;strong&gt;If your code assumes anything persists between transactions on the same connection, transaction pooling will break it — sometimes loudly, sometimes not.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What exactly breaks in transaction mode?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Transaction mode behavior&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Protocol-level named prepared statements&lt;/td&gt;
&lt;td&gt;Collide across clients unless the pooler tracks them&lt;/td&gt;
&lt;td&gt;PgBouncer &lt;code&gt;max_prepared_statements&lt;/code&gt; (1.21+), or disable in the driver&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;SET&lt;/code&gt; outside a transaction&lt;/td&gt;
&lt;td&gt;Leaks to the next client on that server connection&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;SET LOCAL&lt;/code&gt;, or set it in the connection string / server defaults&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;pg_advisory_lock()&lt;/code&gt; (session-scoped)&lt;/td&gt;
&lt;td&gt;Held by an unrelated connection; may never release&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;pg_advisory_xact_lock()&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;LISTEN&lt;/code&gt; / &lt;code&gt;NOTIFY&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Listener silently loses its connection&lt;/td&gt;
&lt;td&gt;Dedicated direct connection, not through the pooler&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Temp tables, &lt;code&gt;WITH HOLD&lt;/code&gt; cursors&lt;/td&gt;
&lt;td&gt;Disappear after &lt;code&gt;COMMIT&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Keep it inside one transaction, or use session mode&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long transactions&lt;/td&gt;
&lt;td&gt;Hold a backend for their whole duration&lt;/td&gt;
&lt;td&gt;Fix the transaction; the pooler can't help&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The driver-side escape hatch is usually one flag. psycopg3:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;psycopg&lt;/span&gt;

&lt;span class="c1"&gt;# Through a transaction-mode pooler: don't let psycopg promote
# frequently-used queries to named prepared statements.
&lt;/span&gt;&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;psycopg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;postgresql://app@pooler.internal:6432/app&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;prepare_threshold&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;asyncpg needs both the statement cache off and unique statement names disabled:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asyncpg&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;asyncpg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;postgresql://app@pooler.internal:6432/app&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;statement_cache_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You pay for this: unprepared queries re-plan on every execution. For short OLTP queries the cost is small; for a complex analytical query in a hot path it isn't. That trade is the reason PgBouncer's prepared-statement support matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before you compare poolers, audit your own code for session state — most "the pooler is broken" tickets are an app holding state the pooler never promised to keep.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  PgBouncer: most control, most of it yours to operate
&lt;/h2&gt;

&lt;p&gt;PgBouncer is the default answer because it's small, battle-tested and runs anywhere. Since 1.21 it tracks protocol-level named prepared statements in transaction mode, which removes the single most common driver breakage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[databases]&lt;/span&gt;
&lt;span class="py"&gt;app&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;host=db.internal port=5432 dbname=app&lt;/span&gt;

&lt;span class="nn"&gt;[pgbouncer]&lt;/span&gt;
&lt;span class="py"&gt;pool_mode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;transaction&lt;/span&gt;
&lt;span class="py"&gt;max_client_conn&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;2000&lt;/span&gt;
&lt;span class="py"&gt;default_pool_size&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;20&lt;/span&gt;
&lt;span class="py"&gt;max_prepared_statements&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;200&lt;/span&gt;
&lt;span class="py"&gt;server_reset_query&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;DISCARD ALL&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Watch the pool from the admin console rather than guessing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- psql "postgresql://pgbouncer@pooler.internal:6432/pgbouncer"&lt;/span&gt;
&lt;span class="k"&gt;SHOW&lt;/span&gt; &lt;span class="n"&gt;POOLS&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="c1"&gt;-- cl_waiting &amp;gt; 0 consistently means default_pool_size is too small&lt;/span&gt;
&lt;span class="c1"&gt;-- or transactions are running too long&lt;/span&gt;
&lt;span class="k"&gt;SHOW&lt;/span&gt; &lt;span class="n"&gt;STATS&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;cl_waiting&lt;/code&gt; is the number that tells you the truth. If it's persistently above zero, clients are queuing on the pooler and your p99 now includes pool wait time that your database metrics will never show you.&lt;/p&gt;

&lt;p&gt;The honest drawback: PgBouncer is essentially single-threaded, so one process saturates a core under heavy traffic and you scale it by running several processes with &lt;code&gt;so_reuseport&lt;/code&gt; or putting instances behind a load balancer — plus you now own a network hop that can page you at 3am. If you want that control and are willing to run it, PgBouncer is the pooler with the fewest surprises about what it's doing to your connections.&lt;/p&gt;

&lt;h2&gt;
  
  
  Supavisor: a clustered pooler you don't hand-tune
&lt;/h2&gt;

&lt;p&gt;Supavisor is Supabase's open-source pooler, written in Elixir, designed to be multi-tenant and to run as a cluster rather than as a single process per host. It exposes transaction mode and session mode on separate ports, so the usual pattern is app traffic on the transaction port and migrations or &lt;code&gt;LISTEN&lt;/code&gt;-style work on the session port. It has added named prepared statement support in transaction mode, but treat that as version-dependent and verify it with your actual driver before relying on it.&lt;/p&gt;

&lt;p&gt;If you're on Supabase, this is not really a choice — the platform's pooler endpoints are Supavisor, and the decision collapses to "which port". Outside Supabase, it's a credible self-hosted option when you want horizontal scaling without assembling it yourself from PgBouncer processes. Supavisor is the option worth considering when you want a pooler that scales across nodes instead of one you shard by hand.&lt;/p&gt;

&lt;p&gt;The drawback is ecosystem maturity: PgBouncer has a decade more field reports, and when something strange happens at 2am, the number of existing answers matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  RDS Proxy: what are you actually paying for?
&lt;/h2&gt;

&lt;p&gt;RDS Proxy is the managed option for RDS and Aurora, and pooling is arguably not its best feature. What you're buying is IAM authentication, credentials pulled from Secrets Manager instead of your app config, and connection handling across failovers — the proxy holds client connections while the underlying instance fails over, which shortens the error window your app sees.&lt;/p&gt;

&lt;p&gt;Its pooling has a specific gotcha: &lt;strong&gt;pinning&lt;/strong&gt;. When RDS Proxy sees session state it can't safely share, it stops multiplexing and dedicates a backend connection to that client for the rest of the session. Your pooler quietly becomes a passthrough. Check it directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CloudWatch → RDS → Proxy metrics
  DatabaseConnectionsCurrentlySessionPinned
  DatabaseConnectionsBorrowLatency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the pinned count tracks your client count, you're paying for a proxy that isn't pooling. As of mid-2026, AWS bills RDS Proxy per vCPU-hour of the underlying database instance rather than per connection or per request, with a floor for small instances — check the current pricing page, but note the shape: cost scales with your database size, not your traffic, so a small instance with bursty Lambda traffic is where it looks best and a large instance with modest connection counts is where it looks worst. RDS Proxy is the right call when IAM auth and failover behavior are requirements, not when raw pooling efficiency is the goal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which one should you run?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;PgBouncer&lt;/th&gt;
&lt;th&gt;Supavisor&lt;/th&gt;
&lt;th&gt;RDS Proxy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Operate it yourself&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes (or managed on Supabase)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Named prepared statements in transaction mode&lt;/td&gt;
&lt;td&gt;Yes, 1.21+&lt;/td&gt;
&lt;td&gt;Yes, version-dependent&lt;/td&gt;
&lt;td&gt;Pinning risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scaling model&lt;/td&gt;
&lt;td&gt;Multiple processes / instances&lt;/td&gt;
&lt;td&gt;Clustered&lt;/td&gt;
&lt;td&gt;Managed by AWS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Auth integration&lt;/td&gt;
&lt;td&gt;Postgres auth, auth_query&lt;/td&gt;
&lt;td&gt;Postgres auth&lt;/td&gt;
&lt;td&gt;IAM + Secrets Manager&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failover handling&lt;/td&gt;
&lt;td&gt;You handle it&lt;/td&gt;
&lt;td&gt;You handle it&lt;/td&gt;
&lt;td&gt;Built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost shape&lt;/td&gt;
&lt;td&gt;Instance you run&lt;/td&gt;
&lt;td&gt;Instance you run, or platform&lt;/td&gt;
&lt;td&gt;Per vCPU-hour of the DB instance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best when&lt;/td&gt;
&lt;td&gt;You want control and predictability&lt;/td&gt;
&lt;td&gt;You're on Supabase, or want a clustered pooler&lt;/td&gt;
&lt;td&gt;You're on RDS/Aurora and need IAM + failover&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What's the difference between session pooling and transaction pooling in PgBouncer?&lt;/strong&gt;&lt;br&gt;
Session pooling assigns a server connection to a client until it disconnects, preserving all session state but providing little multiplexing. Transaction pooling returns the server connection to the pool after each &lt;code&gt;COMMIT&lt;/code&gt;, which is what lets hundreds of clients share a few dozen backends, but it breaks anything that depends on session state — prepared statements, session &lt;code&gt;SET&lt;/code&gt;s, &lt;code&gt;LISTEN/NOTIFY&lt;/code&gt; and session-level advisory locks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why do I get "prepared statement already exists" through PgBouncer?&lt;/strong&gt;&lt;br&gt;
Your driver created a named prepared statement on one server connection, then landed on a different one for the next query. Fix it by enabling &lt;code&gt;max_prepared_statements&lt;/code&gt; in PgBouncer 1.21 or later so the pooler tracks statements per server connection, or by disabling named prepared statements in the driver (&lt;code&gt;prepare_threshold=None&lt;/code&gt; in psycopg3, &lt;code&gt;statement_cache_size=0&lt;/code&gt; in asyncpg).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does RDS Proxy replace PgBouncer?&lt;/strong&gt;&lt;br&gt;
Only partly. RDS Proxy pools connections and adds IAM authentication and failover handling, but it pins connections whenever it detects session state it can't share, which can eliminate the pooling benefit entirely. Watch the &lt;code&gt;DatabaseConnectionsCurrentlySessionPinned&lt;/code&gt; CloudWatch metric before assuming it's pooling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;If you run your own Postgres and want the most predictable behavior, run PgBouncer in transaction mode with &lt;code&gt;max_prepared_statements&lt;/code&gt; set, and watch &lt;code&gt;SHOW POOLS&lt;/code&gt; for &lt;code&gt;cl_waiting&lt;/code&gt;. If you're on Supabase, use Supavisor's transaction port for app traffic and the session port for migrations and listeners — that decision is already made for you. If you're on RDS or Aurora and your real requirements are IAM auth and clean failovers, RDS Proxy earns its cost, provided you check the pinned-connection metric and fix whatever session state is causing it. Whichever you pick, do the app-side audit first: &lt;code&gt;SET LOCAL&lt;/code&gt; instead of &lt;code&gt;SET&lt;/code&gt;, &lt;code&gt;pg_advisory_xact_lock()&lt;/code&gt; instead of &lt;code&gt;pg_advisory_lock()&lt;/code&gt;, and a direct connection for anything that listens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/how-to-test-search-relevance-before-you-ship-a-ranking-change-29o"&gt;How to Test Search Relevance Before You Ship a Ranking Change&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/hotstandbyfeedback-for-a-reporting-replica-query-cancellations-or-primary-bloat-mek"&gt;hot_standby_feedback for a Reporting Replica: Query Cancellations or Primary Bloat?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/should-you-migrate-off-pgvector-run-this-shadow-mode-benchmark-first-3o54"&gt;Should You Migrate Off pgvector? Run This Shadow-Mode Benchmark First&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>devops</category>
      <category>backend</category>
    </item>
    <item>
      <title>Uptime Monitor Says 100% While Users See Errors: What to Check, and When Paid Monitoring Is Worth It</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Thu, 24 Sep 2026 15:58:55 +0000</pubDate>
      <link>https://dev.to/libme/uptime-monitor-says-100-while-users-see-errors-what-to-check-and-when-paid-monitoring-is-worth-it-489p</link>
      <guid>https://dev.to/libme/uptime-monitor-says-100-while-users-see-errors-what-to-check-and-when-paid-monitoring-is-worth-it-489p</guid>
      <description>&lt;p&gt;An uptime monitor that reports 100% during an outage is almost always measuring the wrong thing: a CDN-cached page, a health endpoint that returns 200 whether or not the database answers, or a check that only looks at the status code. The fix is two health endpoints (shallow for the load balancer, deep for the external monitor), &lt;code&gt;Cache-Control: no-store&lt;/code&gt; on both, and an assertion on the response body. Paid monitoring is worth it once your customers notice outages before you do — what you are buying is a tighter check interval, multi-region confirmation, and a way to page a human.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why did the monitor say "up" when the site was down?
&lt;/h2&gt;

&lt;p&gt;The incident that made me rewrite my monitoring: a customer emailed a screenshot of a 500 page while the uptime dashboard showed a flat green line. The app had been throwing &lt;code&gt;ECONNREFUSED&lt;/code&gt; on every database call for about forty minutes.&lt;/p&gt;

&lt;p&gt;The monitor was hitting &lt;code&gt;https://example.com/&lt;/code&gt;, which sat behind a CDN with a long cache TTL. The edge kept serving cached HTML while the origin was on fire: status 200, fast response, green line. The monitor was accurately reporting that the CDN was up.&lt;/p&gt;

&lt;p&gt;Three patterns cause almost every "100% uptime during an outage" report:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The checked URL is cached.&lt;/strong&gt; The CDN answers instead of your origin. &lt;code&gt;curl -sI https://example.com/ | grep -iE 'age|cache|x-cache'&lt;/code&gt; will show a non-zero &lt;code&gt;Age&lt;/code&gt; or a &lt;code&gt;HIT&lt;/code&gt; header if this is happening to you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The health endpoint is shallow.&lt;/strong&gt; &lt;code&gt;/health&lt;/code&gt; returns &lt;code&gt;{"status":"ok"}&lt;/code&gt; as long as the process is alive. Process alive, database gone, still 200.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The check only asserts on the status code.&lt;/strong&gt; Some frameworks render an error page with a 200. Some maintenance pages do too. A status-code-only check cannot tell the difference between "the app works" and "the app produced HTML."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;An uptime check is only as honest as the endpoint it hits, and a cached 200 is the most common lie.&lt;/p&gt;

&lt;h2&gt;
  
  
  How should a health endpoint be built so a monitor can trust it?
&lt;/h2&gt;

&lt;p&gt;You need two endpoints with different jobs, because the load balancer and the external monitor are asking different questions.&lt;/p&gt;

&lt;p&gt;The load balancer asks "should I keep routing to this instance?" That check must be shallow. If it includes the database, one database blip makes every instance fail readiness at the same moment, the load balancer pulls all of them, and a thirty-second hiccup becomes a full outage.&lt;/p&gt;

&lt;p&gt;The external monitor asks "can a user actually get served?" That check should be deep: touch the database, touch the cache, and fail loudly with a 503 if any of them do not answer within a tight timeout. Here is the shape I use, in Express with &lt;code&gt;pg&lt;/code&gt; and &lt;code&gt;redis&lt;/code&gt; (ESM, so top-level &lt;code&gt;await&lt;/code&gt; works):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;express&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;express&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;pg&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pg&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;redis&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;express&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nx"&gt;pg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Pool&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;connectionString&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DATABASE_URL&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createClient&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;REDIS_URL&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;withTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;promise&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;timer&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;timeout&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;reject&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;timer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; timed out after &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;ms`&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
      &lt;span class="nx"&gt;ms&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;race&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nx"&gt;promise&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="k"&gt;finally&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;timer&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Shallow: "is this process alive?" — for the load balancer only.&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/health&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Cache-Control&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no-store&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ok&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Deep: "can this instance serve a real request?" — for the external monitor.&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/health/deep&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Cache-Control&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no-store&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;checks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;postgres&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;withTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SELECT 1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;postgres&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;withTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ping&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;redis&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{};&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;healthy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;check&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;entries&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;checks&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;check&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ok&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="nx"&gt;healthy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;healthy&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;healthy&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ok&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;degraded&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;checks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;results&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three details matter more than the rest. The timeouts are short and explicit, because a health check that hangs for thirty seconds on a dead connection pool is worse than one that fails fast. &lt;code&gt;Cache-Control: no-store&lt;/code&gt; goes on both endpoints, and you then verify the CDN honors it — some edge rules cache regardless of origin headers, in which case you route &lt;code&gt;/health/*&lt;/code&gt; around the cache entirely. And the deep check covers only dependencies you own. Put a third-party API in it and their outage pages you at 3 a.m. for something you cannot fix.&lt;/p&gt;

&lt;p&gt;Then configure the monitor to hit &lt;code&gt;/health/deep&lt;/code&gt; and assert on the body, not just the code. Every serious uptime tool supports a keyword or JSON match; use it to require &lt;code&gt;"status":"ok"&lt;/code&gt; in the response. That single assertion closes the gap where an error page with a 200 slips through.&lt;/p&gt;

&lt;p&gt;The load balancer gets the shallow check and the monitor gets the deep one; giving both the same endpoint is how a database blip turns into a total outage.&lt;/p&gt;

&lt;h2&gt;
  
  
  What are you actually paying for with an uptime monitoring subscription?
&lt;/h2&gt;

&lt;p&gt;Free tiers are generous enough that the paid decision comes down to a handful of capabilities. As of September 2026, pricing models cluster around monitor counts and check intervals, but the numbers move often enough that you should read the current plan page rather than trust any post.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you pay for&lt;/th&gt;
&lt;th&gt;Free tier reality&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Check interval&lt;/td&gt;
&lt;td&gt;Every few minutes&lt;/td&gt;
&lt;td&gt;Interval is your worst-case detection delay; a 3-minute outage can fit inside a 5-minute gap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-region confirmation&lt;/td&gt;
&lt;td&gt;Usually one location&lt;/td&gt;
&lt;td&gt;One vantage point produces false alarms and misses regional DNS or routing failures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Body / JSON assertions&lt;/td&gt;
&lt;td&gt;Often limited&lt;/td&gt;
&lt;td&gt;Without them you are back to status-code-only checks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser (synthetic) checks&lt;/td&gt;
&lt;td&gt;Rare&lt;/td&gt;
&lt;td&gt;The only way to verify "login and checkout still work," not just "the URL answers"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;On-call routing&lt;/td&gt;
&lt;td&gt;Email only&lt;/td&gt;
&lt;td&gt;Phone or SMS escalation is what wakes someone up at night&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Status page&lt;/td&gt;
&lt;td&gt;Basic&lt;/td&gt;
&lt;td&gt;Cuts support email during an incident&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The hidden costs are the ones that scale. Per-monitor pricing feels cheap until every microservice and cron job has its own check. SMS and voice alerts are frequently metered separately. Browser checks are typically priced per run, so a Playwright flow every minute from three regions costs meaningfully more than a ping.&lt;/p&gt;

&lt;p&gt;The four tools I keep coming back to each own a distinct slice of this problem.&lt;/p&gt;

&lt;p&gt;If you just need a free ping every few minutes with email alerts, UptimeRobot is the one most people start with, and for a single side project it is often enough. Its limit is that the free interval is coarse, and assertions and status pages get thin once you need more than the basics.&lt;/p&gt;

&lt;p&gt;If you want the managed version of the whole loop, Better Stack is the one that bundles uptime checks with on-call scheduling and a status page so a failed check pages a human without a second tool. The drawback is that you are paying for incident management you may not need yet, and seats plus phone alerting add up for a small team.&lt;/p&gt;

&lt;p&gt;If your real question is whether a user can still log in and check out, Checkly is the one that runs Playwright scripts on a schedule from multiple regions and keeps the check definitions as code in your repo. The trade-off is maintenance: browser checks break when your UI changes, and per-run pricing punishes aggressive schedules.&lt;/p&gt;

&lt;p&gt;For cron jobs and queue workers that have no URL to hit, Healthchecks.io is the one built around the dead-man's switch: the job pings it on success, and you get alerted when the ping stops arriving. It does not do HTTP monitoring at all, so it complements rather than replaces the others.&lt;/p&gt;

&lt;p&gt;Pay for monitoring when the cost of a missed outage exceeds the subscription, and for most small products that line is crossed the first time a customer reports the outage before your monitor does.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is the free tier genuinely enough?
&lt;/h2&gt;

&lt;p&gt;A free ping monitor is fine when you have one or two public endpoints, a five-minute detection delay is acceptable, email is a reasonable alert channel because nobody is on call anyway, and a deep health endpoint like the one above is what it hits. Solo side projects and internal tools usually qualify.&lt;/p&gt;

&lt;p&gt;Move to a paid tier at the first of these signals: paying customers who would churn over an unnoticed outage, a second person who needs to be paged, a purchase flow that can break while the homepage stays up, or an SLA you signed. Fix the endpoint first either way; a paid monitor pointed at a cached homepage costs more and still lies.&lt;/p&gt;

&lt;p&gt;The free tier stops being free the day a customer becomes your monitoring system.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does my uptime monitor show 100% uptime when my site is down?&lt;/strong&gt;&lt;br&gt;
Because it is checking something other than your application: a CDN-cached page, a shallow health endpoint that returns 200 while the database is unreachable, or a check that only looks at the status code. Point it at a deep health endpoint with &lt;code&gt;Cache-Control: no-store&lt;/code&gt; and assert on the response body.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should a health check endpoint check the database?&lt;/strong&gt;&lt;br&gt;
Yes for the endpoint your external monitor hits, no for the one your load balancer uses. A deep check with short timeouts tells you users are being served; a shallow liveness check keeps a brief database blip from making the load balancer drop every instance at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How often should an uptime monitor check my site?&lt;/strong&gt;&lt;br&gt;
The check interval is your maximum detection delay. Every five minutes is acceptable for hobby projects; every minute with confirmation from a second region is the baseline once customers depend on the service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Fix the endpoint before you buy anything: a deep &lt;code&gt;/health/deep&lt;/code&gt; with dependency timeouts, &lt;code&gt;no-store&lt;/code&gt; headers, and a body assertion in the monitor closes the "100% uptime during an outage" hole on any tool's free tier. Solo projects and internal tools can stay on UptimeRobot's free ping and Healthchecks.io for cron jobs. Teams with paying customers should pay for one-minute multi-region checks and phone escalation, which is where Better Stack fits. If a broken login flow would go unnoticed while the homepage stays green, add Checkly's browser checks for the one or two flows that actually make money.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/the-hidden-costs-of-serverless-what-your-first-big-bill-teaches-you-5d3d"&gt;The Hidden Costs of Serverless: What Your First Big Bill Teaches You&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/cutting-your-side-projects-cloud-bill-a-checklist-that-doesnt-sacrifice-uptime-17kf"&gt;Cutting Your Side Project's Cloud Bill: A Checklist That Doesn't Sacrifice Uptime&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/why-your-p99-looks-fine-while-users-complain-averaged-percentiles-and-histogram-buckets-eej"&gt;Why Your p99 Looks Fine While Users Complain: Averaged Percentiles and Histogram Buckets&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devops</category>
      <category>backend</category>
      <category>saas</category>
      <category>cloud</category>
    </item>
    <item>
      <title>WebSocket Closes Every 60 Seconds With Code 1006: Finding the Proxy Idle Timeout and Fixing It With Heartbeats</title>
      <dc:creator>Libme</dc:creator>
      <pubDate>Wed, 23 Sep 2026 02:51:43 +0000</pubDate>
      <link>https://dev.to/libme/websocket-closes-every-60-seconds-with-code-1006-finding-the-proxy-idle-timeout-and-fixing-it-with-278e</link>
      <guid>https://dev.to/libme/websocket-closes-every-60-seconds-with-code-1006-finding-the-proxy-idle-timeout-and-fixing-it-with-278e</guid>
      <description>&lt;p&gt;If your WebSocket or Server-Sent Events connection dies at a suspiciously round interval (60 seconds is the classic, 55 or 30 also show up) and the browser reports close code &lt;code&gt;1006&lt;/code&gt; with no reason, something between your client and your server has an idle timeout and is cutting the TCP connection when no bytes flow. The fix is not a bigger timeout; it's a server-initiated heartbeat that keeps bytes moving at a shorter interval than the strictest hop, plus a client that reconnects without panicking. This post walks through how to prove which hop is killing you and what the fix looks like on each.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does close code 1006 actually mean?
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;1006&lt;/code&gt; is defined by the WebSocket spec as "abnormal closure": the connection ended without either side sending a close frame. Your server did not say goodbye, and neither did the browser. That is precisely the signature of a middlebox dropping the TCP socket: an L7 proxy tracking idle time decides the connection is dead, closes both sides, and neither endpoint gets a WebSocket-level close.&lt;/p&gt;

&lt;p&gt;In the browser console it looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WebSocket connection to 'wss://app.example.com/ws' failed:
WebSocket is closed before the connection is established.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or, more commonly on an already-open socket, an &lt;code&gt;onclose&lt;/code&gt; event with &lt;code&gt;event.code === 1006&lt;/code&gt; and &lt;code&gt;event.wasClean === false&lt;/code&gt;. On the Node side you often see nothing at all, or a plain &lt;code&gt;ECONNRESET&lt;/code&gt; on the underlying socket. Local development works perfectly because there is no proxy between &lt;code&gt;localhost:3000&lt;/code&gt; and your browser.&lt;/p&gt;

&lt;p&gt;The takeaway: a &lt;code&gt;1006&lt;/code&gt; on a regular interval is a timer in the network path, not a bug in your message handling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which hop has the timer?
&lt;/h2&gt;

&lt;p&gt;Every layer between the browser and your process can have its own idle clock, and only application bytes reset it. TCP keepalive probes do not help here: they are exchanged between adjacent TCP peers only, and an L7 proxy terminates TCP on both sides, so a keepalive from your server never reaches the load balancer's "was there data?" counter.&lt;/p&gt;

&lt;p&gt;The usual suspects, as of mid-2026 (defaults change, so verify against your provider's current docs):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hop&lt;/th&gt;
&lt;th&gt;Default idle behavior&lt;/th&gt;
&lt;th&gt;Where to change it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;nginx (&lt;code&gt;proxy_pass&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Closes if the upstream sends nothing for &lt;code&gt;proxy_read_timeout&lt;/code&gt; (60s)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;proxy_read_timeout&lt;/code&gt; / &lt;code&gt;proxy_send_timeout&lt;/code&gt; in the &lt;code&gt;location&lt;/code&gt; block&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ingress-nginx (Kubernetes)&lt;/td&gt;
&lt;td&gt;Same 60s defaults, inherited from nginx&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;nginx.ingress.kubernetes.io/proxy-read-timeout&lt;/code&gt; annotation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS Application Load Balancer&lt;/td&gt;
&lt;td&gt;Idle timeout, 60s default, applies to WebSockets&lt;/td&gt;
&lt;td&gt;Load balancer attribute &lt;code&gt;idle_timeout.timeout_seconds&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Cloud external HTTP(S) LB&lt;/td&gt;
&lt;td&gt;Backend service timeout (30s default) is treated as the &lt;em&gt;maximum lifetime&lt;/em&gt; of a WebSocket, idle or not&lt;/td&gt;
&lt;td&gt;Backend service &lt;code&gt;timeoutSec&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heroku router&lt;/td&gt;
&lt;td&gt;Terminates after 55s with no data in either direction&lt;/td&gt;
&lt;td&gt;Not configurable; heartbeat is the only option&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloudflare (proxied)&lt;/td&gt;
&lt;td&gt;WebSockets supported; timeouts depend on plan and whether the origin responds&lt;/td&gt;
&lt;td&gt;Check the current docs for your plan; heartbeat regardless&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Google Cloud row is the one that surprises people. Raising it to 3600 does not make the problem go away; it just makes your sockets die once an hour instead of every 30 seconds, which is a harder bug to notice and reproduce. Your client must handle reconnection no matter what, and the heartbeat only prevents the &lt;em&gt;idle&lt;/em&gt; kills.&lt;/p&gt;

&lt;p&gt;To identify the hop quickly: open a socket, send nothing, and time the close. 60s points at nginx or ALB. 55s is Heroku. 30s that ignores your heartbeat entirely is the GCP max-lifetime behavior. If it dies at a different interval after you add a heartbeat, you have found a second timer stacked behind the first.&lt;/p&gt;

&lt;p&gt;The takeaway: measure the time-to-death with a silent connection before touching any config, because the interval tells you which layer to look at.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you keep the connection alive without relying on the client?
&lt;/h2&gt;

&lt;p&gt;Send the heartbeat from the server. Client-side timers are unreliable: Chrome throttles &lt;code&gt;setInterval&lt;/code&gt; in background tabs (intensive throttling aligns timers to once per minute after a tab has been hidden for a while, as of current stable), which is exactly long enough to trip a 60-second proxy timeout. The server has no such constraint and it's the only party that also needs to detect dead clients for cleanup.&lt;/p&gt;

&lt;p&gt;With the &lt;code&gt;ws&lt;/code&gt; library in Node, the pattern from its own README is a ping every N seconds and a &lt;code&gt;terminate()&lt;/code&gt; for any client that did not pong since the last round:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;WebSocketServer&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ws&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;wss&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;WebSocketServer&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3000&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;wss&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;connection&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;isAlive&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pong&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;isAlive&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;HEARTBEAT_MS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;25&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// must be &amp;lt; the strictest proxy idle timeout&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;interval&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ws&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;wss&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;clients&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;isAlive&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;terminate&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// no pong since last tick: assume the peer is gone&lt;/span&gt;
      &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;isAlive&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ping&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;HEARTBEAT_MS&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nx"&gt;wss&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;close&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;interval&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details matter. First, the ping frame counts as application data to every L7 proxy in the path, so it resets each idle clock. Second, &lt;code&gt;terminate()&lt;/code&gt; rather than &lt;code&gt;close()&lt;/code&gt; on the unresponsive branch: a peer that cannot pong will not complete a close handshake either, and &lt;code&gt;close()&lt;/code&gt; would leave a zombie for another full interval.&lt;/p&gt;

&lt;p&gt;Pick the interval as roughly half of the shortest timeout you found in the table. 25 seconds is a safe default under a 55–60 second limit; it's also what Socket.IO ships as its &lt;code&gt;pingInterval&lt;/code&gt; default, which is why Socket.IO users rarely hit this class of bug until they move to a stricter proxy. If you'd rather not own any of this, Socket.IO is the library that handles heartbeats, reconnection with backoff, and transport fallback out of the box; the cost is a custom protocol on top of WebSockets, so non-JS clients need a Socket.IO client library rather than a plain WebSocket.&lt;/p&gt;

&lt;p&gt;For Server-Sent Events the same principle applies, and it's even simpler because SSE has a comment syntax that clients ignore:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Express handler for an SSE stream&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/events&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;text/event-stream&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Cache-Control&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;no-cache&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Connection&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;keep-alive&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;retry: 3000&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// tell EventSource how long to wait before reconnecting&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;keepalive&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;: keepalive&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;25&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;close&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;keepalive&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A line starting with &lt;code&gt;:&lt;/code&gt; is a comment per the SSE spec; &lt;code&gt;EventSource&lt;/code&gt; discards it but the proxy sees bytes. Note the &lt;code&gt;retry:&lt;/code&gt; field: &lt;code&gt;EventSource&lt;/code&gt; reconnects on its own, and this lets you set the delay instead of relying on the browser default.&lt;/p&gt;

&lt;p&gt;The takeaway: a server-sent ping every 25 seconds fixes the idle-timeout class of disconnect on every proxy in the table except the one that enforces a maximum lifetime.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you raise the proxy timeout instead?
&lt;/h2&gt;

&lt;p&gt;Raise it &lt;em&gt;as well&lt;/em&gt;, not instead. Heartbeats solve the idle problem, but a 60-second &lt;code&gt;proxy_read_timeout&lt;/code&gt; is also what kills legitimately long silent responses like a slow LLM completion or a large export, so it's worth loosening for the specific WebSocket location:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/ws&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://app_upstream&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_http_version&lt;/span&gt; &lt;span class="mf"&gt;1.1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Upgrade&lt;/span&gt; &lt;span class="nv"&gt;$http_upgrade&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Connection&lt;/span&gt; &lt;span class="s"&gt;"upgrade"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_read_timeout&lt;/span&gt; &lt;span class="s"&gt;3600s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_send_timeout&lt;/span&gt; &lt;span class="s"&gt;3600s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the large value scoped to the upgrade path. Applying it globally means a stuck upstream on a normal HTTP request can hold a worker connection for an hour.&lt;/p&gt;

&lt;p&gt;The one-line rule: heartbeats are mandatory and portable; timeout increases are a per-environment optimization that you'll forget to carry over to the next hosting provider.&lt;/p&gt;

&lt;h2&gt;
  
  
  How should the client reconnect?
&lt;/h2&gt;

&lt;p&gt;The browser's &lt;code&gt;WebSocket&lt;/code&gt; does not reconnect on its own (&lt;code&gt;EventSource&lt;/code&gt; does). A minimal client needs exponential backoff with jitter so that a proxy restart doesn't turn ten thousand clients into a synchronized reconnect storm:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;onMessage&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;ws&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onopen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onmessage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;onMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onclose&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;base&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="nx"&gt;attempt&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;base&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;base&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// jitter&lt;/span&gt;
      &lt;span class="nx"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;open&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onerror&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// ensure onclose fires and schedules a retry&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;ws&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;client shutdown&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What this does not solve is missed messages during the gap. If a client must not lose events, you need a resumable stream: either a &lt;code&gt;Last-Event-ID&lt;/code&gt; (native to SSE) or a sequence number you send on reconnect so the server can replay from its buffer. That is real work, and it's the point at which a managed service earns its fee. If you want the managed version of this, Ably is the one that handles connection state recovery and message replay across reconnects without you building the buffer; the trade-off is per-message and per-connection pricing that makes chatty, high-fanout workloads expensive compared to a socket you own, so run the math on your peak message rate before committing.&lt;/p&gt;

&lt;p&gt;The takeaway: reconnection with jitter is table stakes, and replay after reconnect is where you decide between building a buffer and buying one.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does my WebSocket close after 60 seconds behind nginx?&lt;/strong&gt;&lt;br&gt;
Because nginx's &lt;code&gt;proxy_read_timeout&lt;/code&gt; defaults to 60 seconds and closes the upstream connection when no data has been transmitted in that window. Send a ping from the server at least every 30 seconds, and raise &lt;code&gt;proxy_read_timeout&lt;/code&gt; on the WebSocket location if you also have long silent responses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does TCP keepalive prevent WebSocket idle timeouts on a load balancer?&lt;/strong&gt;&lt;br&gt;
No. TCP keepalive probes only travel between adjacent TCP peers, and an L7 load balancer terminates TCP on each side, so its idle timer only resets on application data. Use WebSocket ping frames or SSE comment lines instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between WebSocket close code 1006 and 1001?&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;1001&lt;/code&gt; means one side sent a close frame because it is "going away" (page navigation, server shutdown). &lt;code&gt;1006&lt;/code&gt; means the connection dropped with no close frame at all, which almost always indicates a proxy, load balancer, or network cut rather than either application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;If you own a WebSocket or SSE endpoint and it runs behind anything other than your own process, ship a server-initiated heartbeat at about 25 seconds and a client that reconnects with jittered backoff; do that before adjusting a single timeout. Then raise the proxy's read timeout on the upgrade path so slow-but-legitimate silences survive, and check whether your load balancer enforces a maximum connection lifetime (Google Cloud does by default) that no heartbeat can defeat. Teams that also need guaranteed delivery across reconnects should either budget the time for a replay buffer or pay a managed realtime provider for one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/why-your-oauth-integration-randomly-returns-invalidgrant-and-how-to-stop-two-workers-from-racing-4ake"&gt;Why Your OAuth Integration Randomly Returns invalid_grant (and How to Stop Two Workers From Racing)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/token-streaming-works-locally-but-arrives-all-at-once-in-production-finding-the-buffer-1o9g"&gt;Token Streaming Works Locally but Arrives All at Once in Production: Finding the Buffer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/libme/the-hidden-costs-of-serverless-what-your-first-big-bill-teaches-you-5d3d"&gt;The Hidden Costs of Serverless: What Your First Big Bill Teaches You&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>webdev</category>
      <category>node</category>
      <category>backend</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
