<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mukul S</title>
    <description>The latest articles on DEV Community by Mukul S (@mukul_sharma_61fc4dd6f9d8).</description>
    <link>https://dev.to/mukul_sharma_61fc4dd6f9d8</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3518766%2F548d981e-ab90-4068-b53c-831456d6aeec.jpg</url>
      <title>DEV Community: Mukul S</title>
      <link>https://dev.to/mukul_sharma_61fc4dd6f9d8</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mukul_sharma_61fc4dd6f9d8"/>
    <language>en</language>
    <item>
      <title>Stop Paying for the Same Tokens Twice: A Practical Guide to Prompt Caching</title>
      <dc:creator>Mukul S</dc:creator>
      <pubDate>Wed, 05 Aug 2026 19:03:27 +0000</pubDate>
      <link>https://dev.to/mukul_sharma_61fc4dd6f9d8/stop-paying-for-the-same-tokens-twice-a-practical-guide-to-prompt-caching-4938</link>
      <guid>https://dev.to/mukul_sharma_61fc4dd6f9d8/stop-paying-for-the-same-tokens-twice-a-practical-guide-to-prompt-caching-4938</guid>
      <description>&lt;p&gt;You've built a chatbot. Every turn, you re-send the whole conversation — the 8,000-token system prompt, the uploaded PDF, the 15 messages of history — just so the model can answer "and what about Mars?"&lt;br&gt;
The model re-reads all of it. Every. Single. Time. You pay full price for all of it. Every. Single. Time.&lt;br&gt;
Prompt caching fixes this. It's roughly one extra line of JSON, and it can cut your input costs by ~90% on the repeated part while making responses noticeably faster.&lt;/p&gt;
&lt;h2&gt;
  
  
  Let's walk through it.
&lt;/h2&gt;
&lt;h2&gt;
  
  
  The one-liner version
&lt;/h2&gt;

&lt;p&gt;Add &lt;code&gt;cache_control&lt;/code&gt; at the top level of your request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cache_control&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;   &lt;span class="c1"&gt;# &amp;lt;-- this is the whole trick
&lt;/span&gt;    &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are an AI assistant tasked with analyzing literary works...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Analyze the major themes in &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Pride and Prejudice&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's &lt;strong&gt;automatic caching&lt;/strong&gt;. The API caches everything up to and including the last cacheable block in your request. Next time you send a request that starts with the same content, that prefix is read from cache instead of reprocessed.&lt;br&gt;
Same thing in curl, if that's more your speed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.anthropic.com/v1/messages &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"content-type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"anthropic-version: 2023-06-01"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "claude-opus-5",
    "max_tokens": 1024,
    "cache_control": {"type": "ephemeral"},
    "system": "You are an AI assistant tasked with analyzing literary works.",
    "messages": [{"role": "user", "content": "Analyze the major themes in Pride and Prejudice."}]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  The mental model: it's a &lt;em&gt;prefix&lt;/em&gt; cache
&lt;/h2&gt;

&lt;p&gt;This is the single most important thing to internalize.&lt;br&gt;
Your prompt is read in a fixed order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tools  →  system  →  messages
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Caching works on &lt;strong&gt;prefixes&lt;/strong&gt; of that sequence. When you mark a block with &lt;code&gt;cache_control&lt;/code&gt;, you're saying: &lt;em&gt;"cache everything from the start of the request up to and including this block."&lt;/em&gt;&lt;br&gt;
Which means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;✅ Static stuff at the &lt;strong&gt;front&lt;/strong&gt; = cacheable.&lt;/li&gt;
&lt;li&gt;❌ Change something early = everything after it is invalidated.
So the golden rule is: &lt;strong&gt;put the boring, unchanging stuff first&lt;/strong&gt;. Tool definitions, system instructions, that 50-page contract, your 20 few-shot examples. Then the volatile stuff — the user's actual question — goes last.
---&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  What it costs (and saves)
&lt;/h2&gt;

&lt;p&gt;Three price tiers instead of one:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Multiplier vs. base input&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;5-minute cache &lt;strong&gt;write&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;1.25×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1-hour cache &lt;strong&gt;write&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;2×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache &lt;strong&gt;read&lt;/strong&gt; (hit)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.1×&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So for Claude Opus 5 ($5/MTok input):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;First request writes the cache: $6.25/MTok&lt;/li&gt;
&lt;li&gt;Every subsequent hit: &lt;strong&gt;$0.50/MTok&lt;/strong&gt;
You pay a 25% premium once, then 90% off forever after. If you reuse a prefix even twice, you're already ahead.
Cache breakpoints themselves are &lt;strong&gt;free&lt;/strong&gt;. You're never charged for &lt;em&gt;having&lt;/em&gt; a breakpoint — only for tokens actually written and read.&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  Reading the usage fields (this trips everyone up)
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"usage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_read_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_creation_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;input_tokens&lt;/code&gt; is &lt;strong&gt;not&lt;/strong&gt; your total input. It's only the tokens &lt;em&gt;after&lt;/em&gt; your last cache breakpoint. Total is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cache_read_input_tokens&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;cache_creation_input_tokens&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;input_tokens&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the example above processed 100,050 tokens, not 50. Handy side effect: cache hits don't count against your rate limits the way fresh input does, so effective throughput goes up too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick sanity check:&lt;/strong&gt; if both &lt;code&gt;cache_creation_input_tokens&lt;/code&gt; and &lt;code&gt;cache_read_input_tokens&lt;/code&gt; are &lt;code&gt;0&lt;/code&gt;, nothing was cached. Most likely you're under the minimum length (see below) — the API won't error, it just silently skips caching.&lt;/p&gt;




&lt;h2&gt;
  
  
  Minimum sizes — don't skip this
&lt;/h2&gt;

&lt;p&gt;Prompts shorter than a per-model floor simply won't cache:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Minimum cacheable tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5, Fable 5, Mythos 5&lt;/td&gt;
&lt;td&gt;512&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.8, Sonnet 5 / 4.6 / 4.5&lt;/td&gt;
&lt;td&gt;1,024&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.7, Haiku 3.5&lt;/td&gt;
&lt;td&gt;2,048&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.6, Opus 4.5, Haiku 4.5&lt;/td&gt;
&lt;td&gt;4,096&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you're &lt;em&gt;just&lt;/em&gt; under the line, it's often worth padding the cached section (more examples, more context) to get over it. Cache reads are cheap enough that the extra tokens pay for themselves.&lt;/p&gt;




&lt;h2&gt;
  
  
  Multi-turn conversations: let it drive
&lt;/h2&gt;

&lt;p&gt;With automatic caching, the breakpoint walks forward on its own as the conversation grows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Request&lt;/th&gt;
&lt;th&gt;What happens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;System + U1 + A1 + **U2**&lt;/code&gt; → everything written to cache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;... + A2 + **U3**&lt;/code&gt; → System→U2 read from cache, A2+U3 written&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;... + A3 + **U4**&lt;/code&gt; → System→U3 read from cache, A3+U4 written&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;No bookkeeping. No moving markers around. Each turn reads the whole prior conversation from cache and only writes the new bit. This is the single best default for chat apps.&lt;/p&gt;




&lt;h2&gt;
  
  
  Explicit breakpoints: when you need the wheel
&lt;/h2&gt;

&lt;p&gt;Put &lt;code&gt;cache_control&lt;/code&gt; on individual blocks when different parts of your prompt change at different rates. You get up to &lt;strong&gt;4 breakpoints&lt;/strong&gt;.&lt;br&gt;
Classic RAG-agent layout:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/*&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;*/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"get_document"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cache_control"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ephemeral"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"system"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"You are a research assistant..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cache_control"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ephemeral"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"# Knowledge Base&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;## Doc 1..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cache_control"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ephemeral"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Tell me about Perseverance."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cache_control"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ephemeral"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four independent segments:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Tools&lt;/strong&gt; — basically never change&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instructions&lt;/strong&gt; — change on deploys&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAG documents&lt;/strong&gt; — change daily&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversation&lt;/strong&gt; — changes every turn
Swap the RAG docs and you keep segments 1 and 2. Add a turn and you keep 1, 2, and 3. Changes only invalidate their own segment and everything downstream.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  The mistake literally everyone makes
&lt;/h2&gt;

&lt;p&gt;Here's the bug I want you to remember, because it's expensive and silent.&lt;br&gt;
Your prompt: blocks 1–5 are a big static system context. Block 6 is &lt;code&gt;f"[{timestamp}] {user_message}"&lt;/code&gt;. You put &lt;code&gt;cache_control&lt;/code&gt; on block 6, because it's the end and that seems right.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Request 1:&lt;/strong&gt; cache written at block 6. The hash includes the timestamp.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Request 2:&lt;/strong&gt; different timestamp → different hash → miss. The system walks back through blocks 5, 4, 3, 2, 1 looking for entries... but &lt;em&gt;no request ever wrote an entry there&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Result:&lt;/strong&gt; a fresh cache write every single request. You pay the 1.25× premium forever and never get a single read.
The lookback &lt;strong&gt;does not&lt;/strong&gt; find stable content behind your breakpoint and cache it for you. It only finds entries that earlier requests wrote — and &lt;strong&gt;writes happen only at breakpoints.&lt;/strong&gt;
The fix is one line: move &lt;code&gt;cache_control&lt;/code&gt; to &lt;strong&gt;block 5&lt;/strong&gt;, the last block that's identical across requests.
&amp;gt; &lt;strong&gt;Rule of thumb:&lt;/strong&gt; put the breakpoint on the last block whose prefix is &lt;em&gt;identical&lt;/em&gt; across the requests you want to share a cache.
(Note: automatic caching falls into the same trap here, since it targets the last block. If your final block has a per-request timestamp, use an explicit breakpoint on the static prefix instead.)&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  The 20-block lookback window
&lt;/h2&gt;

&lt;p&gt;A related gotcha. When looking for a cache hit, the system checks your breakpoint's position and then walks backward — but only &lt;strong&gt;20 blocks&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Turn 1:&lt;/strong&gt; 10 blocks, breakpoint at 10. Entry written at 10.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn 2:&lt;/strong&gt; 15 blocks, breakpoint at 15. Walks back to 10, finds turn 1's entry. Hit! Only blocks 11–15 processed fresh.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn 3:&lt;/strong&gt; 35 blocks, breakpoint at 35. Checks blocks 35 down to 16, finds nothing. The turn-2 entry at block 15 is &lt;em&gt;one position outside the window&lt;/em&gt;. &lt;strong&gt;Miss.&lt;/strong&gt; Full reprocess.
If your conversation can jump by 20+ blocks in a single turn, add a second breakpoint further back so a write accumulates there before you need it.&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  The 5-minute vs 1-hour decision
&lt;/h2&gt;

&lt;p&gt;Default TTL is &lt;strong&gt;5 minutes&lt;/strong&gt;, and it refreshes for free every time you hit the cache. An active chat session basically keeps itself warm.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"cache_control"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ephemeral"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"ttl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1h"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reach for 1h when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Follow-ups are likely to land &lt;em&gt;after&lt;/em&gt; 5 minutes but &lt;em&gt;within&lt;/em&gt; an hour (a user who steps away; an agent sub-task that runs long)&lt;/li&gt;
&lt;li&gt;Latency matters on those delayed follow-ups&lt;/li&gt;
&lt;li&gt;You're batching, where jobs commonly take 5–60 minutes
Stick with 5m when your prompt is used more often than every 5 minutes — refreshes are free, so you'd be paying 2× writes for nothing.
Mixing TTLs in one request is allowed, with one rule: &lt;strong&gt;longer TTLs must come first&lt;/strong&gt;. 1-hour blocks before 5-minute blocks.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Bonus: pre-warm the cache
&lt;/h2&gt;

&lt;p&gt;Latency-sensitive app? The first user of the day eats the cache-miss penalty. Unless you warm it up first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are an expert software engineer...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;prewarm_cache&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                                  &lt;span class="c1"&gt;# no output generated
&lt;/span&gt;        &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;warmup&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;max_tokens: 0&lt;/code&gt; reads your prompt in, writes the cache at your breakpoint, and returns immediately with an empty &lt;code&gt;content&lt;/code&gt; array and &lt;code&gt;stop_reason: "max_tokens"&lt;/code&gt;. Zero output tokens billed. (You still pay the cache write, naturally.)&lt;br&gt;
Two things to get right:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Put the breakpoint on the &lt;strong&gt;shared&lt;/strong&gt; content (your system prompt), &lt;em&gt;not&lt;/em&gt; on the &lt;code&gt;"warmup"&lt;/code&gt; placeholder — otherwise the entry is keyed to the placeholder and real traffic never hits it. This is why pre-warming needs an explicit breakpoint rather than automatic caching.&lt;/li&gt;
&lt;li&gt;Use the &lt;strong&gt;same&lt;/strong&gt; thinking config and &lt;code&gt;effort&lt;/code&gt; setting as your real requests. Those get rendered into the prompt, so a mismatched pre-warm writes an entry nobody uses.
&lt;code&gt;max_tokens: 0&lt;/code&gt; is rejected with &lt;code&gt;stream: true&lt;/code&gt;, extended thinking, structured outputs, forced &lt;code&gt;tool_choice&lt;/code&gt;, or inside a Batches request.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What breaks the cache
&lt;/h2&gt;

&lt;p&gt;Cache hits need a &lt;strong&gt;100% byte-identical&lt;/strong&gt; prefix. Things that invalidate:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;th&gt;Blast radius&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Tool definitions&lt;/td&gt;
&lt;td&gt;Everything&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toggling web search / citations&lt;/td&gt;
&lt;td&gt;System + messages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Switching fast mode&lt;/td&gt;
&lt;td&gt;System + messages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tool_choice&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Messages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Adding/removing images anywhere&lt;/td&gt;
&lt;td&gt;Messages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking config or &lt;code&gt;effort&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Messages (and sometimes more)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One sneaky one: some languages (&lt;strong&gt;Go, Swift&lt;/strong&gt;) randomize map key order when serializing JSON. If your &lt;code&gt;tool_use&lt;/code&gt; blocks come out with shuffled keys, your cache never hits and you'll have no idea why. Pin the ordering.&lt;br&gt;
Also: caches are isolated per organization, and per workspace on the Claude API. And a cache entry only becomes available &lt;em&gt;after the first response begins&lt;/em&gt; — so firing 10 parallel requests with the same prefix gives you 10 misses. Send one, wait, then fan out.&lt;/p&gt;




&lt;h2&gt;
  
  
  Troubleshooting checklist
&lt;/h2&gt;

&lt;p&gt;Cache not hitting? Run down this list:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Is the cached section &lt;strong&gt;byte-identical&lt;/strong&gt; across calls?&lt;/li&gt;
&lt;li&gt;[ ] Are you over the &lt;strong&gt;minimum token count&lt;/strong&gt; for your model?&lt;/li&gt;
&lt;li&gt;[ ] Is the breakpoint on a block that &lt;strong&gt;stays the same&lt;/strong&gt; (no timestamps, no user input)?&lt;/li&gt;
&lt;li&gt;[ ] Are calls landing &lt;strong&gt;within the TTL&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;[ ] Are &lt;code&gt;tool_choice&lt;/code&gt;, image presence, thinking config, and &lt;code&gt;effort&lt;/code&gt; &lt;strong&gt;consistent&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;[ ] Is your JSON &lt;strong&gt;key order stable&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;[ ] Has a growing conversation pushed you past the &lt;strong&gt;20-block lookback&lt;/strong&gt;?&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Got a caching setup that surprised you — good or bad? Drop it in the comments.&lt;/em&gt; 🚀&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>performance</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Skills as Sub-Agents: Orchestrating Complex work with Claude Skills</title>
      <dc:creator>Mukul S</dc:creator>
      <pubDate>Fri, 31 Jul 2026 17:45:24 +0000</pubDate>
      <link>https://dev.to/mukul_sharma_61fc4dd6f9d8/skills-as-sub-agents-orchestrating-complex-work-with-claude-skills-5d7m</link>
      <guid>https://dev.to/mukul_sharma_61fc4dd6f9d8/skills-as-sub-agents-orchestrating-complex-work-with-claude-skills-5d7m</guid>
      <description>&lt;p&gt;If you've built anything with a coding agent, you've hit the wall: the task is too big for one prompt. You ask it to "find out what's causing this bug," and it starts strong — then drowns. Half its context is the raw output of files it dumped, it's lost the thread of which theory it was testing, and the answer it finally gives is confidently wrong.&lt;br&gt;
The problem isn't the model. It's that you asked &lt;em&gt;one&lt;/em&gt; agent, with &lt;em&gt;one&lt;/em&gt; context window, to do &lt;em&gt;everything&lt;/em&gt; — gather evidence, hold it all in its head, and reason about it at the same time.&lt;br&gt;
There's a cleaner pattern. Split the work into &lt;strong&gt;skills&lt;/strong&gt;, and have one skill &lt;strong&gt;orchestrate the others as sub-agents&lt;/strong&gt;. Let's walk through it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;New to skills or the agent loop? The one-line version: a &lt;strong&gt;skill&lt;/strong&gt; is a packaged set of instructions on disk that the agent loads only when it's relevant. A &lt;strong&gt;sub-agent&lt;/strong&gt; is a fresh agent instance with its own separate context window. This post is about combining the two.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  The core idea: reasoning up top, work down below
&lt;/h2&gt;

&lt;p&gt;Picture two layers.&lt;br&gt;
&lt;strong&gt;The orchestrator&lt;/strong&gt; sits on top. Its only job is to &lt;em&gt;think&lt;/em&gt;: plan the investigation, interpret results, decide what to do next, and conclude. It never opens a file, never greps the codebase, never parses output. The first lines of an orchestrator skill should say exactly that:&lt;/p&gt;

&lt;p&gt;You are a reasoning and coordination layer. You think, you plan, you interpret, you propose — but you never dig into the data yourself. That work belongs in sub-agents.&lt;br&gt;
&lt;strong&gt;The workers&lt;/strong&gt; sit below. Each is a small, focused skill that does &lt;em&gt;one&lt;/em&gt; concrete thing — find every usage of a function, read a file's git history, run a specific test. The orchestrator spawns each worker as a &lt;strong&gt;sub-agent&lt;/strong&gt;: a separate agent with its own fresh context window.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Why does this split matter so much? &lt;strong&gt;Context hygiene.&lt;/strong&gt;&lt;br&gt;
When a worker greps a huge codebase and reads a dozen files, all of that lands in the &lt;em&gt;worker's&lt;/em&gt; context — not the orchestrator's. The worker chews through it, extracts the one thing that matters, and reports back a three-line summary. The orchestrator's context stays clean: it accumulates &lt;em&gt;conclusions&lt;/em&gt;, not &lt;em&gt;raw data&lt;/em&gt;. That's the whole trick, and it's why the orchestrator can run for dozens of steps without falling over.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  Anatomy of a worker skill
&lt;/h2&gt;

&lt;p&gt;A worker skill is just a folder with a &lt;code&gt;SKILL.md&lt;/code&gt; file. The front matter does the heavy lifting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;find-usages&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Find&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;every&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;place&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;symbol,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;function,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pattern&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;used&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;across&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;codebase.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;when&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tracing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;how&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;change&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ripples&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;through&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;code."&lt;/span&gt;
&lt;span class="na"&gt;context&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fork&lt;/span&gt;
&lt;span class="na"&gt;allowed-tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Grep, Read&lt;/span&gt;
&lt;span class="na"&gt;argument-hint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;&amp;lt;symbol&amp;gt; [--path=&amp;lt;dir&amp;gt;]&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="gh"&gt;# Find Usages&lt;/span&gt;
&lt;span class="gu"&gt;## Input&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Symbol**&lt;/span&gt; (required): the function, class, or pattern to trace.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Path**&lt;/span&gt; (optional): limit the search to a directory.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Output file**&lt;/span&gt;: write your full report here.
&lt;span class="gu"&gt;## Procedure&lt;/span&gt;
&lt;span class="p"&gt;1.&lt;/span&gt; Search the codebase for the symbol.
&lt;span class="p"&gt;2.&lt;/span&gt; For each hit, read enough surrounding code to classify it
   (definition, call site, test, re-export).
&lt;span class="p"&gt;3.&lt;/span&gt; Write the full annotated list to the output file.
&lt;span class="p"&gt;4.&lt;/span&gt; Return a 3-line summary: total count, and the 1–2 most relevant sites.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four things are worth calling out, because they're what make it composable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;description&lt;/code&gt;&lt;/strong&gt; — this is the skill's résumé. It's how the orchestrator decides whether to call this worker at all (more on that below). Write it for &lt;em&gt;selection&lt;/em&gt;, not just documentation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;context: fork&lt;/code&gt;&lt;/strong&gt; — run this skill in an isolated context. This is what makes it a sub-agent instead of inline instructions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;allowed-tools&lt;/code&gt;&lt;/strong&gt; — scope the worker down to only what it needs. A read-only search worker gets &lt;code&gt;Grep, Read&lt;/code&gt; and nothing that can modify files. This is your safety boundary, enforced per-skill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The output contract&lt;/strong&gt; — the worker writes its &lt;em&gt;full&lt;/em&gt; findings to a file and returns only a &lt;em&gt;summary&lt;/em&gt;. This is the single most important convention in the whole system. Get it right and everything else falls into place.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The summary-vs-report split
&lt;/h2&gt;

&lt;p&gt;Say it plainly, because it's the load-bearing idea:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A sub-agent writes its full report to a &lt;strong&gt;file&lt;/strong&gt;, and returns only a short &lt;strong&gt;summary&lt;/strong&gt; to the orchestrator.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The orchestrator reads the summary immediately. If it needs the details — the exact line, the full list — it reads the file &lt;em&gt;on demand&lt;/em&gt;. Most of the time it doesn't need to. So the expensive raw data lives on disk, referenced by path, instead of clogging the reasoning layer's context.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the orchestrator picks workers (without reading them all)
&lt;/h2&gt;

&lt;p&gt;Here's a subtlety that trips people up. If the orchestrator has twenty worker skills available, you might think it needs to read all twenty &lt;code&gt;SKILL.md&lt;/code&gt; files to know what they do. That would blow its context before the work even starts.&lt;br&gt;
It doesn't. The agent already sees every skill's &lt;strong&gt;name and one-line description&lt;/strong&gt; — that's the progressive-disclosure model skills are built on. So the orchestrator picks workers &lt;em&gt;by description alone&lt;/em&gt;. Make this an explicit rule in the orchestrator:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do NOT read the full worker SKILL.md files — they are large and will pollute your context. The skill name and description are enough to pick.&lt;br&gt;
This is why that &lt;code&gt;description&lt;/code&gt; field is so important. It's not documentation for humans; it's the interface the orchestrator selects against.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  The orchestration loop
&lt;/h2&gt;

&lt;p&gt;Put it together and the orchestrator runs a loop that should feel familiar if you've seen the basic agent loop — just one level up. Underneath any specific task it's the same four beats:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PLAN (once)
  Confirm the request, scope it, restate the goal.
repeat:
  DISPATCH    → spawn worker sub-agents in parallel to do the next chunk
  INTEGRATE   → read each summary as it returns, update your working picture
  DECIDE      → goal met? finish. gaps remain? dispatch more.
DELIVER
  Present the result (with alternatives called out if you're unsure) and stop.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the general shape. Each kind of task just fills in the blanks differently:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task type&lt;/th&gt;
&lt;th&gt;DISPATCH gathers…&lt;/th&gt;
&lt;th&gt;The "working picture" is…&lt;/th&gt;
&lt;th&gt;DELIVER produces…&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Debugging&lt;/td&gt;
&lt;td&gt;evidence about a failure&lt;/td&gt;
&lt;td&gt;competing hypotheses&lt;/td&gt;
&lt;td&gt;root cause + fix&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Research&lt;/td&gt;
&lt;td&gt;facts from sources&lt;/td&gt;
&lt;td&gt;an outline of the answer&lt;/td&gt;
&lt;td&gt;a synthesized report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Migration&lt;/td&gt;
&lt;td&gt;which files need changing&lt;/td&gt;
&lt;td&gt;a checklist of edits&lt;/td&gt;
&lt;td&gt;a completed, verified change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audit&lt;/td&gt;
&lt;td&gt;findings per rule/area&lt;/td&gt;
&lt;td&gt;a running list of issues&lt;/td&gt;
&lt;td&gt;a prioritized report&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The vocabulary changes; the loop doesn't. Whenever a task is "break it into independent chunks, do them, combine the results, repeat until done," this is the pattern.&lt;br&gt;
A few things make it work well in practice:&lt;br&gt;
&lt;strong&gt;Fan out, then integrate as results arrive.&lt;/strong&gt; The orchestrator spawns workers as &lt;em&gt;background&lt;/em&gt; sub-agents and doesn't block waiting for all of them. As each summary comes back, it folds it into the picture and — crucially — can dispatch a follow-up immediately. A usage search turns up a suspicious call site? Fire off a git-history worker on that file &lt;em&gt;now&lt;/em&gt;, while the other workers are still running. The loop is interleaved, not batch-then-wait.&lt;br&gt;
&lt;strong&gt;Cap the expensive workers.&lt;/strong&gt; Cheap, read-only workers (a quick search, a file read) have no concurrency limit — the cost of running one you didn't need is near zero. Heavier workers (running the full test suite, deep analysis) get a hard cap, e.g. "at most 2 at a time." Match the limit to the cost.&lt;br&gt;
&lt;strong&gt;Refresh the instructions each phase.&lt;/strong&gt; On a long run, the orchestrator's own SKILL.md scrolls out of its context. So each phase's detailed rules live in a &lt;em&gt;separate&lt;/em&gt; file (&lt;code&gt;phase-collect.md&lt;/code&gt;, &lt;code&gt;phase-synthesize.md&lt;/code&gt;…) that the orchestrator &lt;strong&gt;re-reads every time it enters that phase&lt;/strong&gt;. Don't rely on the agent remembering a rule it read forty steps ago — put the rule back in front of it when it's about to act on it.&lt;br&gt;
&lt;strong&gt;Keep an append-only log.&lt;/strong&gt; One file is the run's source of truth: every dispatch, every result, every decision. Append-only — if an earlier note was wrong, you add a correction rather than editing history. It's the audit trail &lt;em&gt;and&lt;/em&gt; the memory that survives context churn.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this beats one big prompt
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context stays clean.&lt;/strong&gt; Raw data is quarantined in worker contexts. The reasoning layer only ever sees distilled conclusions, so it can go deep and long without degrading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real parallelism.&lt;/strong&gt; Ten independent checks run at once instead of one after another. A job that would take an agent twenty sequential steps finishes in a handful of rounds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specialization.&lt;/strong&gt; Each worker is small, testable, and good at exactly one thing. You can improve the usage-finder without touching anything else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety per worker.&lt;/strong&gt; &lt;code&gt;allowed-tools&lt;/code&gt; means your read-only workers &lt;em&gt;cannot&lt;/em&gt; modify files, no matter what the model decides. The boundary is structural, not a polite request in a prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Composability.&lt;/strong&gt; Adding a capability means adding a skill folder. The orchestrator picks it up by description — no changes to the loop.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Setting up your own
&lt;/h2&gt;

&lt;p&gt;The pattern transfers to any complex task — debugging, research, migrations, audits. A starting recipe:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Write the workers first.&lt;/strong&gt; For each concrete sub-task, make a skill folder with a &lt;code&gt;SKILL.md&lt;/code&gt;. Nail the &lt;code&gt;description&lt;/code&gt; (for selection), set &lt;code&gt;context: fork&lt;/code&gt;, scope &lt;code&gt;allowed-tools&lt;/code&gt;, and define the input params + output-file contract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforce summary-vs-report everywhere.&lt;/strong&gt; Every worker writes full output to a file, returns a short summary. Non-negotiable — it's what keeps the orchestrator's context alive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the orchestrator as pure reasoning.&lt;/strong&gt; Tell it, in the strongest terms, to delegate all the actual work. Give it a loop (gather → interpret → decide) and the rule to pick workers by description, not by reading them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Externalize per-phase rules.&lt;/strong&gt; Put detailed checklists in phase files the orchestrator re-reads on entry, so they stay in its attention on long runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cap the expensive workers and log everything.&lt;/strong&gt; Concurrency limits by cost; an append-only log as the source of truth.
That's the architecture. One agent that thinks, many that fetch — each in its own context, each doing one job well. Once you stop cramming everything into a single prompt and start treating skills as sub-agents, the ceiling on what an agent can reliably handle goes way, way up.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
    </item>
    <item>
      <title>How Claude Code Actually Works: The Agentic Loop</title>
      <dc:creator>Mukul S</dc:creator>
      <pubDate>Wed, 29 Jul 2026 22:26:41 +0000</pubDate>
      <link>https://dev.to/mukul_sharma_61fc4dd6f9d8/how-claude-code-actually-works-the-agentic-loop-331i</link>
      <guid>https://dev.to/mukul_sharma_61fc4dd6f9d8/how-claude-code-actually-works-the-agentic-loop-331i</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Coding agents feel like magic. They aren't. Here's the tiny loop at the heart of it -- the one you could sketch on a napkin.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You type "fix the failing test." Claude Code reads a file, runs &lt;code&gt;pytest&lt;/code&gt;, sees a red failure, edits the code, runs &lt;code&gt;pytest&lt;/code&gt; again, sees green, and tells you it's done.&lt;/p&gt;

&lt;p&gt;The first time you watch this happen, it feels like magic.&lt;/p&gt;

&lt;p&gt;It isn't. Underneath, there's a loop so small you could sketch it on a napkin. Once you see it, coding agents stop being mysterious -- and you'll understand &lt;em&gt;why&lt;/em&gt; they behave the way they do.&lt;/p&gt;

&lt;p&gt;Let's build the mental model from scratch.Here's the unlock that makes everything else click:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A language model can't &lt;em&gt;do&lt;/em&gt; everything. It can only produce text.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model can't open your files. It can't run a command. It can't touch the network. Drop it on a desert island and it will happily generate paragraphs forever, but it will never actually &lt;em&gt;run&lt;/em&gt; your tests.&lt;/p&gt;

&lt;p&gt;So how does Claude Code edit files and run commands?&lt;/p&gt;

&lt;p&gt;It doesn't. A regular program does -- and the model just tells it what to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enter the harness
&lt;/h2&gt;

&lt;p&gt;Around the model sits an ordinary program. Call it the &lt;strong&gt;harness&lt;/strong&gt; (that's the CLI you're running).&lt;/p&gt;

&lt;p&gt;The deal between them is simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The harness tells the model, up front: &lt;em&gt;"Here are the tools you can use."&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;The model replies with a &lt;strong&gt;tool call&lt;/strong&gt; -- not prose, but a structured request like &lt;em&gt;"run &lt;code&gt;pytest -q&lt;/code&gt;"&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;The harness actually runs it, captures the output, and hands the result back to the model.&lt;/li&gt;
&lt;li&gt;The model reads that result and decides what to do next.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model is the brain. The harness is the hands. Neither works without the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tools are just a menu + a contract
&lt;/h2&gt;

&lt;p&gt;When a session starts, the harness gives the model a menu of tools, each with a name, a description, and a schema for its inputs. Simplified, it looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"run_bash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Execute a shell command and return its output"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"input_schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now, when the model wants to run the tests, it doesn't &lt;em&gt;say&lt;/em&gt; "you should run pytest." It emits a structured tool call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"run_bash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pytest -q"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The harness executes that command and appends the result to the conversation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tool_result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1 failed, 4 passed&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;E   assert add(2, 2) == 5"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That result goes back into the model on the next turn. Now the model can &lt;em&gt;react&lt;/em&gt; -- it knows the test failed and exactly why.&lt;/p&gt;

&lt;p&gt;This back-and-forth is the whole game.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop
&lt;/h2&gt;

&lt;p&gt;Here's the engine. This is the napkin sketch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;conversation&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;# the model thinks
&lt;/span&gt;    &lt;span class="n"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;has_tool_calls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# the harness acts
&lt;/span&gt;        &lt;span class="n"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="c1"&gt;# loop again -- the model sees the results and continues
&lt;/span&gt;    &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;   &lt;span class="c1"&gt;# no tool calls == the model is done
&lt;/span&gt;
&lt;span class="nf"&gt;show&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# final answer to the user
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. That's the agentic loop.&lt;/p&gt;

&lt;p&gt;Every turn, the model sees the &lt;em&gt;entire&lt;/em&gt; conversation so far -- your request, every command it has run, and every result. It picks the next action. The harness carries it out. Repeat until the model stops asking for tools.&lt;/p&gt;

&lt;p&gt;Let's trace our "fix the failing test" example through it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;You:&lt;/strong&gt; "fix the failing test"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; calls &lt;code&gt;run_bash("pytest -q")&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harness:&lt;/strong&gt; returns &lt;code&gt;1 failed ... assert add(2, 2) == 5&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; calls &lt;code&gt;read_file("math_utils.py")&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harness:&lt;/strong&gt; returns the file contents&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; calls &lt;code&gt;edit_file(...)&lt;/code&gt; to fix the bug&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harness:&lt;/strong&gt; applies the edit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; calls &lt;code&gt;run_bash("pytest -q")&lt;/code&gt; again&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harness:&lt;/strong&gt; returns &lt;code&gt;5 passed&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; no more tool calls -- replies &lt;em&gt;"Fixed it. The bug was..."&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Ten steps, one loop. No magic -- just a model reacting to real output, one turn at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this design is so powerful
&lt;/h2&gt;

&lt;p&gt;Once you see the loop, a lot of things make sense:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It reacts to reality.&lt;/strong&gt; The model doesn't guess whether the test passed -- it &lt;em&gt;sees&lt;/em&gt; the output and adjusts. That feedback is what makes it feel smart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tools are pluggable.&lt;/strong&gt; Want the agent to query a database or hit an API? Add a tool to the menu. The loop doesn't change. (This is exactly what MCP servers do -- they extend the menu.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety lives in the harness, not the model.&lt;/strong&gt; Before step 7 actually writes to disk, the harness can pause and ask &lt;em&gt;you&lt;/em&gt; for permission. The model can &lt;em&gt;request&lt;/em&gt; anything; the harness decides what's &lt;em&gt;allowed&lt;/em&gt;. That separation is your safety layer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Skills: giving the agent a playbook without bloating its brain
&lt;/h2&gt;

&lt;p&gt;So far the model knows two kinds of things: what's baked into its training, and whatever's in the current conversation. But real work needs &lt;em&gt;specifics&lt;/em&gt; -- your team's deploy steps, a tricky migration checklist, the exact way your repo wants commits formatted.&lt;/p&gt;

&lt;p&gt;You &lt;em&gt;could&lt;/em&gt; paste all that into every conversation. But then it's competing for space in that limited context window, every single session, whether you need it or not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skills&lt;/strong&gt; are the fix. A skill is just a packaged set of instructions for a particular kind of task, sitting on disk until it's needed. Think of it as a playbook the agent can pull off the shelf.&lt;/p&gt;

&lt;p&gt;The trick is &lt;em&gt;how&lt;/em&gt; they load, and it's a pattern worth knowing: &lt;strong&gt;progressive disclosure.&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;At startup, the harness shows the model only a one-line description of each skill -- a table of contents, not the book.&lt;/li&gt;
&lt;li&gt;When your request matches one ("deploy the service"), the model calls a tool to &lt;em&gt;open&lt;/em&gt; that skill.&lt;/li&gt;
&lt;li&gt;Only now do the full instructions load into the conversation -- right when they're relevant, and not a moment before.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So the agent can "know" about a hundred skills while paying the context cost of only the one it's actually using. It's the same pluggable idea from earlier, applied to &lt;em&gt;knowledge&lt;/em&gt; instead of actions: the menu stays cheap, and the details arrive on demand.&lt;/p&gt;

&lt;p&gt;And the best part -- a skill is just files. Here's a whole one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deploy-service&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Deploy&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;web&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;staging&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;run&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;smoke&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tests"&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="p"&gt;1.&lt;/span&gt; Run &lt;span class="sb"&gt;`make build`&lt;/span&gt;.
&lt;span class="p"&gt;2.&lt;/span&gt; Push the image with &lt;span class="sb"&gt;`make push TAG=staging`&lt;/span&gt;.
&lt;span class="p"&gt;3.&lt;/span&gt; Wait for the rollout: &lt;span class="sb"&gt;`kubectl rollout status deploy/web`&lt;/span&gt;.
&lt;span class="p"&gt;4.&lt;/span&gt; Hit &lt;span class="sb"&gt;`/healthz`&lt;/span&gt; and confirm a 200 before reporting success.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Drop that markdown file in, and you've taught the agent a new procedure -- no retraining, no changes to the loop itself.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Skills vs MCP, quickly:&lt;/strong&gt; MCP servers add new &lt;em&gt;tools&lt;/em&gt; (new actions the agent can take). Skills add new &lt;em&gt;instructions&lt;/em&gt; (know-how for using the tools it already has). One grows the hands; the other grows the playbook.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The one catch: context isn't infinite
&lt;/h2&gt;

&lt;p&gt;Notice that every tool result gets appended to the conversation. Read a big file? It's in there. Run a chatty command? That's in there too.&lt;/p&gt;

&lt;p&gt;But the model can only look at a limited amount of text at once -- its &lt;strong&gt;context window&lt;/strong&gt;. On a long session, all those results pile up and eventually won't fit. (This is the very problem skills are designed to sidestep.)&lt;/p&gt;

&lt;p&gt;The fix is &lt;strong&gt;compaction&lt;/strong&gt;: when the conversation gets too big, the harness summarizes the older parts to keep the important bits and drop the noise. It's the tradeoff that lets a session run for hours without the model forgetting why it started.&lt;/p&gt;

&lt;p&gt;(That's a whole post of its own -- maybe part 2.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The napkin, one more time
&lt;/h2&gt;

&lt;p&gt;Strip away everything and a coding agent is a handful of pieces:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A model that &lt;strong&gt;only produces text&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A set of &lt;strong&gt;tools&lt;/strong&gt; it can request, with a clear contract&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;loop&lt;/strong&gt; that runs those requests and feeds results back&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;permission gate&lt;/strong&gt; in the harness that decides what's allowed&lt;/li&gt;
&lt;li&gt;Optional &lt;strong&gt;skills&lt;/strong&gt; -- playbooks the agent loads only when they're relevant&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's the whole machine. The intelligence is in the model; the &lt;em&gt;agency&lt;/em&gt; is in the loop.&lt;/p&gt;

&lt;p&gt;Next time you watch Claude Code fix a test on its own, you'll know exactly what's happening under the hood -- one turn at a time.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;If this was useful, next part will dig into context management: how an agent keeps working long after the conversation outgrows its memory. Follow along.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
    </item>
    <item>
      <title>The Hidden Cost of a Log Line : Sync/Async Flush and everything in Between</title>
      <dc:creator>Mukul S</dc:creator>
      <pubDate>Wed, 29 Jul 2026 06:37:15 +0000</pubDate>
      <link>https://dev.to/mukul_sharma_61fc4dd6f9d8/the-hidden-cost-of-a-log-line-syncasync-flush-and-everything-in-between-13of</link>
      <guid>https://dev.to/mukul_sharma_61fc4dd6f9d8/the-hidden-cost-of-a-log-line-syncasync-flush-and-everything-in-between-13of</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;code&gt;log.info("user logged in")&lt;/code&gt; looks free. It isn't. Behind that one line is a chain of decisions — buffer or not, flush or not, block or drop, same thread or another — and each one trades &lt;strong&gt;latency&lt;/strong&gt;, &lt;strong&gt;throughput&lt;/strong&gt;, and &lt;strong&gt;durability&lt;/strong&gt; against the others. This post walks the whole chain, from the method call down to the bytes hitting the disk platter.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you've ever wondered why your p99 latency has a mysterious spike, why logs vanish after a crash, or what "async logging" actually buys you, this is for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, the map: facade vs. implementation
&lt;/h2&gt;

&lt;p&gt;Java logging is a two-layer cake, and mixing up the layers is the #1 source of confusion.&lt;br&gt;
&lt;strong&gt;The facade&lt;/strong&gt; is the API your code calls. &lt;strong&gt;The implementation&lt;/strong&gt; is what actually writes the bytes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;your code
   │  log.info(...)
   ▼
┌───────────────────────────────┐
│  Facade: SLF4J (or Log4j2 API)│   ← the interface you compile against
└──────────────┬────────────────┘
               │ bound at runtime
   ┌───────────┼────────────┬──────────────┐
   ▼           ▼            ▼              ▼
Logback   Log4j2 Core   java.util.logging  ...
(the engine that buffers, formats, and flushes)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SLF4J&lt;/strong&gt; — the de-facto standard facade. Your app should log against this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logback&lt;/strong&gt; — the reference SLF4J implementation. Solid, widely deployed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log4j2&lt;/strong&gt; — the performance-focused implementation, famous for its lock-free async loggers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;java.util.logging (JUL)&lt;/strong&gt; — built into the JDK, rarely chosen on purpose.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Why the split? So you can swap engines without touching a single &lt;code&gt;log.&lt;/code&gt; call. Everything interesting in this post — the buffering, the flushing, the async magic — happens in the &lt;strong&gt;implementation&lt;/strong&gt; layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  The anatomy of a single log call
&lt;/h2&gt;

&lt;p&gt;Before we talk flushing, let's see what one &lt;code&gt;log.info(...)&lt;/code&gt; actually does. There are five stages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Level check      → is INFO enabled for this logger? (cheap, often the fastest bail-out)
2. Build LogEvent   → capture message, timestamp, thread, MDC context, maybe a stack trace
3. Filter           → run any configured filters
4. Layout / encode  → turn the event into bytes ("2026-07-28 12:00:01 INFO ...")
5. Append           → write those bytes to the destination (file, console, socket)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Stage 5 — the &lt;strong&gt;append&lt;/strong&gt; — is where sync vs. async and flush-vs-no-flush live. It's also, by far, the most expensive stage, because it may touch the disk.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Pro tip already:&lt;/strong&gt; Stage 1 is why you sometimes see &lt;code&gt;if (log.isDebugEnabled())&lt;/code&gt;. It skips stages 2–5 when the level is off. With modern parameterized logging (&lt;code&gt;log.debug("x={}", x)&lt;/code&gt;) the framework does this check for you &lt;em&gt;before&lt;/em&gt; building the string — so &lt;code&gt;log.debug("x=" + x)&lt;/code&gt; is the real anti-pattern, because the string concatenation happens regardless.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Sync logging: the default, and its two hidden layers
&lt;/h2&gt;

&lt;p&gt;"Synchronous" means the append happens &lt;strong&gt;on your application thread&lt;/strong&gt;. Your thread doesn't return from &lt;code&gt;log.info(...)&lt;/code&gt; until the write is done. Simple, predictable — and where all the flush nuance lives.&lt;br&gt;
Here's the subtlety most people miss: between your log statement and the actual disk, there are &lt;strong&gt;two separate buffers&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; your app thread
      │  writes formatted bytes
      ▼
┌─────────────────────┐
│ App-level buffer    │   ← BufferedOutputStream inside the appender
│ (e.g. 4 KB)         │      flush() empties THIS
└──────────┬──────────┘
           │ flush()
           ▼
┌─────────────────────┐
│ OS page cache       │   ← the kernel's copy, still in RAM
│ (kernel memory)     │      fsync() empties THIS
└──────────┬──────────┘
           │ fsync()
           ▼
┌─────────────────────┐
│ Physical disk       │   ← now it survives a power loss
└─────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two different operations empty two different buffers, and people constantly confuse them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;flush()&lt;/code&gt;&lt;/strong&gt; pushes bytes from the &lt;em&gt;application&lt;/em&gt; buffer into the &lt;em&gt;OS page cache&lt;/em&gt;. After a flush, another process (like &lt;code&gt;tail -f&lt;/code&gt;) can see your log line. &lt;strong&gt;But it is still in RAM&lt;/strong&gt; — a kernel panic or power loss loses it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;fsync()&lt;/code&gt;&lt;/strong&gt; (via &lt;code&gt;FileChannel.force()&lt;/code&gt;) forces the OS to write the page cache to &lt;em&gt;physical storage&lt;/em&gt;. This is what actually makes a log line survive a crash — and it's dramatically slower.
Most logging frameworks give you knobs for the first buffer and, effectively, never touch the second by default. That's the durability tradeoff hiding in plain sight.
### &lt;code&gt;immediateFlush&lt;/code&gt;: the knob that matters most
Both Logback and Log4j2 file appenders expose &lt;code&gt;immediateFlush&lt;/code&gt;:
&lt;strong&gt;Logback:&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;appender&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"FILE"&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.core.FileAppender"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;file&amp;gt;&lt;/span&gt;app.log&lt;span class="nt"&gt;&amp;lt;/file&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;immediateFlush&amp;gt;&lt;/span&gt;true&lt;span class="nt"&gt;&amp;lt;/immediateFlush&amp;gt;&lt;/span&gt;  &lt;span class="c"&gt;&amp;lt;!-- default: true --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;encoder&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;pattern&amp;gt;&lt;/span&gt;%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n&lt;span class="nt"&gt;&amp;lt;/pattern&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/encoder&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/appender&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Log4j2:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;File&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"FILE"&lt;/span&gt; &lt;span class="na"&gt;fileName=&lt;/span&gt;&lt;span class="s"&gt;"app.log"&lt;/span&gt; &lt;span class="na"&gt;immediateFlush=&lt;/span&gt;&lt;span class="s"&gt;"true"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;PatternLayout&lt;/span&gt; &lt;span class="na"&gt;pattern=&lt;/span&gt;&lt;span class="s"&gt;"%d{HH:mm:ss.SSS} [%t] %-5level %logger{36} - %msg%n"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/File&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What it does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;immediateFlush=true&lt;/code&gt;&lt;/strong&gt; (the default): call &lt;code&gt;flush()&lt;/code&gt; after &lt;strong&gt;every&lt;/strong&gt; log event. Your line is in the OS cache the instant the call returns. Safe if the JVM crashes (the OS still has the bytes and will write them). &lt;strong&gt;But&lt;/strong&gt; it means a syscall per log line — the throughput killer under load.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;immediateFlush=false&lt;/code&gt;&lt;/strong&gt;: let the &lt;code&gt;BufferedOutputStream&lt;/code&gt; fill up (typically 8 KB) and flush only when it's full. Far fewer syscalls, much higher throughput. The cost: if the JVM dies, whatever's sitting in the app buffer (up to 8 KB of your most recent, most interesting logs) is &lt;strong&gt;gone&lt;/strong&gt;.
That's the fundamental sync-flush tradeoff in one sentence: &lt;strong&gt;flush every line for safety, or buffer for speed.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Note: Log4j2 automatically forces &lt;code&gt;immediateFlush=false&lt;/code&gt; behavior for its &lt;strong&gt;async&lt;/strong&gt; loggers — because in the async world, the fix for durability is different (more on that below).&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  So does sync logging survive a crash?
&lt;/h3&gt;

&lt;p&gt;Depends on &lt;em&gt;which&lt;/em&gt; crash and &lt;em&gt;which&lt;/em&gt; flush setting:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure&lt;/th&gt;
&lt;th&gt;&lt;code&gt;immediateFlush=true&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;immediateFlush=false&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;With &lt;code&gt;fsync&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;JVM crash (exception, OOM)&lt;/td&gt;
&lt;td&gt;✅ safe (OS has it)&lt;/td&gt;
&lt;td&gt;⚠️ lose app buffer&lt;/td&gt;
&lt;td&gt;✅ safe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kill -9 the process&lt;/td&gt;
&lt;td&gt;✅ safe (OS has it)&lt;/td&gt;
&lt;td&gt;⚠️ lose app buffer&lt;/td&gt;
&lt;td&gt;✅ safe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OS kernel panic&lt;/td&gt;
&lt;td&gt;❌ lose page cache&lt;/td&gt;
&lt;td&gt;❌ lose more&lt;/td&gt;
&lt;td&gt;✅ safe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power loss&lt;/td&gt;
&lt;td&gt;❌ lose page cache&lt;/td&gt;
&lt;td&gt;❌ lose more&lt;/td&gt;
&lt;td&gt;✅ safe&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;Almost nobody does &lt;code&gt;fsync&lt;/code&gt; per log line — it's punishingly slow (milliseconds per call). Logs are treated as "best effort durable," and that's usually the right call. Just know the guarantee you're actually getting.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Async logging: get off the hot path
&lt;/h2&gt;

&lt;p&gt;Sync logging's real problem isn't correctness — it's that &lt;strong&gt;your request thread pays the I/O bill&lt;/strong&gt;. If the disk hiccups, if a log rotation stalls, if the buffer flushes at the wrong moment, that latency lands directly on the user request that happened to log at that instant. This is a classic source of mysterious p99 spikes.&lt;br&gt;
Async logging fixes this by handing the log event to &lt;strong&gt;another thread&lt;/strong&gt; and returning immediately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; app thread                         background thread
     │  log.info(...)                     │
     │──── enqueue event ────►  [ queue ] │
     │  returns instantly                 │──► format + write + flush
     ▼                                    ▼
 keep serving the request         does the slow I/O
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your app thread's job shrinks to "drop the event in a queue and move on." The slow work — formatting, writing, flushing — happens on a dedicated logging thread that nobody's waiting on.&lt;br&gt;
There are two very different implementations of this idea, and the difference is the whole ballgame.&lt;/p&gt;
&lt;h3&gt;
  
  
  Flavor 1: Logback / Log4j2 &lt;code&gt;AsyncAppender&lt;/code&gt; (a blocking queue)
&lt;/h3&gt;

&lt;p&gt;This is the classic approach: wrap a real appender in an async one backed by a &lt;code&gt;BlockingQueue&lt;/code&gt; (an &lt;code&gt;ArrayBlockingQueue&lt;/code&gt;).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="c"&gt;&amp;lt;!-- Logback --&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;appender&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"ASYNC"&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;"ch.qos.logback.classic.AsyncAppender"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;queueSize&amp;gt;&lt;/span&gt;256&lt;span class="nt"&gt;&amp;lt;/queueSize&amp;gt;&lt;/span&gt;              &lt;span class="c"&gt;&amp;lt;!-- default: 256 --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;discardingThreshold&amp;gt;&lt;/span&gt;51&lt;span class="nt"&gt;&amp;lt;/discardingThreshold&amp;gt;&lt;/span&gt;  &lt;span class="c"&gt;&amp;lt;!-- default: 20% of queueSize --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;neverBlock&amp;gt;&lt;/span&gt;false&lt;span class="nt"&gt;&amp;lt;/neverBlock&amp;gt;&lt;/span&gt;          &lt;span class="c"&gt;&amp;lt;!-- default: false --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;appender-ref&lt;/span&gt; &lt;span class="na"&gt;ref=&lt;/span&gt;&lt;span class="s"&gt;"FILE"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/appender&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The three knobs that decide its behavior under stress:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;queueSize&lt;/code&gt;&lt;/strong&gt; — how many events can wait in line. Bigger = absorbs bigger bursts, uses more memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;discardingThreshold&lt;/code&gt;&lt;/strong&gt; — when the queue gets this full, &lt;strong&gt;drop lower-priority events&lt;/strong&gt; (TRACE/DEBUG/INFO) and keep WARN/ERROR. Logback defaults this to 20% of the queue remaining. Set it to &lt;code&gt;0&lt;/code&gt; to &lt;em&gt;never&lt;/em&gt; discard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;neverBlock&lt;/code&gt;&lt;/strong&gt; — what to do when the queue is completely full:

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;false&lt;/code&gt; (default): the app thread &lt;strong&gt;blocks&lt;/strong&gt; until space frees up. You keep every log, but now you've reintroduced the exact latency you were trying to avoid — async silently becomes sync under load.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;true&lt;/code&gt;: drop the event and move on. Latency stays flat, but you lose logs during bursts.
&lt;strong&gt;There is no free lunch here.&lt;/strong&gt; A bounded queue under sustained overload has only three options: block the producer, drop events, or grow unbounded (OOM). Every async logger is just choosing which of these to do, and when.
### Flavor 2: Log4j2 Async Loggers (the LMAX Disruptor)
Log4j2's headline feature is &lt;strong&gt;Async Loggers&lt;/strong&gt;, built on the &lt;strong&gt;LMAX Disruptor&lt;/strong&gt; — a lock-free ring buffer that came out of high-frequency trading. This is a genuinely different beast from a blocking queue, and it's why Log4j2 benchmarks so aggressively.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="c"&gt;&amp;lt;!-- Make ALL loggers async (system property) --&amp;gt;&lt;/span&gt;
&lt;span class="c"&gt;&amp;lt;!-- -Dlog4j2.contextSelector=org.apache.logging.log4j.core.async.AsyncLoggerContextSelector --&amp;gt;&lt;/span&gt;
&lt;span class="c"&gt;&amp;lt;!-- Or mix: specific loggers async, rest sync --&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;AsyncLogger&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"com.myapp"&lt;/span&gt; &lt;span class="na"&gt;level=&lt;/span&gt;&lt;span class="s"&gt;"info"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;Root&lt;/span&gt; &lt;span class="na"&gt;level=&lt;/span&gt;&lt;span class="s"&gt;"info"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;AppenderRef&lt;/span&gt; &lt;span class="na"&gt;ref=&lt;/span&gt;&lt;span class="s"&gt;"FILE"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/Root&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why a ring buffer instead of a queue? Two reasons:&lt;br&gt;
&lt;strong&gt;1. Lock-free = no contention.&lt;/strong&gt; A &lt;code&gt;BlockingQueue&lt;/code&gt; uses locks; when many threads log at once, they fight over that lock and serialize. The Disruptor uses a pre-allocated ring buffer and atomic sequence counters (CAS) instead of locks. Producers and the consumer coordinate via cursor positions, not mutual exclusion. At high thread counts this is the difference — Log4j2's own numbers show async loggers sustaining &lt;strong&gt;millions&lt;/strong&gt; of messages/sec where a blocking queue plateaus.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;        ring buffer (pre-allocated slots, reused forever)
        ┌───┬───┬───┬───┬───┬───┬───┬───┐
        │ 5 │ 6 │ 7 │   │   │ 1 │ 2 │ 3 │
        └───┴───┴───┴─▲─┴───┴───┴───┴─▲─┘
                      │               │
              consumer cursor    producer cursor
              (writes to disk)   (app threads publish here)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Garbage-free.&lt;/strong&gt; The slots in the ring buffer are allocated once and &lt;strong&gt;reused&lt;/strong&gt;. A queue allocates a new node per event, feeding the garbage collector. The Disruptor pre-allocates, so in steady state it creates (almost) no garbage — which means no GC pauses caused by logging. Log4j2 pairs this with a "garbage-free" layout mode that reuses &lt;code&gt;StringBuilder&lt;/code&gt;s and buffers, so you can log hard with near-zero allocation.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens when the ring buffer fills? &lt;code&gt;AsyncQueueFullPolicy&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Same fundamental problem as before — a bounded buffer can overflow — and Log4j2 makes the policy explicit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Default&lt;/code&gt; policy&lt;/strong&gt;: the producing thread &lt;strong&gt;blocks (busy-spins/waits)&lt;/strong&gt; until a slot frees up. No lost logs, but backpressure hits your app thread.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Discard&lt;/code&gt; policy&lt;/strong&gt;: drop events at or below a configured level (default: drop INFO and below) when full. Keeps latency flat, loses low-priority logs.&lt;/li&gt;
&lt;li&gt;You can also plug in a &lt;strong&gt;custom&lt;/strong&gt; policy.
And there's a sharp edge worth knowing: if a log call happens &lt;em&gt;on the background consumer thread itself&lt;/em&gt; (e.g. logging from inside a layout or an exception handler), blocking would deadlock — so Log4j2 detects this and routes it synchronously instead.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The durability twist: async can lose logs a sync setup wouldn't
&lt;/h2&gt;

&lt;p&gt;Here's the tradeoff people forget. With async logging, when your app thread returns from &lt;code&gt;log.error("about to crash")&lt;/code&gt;, that event is &lt;strong&gt;just sitting in a queue in memory&lt;/strong&gt;. It hasn't been formatted, let alone written or flushed.&lt;br&gt;
If the JVM crashes in the next millisecond, &lt;strong&gt;that error log — the one explaining the crash — is gone.&lt;/strong&gt; With synchronous &lt;code&gt;immediateFlush=true&lt;/code&gt; logging, the same line would have been in the OS cache and survived.&lt;br&gt;
This is the cruel irony of async logging: it's fastest exactly when you're logging the most, which is often right before something goes wrong.&lt;br&gt;
Mitigations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Shutdown hooks / graceful drain.&lt;/strong&gt; Both frameworks try to flush the queue on orderly shutdown. This handles clean exits, not &lt;code&gt;kill -9&lt;/code&gt; or hard crashes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log fatal errors synchronously.&lt;/strong&gt; A common pattern: async for INFO/DEBUG, but route ERROR/FATAL through a synchronous, immediate-flush appender so the important stuff is durable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accept the loss.&lt;/strong&gt; For high-volume, low-value logs (access logs, debug traces), losing the last few hundred on a hard crash is fine. Match the guarantee to the value of the log.&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  Putting numbers to it (and the benchmarking trap)
&lt;/h2&gt;

&lt;p&gt;Rough ordering of throughput, slowest to fastest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sync + immediateFlush=true        ▓▓                      (syscall per line)
sync + immediateFlush=false       ▓▓▓▓▓▓                  (buffered)
async AsyncAppender (queue)       ▓▓▓▓▓▓▓▓▓▓
async Loggers (Disruptor)         ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But — and every honest benchmark says this — &lt;strong&gt;your mileage will vary wildly&lt;/strong&gt;, because:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Peak vs. sustained throughput are different questions.&lt;/strong&gt; Async is fantastic at absorbing &lt;em&gt;bursts&lt;/em&gt;. But your disk's real write bandwidth is a hard ceiling. If you sustain more log volume than the disk can drain, the queue fills and async degrades to whatever its full-policy is (blocking or dropping). Async doesn't make your disk faster — it just decouples your app from disk &lt;em&gt;jitter&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layout cost is often the real bottleneck.&lt;/strong&gt; A fancy pattern with caller location (&lt;code&gt;%class&lt;/code&gt;, &lt;code&gt;%line&lt;/code&gt;, &lt;code&gt;%method&lt;/code&gt;) forces a stack-trace walk on &lt;strong&gt;every&lt;/strong&gt; log line — that can cost more than the write itself. Avoid location info in hot paths.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Microbenchmarks lie.&lt;/strong&gt; Logging in a tight loop with nothing else running doesn't reflect a real app where the GC, the CPU cache, and other threads are all contending.
The practical takeaway isn't a number — it's a decision tree.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  A practical decision guide
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default for most services:&lt;/strong&gt;&lt;br&gt;
Async logging (Log4j2 Async Loggers if you can, Logback AsyncAppender otherwise), with a bounded queue sized to absorb your typical burst, and low-priority events discarded under pressure.&lt;br&gt;
&lt;strong&gt;When you need every log line (audit, compliance, financial):&lt;/strong&gt;&lt;br&gt;
Synchronous + &lt;code&gt;immediateFlush=true&lt;/code&gt;. Accept the throughput hit; you're buying durability. Add &lt;code&gt;fsync&lt;/code&gt; only if you truly can't lose data on a power cut — and know it'll cost you.&lt;br&gt;
&lt;strong&gt;When latency is sacred (trading, real-time):&lt;/strong&gt;&lt;br&gt;
Log4j2 Async Loggers with garbage-free layout and a discard policy — never let logging block a request thread, and never let it trigger GC.&lt;br&gt;
&lt;strong&gt;A solid hybrid that covers most people:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INFO/DEBUG  → async, discard-under-pressure  (high volume, low value)
WARN/ERROR  → sync, immediateFlush=true       (low volume, high value)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You get speed for the noise and durability for the signal.&lt;/p&gt;




&lt;h2&gt;
  
  
  The five things to actually remember
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A log call has two buffers below it.&lt;/strong&gt; &lt;code&gt;flush()&lt;/code&gt; empties the app buffer to the OS; &lt;code&gt;fsync()&lt;/code&gt; empties the OS to disk. They are not the same, and only the second survives a power loss.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;immediateFlush&lt;/code&gt; is the core sync knob.&lt;/strong&gt; &lt;code&gt;true&lt;/code&gt; = safe but a syscall per line; &lt;code&gt;false&lt;/code&gt; = fast but you lose the app buffer on a crash.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Async moves I/O off your request thread&lt;/strong&gt; — killing latency jitter — but every bounded queue must eventually &lt;strong&gt;block, drop, or OOM&lt;/strong&gt; under sustained overload. Know which one yours does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Disruptor wins by being lock-free and garbage-free&lt;/strong&gt;, not by magic. It beats blocking queues under high thread counts and avoids GC pauses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Async trades durability for speed.&lt;/strong&gt; The log explaining your crash is the one most likely to be lost in the queue. Route critical logs synchronously.
Logging feels like the most boring line in your codebase. It's also one of the few that quietly touches concurrency, the memory hierarchy, syscalls, GC, and durability all at once. Now you know what that one line is really doing. 🚀
---&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;What's your logging setup — sync-and-safe, or async-and-fast? And has "async logging ate my crash log" ever bitten you? Tell me in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>java</category>
      <category>performance</category>
    </item>
    <item>
      <title>NGINX, Actually Explained: Architecture, Config, and the Mental Model That Makes It Click</title>
      <dc:creator>Mukul S</dc:creator>
      <pubDate>Wed, 29 Jul 2026 06:06:37 +0000</pubDate>
      <link>https://dev.to/mukul_sharma_61fc4dd6f9d8/nginx-actually-explained-architecture-config-and-the-mental-model-that-makes-it-click-10bf</link>
      <guid>https://dev.to/mukul_sharma_61fc4dd6f9d8/nginx-actually-explained-architecture-config-and-the-mental-model-that-makes-it-click-10bf</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Most people meet NGINX as a magic file they copy-paste until the 502 goes away. This post is the opposite: a tour of the &lt;em&gt;model&lt;/em&gt; behind that file, drawn straight from the official docs at &lt;a href="https://nginx.org/en/docs/" rel="noopener noreferrer"&gt;nginx.org/en/docs&lt;/a&gt;. Once the model clicks, the config stops being scary.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why NGINX is fast (and why the answer matters)
&lt;/h2&gt;

&lt;p&gt;Classic web servers spawn a process or thread &lt;strong&gt;per connection&lt;/strong&gt;. That's fine at 200 connections and catastrophic at 200,000 — each thread eats memory and forces the kernel to context-switch constantly. This is the famous &lt;strong&gt;C10K problem&lt;/strong&gt;.&lt;br&gt;
NGINX was built to dodge it. Instead of "one connection = one thread," it uses a small, fixed number of &lt;strong&gt;worker processes&lt;/strong&gt;, each running a single-threaded, &lt;strong&gt;event-driven, non-blocking&lt;/strong&gt; loop. One worker juggles thousands of concurrent connections by never sitting idle waiting on I/O.&lt;/p&gt;
&lt;h2&gt;
  
  
  That one design decision explains almost everything else about NGINX. Keep it in mind as we go.
&lt;/h2&gt;
&lt;h2&gt;
  
  
  The process model: master + workers
&lt;/h2&gt;

&lt;p&gt;When NGINX starts, you get &lt;strong&gt;one master process&lt;/strong&gt; and &lt;strong&gt;several worker processes&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;          ┌─────────────────┐
          │  master process │   (runs as root, reads config, binds ports)
          └───────┬─────────┘
                  │ manages
      ┌───────────┼───────────┐
      ▼           ▼           ▼
  ┌────────┐  ┌────────┐  ┌────────┐
  │worker 0│  │worker 1│  │worker 2│   (do the actual request handling)
  └────────┘  └────────┘  └────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The master process&lt;/strong&gt; does the privileged, one-time work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reads and validates configuration&lt;/li&gt;
&lt;li&gt;Binds to privileged ports (like &lt;code&gt;:80&lt;/code&gt; / &lt;code&gt;:443&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Spawns, monitors, and restarts workers&lt;/li&gt;
&lt;li&gt;Handles signals for graceful reloads and upgrades
&lt;strong&gt;Worker processes&lt;/strong&gt; do everything else — accepting connections, reading requests, talking to upstreams, sending responses. They're the hot path.
A sane default is one worker per CPU core:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;worker_processes&lt;/span&gt; &lt;span class="s"&gt;auto&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# matches the number of available cores&lt;/span&gt;
&lt;span class="k"&gt;events&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;worker_connections&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# max simultaneous connections *per worker*&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your theoretical connection ceiling is roughly &lt;code&gt;worker_processes × worker_connections&lt;/code&gt;. With 8 cores and 1024 connections each, that's ~8,000 connections without breaking a sweat — and that's a conservative default.&lt;/p&gt;

&lt;h3&gt;
  
  
  The trick that makes reloads zero-downtime
&lt;/h3&gt;

&lt;p&gt;Run &lt;code&gt;nginx -s reload&lt;/code&gt; and NGINX doesn't drop a single connection. The master:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Re-reads and validates the new config&lt;/li&gt;
&lt;li&gt;Starts &lt;strong&gt;new&lt;/strong&gt; workers with the new config&lt;/li&gt;
&lt;li&gt;Tells &lt;strong&gt;old&lt;/strong&gt; workers to stop accepting new connections&lt;/li&gt;
&lt;li&gt;Old workers finish their in-flight requests, then exit
This graceful choreography is why you can ship config changes to a busy production server in the middle of the day.
---
## The event loop: how one worker serves thousands
Here's the part that trips people up. A single worker is &lt;em&gt;single-threaded&lt;/em&gt;, yet handles thousands of connections. How?
Instead of blocking on I/O, each worker asks the kernel: &lt;em&gt;"tell me when any of these connections has something ready."&lt;/em&gt; On Linux that mechanism is &lt;strong&gt;&lt;code&gt;epoll&lt;/code&gt;&lt;/strong&gt; (it's &lt;code&gt;kqueue&lt;/code&gt; on BSD/macOS). The worker then loops:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;events&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;wait_for_ready_connections&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;   &lt;span class="c1"&gt;// epoll_wait — the only place we "block"&lt;/span&gt;
    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nf"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                        &lt;span class="c1"&gt;// never blocks; do a slice of work, move on&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule inside a worker is sacred: &lt;strong&gt;never block.&lt;/strong&gt; A slow disk read or a slow upstream must not freeze the whole loop, because that loop is serving thousands of other clients. When work would block, NGINX either offloads it (thread pools for disk I/O) or parks the connection and comes back when the kernel says it's ready.&lt;br&gt;
Choose the connection-processing method explicitly if you want, though NGINX auto-selects the best one for your OS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;events&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="s"&gt;epoll&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;              &lt;span class="c1"&gt;# Linux&lt;/span&gt;
    &lt;span class="kn"&gt;worker_connections&lt;/span&gt; &lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;multi_accept&lt;/span&gt; &lt;span class="no"&gt;on&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;        &lt;span class="c1"&gt;# accept as many new connections as possible per event&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  This is the whole "why is NGINX so fast" answer in one sentence: &lt;strong&gt;it turns I/O waiting into an event notification instead of a parked thread.&lt;/strong&gt;
&lt;/h2&gt;

&lt;h2&gt;
  
  
  How NGINX processes a request
&lt;/h2&gt;

&lt;p&gt;The docs have a dedicated page on this ("How nginx processes a request"), and it's worth internalizing. Requests are routed in two steps: &lt;strong&gt;which &lt;code&gt;server&lt;/code&gt; block&lt;/strong&gt;, then &lt;strong&gt;which &lt;code&gt;location&lt;/code&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: pick the &lt;code&gt;server&lt;/code&gt; block
&lt;/h3&gt;

&lt;p&gt;NGINX matches the request against &lt;code&gt;listen&lt;/code&gt; directives and the &lt;code&gt;Host&lt;/code&gt; header (via &lt;code&gt;server_name&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;example.com&lt;/span&gt; &lt;span class="s"&gt;www.example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="c1"&gt;# ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt; &lt;span class="s"&gt;default_server&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# fallback when no server_name matches&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;444&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                 &lt;span class="c1"&gt;# drop connections to unknown hosts&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;server_name&lt;/code&gt; matching order: exact name → leading wildcard (&lt;code&gt;*.example.com&lt;/code&gt;) → trailing wildcard (&lt;code&gt;www.example.*&lt;/code&gt;) → regex. First match wins.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: pick the &lt;code&gt;location&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Within the chosen server, NGINX matches the URI against &lt;code&gt;location&lt;/code&gt; blocks. The matching rules are precise and worth memorizing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;/exact&lt;/span&gt;      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;# 1. exact match — highest priority, stops search&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="s"&gt;^~&lt;/span&gt; &lt;span class="n"&gt;/assets/&lt;/span&gt;   &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;# 2. prefix match that *skips* regex if it wins&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="p"&gt;~&lt;/span&gt; &lt;span class="sr"&gt;\.php$&lt;/span&gt;      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;# 3. case-sensitive regex (first match wins)&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="p"&gt;~&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt; &lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="s"&gt;.(jpg|png)&lt;/span&gt;$ &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="c1"&gt;# 4. case-insensitive regex&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt;             &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;# 5. plain prefix — longest match wins as fallback&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The priority order is: &lt;code&gt;=&lt;/code&gt; beats &lt;code&gt;^~&lt;/code&gt; beats regex (&lt;code&gt;~&lt;/code&gt;/&lt;code&gt;~*&lt;/code&gt;) beats plain prefixes. Getting a 404 or serving the wrong file is &lt;em&gt;almost always&lt;/em&gt; a &lt;code&gt;location&lt;/code&gt; precedence surprise — this table is the cure.&lt;/p&gt;




&lt;h2&gt;
  
  
  The config blocks you'll actually use
&lt;/h2&gt;

&lt;p&gt;NGINX config is a tree of &lt;strong&gt;contexts&lt;/strong&gt;. Directives are only valid in certain contexts. Here's the shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="c1"&gt;# main context&lt;/span&gt;
&lt;span class="k"&gt;worker_processes&lt;/span&gt; &lt;span class="s"&gt;auto&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;events&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kn"&gt;...&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;               &lt;span class="c1"&gt;# connection processing&lt;/span&gt;
&lt;span class="s"&gt;http&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;                       &lt;span class="c1"&gt;# everything HTTP lives here&lt;/span&gt;
    &lt;span class="kn"&gt;include&lt;/span&gt; &lt;span class="s"&gt;mime.types&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;sendfile&lt;/span&gt; &lt;span class="no"&gt;on&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;             &lt;span class="c1"&gt;# zero-copy file serving via the kernel&lt;/span&gt;
    &lt;span class="kn"&gt;upstream&lt;/span&gt; &lt;span class="s"&gt;app&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kn"&gt;...&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;     &lt;span class="c1"&gt;# a pool of backend servers&lt;/span&gt;
    &lt;span class="s"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kn"&gt;...&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;           &lt;span class="c1"&gt;# a virtual host&lt;/span&gt;
&lt;span class="err"&gt;}&lt;/span&gt;
&lt;span class="s"&gt;stream&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kn"&gt;...&lt;/span&gt; &lt;span class="err"&gt;}&lt;/span&gt;               &lt;span class="c1"&gt;# raw TCP/UDP proxying (databases, gRPC, etc.)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Serving static files
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;static.example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;root&lt;/span&gt; &lt;span class="n"&gt;/var/www/html&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;try_files&lt;/span&gt; &lt;span class="nv"&gt;$uri&lt;/span&gt; &lt;span class="nv"&gt;$uri&lt;/span&gt;&lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# try file, then dir, else 404&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/assets/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;expires&lt;/span&gt; &lt;span class="s"&gt;30d&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                  &lt;span class="c1"&gt;# aggressive caching for hashed assets&lt;/span&gt;
        &lt;span class="kn"&gt;add_header&lt;/span&gt; &lt;span class="s"&gt;Cache-Control&lt;/span&gt; &lt;span class="s"&gt;"public,&lt;/span&gt; &lt;span class="s"&gt;immutable"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;try_files&lt;/code&gt; is the workhorse here — it's how SPAs get their "always fall back to index.html" behavior:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;try_files&lt;/span&gt; &lt;span class="nv"&gt;$uri&lt;/span&gt; &lt;span class="nv"&gt;$uri&lt;/span&gt;&lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="n"&gt;/index.html&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# client-side routing friendly&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Reverse proxy to an app server
&lt;/h3&gt;

&lt;p&gt;This is probably why you're here. NGINX sits in front of your Node/Python/Go app and forwards requests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;upstream&lt;/span&gt; &lt;span class="s"&gt;backend&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;127.0.0.1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="nf"&gt;127.0.0.1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;3001&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="c1"&gt;# load balancing: round-robin by default&lt;/span&gt;
    &lt;span class="c1"&gt;# least_conn;                # send to the worker with fewest active conns&lt;/span&gt;
    &lt;span class="c1"&gt;# ip_hash;                   # sticky sessions by client IP&lt;/span&gt;
    &lt;span class="kn"&gt;keepalive&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                &lt;span class="c1"&gt;# reuse upstream connections&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;api.example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://backend&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Host&lt;/span&gt;              &lt;span class="nv"&gt;$host&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;X-Real-IP&lt;/span&gt;         &lt;span class="nv"&gt;$remote_addr&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;X-Forwarded-For&lt;/span&gt;   &lt;span class="nv"&gt;$proxy_add_x_forwarded_for&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;X-Forwarded-Proto&lt;/span&gt; &lt;span class="nv"&gt;$scheme&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_http_version&lt;/span&gt; &lt;span class="mf"&gt;1.1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Connection&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;    &lt;span class="c1"&gt;# required for upstream keepalive&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those &lt;code&gt;proxy_set_header&lt;/code&gt; lines matter more than they look. Without them, your backend sees NGINX's IP as the client, loses the original scheme, and can generate broken redirects. This four-line block is the single most copy-pasted (and most misunderstood) snippet in NGINX land.&lt;/p&gt;

&lt;h3&gt;
  
  
  Load balancing methods, at a glance
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Directive&lt;/th&gt;
&lt;th&gt;Use when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Round robin&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(default)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;Backends are roughly equal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Least connections&lt;/td&gt;
&lt;td&gt;&lt;code&gt;least_conn;&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Requests have uneven duration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IP hash&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ip_hash;&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;You need sticky sessions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weighted&lt;/td&gt;
&lt;td&gt;&lt;code&gt;server ... weight=3;&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Backends have different capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  HTTPS in about ten lines
&lt;/h2&gt;

&lt;p&gt;The docs' "Configuring HTTPS servers" page boils down to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;443&lt;/span&gt; &lt;span class="s"&gt;ssl&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;ssl_certificate&lt;/span&gt;     &lt;span class="n"&gt;/etc/nginx/ssl/example.com.crt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;ssl_certificate_key&lt;/span&gt; &lt;span class="n"&gt;/etc/nginx/ssl/example.com.key&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;ssl_protocols&lt;/span&gt;       &lt;span class="s"&gt;TLSv1.2&lt;/span&gt; &lt;span class="s"&gt;TLSv1.3&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;ssl_ciphers&lt;/span&gt;         &lt;span class="s"&gt;HIGH:!aNULL:!MD5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;ssl_session_cache&lt;/span&gt;   &lt;span class="s"&gt;shared:SSL:10m&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# reuse handshakes across connections&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;# Redirect all HTTP to HTTPS&lt;/span&gt;
&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;301&lt;/span&gt; &lt;span class="s"&gt;https://&lt;/span&gt;&lt;span class="nv"&gt;$host$request_uri&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  &lt;code&gt;ssl_session_cache&lt;/code&gt; is the performance sleeper: TLS handshakes are expensive, and caching sessions across connections cuts CPU noticeably on busy sites.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Rate limiting: cheap insurance
&lt;/h2&gt;

&lt;p&gt;NGINX can shield your backend from bursts and abuse before a request ever reaches your app. It uses a &lt;strong&gt;leaky bucket&lt;/strong&gt; algorithm:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;http&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;# define a zone: 10MB of state, 10 requests/sec per client IP&lt;/span&gt;
    &lt;span class="kn"&gt;limit_req_zone&lt;/span&gt; &lt;span class="nv"&gt;$binary_remote_addr&lt;/span&gt; &lt;span class="s"&gt;zone=api:10m&lt;/span&gt; &lt;span class="s"&gt;rate=10r/s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/api/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="kn"&gt;limit_req&lt;/span&gt; &lt;span class="s"&gt;zone=api&lt;/span&gt; &lt;span class="s"&gt;burst=20&lt;/span&gt; &lt;span class="s"&gt;nodelay&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;# allow short bursts of 20&lt;/span&gt;
            &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://backend&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  There's a sibling directive, &lt;code&gt;limit_conn&lt;/code&gt;, for capping &lt;em&gt;concurrent&lt;/em&gt; connections per client. Both live entirely in NGINX — no app code, no extra service.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  The commands worth memorizing
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nginx &lt;span class="nt"&gt;-t&lt;/span&gt;                 &lt;span class="c"&gt;# test config syntax before applying — do this ALWAYS&lt;/span&gt;
nginx &lt;span class="nt"&gt;-s&lt;/span&gt; reload          &lt;span class="c"&gt;# graceful reload, zero dropped connections&lt;/span&gt;
nginx &lt;span class="nt"&gt;-s&lt;/span&gt; quit            &lt;span class="c"&gt;# graceful shutdown (finish in-flight requests)&lt;/span&gt;
nginx &lt;span class="nt"&gt;-s&lt;/span&gt; stop            &lt;span class="c"&gt;# fast shutdown (drop everything now)&lt;/span&gt;
nginx &lt;span class="nt"&gt;-T&lt;/span&gt;                 &lt;span class="c"&gt;# dump the full, resolved config (includes all `include`s)&lt;/span&gt;
nginx &lt;span class="nt"&gt;-v&lt;/span&gt;                 &lt;span class="c"&gt;# version&lt;/span&gt;
nginx &lt;span class="nt"&gt;-V&lt;/span&gt;                 &lt;span class="c"&gt;# version + build flags + configured modules&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Rule of thumb: &lt;strong&gt;never&lt;/strong&gt; &lt;code&gt;reload&lt;/code&gt; without running &lt;code&gt;nginx -t&lt;/code&gt; first. A syntax error caught by &lt;code&gt;-t&lt;/code&gt; is a non-event; the same error hitting &lt;code&gt;reload&lt;/code&gt; on some setups can take the server down.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Debugging: read the logs like a local
&lt;/h2&gt;

&lt;p&gt;Two logs, two jobs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;nginx&lt;/span&gt;
&lt;span class="s"&gt;http&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;log_format&lt;/span&gt; &lt;span class="s"&gt;main&lt;/span&gt; &lt;span class="s"&gt;'&lt;/span&gt;&lt;span class="nv"&gt;$remote_addr&lt;/span&gt; &lt;span class="s"&gt;-&lt;/span&gt; &lt;span class="nv"&gt;$remote_user&lt;/span&gt; &lt;span class="s"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;$time_local&lt;/span&gt;&lt;span class="s"&gt;]&lt;/span&gt; &lt;span class="s"&gt;'&lt;/span&gt;
                    &lt;span class="s"&gt;'"&lt;/span&gt;&lt;span class="nv"&gt;$request&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt; &lt;span class="nv"&gt;$status&lt;/span&gt; &lt;span class="nv"&gt;$body_bytes_sent&lt;/span&gt; &lt;span class="s"&gt;'&lt;/span&gt;
                    &lt;span class="s"&gt;'"&lt;/span&gt;&lt;span class="nv"&gt;$http_referer&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt; &lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$http_user_agent&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt; &lt;span class="s"&gt;'&lt;/span&gt;
                    &lt;span class="s"&gt;'rt=&lt;/span&gt;&lt;span class="nv"&gt;$request_time&lt;/span&gt; &lt;span class="s"&gt;uct="&lt;/span&gt;&lt;span class="nv"&gt;$upstream_connect_time&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt; &lt;span class="s"&gt;'&lt;/span&gt;
                    &lt;span class="s"&gt;'urt="&lt;/span&gt;&lt;span class="nv"&gt;$upstream_response_time&lt;/span&gt;&lt;span class="s"&gt;"'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;access_log&lt;/span&gt; &lt;span class="n"&gt;/var/log/nginx/access.log&lt;/span&gt; &lt;span class="s"&gt;main&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;error_log&lt;/span&gt;  &lt;span class="n"&gt;/var/log/nginx/error.log&lt;/span&gt; &lt;span class="s"&gt;warn&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adding &lt;code&gt;$request_time&lt;/code&gt; and &lt;code&gt;$upstream_response_time&lt;/code&gt; to your access log is the fastest way to answer "is it NGINX or is it the backend that's slow?" — a question you &lt;em&gt;will&lt;/em&gt; be asked during an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  For deep debugging, a build with &lt;code&gt;--with-debug&lt;/code&gt; unlocks &lt;code&gt;error_log ... debug;&lt;/code&gt;, which traces the request lifecycle in detail. That's covered in the docs' "A debugging log" page.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  The mental model, in five lines
&lt;/h2&gt;

&lt;p&gt;If you remember nothing else:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Master configures, workers serve.&lt;/strong&gt; Privilege separation + zero-downtime reloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One worker, one event loop, thousands of connections.&lt;/strong&gt; Never block the loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Requests route in two steps:&lt;/strong&gt; &lt;code&gt;server&lt;/code&gt; block, then &lt;code&gt;location&lt;/code&gt; — and precedence is not intuitive, so learn the table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contexts nest:&lt;/strong&gt; &lt;code&gt;main → http → server → location&lt;/code&gt;. Directives are context-scoped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;nginx -t&lt;/code&gt; before every &lt;code&gt;reload&lt;/code&gt;.&lt;/strong&gt; Always.
Everything in the config file is a consequence of these five ideas. Once you see the event loop behind the directives, NGINX stops being a black box and starts being the most predictable thing in your stack.
---&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Where to go next in the docs
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Beginner's Guide&lt;/strong&gt; — a hands-on first config&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How nginx processes a request&lt;/strong&gt; — the routing internals in depth&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Admin Guide&lt;/strong&gt; — proxying, load balancing, compression, caching&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Module reference&lt;/strong&gt; — every directive, grouped by module (&lt;code&gt;ngx_http_core_module&lt;/code&gt;, &lt;code&gt;ngx_http_proxy_module&lt;/code&gt;, &lt;code&gt;ngx_http_upstream_module&lt;/code&gt;, …)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All at &lt;a href="https://nginx.org/en/docs/" rel="noopener noreferrer"&gt;nginx.org/en/docs&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Found this useful? Drop your gnarliest NGINX config gotcha in the comments — the &lt;code&gt;location&lt;/code&gt; precedence ones are always a good time.&lt;/em&gt; 🚀&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>devops</category>
      <category>performance</category>
    </item>
  </channel>
</rss>
