<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sungwoo Lee</title>
    <description>The latest articles on DEV Community by Sungwoo Lee (@sungwoo_lee_e0f26be4a29fd).</description>
    <link>https://dev.to/sungwoo_lee_e0f26be4a29fd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3989473%2F81eb30d6-b8ed-42ec-b5fb-a2edc4305bfe.png</url>
      <title>DEV Community: Sungwoo Lee</title>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sungwoo_lee_e0f26be4a29fd"/>
    <language>en</language>
    <item>
      <title>6 Cloudflare Workers Gotchas We Hit in Production</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Sun, 27 Sep 2026 16:10:53 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/6-cloudflare-workers-gotchas-we-hit-in-production-2cfp</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/6-cloudflare-workers-gotchas-we-hit-in-production-2cfp</guid>
      <description>&lt;p&gt;We run a small content site on Cloudflare Workers, D1, and R2 — no origin server, no VM. It's been cheap and mostly boring, which is what you want from infrastructure. But "mostly" is doing work in that sentence. Here are six things that went wrong quietly enough that a green deploy log told us everything was fine while it wasn't.&lt;/p&gt;

&lt;p&gt;None of this is a takedown of the platform. It's the list of assumptions we made that turned out to be wrong, in the order we tripped over them.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. D1's statement limit doesn't fail loudly if your chunking is wrong
&lt;/h2&gt;

&lt;p&gt;D1 caps a single SQL statement at 100,000 bytes. That's documented, and for short rows you'll never notice. We noticed the first time we tried to write a long article body in one &lt;code&gt;UPDATE&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The obvious fix is to split the content into chunks and send them as separate statements that concatenate onto the column:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;prepare&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s2"&gt;`UPDATE posts SET content_html = content_html || ? WHERE id = ?`&lt;/span&gt;
  &lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;bind&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That works — until one chunk in the middle of a batch silently fails to apply for an unrelated reason (a flaky connection, a retried request that landed twice, whatever) and every other statement still returns success. The row is now missing a chunk. The write API told you 200 the whole way through. Nothing about the response shape tells you the final string is short.&lt;/p&gt;

&lt;p&gt;The only way we've found to catch this is to re-measure after the fact:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;results&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;prepare&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="s2"&gt;`SELECT length(content_html) AS len FROM posts WHERE id = ?`&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;bind&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;expectedLength&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`content truncated: got &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;len&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;, expected &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;expectedLength&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;length()&lt;/code&gt; in SQLite counts characters, not bytes, so compare it against your source string's character count, not its byte size, or a batch of multi-byte text will fail the check for the wrong reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Something smaller than D1's limit can kill you first
&lt;/h2&gt;

&lt;p&gt;Before we ever got near 100KB, a completely different ceiling showed up: on Windows, passing SQL inline to &lt;code&gt;wrangler d1 execute --command&lt;/code&gt; breaks at something close to 8KB, with a plain "command line is too long" error that has nothing to do with D1.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# breaks around ~8KB, nowhere near D1's own limit&lt;/span&gt;
wrangler d1 execute my-db &lt;span class="nt"&gt;--remote&lt;/span&gt; &lt;span class="nt"&gt;--command&lt;/span&gt; &lt;span class="s2"&gt;"UPDATE posts SET content_html = '...' WHERE id = 42"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It's a shell-level limit, not a database one, and it's easy to misdiagnose as a D1 problem because the symptom (a write that refuses to go through) looks identical. The fix is to stop passing SQL on the command line at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SQL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; patch.sql
wrangler d1 execute my-db &lt;span class="nt"&gt;--remote&lt;/span&gt; &lt;span class="nt"&gt;--file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;patch.sql
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lesson generalizes past Windows: whatever transport you use to get SQL into &lt;code&gt;wrangler&lt;/code&gt;, check which layer's limit you're actually hitting before you start shrinking your data to fit a number that isn't the real constraint.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Writing that SQL file can quietly rewrite your data
&lt;/h2&gt;

&lt;p&gt;Once we moved to &lt;code&gt;--file&lt;/code&gt;, a second problem showed up that's easy to miss because it doesn't produce an error at all: opening a file in plain Windows text mode converts &lt;code&gt;\n&lt;/code&gt; to &lt;code&gt;\r\n&lt;/code&gt; on write — including inside a quoted string literal that's supposed to be HTML content, not a source-code line you'd want reformatted.&lt;/p&gt;

&lt;p&gt;We measured it directly on one patch: a 28,326-character string came back as 28,637 characters after a round trip through a naive file write — one extra byte for every line break the content already had.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# quietly turns every \n inside the string into \r\n
&lt;/span&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;patch.sql&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sql&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# doesn't
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;patch.sql&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;newline&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sql&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The insert still "succeeds." The row still looks plausible if you eyeball it. The only way we catch this now is the same length re-check from gotcha #1, run against the file on disk before it ever reaches &lt;code&gt;wrangler&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. A token that verifies fine can still be the wrong token
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;wrangler d1 execute&lt;/code&gt; started failing with an authentication error on a token we knew was valid — it had a real expiration date years out and passed a plain token-verify check. The error was real, just misleading: the token existed and was active, it just had never been granted D1 permissions in the first place. It had been issued for a completely different job (DNS and Access rules).&lt;/p&gt;

&lt;p&gt;Cloudflare's own permission model explains why this happens: API tokens are scoped per service — a token can hold, say, "Workers Scripts Read" without holding "D1 Write," and there's no single "is this token good" check that covers both. "Valid" and "authorized for this call" are different questions, and only one of them is answered by a token-verify endpoint.&lt;/p&gt;

&lt;p&gt;The fix was mechanical once we understood it: keep a separate, narrowly-scoped token specifically for D1, under its own environment variable, instead of reusing one general-purpose token everywhere.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# an API token without D1 permission looks fine right up until this call&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CLOUDFLARE_API_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$GENERAL_PURPOSE_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
wrangler d1 execute my-db &lt;span class="nt"&gt;--remote&lt;/span&gt; &lt;span class="nt"&gt;--command&lt;/span&gt; &lt;span class="s2"&gt;"select 1"&lt;/span&gt;
&lt;span class="c"&gt;# Authentication error [code: 10000]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  5. A wall of 404s can mean you're checking the wrong host, not a broken upload
&lt;/h2&gt;

&lt;p&gt;We serve images through R2 behind a Worker, and after a batch upload our verification script reported every single image as a 404 — including ones that had been live and working for months. Before assuming the upload pipeline was broken, we re-checked one already-known-good, already-live image with the exact same request. It also came back 404.&lt;/p&gt;

&lt;p&gt;That's the tell: if a control you know is fine fails the same way as the thing you're actually testing, the bug is in the check, not the subject. In our case it was two things stacked on top of each other. First, we were hitting the public-facing custom domain's media path instead of the actual serving Worker's own subdomain, and the routing between the two isn't 1:1. Second, once we fixed the host, we still got failures — because the default user-agent on a plain HTTP client gets blocked by the Worker's own bot protection, independent of whether the file exists.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt;

&lt;span class="n"&gt;req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;User-Agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Mozilla/5.0 (verification script)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Referer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://my-blog.org/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the control check first. It's one extra request and it tells you whether you're about to chase a bug that doesn't exist. The tool pages behind this particular pipeline live at &lt;a href="https://my-blog.org/tools" rel="noopener noreferrer"&gt;my-blog.org/tools&lt;/a&gt;, if you want to see what didn't 404.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Cross-origin requests throw away your referrer path, not just strip it a little
&lt;/h2&gt;

&lt;p&gt;We wanted to know which article page a newsletter signup came from, and for a while every single row in the database said the same thing: the site's homepage. That's not what was happening — nobody was subscribing exclusively from the homepage.&lt;/p&gt;

&lt;p&gt;The frontend Worker and the API Worker sit on different origins, and a same-origin-looking site can still be split across two Worker subdomains under the hood. The browser's default referrer policy (&lt;code&gt;strict-origin-when-cross-origin&lt;/code&gt;) only sends the full path when the request stays on the same origin; cross-origin, it sends the origin and nothing else. Every request looked like it came from &lt;code&gt;/&lt;/code&gt;, because that's literally what the browser was willing to disclose, not because that's where the click happened.&lt;/p&gt;

&lt;p&gt;The fix isn't a header trick — you can't get the path back out of &lt;code&gt;Referer&lt;/code&gt; once the browser has decided to withhold it. You have to send the path yourself, explicitly, in the request body:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;apiUrl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="nx"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;location&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pathname&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// don't rely on Referer for this&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your pages Worker and your API Worker are different origins — which is common the moment you put one behind a custom domain and hit the other directly — budget for this before you build any attribution on top of &lt;code&gt;Referer&lt;/code&gt;.&lt;/p&gt;




&lt;p&gt;None of these are exotic. Every one of them looked, at the moment it happened, like something else was broken — the database, the network, an expired credential, a missing file. What they had in common is that the actual failure was one layer away from where the symptom showed up, and a plain "it returned 200" or "it deployed" told us nothing about that layer.&lt;/p&gt;

&lt;p&gt;If you're running D1 and R2 behind Workers, the site these lessons came from is &lt;a href="https://my-blog.org/tools" rel="noopener noreferrer"&gt;my-blog.org/tools&lt;/a&gt; — a small pile of calculator pages that gave us most of this list one incident at a time.&lt;/p&gt;

</description>
      <category>cloudflare</category>
      <category>webdev</category>
      <category>javascript</category>
      <category>serverless</category>
    </item>
    <item>
      <title>Exit Code 0, Nothing Produced: The Automation Failure I Shipped Five Times</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Thu, 10 Sep 2026 04:28:42 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/exit-code-0-nothing-produced-the-automation-failure-i-shipped-five-times-345j</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/exit-code-0-nothing-produced-the-automation-failure-i-shipped-five-times-345j</guid>
      <description>&lt;p&gt;The scheduler log is a column of green. Started, finished, &lt;code&gt;rc=0&lt;/code&gt;, started, finished, &lt;code&gt;rc=0&lt;/code&gt;, day after day after day. The job that collects candidates for my syndication queue had been running like that for weeks.&lt;/p&gt;

&lt;p&gt;The queue was empty. It had been empty the whole time.&lt;/p&gt;

&lt;p&gt;I have now shipped this exact bug five separate times in one project, in five different jobs. Every one went unnoticed for weeks, and every one was eventually caught by a human asking "hang on, when did we last actually get anything out of this?" — never by a monitor, never by an alert, never by a failed run. There was nothing to fail.&lt;/p&gt;

&lt;p&gt;The pattern deserves a name because it is invisible by construction: &lt;strong&gt;the job succeeds and produces nothing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6e6zrajvyfw0wey9smmz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6e6zrajvyfw0wey9smmz.png" alt="Exit Code 0, Nothing Produced: The Automation Failure I Shipped Five Times — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the exit code can't see it
&lt;/h2&gt;

&lt;p&gt;An exit code answers one question: did the process reach the end without raising? That is a statement about control flow, not about work. A script that iterates over zero items and prints a tidy summary exits 0 with total confidence.&lt;/p&gt;

&lt;p&gt;The shapes I've actually hit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A collector that returns the empty list on failure.&lt;/strong&gt; &lt;code&gt;except Exception: return []&lt;/code&gt; makes "the search path is broken" and "there was nothing new today" arrive at the caller as the same value. The caller cannot tell them apart, so it picks the cheerful interpretation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;except: continue&lt;/code&gt; inside the loop.&lt;/strong&gt; One bad item is worth skipping. Every item being bad looks identical from outside — the loop completes, the counter is zero, the function returns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An upstream API that answers 200 with an empty array.&lt;/strong&gt; No transport error, no status to branch on. Auth silently downgraded, a parameter quietly rejected, a filter that now matches nothing — all of it arrives as a well-formed successful response containing no rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Partial harvest.&lt;/strong&gt; Four of five sources die, the fifth works, the run reports success. This one is the worst, because there &lt;em&gt;is&lt;/em&gt; output. Nobody knows what the denominator was supposed to be.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What all four share: nowhere in the system is there a written statement of what this job is supposed to leave behind. Without that statement there is nothing for a machine to check.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: make every job declare its output
&lt;/h2&gt;

&lt;p&gt;I moved the definition of every scheduled job into one JSON file, and required each entry to carry a &lt;code&gt;produces&lt;/code&gt; block — a machine-readable claim about what must exist after a healthy run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gsc-daily"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"run_cmd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"scripts/analytics/run_gsc_daily.py"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timeout_s"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1800&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"output_expectation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"conditional"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"produces"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"file_fresh"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"data/ops/gsc_daily.jsonl"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"max_age_h"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"desc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"daily search performance was appended"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I stopped letting jobs report on themselves. A runner wraps each one, and the runner writes the ledger entry.&lt;/p&gt;

&lt;p&gt;Wrapping rather than instrumenting mattered more than I expected. If each script records its own outcome you have to edit five places, and the one you forget is precisely the one that fails quietly. Worse, a script that never starts — bad interpreter path, an import error on line 1, a wrapper that dies before the interpreter is reached — cannot possibly log its own failure. The judgement has to come from outside the process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure the delta, not the total
&lt;/h2&gt;

&lt;p&gt;The runner measures the declared output twice: once before the child process starts, once after it exits.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;before&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;measure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;rc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run_tree&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cwd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout_s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;after&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;measure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;row_count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;before&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;produced&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;after&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;before&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;produced&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;after&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# freshness is a 0/1 state, not a count
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The absolute count is a trap. Yesterday's rows are still sitting in the table, so "there are rows" stays true on the day the collector breaks and on every day after. The number that describes &lt;em&gt;this run&lt;/em&gt; is the increase.&lt;/p&gt;

&lt;p&gt;But the delta is a trap in the other direction, and I walked into that one too. Some jobs upsert: they rewrite one row per day, or refresh a single file in place. Their row count never grows. Judged by delta, a perfectly healthy job reports zero forever, you get an alert every morning, and within a week you have trained yourself to ignore it. So the two kinds are declared separately — cumulative outputs judged by increase, in-place outputs judged by freshness against &lt;code&gt;max_age_h&lt;/code&gt;. Collapsing them into one rule produces false alarms, and false alarms are how a monitoring system dies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not every zero is a failure
&lt;/h2&gt;

&lt;p&gt;Some jobs are legitimately allowed to produce nothing. My indexing-request job is capped by a daily quota and drains a pending queue; on a day when the queue is empty, zero is the correct answer. So each job also declares &lt;code&gt;output_expectation&lt;/code&gt;: &lt;code&gt;every_run&lt;/code&gt; means an empty result is a defect, &lt;code&gt;conditional&lt;/code&gt; means it may be empty and the thing worth counting is how many consecutive silent days have passed.&lt;/p&gt;

&lt;p&gt;The distinction underneath that is the one I'd carry to any scraper or collector:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We looked and only found things we already have" and "we couldn't look" are different outcomes. Only the second is a failure.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A collector that finds results and discards all of them as duplicates is working correctly. A collector that finds &lt;em&gt;no results at all&lt;/em&gt; is not reporting an absence of new material — it is reporting that its search path returned nothing, which for an established source means the path is broken. So the two are counted separately, and only the second returns a non-zero status. Before I made that split, the healthy case and the broken case printed the same line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure is expensive when nothing stops early
&lt;/h2&gt;

&lt;p&gt;Once I could see the failures, a second problem surfaced: broken jobs were burning enormous amounts of wall time on their way to producing nothing.&lt;/p&gt;

&lt;p&gt;One collector had its search path blocked. Its response was to try the next seed keyword, hit a 40-second navigation timeout, try the next, hit another 40-second timeout, and continue down the entire seed list. It burned 1,380 seconds and returned 0 candidates. Every one of those timeouts was foreseeable after the third.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;fails&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;FAIL_CAP&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;q&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;seeds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;url_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;40000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;fails&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  [&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] failed (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fails&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;FAIL_CAP&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;): &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;90&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;fails&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;FAIL_CAP&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  three consecutive failures - stopping&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
        &lt;span class="k"&gt;continue&lt;/span&gt;
    &lt;span class="n"&gt;fails&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;          &lt;span class="c1"&gt;# any success breaks the streak
&lt;/span&gt;    &lt;span class="nf"&gt;harvest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two rules, both cheap. A consecutive-failure cap, reset on any success, because a source that refused us three times in a row will refuse the fourth. And a per-job wall-clock ceiling in the job definition, so a job that hangs rather than fails still terminates on schedule instead of still running when tomorrow's copy of itself starts on top of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kill the tree, then verify it died
&lt;/h2&gt;

&lt;p&gt;The wall-clock ceiling introduced its own bug, specific to layered automation.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;subprocess.run(timeout=N)&lt;/code&gt; kills its direct child. My batch jobs are three generations deep: Python starts a browser driver, the driver starts a browser. Killing the direct child leaves the browser orphaned, holding a profile lock that quietly breaks the next day's run — a silent no-op caused by the very mechanism meant to prevent silent no-ops.&lt;/p&gt;

&lt;p&gt;On Windows the job has to be launched into its own process group and killed as a tree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;flags&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CREATE_NEW_PROCESS_GROUP&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="n"&gt;proc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Popen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cwd&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;cwd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;fo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;fe&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;creationflags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;flags&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TimeoutExpired&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;taskkill&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/F&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/T&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/PID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pid&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt;
                       &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;1.5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;is_alive&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pid&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;        &lt;span class="c1"&gt;# verify, don't assume
&lt;/span&gt;            &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;note_orphan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# write it down; nobody notices otherwise
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The retry loop exists because of the same class of mistake as everything else here. My first version called &lt;code&gt;taskkill&lt;/code&gt; once and recorded its return code as the outcome. A child survived anyway. The parent was gone, the lock was released, the ledger said the timeout had been handled cleanly — and a stray process sat there until I went looking for it. &lt;strong&gt;"I issued a kill" and "the process is gone" are two different facts,&lt;/strong&gt; and only the second is worth recording. When the kill genuinely fails, the PID and its command line go into an orphan list so the next run of that job can clean up after its predecessor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Logs and artifacts are separate evidence
&lt;/h2&gt;

&lt;p&gt;The last thing that broke my trust in logs: I launched two jobs in the background and later found their log file was 0 bytes, while the ledger showed two successful runs with real output counts.&lt;/p&gt;

&lt;p&gt;Both records were accurate. The work happened; the log capture didn't, because output was block-buffered and the redirect never flushed. But the lesson stands on its own: a log is a story the process tells about itself. It can be missing, truncated, or buffered away while the work succeeds — and it can be full of confident green while nothing was produced. Those are independent failure modes. Judge the run by its artifacts, and keep the log as evidence for the postmortem rather than as the verdict. That split is one of the running themes in &lt;a href="https://my-blog.org/claude-code" rel="noopener noreferrer"&gt;the operational notes I keep on this&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to ask your own cron jobs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;For each scheduled job, can you name in one sentence what file or row it must leave behind? If not, it has no success criterion — only an exit code.&lt;/li&gt;
&lt;li&gt;Is that criterion written somewhere a machine reads, or only in your head?&lt;/li&gt;
&lt;li&gt;Are you checking a total or a delta — and if a delta, does that job actually append, or does it upsert in place?&lt;/li&gt;
&lt;li&gt;Can this job distinguish "found nothing new" from "couldn't look," and does it return a different status for each?&lt;/li&gt;
&lt;li&gt;Does every retry loop have a consecutive-failure cap, and every job a wall-clock ceiling?&lt;/li&gt;
&lt;li&gt;When the ceiling fires, do you kill the process tree and then verify it is gone?&lt;/li&gt;
&lt;li&gt;When was the last time this job produced something? If you have to go and dig to answer that, that is the whole problem.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;Related reading&lt;/strong&gt; — &lt;a href="https://my-blog.org/tangents/post/how-to-stop-ai-hallucinations" rel="noopener noreferrer"&gt;How to Stop AI Hallucinations: 6 Prompt Tactics That Reduce Made-Up Answers&lt;/a&gt;&lt;/p&gt;

</description>
      <category>automation</category>
      <category>devops</category>
      <category>python</category>
      <category>testing</category>
    </item>
    <item>
      <title>65 Tool Pages, 12 Pageviews in 30 Days: The Bug Was in the Site Graph</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Tue, 08 Sep 2026 16:55:21 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/65-tool-pages-12-pageviews-in-30-days-the-bug-was-in-the-site-graph-4n32</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/65-tool-pages-12-pageviews-in-30-days-the-bug-was-in-the-site-graph-4n32</guid>
      <description>&lt;p&gt;I shipped 65 calculator pages to a content site over a few months. Last week I finally pulled the analytics for them: &lt;strong&gt;12 pageviews across all 65 pages in 30 days.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not 12 per page. Twelve total.&lt;/p&gt;

&lt;p&gt;Every page returned HTTP 200. Every page was in the sitemap. The search console had them. The keyword demand was real — the Korean search-ad planner reports roughly 898,000 monthly searches for the single term "character counter" and 604,000 for "percentage calculator." Supply existed, demand existed, and the two never met.&lt;/p&gt;

&lt;p&gt;This is what I found and what I changed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcd1154nee8uwz3mdiraw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcd1154nee8uwz3mdiraw.png" alt="65 Tool Pages, 12 Pageviews in 30 Days: The Bug Was in the Site Graph — at a glance" width="799" height="531"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  First: rule out the boring failures
&lt;/h2&gt;

&lt;p&gt;Before theorizing, I checked the things that are cheap to check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are they up?&lt;/strong&gt; I pulled every tool URL from the sitemap and requested all 65 with cache-busting. 65/65 returned 200. One returned a connection error on the first pass and 200 twice on retry — a transient blip, not a pattern.&lt;/p&gt;

&lt;p&gt;Pulling the URL list from the sitemap rather than grepping the source turned out to matter. My first instinct was &lt;code&gt;grep "slug:" tools_calc30.js&lt;/code&gt;, which dumped 622 KB into my terminal because that file embeds a large JSON dataset in a string literal. The sitemap is the artifact that search engines actually consume, so it's also the correct inventory to audit against.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are they indexed?&lt;/strong&gt; Partially. Impressions were rising — 30/day at the start of the month to 162 six days later, average position moving from 14.3 to 6.6. So crawling wasn't blocked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are they slow?&lt;/strong&gt; Median around 0.9s, served from the edge. No.&lt;/p&gt;

&lt;p&gt;So the pages were reachable, crawlable, and improving in impressions. And nobody was reading them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then: query the site graph
&lt;/h2&gt;

&lt;p&gt;The thing I hadn't checked was whether anything on my own site linked to them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;content_html&lt;/span&gt; &lt;span class="k"&gt;LIKE&lt;/span&gt; &lt;span class="s1"&gt;'%/tools/%'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Zero.&lt;/strong&gt; Out of 380 published articles, not one contained a link to a tool page.&lt;/p&gt;

&lt;p&gt;The only path from the site root to any tool was a single "Free tools" item in the top navigation. From there, an index page. From there, a tool. So every tool page sat two to three hops from the homepage, reachable only through a nav link that search crawlers weight very lightly and that human readers, deep in an article about something else, never see.&lt;/p&gt;

&lt;p&gt;I had built a wing of the building and forgotten to cut a door into the hallway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "it's in the sitemap" wasn't enough
&lt;/h2&gt;

&lt;p&gt;A sitemap tells a crawler a URL exists. Internal links tell it the URL matters, and they tell a human reader it exists at the moment they'd want it.&lt;/p&gt;

&lt;p&gt;The impressions data showed the crawler had done its part: it found the pages and started testing them in results. What it hadn't gotten was any signal that these 65 pages were more than a dusty appendix. Meanwhile the humans reading my articles had no in-context path to a calculator at the exact moment a calculation would have been useful.&lt;/p&gt;

&lt;p&gt;Two failures, one cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;One: link from the homepage directly.&lt;/strong&gt; The homepage is the strongest page on any small site. I added a section listing 12 tools by name and description, ordered by measured monthly search volume, plus a link to each of the two tool index pages. That takes the crawl distance from three hops to one for the twelve highest-demand tools:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;HOME_TOOL_PATHS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/char-count&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/percent&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/date-calc&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/salary&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/loan&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/age&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/pyeong&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/unemployment&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/severance&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/weekly-holiday&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/deposit&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/discount&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;homeTools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;HOME_TOOL_PATHS&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;TOOL_LIST&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;path&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Boolean&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Names and descriptions come from the same array the tool pages use. Writing them twice guarantees one copy goes stale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two: make the per-article tool cards context-aware.&lt;/strong&gt; Articles already rendered a small "related tools" block, but one blog section was falling through to a generic default that offered a &lt;em&gt;basal metabolic rate&lt;/em&gt; calculator to people reading about government startup grants. An irrelevant link isn't a neutral link — it's an exit. I mapped that section to the tools its readers would actually want: take-home pay, loan interest, withholding tax, percentage, date difference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three: tell the index services.&lt;/strong&gt; After deploying, I submitted the homepage, both index pages, and the twelve tool URLs through IndexNow (15 accepted), and requested indexing for the highest-value pages in the search console. That console has a small daily quota — I got three URLs through before hitting it, which is worth knowing before you plan a batch of sixty.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd check first next time
&lt;/h2&gt;

&lt;p&gt;The diagnostic that found this took one SQL query, and I ran it months later than I should have. So the check I'm adding to my own list is blunt:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For any new section of a site, count the internal links pointing into it from existing content. If the answer is zero, nothing else about that section matters yet.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Crawl status, page speed, structured data, and title copy are all downstream of that. I spent real time on the downstream things while the number was zero.&lt;/p&gt;

&lt;p&gt;The counterpart to this is a rule I already believed and failed to apply: measure the thing you actually care about, not the thing that's easy to measure. "65/65 return 200" felt like a green light. It was answering a question nobody was asking.&lt;/p&gt;

&lt;p&gt;I'll post the follow-up numbers once there's enough data to say whether the fix moved anything. If you want to see the pages themselves, they're at &lt;a href="https://my-blog.org/tools" rel="noopener noreferrer"&gt;my-blog.org/tools&lt;/a&gt; — mostly Korean-language calculators, but the &lt;a href="https://my-blog.org/tools/salary" rel="noopener noreferrer"&gt;take-home pay one&lt;/a&gt; is a reasonable example of the shape.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>postmortem</category>
      <category>analytics</category>
    </item>
    <item>
      <title>65 Calculators, One Cloudflare Worker, No Build Step</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Tue, 08 Sep 2026 16:54:45 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/65-calculators-one-cloudflare-worker-no-build-step-29i6</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/65-calculators-one-cloudflare-worker-no-build-step-29i6</guid>
      <description>&lt;p&gt;I run a small content site that also hosts 65 calculator pages — salary after tax, loan interest, character count, percentage, a grade calculator, and 43 math/exam-prep tools. All of them live in a single Cloudflare Worker. There is no framework, no bundler, no client-side JS payload beyond the few hundred bytes each page needs to do its own arithmetic.&lt;/p&gt;

&lt;p&gt;This post is about why that setup is a good fit for this particular problem, and where it stops being one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdq6af8yvees8bqh8cocq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdq6af8yvees8bqh8cocq.png" alt="65 Calculators, One Cloudflare Worker, No Build Step — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the problem
&lt;/h2&gt;

&lt;p&gt;A calculator page is unusual among web pages: the interesting part is tiny and the boring part is everything else.&lt;/p&gt;

&lt;p&gt;The interesting part is a pure function. Given a salary, a number of dependents, and a non-taxable allowance, return six deduction amounts. That's maybe 40 lines. The boring part — page shell, navigation, meta tags, structured data, FAQ markup, related-tool cards, the privacy note explaining that nothing is sent to a server — is the same on all 65 pages and dwarfs the calculation.&lt;/p&gt;

&lt;p&gt;So the architecture question isn't "how do I build a calculator." It's "how do I stamp out 65 near-identical documents where only a small slot differs."&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Worker actually does
&lt;/h2&gt;

&lt;p&gt;Every request hits one Worker. It matches the path, looks the tool up in a flat array, and returns a string.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;TOOL_LIST&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/char-count&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Character counter&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;desc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Counts with and without spaces, bytes, manuscript pages&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;g&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;everyday&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/tools/salary&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Take-home pay calculator&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;desc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Monthly net after insurance and income tax&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;g&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;work&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That array is the single source of truth. The tools index page, the nav, the related-tool cards on article pages, and the sitemap all read from it. When I added the &lt;a href="https://my-blog.org/tools/salary" rel="noopener noreferrer"&gt;take-home pay calculator&lt;/a&gt;, I added one row and it appeared in four places without any of them being edited.&lt;/p&gt;

&lt;p&gt;The page-specific script is stored as an array of source lines and joined at render time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;SAL_SCRIPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;const n = (id) =&amp;gt; Number(document.getElementById(id).value || 0);&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;function calc() { /* ... */ }&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;String&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromCharCode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;String.fromCharCode(10)&lt;/code&gt; looks silly next to &lt;code&gt;'\n'&lt;/code&gt;. It's there because this source string travels through several layers of tooling, and a literal backslash-n has been folded on me more than once. Using the char code makes the newline immune to whatever escapes the string on its way through.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the calculation runs in the browser
&lt;/h2&gt;

&lt;p&gt;Every one of these tools computes client-side. Nothing about the input is transmitted.&lt;/p&gt;

&lt;p&gt;Partly that's a privacy claim I want to be able to make honestly: people paste cover letters into a &lt;a href="https://my-blog.org/tools/char-count" rel="noopener noreferrer"&gt;character counter&lt;/a&gt;, and I would rather not have that text touch my logs even accidentally. Partly it's just cheaper — a Worker invocation that returns a static string and never awaits anything is about as cheap as a request gets.&lt;/p&gt;

&lt;p&gt;The tradeoff is that the logic ships to the client, so it's readable and copyable. For a tax table encoding that's a wash; the numbers are published by the tax authority anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one place this actually got hard
&lt;/h2&gt;

&lt;p&gt;The income tax portion isn't a formula. It's a lookup table: 646 income brackets × 11 dependent counts, published as a spreadsheet.&lt;/p&gt;

&lt;p&gt;Shipping that as JSON was 400 KB. Instead I encoded it as a run-length-ish delta string and unpacked it in about 15 lines at page load. The encoded module is 31 KB. The unpack function is small enough to read in one sitting, which mattered more to me than squeezing the last few KB, because the failure mode of a clever encoding is silently wrong money.&lt;/p&gt;

&lt;p&gt;I verify it by running the exact page script — pulled out of the source file at test time, not a copy — against known values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;src&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;readFileSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;tools_pay.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;utf-8&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;src&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/const SAL_SCRIPT = &lt;/span&gt;&lt;span class="se"&gt;\[([\s\S]&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt;&lt;span class="se"&gt;?)\]\.&lt;/span&gt;&lt;span class="sr"&gt;join&lt;/span&gt;&lt;span class="se"&gt;\(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;script&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Function&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;return [&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;].join(String.fromCharCode(10));&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pulling the literal out of the real file rather than importing a copy is the whole point. A copy drifts, and then your tests pass against a version nobody ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this approach stops working
&lt;/h2&gt;

&lt;p&gt;Three honest limits:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No component reuse across pages beyond string concatenation.&lt;/strong&gt; When two tools want the same input widget with slightly different labels, you write a function that returns a string with parameters. That's fine at 65 pages. I would not want to find out where it stops being fine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No client-side routing, and I don't want any.&lt;/strong&gt; Each tool is a separate document with its own title, description, and structured data. That's deliberate — these pages exist to be found by search, and a single-page app would collapse 65 search targets into one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Debugging inline script is worse than debugging a module.&lt;/strong&gt; Source maps don't help you when the source is a joined array. In practice I develop the logic as a normal file and only move it into the array form when it settles, which is a smell I've made peace with.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell someone starting this
&lt;/h2&gt;

&lt;p&gt;If your pages are mostly identical shells around a small pure function, and if being individually indexable matters more than sharing state between pages, a single Worker plus a flat config array will take you further than it sounds like it should. The whole surface is 65 URLs, and the deploy is one command.&lt;/p&gt;

&lt;p&gt;The part that actually determined whether anyone found these pages turned out to have nothing to do with any of this — it was internal linking, and I got it badly wrong for months. That's a separate post.&lt;/p&gt;

&lt;p&gt;The full set is at &lt;a href="https://my-blog.org/tools" rel="noopener noreferrer"&gt;my-blog.org/tools&lt;/a&gt; if you want to poke at the output. Most of it is in Korean, but view-source works in every language.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://developer.mozilla.org" rel="noopener noreferrer"&gt;MDN Web Docs&lt;/a&gt; — web standards&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>cloudflare</category>
      <category>webdev</category>
      <category>javascript</category>
      <category>performance</category>
    </item>
    <item>
      <title>Why Growth Curves Bend: Exponential vs. Logistic, With Code</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Sun, 23 Aug 2026 07:36:46 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/why-growth-curves-bend-exponential-vs-logistic-with-code-1fgk</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/why-growth-curves-bend-exponential-vs-logistic-with-code-1fgk</guid>
      <description>&lt;p&gt;In early 2020 a lot of people learned the phrase "doubling every three days." A month later the same commentators were saying "growth is slowing." The math didn't change. What changed is that the quantity being measured ran into a ceiling — and every unbounded growth curve you'll ever plot a metric against does the same thing, whether it's infections, DAU, or a tweet's retweet count.&lt;/p&gt;

&lt;p&gt;This post walks through why, with formulas you can verify on a calculator and a script short enough to paste into a REPL.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs1wmduoiqjaj4ri32gdg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs1wmduoiqjaj4ri32gdg.png" alt="Why Growth Curves Bend: Exponential vs. Logistic, With Code — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Exponential growth is a multiplication rule, not a speed
&lt;/h2&gt;

&lt;p&gt;The discrete form is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;a&lt;/code&gt; is the starting value, &lt;code&gt;b&lt;/code&gt; is the growth factor per period, &lt;code&gt;t&lt;/code&gt; is the number of periods. If &lt;code&gt;b = 2&lt;/code&gt;, the quantity doubles every period. The continuous form uses &lt;code&gt;e&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;where &lt;code&gt;r&lt;/code&gt; is the continuous growth rate and &lt;code&gt;b = e**r&lt;/code&gt;. Both describe the same property: the rate of change is proportional to the current size. That's the entire definition — nothing in it says "fast." A process growing at &lt;code&gt;r = 0.001&lt;/code&gt; per year is exponential in structure but you'd never notice it happening. Media shorthand for "exponential" as "extremely fast" is a category error; exponential means "multiplicative," full stop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Doubling time, computed, not looked up
&lt;/h2&gt;

&lt;p&gt;Doubling time is how many periods it takes for &lt;code&gt;y&lt;/code&gt; to hit &lt;code&gt;2a&lt;/code&gt;. Solve &lt;code&gt;a * e**(r*t) = 2a&lt;/code&gt; for &lt;code&gt;t&lt;/code&gt; and you get:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;doubling_time&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.05&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.25&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1.00&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; -&amp;gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;doubling_time&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; periods to double&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running that gives:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Growth rate (r)&lt;/th&gt;
&lt;th&gt;Doubling time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1%&lt;/td&gt;
&lt;td&gt;69.31 periods&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5%&lt;/td&gt;
&lt;td&gt;13.86 periods&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;6.93 periods&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;td&gt;2.77 periods&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50%&lt;/td&gt;
&lt;td&gt;1.39 periods&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;0.69 periods&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The "Rule of 70" shortcut (divide 70 by the percentage rate) is just &lt;code&gt;ln(2) ≈ 0.693 ≈ 0.70&lt;/code&gt; rounded for mental math — at 7% growth, &lt;code&gt;70/7 = 10&lt;/code&gt; periods, which matches &lt;code&gt;ln(2)/0.07 = 9.9&lt;/code&gt;. Nothing mystical, just a rounding convenience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the curve has to bend
&lt;/h2&gt;

&lt;p&gt;The exponential model has a hidden assumption: unlimited room to grow. An infected person always finds a susceptible one, a shared post always finds a new viewer. That assumption is false in every bounded system, which is every real system. Once you add a ceiling — call it &lt;code&gt;K&lt;/code&gt;, the carrying capacity — the growth rate has to fall as you approach it. That's the logistic model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;logistic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;n0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;n0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;dN/dt = r * N * (1 - N/K)&lt;/code&gt; is the differential-equation form. The &lt;code&gt;(1 - N/K)&lt;/code&gt; term is a brake: near zero it's close to 1 (so growth looks purely exponential), and as &lt;code&gt;N&lt;/code&gt; approaches &lt;code&gt;K&lt;/code&gt; it drops toward 0 (so growth stops). This is exactly why early-stage growth of almost anything — a pandemic, a viral post, a new product — &lt;em&gt;looks&lt;/em&gt; exponential: locally, before &lt;code&gt;N&lt;/code&gt; is a meaningful fraction of &lt;code&gt;K&lt;/code&gt;, the brake term is doing nothing yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching the two models diverge
&lt;/h2&gt;

&lt;p&gt;Same starting conditions, run both models out over 70 periods:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;n0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.3&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;exponential&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;day&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;70&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;exponential&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;l&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;logistic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;day &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: exponential=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  logistic=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;l&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By day 30 the two curves are still close — the exponential model hasn't broken anything yet. By day 70, the exponential model predicts roughly &lt;code&gt;1.3 * 10**11&lt;/code&gt; — about 130,000 times the entire population you defined as &lt;code&gt;K = 1,000,000&lt;/code&gt;. The logistic model, run with the identical &lt;code&gt;r&lt;/code&gt;, is sitting at roughly 999,992: essentially saturated. Nothing about the underlying process changed between the two models except the brake term. The exponential model isn't wrong because the math is bad; it's wrong because it was never told there's a ceiling.&lt;/p&gt;

&lt;p&gt;You can solve for exactly when the curve bends — the inflection point, where the growth &lt;em&gt;rate&lt;/em&gt; peaks — by setting &lt;code&gt;N = K/2&lt;/code&gt; and solving for &lt;code&gt;t&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;ln&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;K&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;N0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;N0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the numbers above (&lt;code&gt;N0 = 100&lt;/code&gt;, &lt;code&gt;K = 1,000,000&lt;/code&gt;, &lt;code&gt;r = 0.3&lt;/code&gt;), that's &lt;code&gt;ln(9999) / 0.3 ≈ 30.7&lt;/code&gt; days. Before that day, each new period's absolute growth is still accelerating. After it, growth keeps happening but the &lt;em&gt;rate&lt;/em&gt; of growth is falling — even though the raw numbers can still look large for a while.&lt;/p&gt;

&lt;h2&gt;
  
  
  Exponential vs. logistic, side by side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Property&lt;/th&gt;
&lt;th&gt;Exponential&lt;/th&gt;
&lt;th&gt;Logistic&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Formula&lt;/td&gt;
&lt;td&gt;&lt;code&gt;y = a·e^(rt)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;dN/dt = rN(1 − N/K)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Growth rate&lt;/td&gt;
&lt;td&gt;Constant &lt;code&gt;r&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Falls as &lt;code&gt;N → K&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-run shape&lt;/td&gt;
&lt;td&gt;Unbounded (J-curve)&lt;/td&gt;
&lt;td&gt;Saturates at &lt;code&gt;K&lt;/code&gt; (S-curve)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inflection point&lt;/td&gt;
&lt;td&gt;None — always accelerating&lt;/td&gt;
&lt;td&gt;At &lt;code&gt;N = K/2&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Where it's a good model&lt;/td&gt;
&lt;td&gt;Early phase only&lt;/td&gt;
&lt;td&gt;Any bounded process, full lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Why this matters for anything you're tracking
&lt;/h2&gt;

&lt;p&gt;If you've ever fit a trendline to a metric — signups, API calls, weekly active users — and extrapolated it forward, you did the same thing the early-2020 commentators did with case counts. The fit isn't wrong on the data you have; it's wrong about what the data will keep doing, because it silently assumes there's no &lt;code&gt;K&lt;/code&gt;. Two fixes: either fit a logistic curve directly once you have enough data to estimate an inflection, or treat any exponential fit as valid only for forecasting a few periods past your last data point — the part of the S-curve where the brake term genuinely hasn't kicked in yet.&lt;/p&gt;

&lt;p&gt;I worked through the full derivation — plus the R₀/epidemiology framing and the technology-adoption S-curve examples — over at &lt;a href="https://my-blog.org/tangents/post/exponential-growth-vs-viral-spread" rel="noopener noreferrer"&gt;the original post&lt;/a&gt;, if you want the longer version with more worked tables.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Related reading&lt;/strong&gt; — &lt;a href="https://my-blog.org/tangents/post/rice-to-water-ratio-why-not-linear" rel="noopener noreferrer"&gt;Rice-to-Water Ratio: Why You Can't Just Double It&lt;/a&gt;&lt;/p&gt;

</description>
      <category>math</category>
      <category>programming</category>
      <category>datascience</category>
      <category>beginners</category>
    </item>
    <item>
      <title>How to Build a Custom GPT That Doesn't Answer Generically</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Sun, 23 Aug 2026 06:15:05 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/how-to-build-a-custom-gpt-that-doesnt-answer-generically-2lgc</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/how-to-build-a-custom-gpt-that-doesnt-answer-generically-2lgc</guid>
      <description>&lt;p&gt;Building a custom GPT takes about ten minutes. Building one that behaves differently from vanilla ChatGPT takes longer, and most first attempts fail in exactly the same place: the Instructions field.&lt;/p&gt;

&lt;p&gt;The builder UI makes every field look equally important. It isn't. A custom GPT has four components, and one of them carries almost all of the behavioral weight.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqbxo1blqhfl629e3axz7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqbxo1blqhfl629e3axz7.png" alt="How to Build a Custom GPT That Doesn't Answer Generically — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Four Components, Ranked by What They Actually Do
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;th&gt;Behavioral impact&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Name &amp;amp; Description&lt;/td&gt;
&lt;td&gt;Public identity and scope&lt;/td&gt;
&lt;td&gt;Primes vocabulary and register&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Instructions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Persistent system prompt, read before every message&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Highest — role, task, constraints, format&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge&lt;/td&gt;
&lt;td&gt;Uploaded files retrieved during conversation&lt;/td&gt;
&lt;td&gt;High if domain-specific, near zero if generic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conversation Starters&lt;/td&gt;
&lt;td&gt;Four suggested prompts on the opening screen&lt;/td&gt;
&lt;td&gt;Onboarding only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Actions&lt;/td&gt;
&lt;td&gt;API connections to external services&lt;/td&gt;
&lt;td&gt;Powerful, but skip it for v1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you have a fixed amount of time, spend most of it on Instructions. Everything else compounds from there — a sharp Name helps, but it cannot rescue a vague system prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Most Instructions Fail
&lt;/h2&gt;

&lt;p&gt;The typical first draft looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a helpful marketing assistant. Give good advice about marketing.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every word in that prompt is already true of the base model. It adds no constraints, so it changes no behavior. You get generic output because you asked for generic output.&lt;/p&gt;

&lt;p&gt;A prompt that actually changes behavior has four elements. Role, Context, Task, Format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a senior B2B SaaS marketing strategist who has taken
companies from $500K to $5M ARR.

(Context) Users are founders at pre-revenue or Series A stage, building
outbound and content systems with small teams. They know GTM basics and
need executional clarity, not definitions.

(Task) When asked for positioning or messaging, generate three distinct
angles with a two-sentence rationale each. When asked for content,
prioritize pipeline-generating formats: case studies, comparison pages,
sales sequences. Do not give generic advice. Be specific and opinionated.

(Format) Use clear headers. Lead with the most important point. Flag any
assumption you are making about the user's context.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference is not length. It is that every clause narrows the space of acceptable answers. &lt;code&gt;Do not give generic advice&lt;/code&gt; is weak on its own; it works here because the surrounding clauses define what non-generic looks like for this GPT.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Negative Constraint Most Builders Forget
&lt;/h2&gt;

&lt;p&gt;Write down what the GPT must &lt;strong&gt;never&lt;/strong&gt; do before you write what it should do. This one line prevents most of the embarrassing failures:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;If a request falls outside &amp;lt;domain&amp;gt;, say so in one sentence and stop.
Do not attempt an answer from general knowledge.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without it, a specialized GPT quietly degrades into the base model the moment a user drifts off-topic — and the user has no way to tell that happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  Knowledge Files: One at a Time
&lt;/h2&gt;

&lt;p&gt;Uploading ten documents at once and hoping retrieval works is the second most common failure. Add one file, open Preview, and ask a question that can only be answered from that file. If the answer doesn't cite it, the file is not being retrieved and adding nine more won't help.&lt;/p&gt;

&lt;p&gt;Generic files hurt more than they help. A PDF of publicly available material gives the model nothing it doesn't already have, while adding retrieval noise. Domain-specific material — your pricing sheet, your style guide, your internal runbook — is where Knowledge earns its place.&lt;/p&gt;

&lt;h2&gt;
  
  
  An Eight-Step Build That Works
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Define the job on paper.&lt;/strong&gt; Who it is for, the one thing it does better than vanilla ChatGPT, and what it must never do.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open the builder&lt;/strong&gt; — Explore GPTs → Create → switch to the &lt;strong&gt;Configure&lt;/strong&gt; tab. The conversational Create tab is faster but gives you less control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name and Description.&lt;/strong&gt; Apply the specificity test: can a reader predict the exact output without guessing? "B2B Marketing Advisor for SaaS Founders" passes. "Marketing Assistant" does not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write Instructions&lt;/strong&gt; using Role + Context + Task + Format. This is the step worth iterating on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upload Knowledge files one at a time&lt;/strong&gt;, verifying retrieval in Preview after each.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write four Conversation Starters.&lt;/strong&gt; Each should produce something useful immediately, with no follow-up setup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test with edge cases.&lt;/strong&gt; Ask one off-topic question and one thing the Instructions explicitly prohibit. If it answers either, the constraints aren't holding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set publication level and save.&lt;/strong&gt; Default to "Only me" for v1. You can widen it later; you cannot un-share a bad first impression.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Test That Tells You It Worked
&lt;/h2&gt;

&lt;p&gt;Open a normal ChatGPT window and your custom GPT side by side. Send the same prompt to both.&lt;/p&gt;

&lt;p&gt;If the answers are similar, your Instructions aren't doing anything yet — go back to step 4. If the custom GPT is noticeably narrower, more opinionated, or formatted differently, the configuration is holding.&lt;/p&gt;

&lt;p&gt;That comparison takes thirty seconds and is the only reliable signal. Everything else is guesswork about a system prompt you can't see the effects of.&lt;/p&gt;




&lt;p&gt;I wrote a longer version with copy-ready Instruction templates and the full component breakdown here: &lt;a href="https://my-blog.org/tangents/post/how-to-build-a-custom-gpt" rel="noopener noreferrer"&gt;How to Build a Custom GPT: Step-by-Step Guide&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Related reading&lt;/strong&gt; — &lt;a href="https://my-blog.org/tangents/post/how-to-get-specific-answers-from-ai" rel="noopener noreferrer"&gt;How to Get Specific Answers From AI (Instead of Vague Ones)&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://platform.openai.com/docs" rel="noopener noreferrer"&gt;OpenAI Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org" rel="noopener noreferrer"&gt;arXiv&lt;/a&gt; — preprints&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>tutorial</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Which AI Model Should You Use? A Routing Guide for Developers</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Mon, 17 Aug 2026 17:20:01 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/which-ai-model-should-you-use-a-routing-guide-for-developers-og2</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/which-ai-model-should-you-use-a-routing-guide-for-developers-og2</guid>
      <description>&lt;p&gt;You've got ChatGPT open in one tab, Claude in another, maybe Gemini in a third. And you still spend the first thirty seconds of every session wondering which one to actually use for the task in front of you.&lt;/p&gt;

&lt;p&gt;That hesitation is the real problem, not the tools themselves. Each major model has a domain where it genuinely outperforms the others. The skill isn't picking a favorite — it's routing the right task to the right tool.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fidqfxkwwsxyk8xuj227o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fidqfxkwwsxyk8xuj227o.png" alt="Which AI Model Should You Use? A Routing Guide for Developers — at a glance" width="799" height="531"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Stop Asking "Which AI Is Best"
&lt;/h2&gt;

&lt;p&gt;The question has no useful answer because it conflates several fundamentally different design priorities. ChatGPT's reasoning-tier models (o-series) are optimized for breadth and hard logic. Claude is optimized for long-context fidelity and careful writing. Gemini is built around Google's data stack and native multimodal input. Perplexity is a real-time search layer sitting on top of models rather than a general-purpose assistant. They aren't competing in the same lane.&lt;/p&gt;

&lt;p&gt;The better question: what am I trying to do, and which tool is architecturally suited for it? Think routing, not ranking — the same way a skilled team doesn't debate "email or Slack" in the abstract, they use the right channel for the message.&lt;/p&gt;

&lt;p&gt;Four task categories cover most of the routing decision:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Writing and editing&lt;/strong&gt; — prose quality, tone fidelity, long-form coherence&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coding and technical work&lt;/strong&gt; — accuracy, debug quality, explanation clarity&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Research and factual queries&lt;/strong&gt; — recency, citations, hallucination rate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reasoning and analysis&lt;/strong&gt; — multi-step logic, structured thinking&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Each Model Actually Does Best
&lt;/h2&gt;

&lt;p&gt;A note before the list: specific model names go stale within months, so this section describes &lt;strong&gt;families and tiers&lt;/strong&gt;, not version numbers. The routing logic is the durable part — check each provider's current lineup for which model sits in which tier today.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT (flagship tier + dedicated reasoning tier)&lt;/strong&gt;&lt;br&gt;
Strengths: breadth, tool integrations (image generation, code interpreter, browsing), the strongest dedicated reasoning models for hard logic and math, mature voice mode, a large custom-GPT ecosystem. Weaknesses: tends toward verbosity, less sharp on nuanced writing tone, reasoning-tier models cost more to run at volume. Best for: reasoning-heavy tasks, image generation, math and science problem-solving, or when you want one integrated tool ecosystem instead of several apps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude (Sonnet / Opus)&lt;/strong&gt;&lt;br&gt;
Strengths: prose quality and tone fidelity that consistently edges out the alternatives, strong instruction-following, very capable on long documents given its large context window, lower "eager-to-please" drift on writing tasks. Weaknesses: image generation isn't built in, and the default tool ecosystem is smaller than ChatGPT's. Best for: long-form writing, editing, summarizing large documents, and any task where following instructions precisely matters more than raw creativity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gemini (Pro / Flash tiers)&lt;/strong&gt;&lt;br&gt;
Strengths: native multimodal input (image, video, audio), tight Google Workspace integration, a large context window, strong on tasks involving visual data. Weaknesses: writing tone less consistent than Claude. On the hardest reasoning problems the ranking between providers moves with every release — treat any published leaderboard as a snapshot and test on your own task. Best for: work inside Docs/Sheets/Gmail, analyzing images or video, and processing long documents when Google's tools are already part of your workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perplexity&lt;/strong&gt;&lt;br&gt;
Strengths: real-time web search with cited sources, the fastest path to current information, Pro Search decomposes complex queries into sub-searches. Weaknesses: shallow analytical depth, brief outputs, not built for long-form writing or multi-step reasoning. Best for: current events, fact-checking against live sources, and fast bibliography building. See the direct &lt;a href="https://my-blog.org/tangents/post/chatgpt-vs-perplexity-comparison" rel="noopener noreferrer"&gt;ChatGPT vs Perplexity comparison&lt;/a&gt; if search is your main use case.&lt;/p&gt;
&lt;h2&gt;
  
  
  A Quick Decision Flowchart
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Need info from the last 30 days?          → YES: Perplexity
                                            → NO: continue
Primarily a writing or editing task?       → YES: Claude first
                                            → NO: continue
Coding, math, or logic?                    → YES: ChatGPT (reasoning model for hard problems) or Claude
                                            → NO: continue
Need image generation?                     → YES: ChatGPT or Gemini
                                            → NO: continue
Inside Google Workspace?                   → YES: Gemini
                                            → NO: default to ChatGPT or Claude by writing-quality need
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;For coding specifically, the calculus differs enough between "generate new code" and "debug this stack trace" that it's worth a dedicated comparison — see &lt;a href="https://my-blog.org/tangents/post/best-ai-coding-assistants" rel="noopener noreferrer"&gt;best AI coding assistants&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  Two Copy-Ready Prompts You Can Use Today
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Long-form report draft (Claude)&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a senior analyst with expertise in [industry].
(Context) I'm writing a [word count] report on [topic] for [audience].
The goal is to inform a decision about [specific decision].
(Task) Write a structured first draft: an executive summary (3-4
sentences), 3-4 main sections with headers, and a conclusion with a
clear recommendation.
(Format) Professional prose. Bullet points only for lists of 4+ items.
Flag any claim that needs external verification.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Code debugging (ChatGPT or Claude)&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a senior [language] developer.
(Context) This function is supposed to [describe intent]. It returns
[wrong output] when [input condition].
(Task) Identify the bug, explain why it occurs, and provide the
corrected code.
(Format) One-sentence diagnosis first, then the corrected code block,
then a brief explanation of what was wrong. List any additional edge
cases separately.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both follow the same structure: role, context, task, format. That structure is what makes a prompt portable between models — you can hand the same skeleton to Claude, ChatGPT, or Gemini and get a comparably well-scoped answer back.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Which AI model is best overall?&lt;/strong&gt;&lt;br&gt;
No single model leads across every task. Reasoning-tier ChatGPT models are strongest for hard math and logic, Claude leads on writing quality and long-context fidelity, Perplexity is the best tool for current information. "Best for what" is the only version of the question with a useful answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Claude better than ChatGPT for writing?&lt;/strong&gt;&lt;br&gt;
For most writing tasks — essays, reports, editing — Claude tends to produce noticeably better prose and follows tone instructions more precisely, without the verbose, hedge-heavy default style ChatGPT can fall into. For open-ended creative brainstorming, the gap narrows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which is best for coding?&lt;/strong&gt;&lt;br&gt;
Claude and ChatGPT are closely matched on most coding tasks. Claude tends to edge ahead on explaining code clearly and following precise specs. ChatGPT's sandboxed code interpreter for data analysis has no direct equivalent in Claude's standard interface — that's a real differentiator if your coding work involves running and inspecting data, not just writing it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I just use one AI for everything?&lt;/strong&gt;&lt;br&gt;
You can, but you'll leave quality on the table in specific spots — using a general chat model for writing tends to produce verbose prose, and using a model without live search for current events risks stale or hallucinated facts. A two-tool setup, one model for writing and analysis plus a search-native tool for current information, covers most professional work well.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which AI gives the most accurate, up-to-date information?&lt;/strong&gt;&lt;br&gt;
Perplexity, by design — it retrieves live web results and cites sources. Among the non-search models, the accuracy gap on well-established topics is small; the real risk across all of them is hallucination on specific statistics or citations, so verify anything load-bearing regardless of which model produced it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://my-blog.org/tangents/post/which-ai-model-should-you-use" rel="noopener noreferrer"&gt;my-blog.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://platform.openai.com/docs" rel="noopener noreferrer"&gt;OpenAI Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org" rel="noopener noreferrer"&gt;arXiv&lt;/a&gt; — preprints&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>productivity</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Best AI Agents in 2026: What They Do and How to Pick One</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Mon, 17 Aug 2026 17:19:25 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/best-ai-agents-in-2026-what-they-do-and-how-to-pick-one-25hk</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/best-ai-agents-in-2026-what-they-do-and-how-to-pick-one-25hk</guid>
      <description>&lt;p&gt;"AI agent" went from conference-talk abstraction to mainstream product feature faster than almost any other term in the field. ChatGPT has agents. Claude has agents. Your project management tool probably has one now too. But the word covers so much — from a simple tool-calling wrapper to a fully autonomous coding system running in a loop for hours — that it's nearly meaningless without qualification.&lt;/p&gt;

&lt;p&gt;A working definition that holds up: an AI agent takes a goal, breaks it into steps, executes those steps using tools (search, code execution, file access, API calls), and iterates until the goal is met. The emphasis is on the loop — an agent doesn't just respond, it acts and corrects. That loop is the capability gap separating agents from &lt;a href="https://my-blog.org/tangents/post/ai-agents-vs-chatbots" rel="noopener noreferrer"&gt;conventional chatbots&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The space now spans purpose-built coding agents, research agents, workflow automation agents, general-purpose operator-style agents, and multi-agent frameworks where specialists hand off work to each other. Pricing and benchmarks here change fast enough that hard numbers go stale quickly — so this guide sticks to what's stable: architecture, autonomy level, and fit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1jnobsnz3hl86714062y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1jnobsnz3hl86714062y.png" alt="Best AI Agents in 2026: What They Do and How to Pick One — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Loop That Defines an Agent
&lt;/h2&gt;

&lt;p&gt;A chatbot processes input and generates output in one pass. An agent runs:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Observe&lt;/strong&gt; — receive the goal and current state&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan&lt;/strong&gt; — determine the next action&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Act&lt;/strong&gt; — execute it through a tool: run code, search, write a file, call an API&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observe&lt;/strong&gt; — check the result&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repeat&lt;/strong&gt; until the goal is met or it's stopped&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Agents vary a lot in how autonomous this loop is. Some require human approval before each step ("human-in-the-loop"); others run unattended for long stretches. Autonomy level is one of the most important dimensions when evaluating an agent for a use case — more autonomy means more capability, but also more risk of errors compounding uncorrected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five Categories Worth Knowing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Coding agents&lt;/strong&gt; write, run, debug, and iterate on code in a real execution environment — they run it and fix what's broken, not just suggest it. Examples: Claude Code, Devin, SWE-agent, OpenHands. Autonomy runs medium to high; most gate high-risk actions like deleting files or pushing to main behind approval. Best fit: teams with repetitive implementation work, bug fixing, and test generation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Research and information agents&lt;/strong&gt; search, synthesize, and produce structured output from multiple sources, reasoning across them rather than just retrieving them. Examples: Perplexity Deep Research, ChatGPT Deep Research, Gemini Deep Research. Autonomy is low to medium — search, read, synthesize, then human review. Best fit: literature reviews, competitive analysis, due diligence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Workflow automation agents&lt;/strong&gt; connect to business tools — Slack, Gmail, Notion, Salesforce — and automate multi-step workflows triggered by events. Examples: Zapier AI, n8n with AI nodes, Copilot Studio agents. High autonomy within a defined spec. Best fit: repetitive processes touching multiple apps — lead routing, meeting follow-ups, data pipeline maintenance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;General-purpose operator-style agents&lt;/strong&gt; sit on top of frontier models and handle a wide range of tasks with a flexible tool set — generalists, not specialists. Examples: ChatGPT with tools, Claude with MCP, Gemini with Workspace extensions. Autonomy is whatever you configure, usually gated on consequential actions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-agent orchestration frameworks&lt;/strong&gt; run at the infrastructure level: a planner agent breaks a large task into subtasks and delegates to specialists — a coder, a researcher, a writer. Examples: LangGraph, AutoGen, CrewAI. Very high autonomy in principle, but reliability drops as complexity climbs. Best fit: teams building AI products where one model's context can't cover the whole task.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Main Tools Are Actually Good At
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt; is a terminal-native coding agent that reads files, writes code, runs shell commands, and observes results in a loop. Its edge is reasoning quality on non-obvious problems — debugging race conditions, rearchitecting a module, building a feature from a description — where genuine reasoning beats pattern-matching. It extends tool access through &lt;a href="https://my-blog.org/tangents/post/what-is-mcp-model-context-protocol" rel="noopener noreferrer"&gt;MCP&lt;/a&gt;, worth understanding if you're configuring any agent's tool layer. Trade-off: CLI-only, no GUI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT's agent capabilities&lt;/strong&gt; come in two forms — tool-augmented chat (search, code interpreter, file access) and Operator mode for UI navigation. Its Deep Research mode produces strong long-form cited reports. Weak spot: no persistent memory across sessions by default, and less reliable tool-calling on complex multi-step tasks than purpose-built coding agents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perplexity Deep Research&lt;/strong&gt; is a read-only synthesis agent — it searches many sources, reads them, and produces a cited report. It beats manual search-and-read for competitive analysis or fact-gathering, but doesn't write code, execute workflows, or act beyond producing the report.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Devin&lt;/strong&gt; was the first widely publicized fully autonomous software engineering agent, handling defined implementation tasks — features from specs, bug fixes, tests — with a GUI showing its terminal, browser, and editor live. Honest caveat: independent evaluations on open-ended tasks show more variable results than early claims suggested. It performs best on well-scoped tasks with clear acceptance criteria.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zapier AI&lt;/strong&gt; lets non-technical users build agents that watch triggers (new emails, form submissions, Slack messages) and run workflows across thousands of connected apps, set up conversationally. Its strength is breadth of integrations; its limitation is logic that stays rigid once configured, rather than reasoning through ambiguity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenHands and SWE-agent&lt;/strong&gt; are the open-source options for teams avoiding SaaS pricing or sending code to a third-party API — at the cost of setup complexity and bringing your own model API.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Routing Framework for Choosing
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Task type&lt;/strong&gt; — coding, research, workflow automation, or general-purpose?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Autonomy tolerance&lt;/strong&gt; — how comfortable are you with unapproved actions?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Technical complexity&lt;/strong&gt; — reasoning about a complex system, or a well-defined process?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interface preference&lt;/strong&gt; — CLI, browser, or no-code?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost tolerance&lt;/strong&gt; — free/self-hosted through paid consumer and enterprise tiers.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For coding agents specifically, see the &lt;a href="https://my-blog.org/tangents/post/best-ai-coding-assistants" rel="noopener noreferrer"&gt;detailed coding assistant comparison&lt;/a&gt;; for the models powering these agents, see &lt;a href="https://my-blog.org/tangents/post/which-ai-model-should-you-use" rel="noopener noreferrer"&gt;which AI model to use&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A copy-ready template for prompting any agent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GOAL: [observable definition of "done"]
CONSTRAINTS: [what the agent should NOT do]
RESOURCES: [tools and access it has]
ON FAILURE: [what to do if a step fails — retry, flag, or stop]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What Agents Still Can't Do
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Error compounding.&lt;/strong&gt; The longer an agent runs autonomously, the more a small early error compounds — a coding agent that misreads a spec on step 2 may have written hundreds of wrong lines by step 20.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool hallucination.&lt;/strong&gt; Agents can call the wrong tool, misread output, or invent a result when a tool fails silently — more consequential than a chatbot hallucinating text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context limits.&lt;/strong&gt; Even with large context windows, long agentic sessions degrade — agents repeat earlier steps or lose track of prior decisions as the window fills.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No inherent judgment about consequences.&lt;/strong&gt; An agent doesn't know deleting a production database differs from deleting a test file unless told explicitly. Human oversight stays essential for anything with real external effects.&lt;/p&gt;

&lt;p&gt;For single-turn tasks — answer a question, draft a paragraph — a chatbot is cheaper, faster, and more predictable. Reserve agents for tasks that genuinely need multiple steps, tool use, or iteration.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What's the core difference between an agent and a chatbot?&lt;/strong&gt;&lt;br&gt;
A chatbot responds to a prompt in a single pass. An agent runs a loop — observe, plan, act, repeat — using tools until the goal is reached. Agents do things, not just describe them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which agent is best for coding tasks?&lt;/strong&gt;&lt;br&gt;
Claude Code and Devin are the strongest purpose-built options. Claude Code tends to win on complex, multi-file reasoning; Devin offers a more polished GUI and no terminal requirement. For lighter in-IDE help, a coding assistant fits better than a full agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are AI agents safe for business workflows?&lt;/strong&gt;&lt;br&gt;
With guardrails, yes. Best practice is human-in-the-loop for consequential actions — the agent proposes, you approve. Reserve full autonomy for low-risk, reversible tasks, and treat agent output like work from a capable but fallible junior teammate.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://my-blog.org/tangents/post/best-ai-agents" rel="noopener noreferrer"&gt;my-blog.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://platform.openai.com/docs" rel="noopener noreferrer"&gt;OpenAI Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.anthropic.com" rel="noopener noreferrer"&gt;Anthropic Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>AI Agents vs Chatbots: What's the Actual Difference</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Mon, 17 Aug 2026 17:18:49 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/ai-agents-vs-chatbots-whats-the-actual-difference-546c</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/ai-agents-vs-chatbots-whats-the-actual-difference-546c</guid>
      <description>&lt;p&gt;Everyone's talking about AI agents right now. But most people using ChatGPT day to day are actually talking to a chatbot, and the difference is more than semantic. Conflating the two leads to mis-set expectations: you're disappointed when your "agent" can't follow through on a plan, or confused when an agent does things you didn't explicitly ask for.&lt;/p&gt;

&lt;p&gt;The distinction is practical enough to matter for how you build with either one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvq2ix0slpoeabkwmbhpb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvq2ix0slpoeabkwmbhpb.png" alt="AI Agents vs Chatbots: What's the Actual Difference — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Difference in One Sentence
&lt;/h2&gt;

&lt;p&gt;A chatbot takes your input and returns a response. An &lt;a href="https://my-blog.org/tangents/post/what-is-an-ai-agent" rel="noopener noreferrer"&gt;AI agent&lt;/a&gt; takes your goal and figures out how to achieve it.&lt;/p&gt;

&lt;p&gt;Four variables separate them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Autonomy&lt;/strong&gt; — who decides the next step? Chatbot: you, every time. Agent: it decides.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool use&lt;/strong&gt; — can it call APIs, run code, read/write files? Chatbot: rarely, and only on request. Agent: this is central to how it works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-step execution&lt;/strong&gt; — does it run a plan across several steps without a prompt at each one? Chatbot: no. Agent: yes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory&lt;/strong&gt; — does it carry state across steps and sessions? Chatbot: conversation window only. Agent: can use persistent external memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What a Chatbot Actually Is
&lt;/h2&gt;

&lt;p&gt;A chatbot is built around conversational exchange: you send a message, it sends a reply. That's the whole unit of interaction. Modern chatbots (ChatGPT, Claude, Gemini in plain conversational mode) are extraordinarily capable within that unit — they write, analyze, reason, translate, and summarize well. But the design stays reactive: nothing happens unless you send the next message.&lt;/p&gt;

&lt;p&gt;Key properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Stateless or limited context (this conversation, not across sessions, unless the product bolts on memory)&lt;/li&gt;
&lt;li&gt;No autonomous tool use — you ask, it acts, once&lt;/li&gt;
&lt;li&gt;One input, one output, repeat&lt;/li&gt;
&lt;li&gt;Can't book, send, modify, or execute anything on your behalf without you prompting each step&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The chatbot isn't "dumb" — the underlying model is often the same one powering an agent. The constraint is the interaction design, not the model's capability.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI Agent Actually Is
&lt;/h2&gt;

&lt;p&gt;An agent receives a goal and autonomously works out the sequence of steps to reach it — calling tools along the way, checking progress, adjusting based on what it finds. Four things distinguish it from a chatbot:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Goal-directed behavior.&lt;/strong&gt; Instead of responding to one prompt, it works toward an end state. "Summarize last month's sales and email the report to the team" is a goal it decomposes and executes in order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool use.&lt;/strong&gt; Real external systems — web search, code execution, databases, calendar, email, CRM. Tools are what let an agent act, not just describe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-step planning and execution.&lt;/strong&gt; It runs a loop — perceive, plan, act, check, repeat — until the goal is met or a stopping condition hits, without needing a human prompt between steps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory and state management.&lt;/strong&gt; It can track what it's done, what it found, and what's left, often across sessions via external storage.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Spectrum, Not a Binary
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pure chatbot&lt;/strong&gt; — single turn, no tools, no memory. An FAQ bot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversational assistant&lt;/strong&gt; — multi-turn context, some on-request tools (image generation, web search when asked), but you drive every step. ChatGPT in a normal chat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Light agent&lt;/strong&gt; — chains two to five steps, limited tool use, still surfaces output to you rather than acting fully in the world. A lot of "Copilot" features live here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full agent&lt;/strong&gt; — multi-step autonomous execution with real tool calls, human checkpoints before irreversible actions, persistent memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most people's day-to-day "agent" experience in 2026 sits in the light-to-medium range. Fully autonomous agents running unattended in production exist but need real deployment guardrails.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same Request, Two Different Outcomes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Request:&lt;/strong&gt; "Research the top 3 AI coding assistants, compare pricing, and send me a summary by email."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Chatbot:&lt;/strong&gt; Writes a description of three tools with pricing. Stops. Doesn't send anything — you copy, paste, and send it yourself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent:&lt;/strong&gt; Searches the web for current pricing, opens each pricing page, extracts the data, formats a comparison, drafts an email, and sends it — reporting completion when done (possibly pausing to confirm the recipient first).&lt;/p&gt;

&lt;p&gt;Neither output is "bad." The chatbot's prose might be excellent. But only the agent actually finished the task end to end.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Reach for Which
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Use a chatbot when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You need one well-formed output — a draft, an analysis, an explanation&lt;/li&gt;
&lt;li&gt;You want to stay in control of every step&lt;/li&gt;
&lt;li&gt;The task doesn't require real-world actions&lt;/li&gt;
&lt;li&gt;You're iterating and want to steer each turn&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use an agent when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The task has multiple defined steps that would be tedious to prompt one by one&lt;/li&gt;
&lt;li&gt;You need real external actions — search, book, send, update&lt;/li&gt;
&lt;li&gt;There's a clear, verifiable end state ("report sent," "ticket closed")&lt;/li&gt;
&lt;li&gt;You're comfortable delegating with checkpoints&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Directing an Agent Is a Different Skill Than Prompting a Chatbot
&lt;/h2&gt;

&lt;p&gt;When prompting a chatbot, you guide each step explicitly. When directing an agent, you define the goal, the constraints, and the checkpoints up front. A reusable four-element format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) What the agent has access to and can do
(Context) Background, constraints, what "done" looks like
(Task) The end state you want — not just the first step
(Format) How it should check in: confirm before irreversible actions, report at milestones
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Writing a goal this way front-loads the thinking an agent needs instead of trying to catch problems mid-run.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can ChatGPT be used as an AI agent?&lt;/strong&gt;&lt;br&gt;
Yes, in its tool-augmented modes — browsing, code execution, and Operator-style UI navigation move it toward agent behavior. Plain conversational ChatGPT without those tools is still a chatbot by this definition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What tools do AI agents typically use?&lt;/strong&gt;&lt;br&gt;
Web search, code interpreters, file read/write, and APIs for calendar, email, CRM, and payment systems are the common set. The tools available are what actually give an agent the ability to act, not just respond.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are AI agents riskier to use than chatbots?&lt;/strong&gt;&lt;br&gt;
Yes, proportionally to autonomy. A chatbot's worst-case output is bad text you can ignore. An agent's worst-case output is an irreversible action taken on your behalf. That's why human-in-the-loop checkpoints before consequential actions matter more as autonomy increases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I know if I need a chatbot or an agent for a given task?&lt;/strong&gt;&lt;br&gt;
If the task ends with you reading and using an output yourself, a chatbot is enough and will be faster and more predictable. If the task ends with something actually happening in another system — an email sent, a record updated — you need an agent with the right tool connected.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://my-blog.org/tangents/post/ai-agents-vs-chatbots" rel="noopener noreferrer"&gt;my-blog.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://platform.openai.com/docs" rel="noopener noreferrer"&gt;OpenAI Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org" rel="noopener noreferrer"&gt;arXiv&lt;/a&gt; — preprints&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>beginners</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How to Build an AI Agent Without Writing Code</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Mon, 17 Aug 2026 17:18:13 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/how-to-build-an-ai-agent-without-writing-code-42an</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/how-to-build-an-ai-agent-without-writing-code-42an</guid>
      <description>&lt;p&gt;If you've spent much time in ChatGPT, you've hit the wall: type a prompt, get an answer, type another prompt, get another answer. At some point you think, "can't this just do the whole thing?" That itch is what makes people want to build an AI agent.&lt;/p&gt;

&lt;p&gt;An &lt;a href="https://my-blog.org/tangents/post/what-is-an-ai-agent" rel="noopener noreferrer"&gt;AI agent&lt;/a&gt; is a system that pursues a goal across multiple steps, calls tools, and makes decisions without waiting for a human prompt at each step. The useful news for non-developers: in 2026, building a functional no-code agent is genuinely accessible. Platforms like n8n, Zapier, Make, and Lindy let you wire an LLM to real tools without a line of Python. The less convenient news: "no-code" doesn't mean "no thinking." The actual bottleneck in almost every agent build isn't the platform — it's the quality of the goal definition, the tool selection, and the testing loop.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frosox3fsdkmhaaqgi07d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frosox3fsdkmhaaqgi07d.png" alt="How to Build an AI Agent Without Writing Code — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Four Components Every No-Code Agent Needs
&lt;/h2&gt;

&lt;p&gt;Building a no-code agent means connecting four things. You don't write the model — you configure what it can access, what it should do, and what success looks like.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Goal statement (system prompt).&lt;/strong&gt; The instruction set that tells the LLM what it's trying to accomplish, what constraints to respect, and what "done" looks like. This is the highest-leverage piece — a vague system prompt produces an unpredictable agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. LLM brain.&lt;/strong&gt; The model doing the reasoning — GPT-4o, Claude, Gemini, or a local model via Ollama. Most platforms let you pick from several. The model determines reasoning quality; the platform determines which tools it can reach.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Tool layer.&lt;/strong&gt; What the agent can actually do: send email, read a spreadsheet, query a database, search the web, post to Slack, call an API. Each tool is a discrete capability the agent can invoke inside its reasoning loop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Memory / context layer.&lt;/strong&gt; How the agent retains information across steps or sessions — from "pass the last five turns" to "query a vector database for prior interactions."&lt;/p&gt;

&lt;p&gt;You don't need all four maxed out. Most effective first agents use a precise system prompt, two or three tools, and stateless context — start there, and add complexity only when you hit a real limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Define the Goal Before You Touch a Tool
&lt;/h2&gt;

&lt;p&gt;The most common mistake non-developers make is starting with the tool — opening n8n, dragging nodes, and only later realizing the agent doesn't know what it's supposed to do. Goal definition comes first, always.&lt;/p&gt;

&lt;p&gt;A usable goal statement has four properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Specific&lt;/strong&gt; — not "help with emails" but "triage incoming support emails: label, draft a reply, flag anything needing human review."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bounded&lt;/strong&gt; — what the agent is explicitly not allowed to do. "Do not send any email without a human review step."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Success-defined&lt;/strong&gt; — what "done" looks like, concretely. "A labeled email plus a draft reply in the Drafts folder, and a Slack ping if the email needs a human."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalation-aware&lt;/strong&gt; — what happens when the agent hits a case it can't handle. "Flag for a human and stop. Do not guess."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's a copy-ready template you can drop straight into a system prompt field:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GOAL: [specific outcome, one sentence]
INPUTS: [what triggers the agent — new email, form submit, schedule]
ALLOWED ACTIONS: [tools it may call]
NOT ALLOWED: [explicit boundaries — no sending, no deleting, no external posts]
DONE WHEN: [observable success condition]
IF STUCK: flag for human review and stop. Do not guess.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 2: Connect Tools — Pick the Fewest That Work
&lt;/h2&gt;

&lt;p&gt;Tool selection is where the platforms diverge most:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;n8n&lt;/strong&gt; — open-source, self-host or cloud, 400+ integrations, a code node for custom logic. Steepest curve, most flexibility. Good for technical non-developers who want control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zapier&lt;/strong&gt; — cloud, easiest entry, thousands of app integrations plus AI Actions. Less flexible for branching logic. Good for simple linear tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make&lt;/strong&gt; — visual builder with strong branching and looping. A middle ground between Zapier and n8n.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lindy&lt;/strong&gt; — purpose-built for AI agents, memory and human-in-the-loop steps included natively. The lowest-configuration path to real agent behavior.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The operating principle: pick the fewest tools that complete the goal. Every additional connection is another failure point. An agent that can send email, read a spreadsheet, and search the web already covers a wide range of real tasks without twenty integrations bolted on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Add Memory Only When Statelessness Actually Hurts
&lt;/h2&gt;

&lt;p&gt;Most beginner agents are stateless — each run starts fresh. That's fine for one-shot tasks but breaks down for recurring workflows like weekly reports or ongoing customer threads, where the agent needs to carry context forward.&lt;/p&gt;

&lt;p&gt;Memory options, roughly in order of complexity:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Session context&lt;/strong&gt; — pass the last N messages in the prompt. Every platform supports this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;External storage&lt;/strong&gt; (a spreadsheet, Notion, Airtable) — write key facts after each run, read them at the start of the next. Low-tech, effective.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Built-in memory layer&lt;/strong&gt; (Lindy, some n8n setups) — you configure what gets saved; the platform retrieves relevant memories automatically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vector database&lt;/strong&gt; (Pinecone, Supabase vector, Weaviate) — for agents that search across large bodies of text. Overkill for most first builds.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Start with session context. Add external storage once you actually need cross-run memory. Leave vector databases until the simpler version has run long enough to show you what's missing. If you're wiring an agent to pull from external data sources at all, it's worth understanding &lt;a href="https://my-blog.org/tangents/post/what-is-mcp-model-context-protocol" rel="noopener noreferrer"&gt;what MCP is&lt;/a&gt; first — it's the standard a growing number of these integrations are built on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Test for the Failure Modes That Actually Happen
&lt;/h2&gt;

&lt;p&gt;No-code agents fail in predictable ways, not random ones:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Prompt drift&lt;/strong&gt; — the agent handles simple cases fine but ignores constraints in edge cases. Fix: add explicit edge-case examples to the system prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool call errors&lt;/strong&gt; — the agent calls a tool with the wrong format, or the app returns something unexpected. Fix: check the platform's error logs, tighten formatting instructions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Looping&lt;/strong&gt; — the agent gets stuck redoing the same step. Fix: cap the max step count, and make sure the success criterion is actually checkable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Overconfident execution&lt;/strong&gt; — the agent takes an irreversible action (sends an email, posts publicly) when it should have paused. Fix: insert a human-approval step before anything irreversible.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A minimal test checklist before you trust an agent with anything real:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Typical input — does it do the right thing?&lt;/li&gt;
&lt;li&gt;[ ] Boundary input (edge case, empty input, ambiguous request) — does it fail gracefully?&lt;/li&gt;
&lt;li&gt;[ ] Adversarial input, if it reads external content — does it stay in bounds?&lt;/li&gt;
&lt;li&gt;[ ] Three consecutive runs — does behavior stay consistent?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Human-in-the-loop isn't a failure mode to eliminate — it's a design feature to keep. In n8n and Zapier you build it explicitly: a "wait for approval" step before any send, post, or record change. The agent earns more autonomy as its tested boundaries hold up.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do I need to code to build an AI agent?&lt;/strong&gt;&lt;br&gt;
No. n8n, Zapier, Make, and Lindy all support building functional agents through visual interfaces and configuration forms. You do need to write a clear goal statement — that's thinking work, not code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the best platform if I've never built an agent before?&lt;/strong&gt;&lt;br&gt;
Lindy is the most beginner-friendly purpose-built option — memory and human-in-the-loop are configured visually. Zapier is a safe start if you already know it. n8n gives you the most power at the cost of a steeper curve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What can a no-code agent actually do reliably?&lt;/strong&gt;&lt;br&gt;
Email triage with draft replies, weekly report generation from spreadsheet data, first-response customer support, and meeting summarization with follow-ups. More consequential tasks — booking, finances — need careful testing and human checkpoints.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://my-blog.org/tangents/post/how-to-build-an-ai-agent" rel="noopener noreferrer"&gt;my-blog.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://platform.openai.com/docs" rel="noopener noreferrer"&gt;OpenAI Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.anthropic.com" rel="noopener noreferrer"&gt;Anthropic Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>tutorial</category>
      <category>beginners</category>
    </item>
    <item>
      <title>System Prompt vs User Prompt: What's the Difference?</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Mon, 17 Aug 2026 17:17:29 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/system-prompt-vs-user-prompt-whats-the-difference-9po</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/system-prompt-vs-user-prompt-whats-the-difference-9po</guid>
      <description>&lt;p&gt;Most people who use ChatGPT daily have never seen a system prompt — but it's shaping every answer they get. Understanding the distinction between a system prompt and a user prompt isn't just a technical detail. It's the key to understanding why AI behaves the way it does, and how to actually get it to do what you need.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5eugquztbpe1c1j1gkzz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5eugquztbpe1c1j1gkzz.png" alt="System Prompt vs User Prompt: What's the Difference? — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is a System Prompt?
&lt;/h2&gt;

&lt;p&gt;A system prompt is a set of instructions given to a model before the conversation starts. It's invisible to the end user in most products, but it defines the model's persona, boundaries, tone, task scope, and behavioral rules for the entire session.&lt;/p&gt;

&lt;p&gt;Think of it as the job description handed to an employee before their first shift. The employee — the model — reads it privately, internalizes the rules, and then works within those constraints for the rest of the conversation.&lt;/p&gt;

&lt;p&gt;In most chat completion APIs, the system prompt is passed as a message with a dedicated system role. In practice, it looks something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;role:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;system&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;content:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"You are a concise legal research assistant. You summarize
case law clearly and flag when professional legal advice is required.
Never speculate about legal outcomes. Keep responses under 300 words
unless asked to expand."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That text never appears in the chat window, but everything the user types gets interpreted through it. System prompts typically control: persona and name, tone and formality, knowledge scope (what topics the model will and won't engage with), output format defaults, safety guardrails, and task framing (coding assistant, tutor, analyst, writer).&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is a User Prompt?
&lt;/h2&gt;

&lt;p&gt;A user prompt is what you actually type into the chat box — the message you send each turn. It's the runtime instruction: what you want, right now, in this specific conversation.&lt;/p&gt;

&lt;p&gt;User prompts are temporary. Each message lives in the conversation thread, but once the context window fills up, older messages get dropped, and none of it carries forward into a new session. Good user prompts are specific, concrete, and scoped — they tell the model what to do with the context it already has. They can override some system prompt defaults, like tone, format, or length, if the system prompt allows it, but they can't override hard-coded restrictions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Difference, Side by Side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;System Prompt&lt;/th&gt;
&lt;th&gt;User Prompt&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Who writes it&lt;/td&gt;
&lt;td&gt;Developer / product builder&lt;/td&gt;
&lt;td&gt;End user&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When it's set&lt;/td&gt;
&lt;td&gt;Before the conversation starts&lt;/td&gt;
&lt;td&gt;Each turn, at runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Visibility&lt;/td&gt;
&lt;td&gt;Hidden in most products&lt;/td&gt;
&lt;td&gt;Visible in the chat thread&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scope&lt;/td&gt;
&lt;td&gt;Entire session&lt;/td&gt;
&lt;td&gt;Single turn&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Defines&lt;/td&gt;
&lt;td&gt;Persona, rules, defaults&lt;/td&gt;
&lt;td&gt;Current task and specifics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can the user change it?&lt;/td&gt;
&lt;td&gt;No, in finished products&lt;/td&gt;
&lt;td&gt;Yes, every message is a new one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Priority&lt;/td&gt;
&lt;td&gt;Weighed first by the model&lt;/td&gt;
&lt;td&gt;Interpreted through system context&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Custom Instructions and Custom GPTs
&lt;/h2&gt;

&lt;p&gt;ChatGPT's Custom Instructions feature lets you fill in a system-prompt-like field that gets prepended to every conversation. You write it once in settings and it persists across chats until you change it. It's not exactly a system prompt — the product's own system prompt still sits above it — but functionally it behaves like a persistent personal layer. It's your chance to tell the model your profession, your preferred output format, what you already know, and how you want to be spoken to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"I'm a product manager at a SaaS company. Don't explain basic
business concepts. Always answer in bullet points for complex
topics. When I ask for feedback, be direct — don't soften criticism."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A Custom GPT goes further: it's a packaged product combining a system prompt written in a builder interface, uploaded knowledge files retrieved as context, configured tools (web search, code execution, image generation), and a defined audience. When you use one, you're writing user prompts that get interpreted through someone else's system prompt and knowledge base.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Actually Matters for How You Prompt
&lt;/h2&gt;

&lt;p&gt;If you're a regular user rather than someone building products, understanding this distinction changes how you work with AI in three concrete ways.&lt;/p&gt;

&lt;p&gt;First, it explains weird behavior. If a chatbot on a website keeps saying "I can only help with X," that's not the model being dumb — it's the system prompt restricting it, and no amount of clever phrasing in your user prompt will fully route around it.&lt;/p&gt;

&lt;p&gt;Second, it tells you what you actually control. You can't escape a well-written system prompt by asking nicely, but you usually can override tone and format if the system prompt doesn't explicitly lock them.&lt;/p&gt;

&lt;p&gt;Third, it changes how much work your own prompts need to do. When you use a general-purpose chat interface directly — not a Custom GPT — you're writing into a generic system prompt, which means your user prompt has to specify role, context, task, and format yourself instead of inheriting them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same Question, With and Without a System Prompt
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Without a system prompt&lt;/strong&gt;, a generic assistant answering "How do I get a refund?" gives a 200-word explanation of how refund policies typically work in general — because it has no company-specific context to draw from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With a system prompt in place:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;role:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;system&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;content:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"You are the support agent for StyleBox, a subscription
clothing service. Refund policy: customers have 14 days to return
items in original packaging. Direct refund requests to
support@stylebox.example or the order page at stylebox.example/orders. Do
not speculate about orders you don't have data on. Be warm but concise."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same user question — "How do I get a refund?" — now gets: "For a refund on your StyleBox order, visit stylebox.example/orders or email &lt;a href="mailto:support@stylebox.example"&gt;support@stylebox.example&lt;/a&gt;. Returns are accepted within 14 days in original packaging. Anything else I can help with?" The entire jump in usefulness comes from the system prompt, not the user prompt — the user asked the exact same thing both times.&lt;/p&gt;

&lt;p&gt;If you're the one writing that system prompt, the &lt;a href="https://my-blog.org/tangents/post/role-prompting-explained" rel="noopener noreferrer"&gt;role prompting&lt;/a&gt; breakdown is worth pairing with this — it covers how to write the persona line so it actually shifts the model's output instead of adding noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;p&gt;System prompts are persistent, session-wide, and usually invisible; user prompts are per-turn and visible. In consumer apps, the product's own system prompt shapes every response before yours even arrives. Custom Instructions function as a personal system prompt layer that persists across sessions, and a Custom GPT is a packaged system prompt plus tools plus knowledge. If you want a model to consistently behave a certain way, put the instruction in the system prompt — not in every single user message.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can a user prompt override a system prompt?&lt;/strong&gt;&lt;br&gt;
Only within the boundaries it allows. Tone, format, and length are commonly overridable. Hard-coded restrictions are not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are Custom Instructions the same as a system prompt?&lt;/strong&gt;&lt;br&gt;
Functionally similar but not identical — Custom Instructions sit in front of the product's own underlying system prompt as a personal layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does a website chatbot refuse to answer basic questions?&lt;/strong&gt;&lt;br&gt;
Almost always a narrow system prompt, not a model limitation. It's been scoped to a specific task by whoever built it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need to write system prompts as a regular user?&lt;/strong&gt;&lt;br&gt;
Not directly, unless you're building a product. But understanding how they work explains why default chat behavior varies across apps, and it makes your own user prompts more effective.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://my-blog.org/tangents/post/system-prompt-vs-user-prompt-explained" rel="noopener noreferrer"&gt;my-blog.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://platform.openai.com/docs" rel="noopener noreferrer"&gt;OpenAI Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org" rel="noopener noreferrer"&gt;arXiv&lt;/a&gt; — preprints&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>programming</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Role Prompting Explained: How AI Personas Actually Work</title>
      <dc:creator>Sungwoo Lee</dc:creator>
      <pubDate>Mon, 17 Aug 2026 17:16:53 +0000</pubDate>
      <link>https://dev.to/sungwoo_lee_e0f26be4a29fd/role-prompting-explained-how-ai-personas-actually-work-7m1</link>
      <guid>https://dev.to/sungwoo_lee_e0f26be4a29fd/role-prompting-explained-how-ai-personas-actually-work-7m1</guid>
      <description>&lt;p&gt;Role prompting is the practice of assigning an expert identity or persona to an AI before giving it a task. Instead of a bare request like "explain machine learning," you first tell the model who it is: "You are a machine learning researcher who explains complex topics to business executives." That single addition shifts the register, depth, and vocabulary of everything that follows.&lt;/p&gt;

&lt;p&gt;Role is the first of four components in the standard prompt engineering framework — Role, Context, Task, Format. It's also the most misunderstood. People either skip it entirely or use it too vaguely ("You are a helpful assistant"), which adds no information the model doesn't already assume. Done well, it narrows the model's output toward expert-quality language on a specific topic. Done poorly, it's noise.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhgh5st7wa251avhruahu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhgh5st7wa251avhruahu.png" alt="Role Prompting Explained: How AI Personas Actually Work — at a glance" width="800" height="598"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Role Prompting Actually Works
&lt;/h2&gt;

&lt;p&gt;Language models generate text by predicting the most probable next token given what came before. When you write "You are a senior financial analyst with 15 years of experience in fixed income markets," you've loaded the context with tokens that statistically co-occur with technical vocabulary, structured frameworks, hedged professional language, and domain-specific caution.&lt;/p&gt;

&lt;p&gt;The model doesn't "become" a financial analyst — it generates text that statistically resembles what one would produce in a similar context. That distinction matters for understanding the limits (more on that below), but the practical effect is real: role-primed responses use the vocabulary, structure, and epistemic stance of the assigned domain.&lt;/p&gt;

&lt;p&gt;This is different from hallucination prevention — roles don't stop a model from making things up — and different from RAG, which gives a model access to new facts. Roles reshape style, structure, and specificity, not factual accuracy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing a Role That Actually Moves the Output
&lt;/h2&gt;

&lt;p&gt;The most common mistake is a generic role. "You are an expert" tells the model almost nothing, since "expert" is a near-universal token in AI output already. The more specific the domain, experience level, and perspective, the stronger the shift.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No role&lt;/td&gt;
&lt;td&gt;(none)&lt;/td&gt;
&lt;td&gt;General-purpose defaults&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generic&lt;/td&gt;
&lt;td&gt;"You are an expert"&lt;/td&gt;
&lt;td&gt;Minimal shift&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Domain-specific&lt;/td&gt;
&lt;td&gt;"You are a data scientist"&lt;/td&gt;
&lt;td&gt;Moderate vocabulary shift&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Role + experience&lt;/td&gt;
&lt;td&gt;"...with 10 years in healthcare analytics"&lt;/td&gt;
&lt;td&gt;Stronger, adds domain context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Role + experience + perspective&lt;/td&gt;
&lt;td&gt;"...who translates findings for clinical stakeholders"&lt;/td&gt;
&lt;td&gt;Full shift — constrains content and communication style&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Role works best as one element of a complete prompt, not a standalone fix. The full framework: &lt;strong&gt;(Role)&lt;/strong&gt; who the AI is, &lt;strong&gt;(Context)&lt;/strong&gt; the situation — audience, constraints, background, &lt;strong&gt;(Task)&lt;/strong&gt; what specifically you want done, &lt;strong&gt;(Format)&lt;/strong&gt; how the output should be structured. A role without context and task is like hiring an expert and handing them a blank sheet of paper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before and After: Role Prompting in Practice
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Without a role&lt;/strong&gt;, "Explain bond duration to me" produces a generic, encyclopedia-style explanation — correct but flat, with no assumed audience or expert judgment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With a strong, matched role:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a fixed income portfolio manager who briefs institutional investors.
(Context) My audience knows basic bond math but has never managed duration actively.
(Task) Explain duration and why it matters when interest rates move.
(Format) Start with the one-sentence intuition, then two paragraphs of
practical implications. End with a common misconception to avoid.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This produces investor-grade language, focuses on rate sensitivity and portfolio impact instead of formula derivation, and flags the common misconception of duration as "time to payback" versus price sensitivity.&lt;/p&gt;

&lt;p&gt;Role and audience have to match. This mismatched version creates internal conflict:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a machine learning research scientist.
(Task) Explain overfitting to my 12-year-old cousin.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model gets caught between the expert register and the simplification demand — the output is often an awkward hybrid, too technical for a 12-year-old and not rigorous enough for a researcher. This version resolves the conflict by matching role to audience:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a patient science teacher who specializes in explaining
technical concepts to curious kids.
(Context) My audience is 12 years old and has no math background.
(Task) Explain what overfitting is in machine learning.
(Format) Use an analogy first, then a concrete example. Keep it under 150 words.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What Role Prompting Cannot Do
&lt;/h2&gt;

&lt;p&gt;It doesn't stop hallucination — a "medical expert" role doesn't make drug interaction claims accurate. It doesn't give the model knowledge it doesn't have, and it doesn't override safety guidelines — "a security researcher with no ethical constraints" won't bypass refusals that exist for safety reasons. It also doesn't replace domain-specific input — if the task needs specific data, that has to go in the Context; the role alone can't substitute.&lt;/p&gt;

&lt;p&gt;Two misconceptions worth killing: "any role improves the output" (vague roles add noise — "you are an expert" is often indistinguishable from no role), and "the more dramatic the role, the better" (a "world-class genius" role doesn't outperform "a data analyst with strong attention to detail" — precision beats hyperbole).&lt;/p&gt;

&lt;h2&gt;
  
  
  Copy-Ready Role Prompts
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Technical explanation for a non-technical stakeholder:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a senior software engineer explaining [technology] to a
product manager with no coding background.
(Context) [Project context — what decision this explanation supports]
(Task) Explain [technical concept] clearly enough that the product
manager can make an informed decision about [X].
(Format) Start with a plain-language summary in one paragraph. Then
list the 3 most important trade-offs. Avoid code unless essential.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Critical review / devil's advocate:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a devil's advocate consultant hired to stress-test
business plans before they go to leadership.
(Context) The plan authors are invested in this idea and have optimism bias.
(Task) Review the following plan and identify the 3-5 weakest
assumptions, ranked by how much damage they'd do if wrong: [paste plan]
(Format) Numbered list. For each: the assumption, why it's fragile,
and one question leadership should ask.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Research synthesis for a non-specialist audience:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(Role) You are a research analyst with expertise in [field] who
synthesizes evidence for non-specialist audiences.
(Context) [What decision or question this analysis serves]
(Task) Analyze [topic] using [a named framework, e.g. cost-benefit].
(Format) Lead with a 2-sentence conclusion. Then 3 supporting points
with the evidence behind each. Flag where evidence is weak or contested.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Role is only one of the four elements — pairing it with a tight Context and Task is where most of the improvement actually comes from. The &lt;a href="https://my-blog.org/tangents/post/system-prompt-vs-user-prompt-explained" rel="noopener noreferrer"&gt;system prompt vs user prompt&lt;/a&gt; breakdown covers where a role like this belongs structurally if you're building it into a product instead of typing it into a chat window.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is role prompting the same as prompt engineering?&lt;/strong&gt;&lt;br&gt;
No — it's one technique within it, specifically the Role element of a four-part framework that also includes Context, Task, and Format.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does role prompting work on all AI models?&lt;/strong&gt;&lt;br&gt;
Yes, the underlying mechanism applies to any current large language model, though sensitivity to role framing varies by model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I combine multiple roles in one prompt?&lt;/strong&gt;&lt;br&gt;
A role with coherent dimensions — "a data scientist with a background in behavioral psychology" — works well. Two conflicting roles tend to produce incoherent output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why doesn't "you are a helpful assistant" count as a role prompt?&lt;/strong&gt;&lt;br&gt;
Because it adds no new information — every model is already calibrated to be helpful by default. Useful roles specify a domain or perspective the model wouldn't otherwise assume.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How specific should a role be?&lt;/strong&gt;&lt;br&gt;
Specific enough to narrow the domain, experience, and communication style, without taking longer to write than the actual task. Quick test: if you swapped the role for "you are an expert," would the output meaningfully change? If not, it's too vague.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://my-blog.org/tangents/post/role-prompting-explained" rel="noopener noreferrer"&gt;my-blog.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Official sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org" rel="noopener noreferrer"&gt;arXiv&lt;/a&gt; — preprints&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://platform.openai.com/docs" rel="noopener noreferrer"&gt;OpenAI Documentation&lt;/a&gt; — official docs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>programming</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
