<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Zhengxin</title>
    <description>The latest articles on DEV Community by Zhengxin (@_94be737e156beb4d74df2).</description>
    <link>https://dev.to/_94be737e156beb4d74df2</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4043494%2F624f59c0-0fbb-4fb6-a54b-92fb2d93c548.jpg</url>
      <title>DEV Community: Zhengxin</title>
      <link>https://dev.to/_94be737e156beb4d74df2</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/_94be737e156beb4d74df2"/>
    <language>en</language>
    <item>
      <title>Meta Connect 2026 in 5 minutes: 7 products, prices, ship dates</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Thu, 24 Sep 2026 05:40:12 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/meta-connect-2026-in-5-minutes-7-products-prices-ship-dates-37ak</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/meta-connect-2026-in-5-minutes-7-products-prices-ship-dates-37ak</guid>
      <description>&lt;p&gt;Meta Connect 2026 was a 55-minute keynote. Here is what actually shipped, with prices and dates, in the order it was announced. The full recap with timestamps into the keynote video is here: &lt;a href="https://summarizevideototext.com/events/meta-connect-2026" rel="noopener noreferrer"&gt;Meta Connect 2026: the 55-minute keynote in 5 minutes&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Muse gets computer use, email and shopping
&lt;/h2&gt;

&lt;p&gt;Muse is the personal agent Meta launched a few weeks ago, and it was the center of the whole show. New today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Computer use.&lt;/strong&gt; With your permission Muse can drive any app on your Mac and keeps working on queued jobs after you walk away. The Mac app itself launched last week.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Its own email address.&lt;/strong&gt; Add Muse to a thread or forward mail and it handles it for you. Launch connectors: Spotify, OpenTable, Plaid, Ticketmaster. The connector platform is open, with 1,500+ apps already.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shopping.&lt;/strong&gt; Deep partnerships with Stripe and Shopify: Muse browses the Shopify catalog and buys with Link and Shop Pay, PayPal coming. Walmart, Best Buy, Gap, Sephora and Wayfair are integrating.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Voice.&lt;/strong&gt; You design your Muse's voice, and a real-time model animates its avatar while you talk.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each Muse runs on its own Muse Secure VM. A Confidential VM is coming so nobody, including Meta, can read your data. Free for a large token allowance; Meta plans to take a small fee on transactions. Just rolled out in Canada.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Ray-Ban Meta Audio: glasses with no camera
&lt;/h2&gt;

&lt;p&gt;A new category. Drop the camera and you get the thinnest, lightest Ray-Ban Meta yet with a full day of battery, built around Muse, music and calls. The reworked electronics fit classic shapes like the Clubmaster.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;$349. Pre-order now, ships October 13.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Ray-Ban Meta Gen 3
&lt;/h2&gt;

&lt;p&gt;The whole Ray-Ban Meta line moves to Gen 3: slimmer frame and temples, adjustable temple tips, better battery, Dolby Atmos spatial audio capture, and microphones that Meta says work on a jet ski. Two new shapes for the first time: Aviator and Zena, plus a 90th-anniversary launch edition Aviator.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;$449. Available now.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Meta Glasses lineup: 51 styles, Adventurer from $249
&lt;/h2&gt;

&lt;p&gt;The Meta Glasses line from June grows to 51 style combinations. New frames Capri (cat-eye) and Nova (panto), an ivory colorway with Kylie Jenner, and a LISA edition with a custom case and unlockable app features. By year end Meta will sell over 100 styles across Display, audio, Ray-Ban, Oakley and Meta Glasses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;From $249. Adventurer ships October 23.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Meta Ray-Ban Display is finally in stock
&lt;/h2&gt;

&lt;p&gt;The Display glasses with the Neural Band sold out for most of the year. Supply has caught up: on sale online at meta.com in the US, Canada and UK, with France, Italy and Germany on October 13. Hologram video calls from the glasses are coming.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Meta VR Glasses: 100 grams, $1,299.99, spring 2027
&lt;/h2&gt;

&lt;p&gt;The big reveal. About 100 grams on your face, a fifth of Quest 3, with compute and battery moved off the face into a tethered puck. Muse is built into the OS. Input is eyes, hands and voice, including typing on any flat surface (Meta's CFO reportedly hit 120 wpm).&lt;/p&gt;

&lt;p&gt;Content: first IMAX Enhanced certified VR device with Dolby Vision and Atmos, 75 hands-only launch titles, the full Quest library with optional controllers, Xbox Cloud Gaming, a Disney+ 3D app, ESPN front-row 180-degree live games at no extra cost, and full-body hologram calling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;$1,299.99 (from the press release). Ships spring 2027.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Muse Charm: the agent on a keychain
&lt;/h2&gt;

&lt;p&gt;The closer. The real-time voice and avatar stack squeezed into a keychain: tap the fingerprint sensor and talk, no phone needed. Materials are still being finalized.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Price not announced. Planned for the December holidays.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Three prices above ($349, $449 and the exact $1,299.99) came from Meta's newsroom posts, not the stage. Everything else is from the keynote, and the &lt;a href="https://summarizevideototext.com/events/meta-connect-2026" rel="noopener noreferrer"&gt;full recap&lt;/a&gt; links each claim to the exact second in the video, with a chapter list and FAQ. It was built by running the keynote through &lt;a href="https://summarizevideototext.com/" rel="noopener noreferrer"&gt;Summarize Video To Text&lt;/a&gt; and hand-checking the numbers.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>wearables</category>
      <category>vr</category>
      <category>news</category>
    </item>
    <item>
      <title>ChatGPT Still Can't Read YouTube Links. Here's the Copy-Paste That Actually Works.</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Wed, 23 Sep 2026 08:01:21 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/chatgpt-still-cant-read-youtube-links-heres-the-copy-paste-that-actually-works-1jl</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/chatgpt-still-cant-read-youtube-links-heres-the-copy-paste-that-actually-works-1jl</guid>
      <description>&lt;p&gt;Someone sends you a 45-minute video. You don't have 45 minutes.&lt;/p&gt;

&lt;p&gt;So you paste the link into ChatGPT and ask what's in it. It answers immediately — five clean bullet points, confident, well organised. You skim them, feel caught up, and move on.&lt;/p&gt;

&lt;p&gt;You just read the video's description. Not the video.&lt;/p&gt;

&lt;p&gt;This is the most common way people waste time with AI and video, and it's invisible, because the wrong answer looks exactly like the right one. r/ChatGPT has been asking about it for three years straight:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://old.reddit.com/r/ChatGPT/comments/11sa120/is_anyone_else_using_chatgpt_to_summarize_youtube/" rel="noopener noreferrer"&gt;Is anyone else using ChatGPT to summarize youtube video transcripts?&lt;/a&gt; — 2023&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://old.reddit.com/r/ChatGPT/comments/18tpxie/struggling_to_summarize_youtube_videos_with/" rel="noopener noreferrer"&gt;Struggling to Summarize YouTube Videos with ChatGPT - Any Tips?&lt;/a&gt; — 2023&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://old.reddit.com/r/ChatGPT/comments/1j6bi6y/why_cant_chatgpt_summarize_youtube_videos_directly/" rel="noopener noreferrer"&gt;Why Can't ChatGPT Summarize YouTube Videos Directly?&lt;/a&gt; — 2025&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://old.reddit.com/r/ChatGPT/comments/1m4mg9s/anybody_what_is_going_on_with_youtube_summarizers/" rel="noopener noreferrer"&gt;Anybody what is going on with Youtube summarizers and ChatGPT?&lt;/a&gt; — 2025&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ChatGPT can't watch a YouTube video. It reads the title and the description — the blurb the creator wrote to get you to click — and summarizes that. No prompt fixes it. It isn't your wording.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to tell it happened to you
&lt;/h2&gt;

&lt;p&gt;Three tells, ten seconds to check:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's suspiciously even.&lt;/strong&gt; Real talks ramble and spend nine minutes on one point. A summary where every section gets equal weight is a summary of an outline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing in it is specific.&lt;/strong&gt; No numbers, no names, no "she argues that…". Just topics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It matches the thumbnail.&lt;/strong&gt; If it's a longer version of the promise on the thumbnail, you got the marketing copy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The real test: take one claim from the summary and ask whether it could &lt;em&gt;only&lt;/em&gt; have come from someone speaking. If nothing passes, you didn't get the video.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: 15 seconds, no account
&lt;/h2&gt;

&lt;p&gt;Go to &lt;a href="https://summarizevideototext.com/?utm_source=devto" rel="noopener noreferrer"&gt;SummarizeVideoToText&lt;/a&gt;, paste the same link, press the button.&lt;/p&gt;

&lt;p&gt;What comes back is what you thought you were getting the first time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The full transcript&lt;/strong&gt; — every word actually spoken, as clean text you can search, copy, or keep.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A structured summary built from that transcript&lt;/strong&gt;, so the specifics survive: the numbers, the names, the part where the argument actually turns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chapter timestamps&lt;/strong&gt;, so when one section matters you jump straight to minute 31 instead of scrubbing for it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No sign-up for any of that. Anonymous use covers two videos a day, up to 15 minutes each — enough to settle most "what's in this" questions on the spot. A free account raises it to an hour per video, which is where lectures, interviews and conference talks live.&lt;/p&gt;

&lt;p&gt;It also works on the videos ChatGPT has no story for at all: upload a Zoom or Teams recording, a lecture file, a video off X or TikTok. Anything you can get to, you can get the words out of.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then take it back to ChatGPT, if you want
&lt;/h2&gt;

&lt;p&gt;The summary is usually where it ends. But if this is something you need to think with — a talk you're responding to, research you're going to cite — copy the transcript out and paste it into ChatGPT with a real prompt. Now it has the actual words, so it can do the thing it's good at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Below is the transcript of a video.

Give me:
1. A 3-sentence summary of the actual argument
2. The 5 most specific claims made, with the
   reasoning given for each
3. Anything stated as fact that I should verify
4. What it does NOT cover that I'd expect it to

Skip the intro and the sponsor read.

Transcript:
[paste]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Asking for "the reasoning given" is what keeps it inside the transcript. Ask for "key points" and it starts filling in from general knowledge, and you're back where you started.&lt;/p&gt;

&lt;p&gt;One thing to watch: an hour of video is around 10,000 words, and a very long paste can get quietly cut off — you'll get a summary of the first third with no warning. Attach it as a file instead of pasting, or do it in two halves.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you'd rather not use anything
&lt;/h2&gt;

&lt;p&gt;You can pull the transcript out of YouTube by hand: under the video, &lt;strong&gt;... → Show transcript&lt;/strong&gt;, then the three dots in that panel → turn off &lt;strong&gt;Toggle timestamps&lt;/strong&gt;, select all, copy. It's free and it works.&lt;/p&gt;

&lt;p&gt;It's also only available when the video has captions, it's missing on mobile web in places, and past about 30 minutes you're dragging a selection down a panel that keeps scrolling out from under you. Fine once. Not something you'll do twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line version
&lt;/h2&gt;

&lt;p&gt;ChatGPT isn't bad at summarizing videos. It's never seen one. Give it the words — or skip that step and get the summary directly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://summarizevideototext.com/?utm_source=devto" rel="noopener noreferrer"&gt;Paste your link here.&lt;/a&gt; Transcript, summary and timestamps, no account, about fifteen seconds.&lt;/p&gt;

&lt;p&gt;Comparing tools for this? We keep &lt;a href="https://summarizevideototext.com/blog/youtube-video-summarizer?utm_source=devto" rel="noopener noreferrer"&gt;a rundown&lt;/a&gt; separately.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>productivity</category>
      <category>learning</category>
    </item>
    <item>
      <title>Studying a YouTube course in Obsidian: playlist in, notes and flashcards out</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Sun, 20 Sep 2026 11:01:03 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/studying-a-youtube-course-in-obsidian-playlist-in-notes-and-flashcards-out-1ell</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/studying-a-youtube-course-in-obsidian-playlist-in-notes-and-flashcards-out-1ell</guid>
      <description>&lt;p&gt;A single conference talk is easy to deal with. A course is the thing that defeats me.&lt;/p&gt;

&lt;p&gt;Forty videos in a playlist, and by video six the notes have stopped. The material isn't the problem. The bookkeeping is. Which ones have I done. What did video three establish that video nine now assumes I remember. Where did I put any of it.&lt;/p&gt;

&lt;p&gt;I've spent the last few months building &lt;a href="https://summarizevideototext.com" rel="noopener noreferrer"&gt;SummarizeVideoToText&lt;/a&gt; and its &lt;a href="https://community.obsidian.md/plugins/summarize-video-to-text" rel="noopener noreferrer"&gt;Obsidian plugin&lt;/a&gt;, and I wrote up &lt;a href="https://dev.to/_94be737e156beb4d74df2/turning-youtube-videos-into-obsidian-notes-and-what-shipping-the-plugin-taught-me-6m2"&gt;the plugin internals&lt;/a&gt; a while back — markers, tokens, the review rules. This post is the other half. Not how it's built, but what the two halves are for when you point them at a whole course instead of one video. The annoying parts are still in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The playlist becomes a queue, not a folder
&lt;/h2&gt;

&lt;p&gt;Paste a YouTube playlist URL into the web app and it imports as a collection, up to 500 videos.&lt;/p&gt;

&lt;p&gt;You get a preview step first: the title, the count, the first handful of videos, and then a confirm. That step exists because an import that silently grabs the wrong 200 videos is worse than no import at all, and you don't find out until you're three summaries deep.&lt;/p&gt;

&lt;p&gt;Everything lands marked "to learn". Nothing is summarized yet.&lt;/p&gt;

&lt;p&gt;That last part is the whole design. Summarizing 40 videos up front gives you 40 documents nobody opens. What you want is a queue with a cursor in it. The collection row on the dashboard carries one button — &lt;code&gt;Continue · Part 7&lt;/code&gt; — and that button is the only decision you have to make when you sit down.&lt;/p&gt;

&lt;h2&gt;
  
  
  One video becomes one note
&lt;/h2&gt;

&lt;p&gt;From inside Obsidian: one command, one link, one note.&lt;/p&gt;

&lt;p&gt;What lands in the vault is a TL;DR, key insights, timestamped chapters, frontmatter shaped for Dataview, and an embedded player at the top so a timestamp seeks in place instead of throwing you back to YouTube. Optional extras are the full transcript in a collapsed callout, a quiz, and a longer summary in one of several templates.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwhk4cdobwsks9yqh5k0u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwhk4cdobwsks9yqh5k0u.png" alt="An Obsidian note generated from a 3Blue1Brown video: breadcrumb, links back to the source and to summarizevideototext.com, an embedded YouTube player, and a TL;DR callout" width="799" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two things about it that matter over a long course rather than a single video.&lt;/p&gt;

&lt;p&gt;The note has a generated region between &lt;code&gt;%% svt:start %%&lt;/code&gt; and &lt;code&gt;%% svt:end %%&lt;/code&gt;, and a "My notes" section outside it. Re-running the command replaces the generated region and leaves your writing alone. On video nine, when you realise your frontmatter tags were wrong on all eight previous notes, you can re-run all of them without losing a word you wrote yourself.&lt;/p&gt;

&lt;p&gt;And re-opening a video that's already been processed never costs a credit. Refreshing a note you made last week is free.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblyubl8kluw1eoujs1ni.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblyubl8kluw1eoujs1ni.png" alt="The same note further down: a Key insights list and timestamped chapters, each timestamp a clickable link that seeks the embedded player" width="799" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you don't use Obsidian, the same note exports to Notion as a real page, or downloads as Markdown with the timestamps still clickable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The quiz becomes cards
&lt;/h2&gt;

&lt;p&gt;There's a command that turns the quiz in a video note into &lt;a href="https://github.com/st3v3nmw/obsidian-spaced-repetition" rel="noopener noreferrer"&gt;Spaced Repetition&lt;/a&gt; cards in a separate note.&lt;/p&gt;

&lt;p&gt;Run it again after a refresh and it only appends the new questions. Cards you already have, cards you wrote by hand, and the &lt;code&gt;&amp;lt;!--SR:...--&amp;gt;&lt;/code&gt; review scheduling the SR plugin writes back are all left untouched. This sounds like a detail and it is the reason the command is safe to press. A "regenerate my flashcards" button that resets three weeks of review intervals is a button you press exactly once.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxnznaycy01di458v958t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxnznaycy01di458v958t.png" alt="A separate flashcards note in Obsidian, tagged #flashcards, with question and answer pairs and a timestamp link under each answer" width="799" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A quiz across the whole collection
&lt;/h2&gt;

&lt;p&gt;Per-video quizzes test whether you were awake. The thing I actually wanted was a quiz drawn across the collection, because that's where a course either connected for you or didn't — video three's definition showing up inside video nine's argument.&lt;/p&gt;

&lt;p&gt;That one runs on the web side, not in the vault, and it's the feature I'd point at if you asked me what a "collection" is for beyond grouping.&lt;/p&gt;

&lt;p&gt;Here's one of mine, published: &lt;a href="https://summarizevideototext.com/u/diaozxin/3blue1brown" rel="noopener noreferrer"&gt;3Blue1Brown's neural network series&lt;/a&gt; — eight videos, eight notes, one page. That's the shape the vault ends up in too, minus the Obsidian chrome.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa7aeieo8h4xxspsqu7hg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa7aeieo8h4xxspsqu7hg.png" alt="A published collection page on summarizevideototext.com: 3Blue1Brown Neural Networks and Deep Learning, 8 videos and 8 notes, a " width="800" height="570"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No account at all:&lt;/strong&gt; 2 summaries a day, videos up to 15 minutes, plus one TikTok / Instagram / X video a day up to 5 minutes. Transcripts are readable, copyable and downloadable without signing in, five a day. No email, no signup wall to get a look at the output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free, signed in:&lt;/strong&gt; 10 credits a month, which is roughly ten videos, with videos up to an hour, plus saved history, collections, and the Obsidian and Notion exports. Playlist import needs this tier — 500 videos of state has to live somewhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro:&lt;/strong&gt; $9/month (launch price, normally $12) or $54/year. Unlimited summaries under fair use, videos up to 6 hours, and 300 credits a month for the transcription cases. The 6 hours is the reason this tier exists at all: lecture recordings and conference livestreams are absurdly long.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Top-ups, if you don't want a subscription:&lt;/strong&gt; $1 for 20 credits, once per account, or $5 for 100. Bought credits don't expire, and they're only spent once the monthly ones run out.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;TikTok, Instagram, X and uploaded files cost more per video than YouTube does, because those have to be transcribed rather than read.&lt;/p&gt;

&lt;p&gt;Output comes in 14 languages, and the output language is independent of the video's. Watching a Spanish lecture and keeping notes in English is the normal case, not a workaround.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't do
&lt;/h2&gt;

&lt;p&gt;Everything downstream is derived from the words. Captions when YouTube has them, speech-to-text when it doesn't. That means a lecture where the argument lives on the whiteboard and the speaker says "so this equals that" produces a summary that is technically accurate and completely useless. I have no good answer for that one yet. Diagram-heavy talks are where I still take notes by hand.&lt;/p&gt;

&lt;p&gt;The plugin also holds no model keys and does no inference locally. It's a client for an account on the server, and it needs the network. If your requirement is that your vault never talks to anything, this is the wrong tool and I'd rather you know that from the third paragraph than from the settings screen.&lt;/p&gt;

&lt;p&gt;It needs Obsidian 1.13.0 or newer. The declarative settings API the community review pushed me toward doesn't exist before that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Web app: &lt;a href="https://summarizevideototext.com" rel="noopener noreferrer"&gt;summarizevideototext.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;A finished collection, if you want to see the output before installing anything: &lt;a href="https://summarizevideototext.com/u/diaozxin/3blue1brown" rel="noopener noreferrer"&gt;3Blue1Brown, 8 chapters with notes&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Obsidian plugin: &lt;a href="https://community.obsidian.md/plugins/summarize-video-to-text" rel="noopener noreferrer"&gt;community.obsidian.md/plugins/summarize-video-to-text&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Source, MIT: &lt;a href="https://github.com/diaozxin007/obsidian-summarize-video-to-text" rel="noopener noreferrer"&gt;github.com/diaozxin007/obsidian-summarize-video-to-text&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Setup walkthrough: &lt;a href="https://summarizevideototext.com/obsidian-plugin" rel="noopener noreferrer"&gt;summarizevideototext.com/obsidian-plugin&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;If your material is lectures specifically: &lt;a href="https://summarizevideototext.com/ai-lecture-summarizer" rel="noopener noreferrer"&gt;summarizevideototext.com/ai-lecture-summarizer&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The playlist-as-a-queue idea is the part I'd keep if I rebuilt all of it. Summarizing everything up front feels productive and produces an archive. Summarizing the next one produces a course you finished.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>productivity</category>
      <category>learning</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Two years of pelicans on bicycles: 103 SVGs, 66 models, and what the guy who invented the benchmark says about it</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Thu, 17 Sep 2026 07:24:24 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/two-years-of-pelicans-on-bicycles-103-svgs-66-models-and-what-the-guy-who-invented-the-benchmark-4c7e</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/two-years-of-pelicans-on-bicycles-103-svgs-66-models-and-what-the-guy-who-invented-the-benchmark-4c7e</guid>
      <description>&lt;p&gt;In October 2024, Simon Willison gave sixteen LLMs the same eight words:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Generate an SVG of a pelican riding a bicycle.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He has been running it on every new model since. Twenty-three months later there are 103 of these drawings scattered across dozens of his posts, and he has written a section explaining why you should not use his benchmark to compare models.&lt;/p&gt;

&lt;p&gt;We pulled all of them into one page — &lt;a href="https://pelicanzoo.ai" rel="noopener noreferrer"&gt;pelicanzoo.ai&lt;/a&gt; — because reading them in order turns out to be a surprisingly good history of the last two years. This post is what we learned putting it together, and how the thing is built.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everything here is Simon's work.&lt;/strong&gt; The images, the commentary and the judgement calls are his; we mirrored and organised them, and every specimen links back to the post it came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  The prompt is deliberately unfair
&lt;/h2&gt;

&lt;p&gt;His stated reasons are mundane: he likes pelicans (he lives in Half Moon Bay, they're on the beach), and he was confident no SVG of a pelican on a bicycle existed in any training set.&lt;/p&gt;

&lt;p&gt;The reasons it &lt;em&gt;works&lt;/em&gt; are better. LLMs are text models that can't see, but they can write code, and SVG is code. So the task is to assemble a picture out of coordinates and radii with no feedback loop — the model never sees what it drew. Bicycles are hard for humans to draw from memory (try it: you'll get the frame wrong). And, in his words:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Pelicans can't ride bicycles. They don't have the right shape.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There's a bonus that matters more than it sounds: models leave comments in the SVG. &lt;code&gt;&amp;lt;!-- wheel --&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;!-- beak --&amp;gt;&lt;/code&gt;. You can read what it &lt;em&gt;meant&lt;/em&gt; to draw and compare it against what came out. The gap is usually the joke.&lt;/p&gt;

&lt;h2&gt;
  
  
  2024: it's mostly shapes
&lt;/h2&gt;

&lt;p&gt;The early attempts are bad in specific, funny ways. Claude 3.5 Sonnet produced two circles, a triangle, and a beak Simon described as a yellow banana smile. Llama 3.3 70B drew a small circle, a vertical line, and something shaped like a sink. Qwen2.5-Coder's bicycle is a pile of brown shapes that reads as a tractor. Mistral Small 3 delivered a squat white duck squatting on a barbell.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvz7ubjiq4fap1xnpp3l9.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvz7ubjiq4fap1xnpp3l9.jpeg" alt="Mistral Small 3's duck on a barbell" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then DeepSeek R1 landed in January 2025, and shortly afterwards Nvidia lost about $600bn in market cap. Simon later singled its pelican out as "the pelican that crashed the stock market" — the best one so far: a recognisable bicycle, a bird next to it that's arguably a pelican, not actually riding it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbrlgszp8vrqc9ftoh2nb.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbrlgszp8vrqc9ftoh2nb.jpeg" alt="DeepSeek R1's pelican" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For that first year the drawings tracked real capability fairly well. Models that drew better pelicans were better at writing code. A joke about how silly model comparisons are started behaving like a real benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 560-match tournament that cost 18 cents
&lt;/h2&gt;

&lt;p&gt;By June 2025 he had 34 pelicans and wanted a ranking without looking at 34 images himself.&lt;/p&gt;

&lt;p&gt;So he had Claude write a tool that renders any two pelicans side by side and screenshots the pair. 34 images, every pairing, 560 matches. Each match went to GPT-4.1 mini with instructions to pick the better "pelican riding a bicycle" illustration and explain why. Elo scores out of the results.&lt;/p&gt;

&lt;p&gt;Total judging cost: &lt;strong&gt;about 18 cents.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdv3au2cxpjf98a925hia.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdv3au2cxpjf98a925hia.jpeg" alt="Tournament winner and loser" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Winner was a Gemini 2.5 Pro preview; last place, Llama 3.3 70B. The judge's reasoning for that matchup: the left image clearly depicts a pelican riding a bicycle, the right is a few minimal shapes that don't read as anything.&lt;/p&gt;

&lt;p&gt;The same talk includes the cost comparison that stuck with me. o1-pro: &lt;strong&gt;88.755 cents&lt;/strong&gt; for one pelican. Gemini 2.5 Pro: &lt;strong&gt;4.77 cents&lt;/strong&gt;. Gemini's was better.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvhirh1b6qy09o15afh9p.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvhirh1b6qy09o15afh9p.jpeg" alt="o1-pro at 88 cents" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The models start editorialising
&lt;/h2&gt;

&lt;p&gt;August 2025, Simon runs Qwen3-4B-Thinking locally. It declines to draw a bicycle. It returns a blue circle with red text across it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is art - pelicans don't ride bikes!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He made that the headline of the post.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fctp8laihq07hsg0dhhpn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fctp8laihq07hsg0dhhpn.png" alt="Qwen3-4B-Thinking: " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In November, Gemini 3 at low reasoning effort gave its pelican a little hat, with an SVG comment reading &lt;code&gt;hat (optional fun detail)&lt;/code&gt;. At high effort the frame geometry was finally correct.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feoobon0vn98fyjwxb5cm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feoobon0vn98fyjwxb5cm.png" alt="Gemini 3, low effort, wearing a hat" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;He decided the task had gotten too easy and escalated: a California brown pelican, bicycle with spokes and a correct frame, large gular pouch, visible feather texture, unambiguously pedalling, in breeding plumage. He attached a photo he took himself.&lt;/p&gt;

&lt;p&gt;Most models missed the detail that makes it a trick question — the California brown pelican isn't brown.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdt90c8xh24kxbmh7m2yj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdt90c8xh24kxbmh7m2yj.png" alt="Gemini 3 high effort on the harder prompt" width="800" height="667"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The laptop model that beat Opus
&lt;/h2&gt;

&lt;p&gt;April 16, 2026: Qwen3.6-35B-A3B and Claude Opus 4.7 shipped the same day. Simon ran Qwen as a 20.9GB quantised model on his MacBook and Opus through the API.&lt;/p&gt;

&lt;p&gt;He gave it to Qwen. Opus got the frame geometry wrong, and got it wrong again at maximum reasoning effort.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qsw2p1l3xbb3j1q3xy8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qsw2p1l3xbb3j1q3xy8.png" alt="Qwen3.6-35B-A3B, running on a laptop" width="800" height="667"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9p690y1tcdooyhxxgnj8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9p690y1tcdooyhxxgnj8.png" alt="Claude Opus 4.7 — look at the frame" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This result made him suspicious of his own benchmark. There's a persistent theory that labs train against his "stupid test", and he'd been keeping a spare prompt in reserve for exactly this moment. He burned it: &lt;strong&gt;a flamingo riding a unicycle.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Qwen's flamingo wears sunglasses and a bow tie, appears to be smoking, and is flanked by heart emoji and the caption "Flamingo on a Unicycle". The source comment says &lt;code&gt;Give the flamingo sunglasses!&lt;/code&gt;. Opus drew a competent, entirely unremarkable flamingo. That one went to Qwen as well.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftyzecq80byem8c56mkh2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftyzecq80byem8c56mkh2.png" alt="Qwen's flamingo" width="800" height="978"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5fzqpt2c8jw3yxpy7utj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5fzqpt2c8jw3yxpy7utj.png" alt="Opus's flamingo" width="800" height="1000"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;His own framing was careful: he has enormous respect for Qwen, and he does not believe a 21GB quantised model is more useful than Anthropic's latest flagship. But if the thing you want is an SVG of a pelican on a bicycle, the laptop won that day.&lt;/p&gt;

&lt;p&gt;That's roughly where pelican quality and model capability stop tracking each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  The inventor says don't use it
&lt;/h2&gt;

&lt;p&gt;Kimi K3, July 2026. Its pelican burned 13,241 reasoning tokens and cost 25 cents. Simon used the post to write a section on what the benchmark is still good for.&lt;/p&gt;

&lt;p&gt;His conclusion after 21 months: it was never a good benchmark. The correlation in year one was a surprise and it's basically gone now — GLM-5.2 draws a better pelican than GPT-5.6 or Claude Fable 5, and he doesn't think GLM is in that class. More importantly it tests nothing about what actually matters today, namely whether a model can hold a long conversation and call tools reliably. His words: &lt;strong&gt;don't use pelicans to compare models.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So why keep going? Two reasons, both practical:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It forces him to actually run each new model. Posting a pelican is proof he got the thing working, rather than just reading the launch post.&lt;/li&gt;
&lt;li&gt;Even one prompt leaks information. On the Kimi K3 run he noticed his eight-word prompt was billed as 95 input tokens, and sending just "hi" cost 86 — implying about 85 tokens of hidden system prompt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;September 1, 2026: Claude Fable 5.1, across five reasoning levels. Low and medium barely think and finish in about twenty seconds. Maximum ran &lt;strong&gt;13 minutes 54 seconds&lt;/strong&gt;, emitted 65,927 tokens, cost &lt;strong&gt;$3.30&lt;/strong&gt;, and produced a pelican in a blue cap with a fish in the basket. Best pelican he's seen from Anthropic.&lt;/p&gt;

&lt;p&gt;The benchmark has become a ritual rather than a measurement. It won't tell you which model is strongest. One pelican's invoice and reasoning trace will still tell you a lot about how a model works, and whether it'll decide on its own to put a hat on your bird.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building the zoo
&lt;/h2&gt;

&lt;p&gt;The drawings live across two years of posts, which makes reading them in sequence tedious. So: one page, 103 specimens, 66 models, 12 vendors, October 2024 to now. Each one is tagged with model and date and links back to Simon's original post.&lt;/p&gt;

&lt;p&gt;A few things that turned out to be more interesting than expected:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;24 of them are still vector.&lt;/strong&gt; Where the original SVG survives, the page serves the SVG — zoom in as far as you like, and four of them are animated (wheels turning, pelican bobbing). The rest exist only as screenshots from his conference slides, so they're labelled &lt;em&gt;print&lt;/em&gt; rather than &lt;em&gt;live&lt;/em&gt;. Making that distinction visible mattered more than we expected; a screenshot of an SVG and the SVG are not the same artifact, and the ones that are still code are the ones you can actually inspect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The critic reads source, not pixels.&lt;/strong&gt; There's a button under each specimen that asks a model to review the drawing — and it's given the SVG source, not an image. Half of that is cost. The other half is that it's funnier and more specific: it can complain that the body is an ellipse of radius 104, or that the rear wheel is out of proportion with the frame. Using models to review models closes a loop that probably shouldn't be closed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Submissions have no backend.&lt;/strong&gt; You can paste in an SVG a model drew for you. The page sanitises it in the browser, then generates a pre-filled GitHub pull request. No server, no database — one pelican is one file. It's a static site on Cloudflare.&lt;/p&gt;

&lt;p&gt;The sanitiser is the part I'd actually defend in review. Model-generated SVG pasted in by strangers is untrusted markup, so it strips &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;foreignObject&amp;gt;&lt;/code&gt;, &lt;code&gt;iframe&lt;/code&gt;/&lt;code&gt;object&lt;/code&gt;/&lt;code&gt;embed&lt;/code&gt;, inline &lt;code&gt;on*&lt;/code&gt; handlers, &lt;code&gt;javascript:&lt;/code&gt; URLs, &lt;code&gt;@import&lt;/code&gt;, and any remote &lt;code&gt;href&lt;/code&gt;/&lt;code&gt;src&lt;/code&gt;/&lt;code&gt;url()&lt;/code&gt; reference. Data URIs stay — they're self-contained. Remote references go because they leak the visitor's IP to whoever hosts them and let someone swap the picture out later.&lt;/p&gt;

&lt;p&gt;Two design decisions in there I'd repeat:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// It reports what it removed instead of only returning a cleaned string.&lt;/span&gt;
&lt;span class="c1"&gt;// @returns {{ ok, svg, removed: string[], error }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same function runs in three places — the static build, the browser preview, and CI on incoming pull requests. Because the site strips &lt;em&gt;before&lt;/em&gt; it builds the PR, a submission that still arrives dirty was hand-crafted, so CI stops it for a human to read rather than silently cleaning it. A sanitiser that returns a clean string and nothing else throws away the signal that someone was trying.&lt;/p&gt;

&lt;p&gt;The other one is dull and bit us first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Percentages need something to be a percentage of; without a viewBox an SVG&lt;/span&gt;
&lt;span class="c1"&gt;// whose width/height we strip will stretch to fill whatever box it lands in.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We strip fixed &lt;code&gt;width&lt;/code&gt;/&lt;code&gt;height&lt;/code&gt; so specimens scale in a grid. Do that to an SVG with no &lt;code&gt;viewBox&lt;/code&gt; and it stretches to fill its container. So the sanitiser synthesises one from the original dimensions when it's missing.&lt;/p&gt;

&lt;p&gt;There's also a guessing game — you get a drawing and four model names, 73 rounds over 51 models. Distractors are picked from &lt;em&gt;different&lt;/em&gt; vendors on purpose. Three Gemini variants side by side isn't difficult, it's just unfair. The 2024 ones are easy. The 2026 ones are basically a coin flip, which is its own kind of result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Credit
&lt;/h2&gt;

&lt;p&gt;The images and the original keeper's notes are Simon Willison's, collected from his posts, and the footer links back to &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/" rel="noopener noreferrer"&gt;his pelican-riding-a-bicycle tag&lt;/a&gt;. We built the cage, not the birds.&lt;/p&gt;

&lt;p&gt;If you've got one that's especially good, or especially wrong, the submission box takes it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://pelicanzoo.ai" rel="noopener noreferrer"&gt;pelicanzoo.ai&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/" rel="noopener noreferrer"&gt;Pelicans on a bicycle&lt;/a&gt; (2024-10-25)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://simonwillison.net/2025/Jun/6/six-months-in-llms/" rel="noopener noreferrer"&gt;The last six months in LLMs, illustrated by pelicans on bicycles&lt;/a&gt; (2025-06-06)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://simonwillison.net/2025/Nov/18/gemini-3/" rel="noopener noreferrer"&gt;Trying out Gemini 3 Pro with audio transcription and a new pelican benchmark&lt;/a&gt; (2025-11-18)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://simonwillison.net/2026/Apr/16/qwen-beats-opus/" rel="noopener noreferrer"&gt;Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7&lt;/a&gt; (2026-04-16)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/" rel="noopener noreferrer"&gt;Kimi K3, and what we can still learn from the pelican benchmark&lt;/a&gt; (2026-07-16)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://simonwillison.net/2026/Sep/1/claude-fable-5-1/" rel="noopener noreferrer"&gt;Claude Fable 5.1 made me a really nice animated pelican&lt;/a&gt; (2026-09-01)&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>showdev</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Turning YouTube Videos Into Obsidian Notes — and What Shipping the Plugin Taught Me</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Mon, 14 Sep 2026 05:02:17 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/turning-youtube-videos-into-obsidian-notes-and-what-shipping-the-plugin-taught-me-6m2</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/turning-youtube-videos-into-obsidian-notes-and-what-shipping-the-plugin-taught-me-6m2</guid>
      <description>&lt;p&gt;Every so often I'd watch a 90-minute conference talk, close the tab, and realise the only trace it left in my vault was a URL in a list called "watch later (again)".&lt;/p&gt;

&lt;p&gt;So I built an Obsidian plugin that turns a video into an actual note. It's been in the community directory since September and is on 0.3.0 now. This is what it does, and — more usefully — the handful of decisions and review rules I'd want to know before building another one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What lands in the vault
&lt;/h2&gt;

&lt;p&gt;One command, one link, and you get a note:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwhk4cdobwsks9yqh5k0u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwhk4cdobwsks9yqh5k0u.png" alt="A video note in Obsidian: embedded player, TL;DR callout and key insights" width="799" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a TL;DR callout and key insights&lt;/li&gt;
&lt;li&gt;an embedded player at the top — timestamps seek it in place, without leaving the note&lt;/li&gt;
&lt;li&gt;chapters as timestamped bullets&lt;/li&gt;
&lt;li&gt;optional summary (several templates), quiz, and full transcript in a collapsed callout&lt;/li&gt;
&lt;li&gt;frontmatter with title, source, channel, language and tags, shaped for Dataview&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblyubl8kluw1eoujs1ni.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblyubl8kluw1eoujs1ni.png" alt="Key insights and chapters with timestamp links" width="799" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Plus a right-sidebar chat panel scoped to the note you're reading, and a command that turns the quiz into &lt;a href="https://github.com/st3v3nmw/obsidian-spaced-repetition" rel="noopener noreferrer"&gt;Spaced Repetition&lt;/a&gt; cards.&lt;/p&gt;

&lt;p&gt;Now the parts that were actually decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plugin holds no model keys
&lt;/h2&gt;

&lt;p&gt;An Obsidian plugin runs on the user's machine. You can't ship an API key in it, and asking every user to paste their own is both friction and a support queue you will personally staff forever.&lt;/p&gt;

&lt;p&gt;So the plugin generates nothing. It sends a video link and an account token to a server, and writes down what comes back. That's the whole client.&lt;/p&gt;

&lt;p&gt;The honest trade-off: this means the plugin needs an account, which is a real cost to the user and something I put in the README under a &lt;strong&gt;Before you install&lt;/strong&gt; heading rather than burying. A plugin that silently requires a signup is a plugin that gets one-starred.&lt;/p&gt;

&lt;h2&gt;
  
  
  The connect handshake, and the one line that makes it safe
&lt;/h2&gt;

&lt;p&gt;Settings → &lt;strong&gt;Connect&lt;/strong&gt; opens the browser at &lt;code&gt;/connect/obsidian?state=…&amp;amp;vault=…&lt;/code&gt;. You sign in, the server mints a token, and the page bounces you back via &lt;code&gt;obsidian://svt-connect?token=…&amp;amp;state=…&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;state&lt;/code&gt; is generated by the plugin and checked when the redirect comes home. Without that check, &lt;strong&gt;any web page could deep-link a token into your vault&lt;/strong&gt; — the &lt;code&gt;obsidian://&lt;/code&gt; scheme is open to anyone. It's four lines of code and it's the difference between a handshake and a hole.&lt;/p&gt;

&lt;p&gt;Server side the token is stored as a sha256 hash. Client side it lands in &lt;code&gt;.obsidian/plugins/&amp;lt;id&amp;gt;/data.json&lt;/code&gt;, which is worth saying out loud in your docs: if the user syncs their vault, that file goes with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  One note format, two producers
&lt;/h2&gt;

&lt;p&gt;The website also has an &lt;em&gt;Export → Obsidian&lt;/em&gt; button. That's two code paths producing "the same" Markdown, which in my experience drift apart within about a week.&lt;/p&gt;

&lt;p&gt;So the Markdown is assembled in exactly one place — a server endpoint — and the plugin writes whatever it receives verbatim. The website button and the plugin produce identical notes because they are, literally, the same function call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Re-running a command must not eat the user's writing
&lt;/h2&gt;

&lt;p&gt;People add their own notes under the generated stuff. Refreshing a note has to be non-destructive or nobody will ever press it twice.&lt;/p&gt;

&lt;p&gt;The generated region is fenced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;%% svt:start %%
…everything the server produced…
%% svt:end %%

&lt;span class="gu"&gt;## My notes&lt;/span&gt;
whatever you wrote, untouched
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Refresh replaces only what's between the markers. Frontmatter is the subtler half: the server returns an &lt;code&gt;ownedKeys&lt;/code&gt; list alongside the note, so the plugin overwrites the keys it owns and leaves every key you added alone. Writes go through &lt;code&gt;vault.process&lt;/code&gt; rather than read-then-write, which is also what the community review expects.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;requestUrl&lt;/code&gt; doesn't stream
&lt;/h2&gt;

&lt;p&gt;Obsidian's &lt;code&gt;requestUrl&lt;/code&gt; sidesteps CORS, which you want, because your plugin isn't an origin anyone will whitelist. What it doesn't do is stream.&lt;/p&gt;

&lt;p&gt;My summary endpoint is server-sent events. So the plugin waits for the whole response and parses the event stream after the fact. You lose the progressive typewriter effect; you keep mobile support and never fight a preflight. Worth knowing before you architect around a stream that won't arrive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Flashcards, and only appending
&lt;/h2&gt;

&lt;p&gt;0.3.0 turns the quiz into a card note: multi-line format, the right answer and its explanation on the back, and a timestamp link to the moment the question came from.&lt;/p&gt;

&lt;p&gt;Run it again after new questions appear and only the missing cards are appended. Existing cards, their &lt;code&gt;&amp;lt;!--SR:…--&amp;gt;&lt;/code&gt; review schedules, and any cards you wrote yourself are never touched. Rewriting a review schedule is the fastest way to lose a user who has been reviewing for six months.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxnznaycy01di458v958t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxnznaycy01di458v958t.png" alt="A card note ready for the Spaced Repetition plugin" width="799" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the community directory review actually checks
&lt;/h2&gt;

&lt;p&gt;Submitting is no longer a PR against &lt;code&gt;obsidian-releases&lt;/code&gt;. You log in at &lt;a href="https://community.obsidian.md" rel="noopener noreferrer"&gt;community.obsidian.md&lt;/a&gt;, link GitHub, and an automated reviewer runs. Mine came back with a list. The useful items:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No inline styles.&lt;/strong&gt; Anything you'd reach for as &lt;code&gt;el.style.foo = bar&lt;/code&gt; belongs in &lt;code&gt;styles.css&lt;/code&gt;. My auto-growing textarea became the CSS &lt;code&gt;field-sizing: content&lt;/code&gt; instead of a resize handler, which is a better implementation anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Settings must use the declarative 1.13 API&lt;/strong&gt; (&lt;code&gt;getSettingDefinitions()&lt;/code&gt; with groups, &lt;code&gt;getControlValue&lt;/code&gt; / &lt;code&gt;setControlValue&lt;/code&gt;). That forced &lt;code&gt;minAppVersion&lt;/code&gt; up to &lt;code&gt;1.13.0&lt;/code&gt;, because the &lt;em&gt;no-unsupported-api&lt;/em&gt; rule rejects 1.13 APIs while you claim to support older versions. Pick a lane; you can't straddle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attest your build.&lt;/strong&gt; &lt;code&gt;actions/attest-build-provenance&lt;/code&gt; in the release workflow, with &lt;code&gt;id-token&lt;/code&gt; and &lt;code&gt;attestations&lt;/code&gt; write permissions. &lt;code&gt;gh attestation verify main.js --owner &amp;lt;you&amp;gt;&lt;/code&gt; exiting 0 with no output is what success looks like.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The tag must match the manifest version&lt;/strong&gt;, and &lt;code&gt;npm version&lt;/code&gt; writes tags with a &lt;code&gt;v&lt;/code&gt; prefix that Obsidian doesn't want. &lt;code&gt;git tag -d v0.2.2 &amp;amp;&amp;amp; git tag 0.2.2&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One false alarm worth pre-empting: right after I re-cut a release, the checker insisted "No release matches your manifest version." It had scanned during the draft window and cached that. Re-scan, it passes. Don't go rewriting your workflow like I nearly did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Obsidian CSS facts that cost me an evening
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;.view-content&lt;/code&gt; ships with &lt;code&gt;padding: 12px 12px 32px&lt;/code&gt;. If your custom view's layout looks mysteriously inset, that's why. Zero it with a same-specificity selector: &lt;code&gt;.workspace-leaf-content[data-type="…"] .view-content.your-class&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;On desktop the status bar is &lt;code&gt;position: fixed&lt;/code&gt;, 27px tall, and floats &lt;strong&gt;over&lt;/strong&gt; the bottom of the right sidebar. Leave it room or your input box is permanently half-covered.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;button:not(.clickable-icon)&lt;/code&gt; picks up Obsidian's background and shadow. For a text link, use an &lt;code&gt;&amp;lt;a&amp;gt;&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Plugin: &lt;a href="https://community.obsidian.md/plugins/summarize-video-to-text" rel="noopener noreferrer"&gt;community.obsidian.md/plugins/summarize-video-to-text&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Source (MIT): &lt;a href="https://github.com/diaozxin007/obsidian-summarize-video-to-text" rel="noopener noreferrer"&gt;github.com/diaozxin007/obsidian-summarize-video-to-text&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Setup walkthrough: &lt;a href="https://summarizevideototext.com/obsidian-plugin" rel="noopener noreferrer"&gt;summarizevideototext.com/obsidian-plugin&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're building something that writes into other people's vaults, the markers-and-&lt;code&gt;ownedKeys&lt;/code&gt; pattern is the piece I'd steal. Everything else was replaceable; that one is what makes a destructive command safe to press twice.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>showdev</category>
      <category>typescript</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Code Agent Loop Deep Dive (0): From the Chat Window to the Loop</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Wed, 09 Sep 2026 02:50:21 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/claude-code-agent-loop-deep-dive-0-from-the-chat-window-to-the-loop-1e4i</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/claude-code-agent-loop-deep-dive-0-from-the-chat-window-to-the-loop-1e4i</guid>
      <description>&lt;p&gt;Open Claude Code, type “Help me investigate this bug,” and press Enter.&lt;/p&gt;

&lt;p&gt;The window starts scrolling. Claude reads &lt;code&gt;auth.py&lt;/code&gt;, then &lt;code&gt;login.py&lt;/code&gt;, runs a search, edits one line, runs a test, and finally says: “Done—the cause was X.”&lt;/p&gt;

&lt;p&gt;It feels like a conversation with something that remembers what was said, can use tools, and waits for you when it needs input.&lt;/p&gt;

&lt;p&gt;But an LLM API is &lt;strong&gt;stateless&lt;/strong&gt;. Each request is independent; the server does not remember your previous prompt. To make the model appear to remember a conversation, the application must send the full message history again on every call.&lt;/p&gt;

&lt;p&gt;That creates an intermediate layer: the &lt;strong&gt;harness&lt;/strong&gt;. It retains the history and packages it every time it calls the LLM. What looks like a continuous chat is closer to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Turn 1: harness sends [1 message] → LLM returns a response
Turn 2: harness sends [3 messages] → LLM returns a response
Turn 3: harness sends [5 messages] → LLM returns a response
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The message array only grows. The LLM remembers nothing; the harness remembers everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  One user message, how many model calls?
&lt;/h2&gt;

&lt;p&gt;Consider the bug fix above. Claude read two files, searched the codebase, made an edit, and ran a test: five tool calls. After each tool completes, the harness puts its &lt;code&gt;tool_result&lt;/code&gt; back into &lt;code&gt;messages&lt;/code&gt; and calls the LLM again so it can decide the next action.&lt;/p&gt;

&lt;p&gt;One user message can therefore mean &lt;strong&gt;five to ten LLM calls&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User presses Enter
  → Call A: “I should read this file”
  → harness reads the file
  → Call B: “I should inspect another file”
  → ...
  → a response with no tool request
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only then does the user-visible turn finish.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is the loop.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The minimal tool-use loop is deceptively small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;has_tool_use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;execute_tools&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_use&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In plain language:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;while True&lt;/code&gt;: keep going until something explicitly ends the turn.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;call_llm(messages)&lt;/code&gt;: send the complete current history and get the next response.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;has_tool_use&lt;/code&gt;: check whether the model asked to use tools.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;execute_tools&lt;/code&gt; plus two appends: run the requested tools, then retain both the model’s request and the result.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;break&lt;/code&gt;: if there is no tool request, the turn is over.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The core is only five actions: call the model, inspect the response, execute tools when requested, append the results, and stop when no more tools are needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The user is only at the beginning and the end
&lt;/h2&gt;

&lt;p&gt;In this loop, the user appears at two points: pressing Enter at the start and seeing the final answer at the end. The middle—model call, tool execution, result append, and next model call—runs automatically. The harness does not ask the user to approve each ordinary step.&lt;/p&gt;

&lt;p&gt;This is the deepest difference between a chatbot and an agent. A chatbot normally waits for a person every turn. An agent waits for a person only at the boundaries; it drives itself through the middle of the task.&lt;/p&gt;

&lt;p&gt;That self-driving property makes several mechanisms necessary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Interrupts:&lt;/strong&gt; a user needs a way to stop a loop that will otherwise keep running.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Permission approval:&lt;/strong&gt; dangerous actions, such as destructive shell commands, must be able to pause the loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;maxTurns&lt;/code&gt; circuit breaker:&lt;/strong&gt; an accidental loop needs a hard ceiling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hooks:&lt;/strong&gt; custom logic must be inserted through lifecycle hooks because the user is not manually operating the intermediate steps.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without automatic iteration, none of these would be particularly important.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five-line happy path is not the product
&lt;/h2&gt;

&lt;p&gt;The toy loop assumes everything works. A real product must handle a much harsher world:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the context window fills up;&lt;/li&gt;
&lt;li&gt;the API fails due to a network problem, rate limit, or overload;&lt;/li&gt;
&lt;li&gt;the model refuses a request;&lt;/li&gt;
&lt;li&gt;the user interrupts while a tool is running;&lt;/li&gt;
&lt;li&gt;a tool crashes;&lt;/li&gt;
&lt;li&gt;output reaches &lt;code&gt;max_tokens&lt;/code&gt; and is cut off;&lt;/li&gt;
&lt;li&gt;output is too long, requiring a fallback model or recovery path;&lt;/li&gt;
&lt;li&gt;several tool uses may need parallel or serialized execution;&lt;/li&gt;
&lt;li&gt;the user sends another message mid-turn, so it must be queued;&lt;/li&gt;
&lt;li&gt;hooks in &lt;code&gt;.claude/settings.json&lt;/code&gt; must run at session and tool boundaries;&lt;/li&gt;
&lt;li&gt;a first-use permission prompt blocks the loop until the UI answers;&lt;/li&gt;
&lt;li&gt;a subagent starts its own loop while remaining distinct from the main one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every item is outside the five-line example.&lt;/p&gt;

&lt;h2&gt;
  
  
  From a loop to an agent runtime
&lt;/h2&gt;

&lt;p&gt;Claude Code weaves these cases into its main loop. Conceptually it looks more like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;while (true) {
    choose a route from state.transition.reason:
        next_turn                    → normal LLM call
        collapse_drain_retry         → context collapse
        reactive_compact_retry       → compact and retry
        max_output_tokens_escalate   → increase the token ceiling
        max_output_tokens_recovery   → inject a “continue” message
        stop_hook_blocking           → a stop hook requires continuation
        token_budget_continuation    → continue under a token budget

    call the LLM and process its stream

    handle stop_reason:
        end_turn                     → complete when there is no tool use
        tool_use                     → runTools
        max_tokens                   → output-token recovery
        refusal                      → explain how to change model settings
        context_window_exceeded      → reactive compaction

    execute tools:
        batch by isConcurrencySafe for parallel or serial execution
        convert every tool failure into an is_error tool_result
        optionally start tools while streaming JSON is still being parsed

    append tool results to messages

    check terminal conditions:
        AbortController.aborted      → aborted_tools
        turnCount &amp;gt; maxTurns         → max_turns
        PostToolUse hook blocks      → hook_stopped
        auto-compact threshold       → compact
        prompt_too_long              → withhold and recover

    continue to the next iteration
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The simple loop has not disappeared. It is still the heart of the system. The production runtime wraps it with recovery paths, safety gates, queueing, streaming, concurrency control, lifecycle hooks, and context management.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this series will examine
&lt;/h2&gt;

&lt;p&gt;This series studies those pieces one by one: how history is retained, how state transitions choose a path, how tools are executed and scheduled, how context is compacted, how approval and interrupts work, and how subagents reuse the same machinery.&lt;/p&gt;

&lt;p&gt;The key idea to keep in mind is simple: Claude Code is not a chat UI with a few tools attached. It is a harness that repeatedly runs an LLM, observes whether it wants an action, performs that action, and feeds the result back into the next decision. Everything else in the product exists to make that loop reliable in the real world.&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Code Tools Deep Dive (14): The Background Mechanism</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Tue, 08 Sep 2026 04:15:34 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-14-the-background-mechanism-332e</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-14-the-background-mechanism-332e</guid>
      <description>&lt;p&gt;This is the fourteenth article in my Claude Code tools series. The first thirteen articles took tools apart one by one. This article is different: it examines a &lt;strong&gt;cross-tool, orthogonal capability&lt;/strong&gt;—the Background mechanism.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This series begins with a prerequisite article explaining what tools are and how Claude uses them. Here I apply the same four lenses—naming, tool descriptions, field descriptions, and schema—to a capability distributed across several tools.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why Background deserves its own article
&lt;/h2&gt;

&lt;p&gt;There is no tool literally named &lt;code&gt;Background&lt;/code&gt; in Claude Code. That is a design choice, not an omission. Background behavior is spread across the ecosystem:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;th&gt;Form&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bash &lt;code&gt;run_in_background: true&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Parameter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent &lt;code&gt;run_in_background: true&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Parameter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monitor&lt;/td&gt;
&lt;td&gt;A tool whose purpose is continuous background listening&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CronCreate&lt;/td&gt;
&lt;td&gt;A future, background trigger&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TaskStop&lt;/td&gt;
&lt;td&gt;A tool for stopping background work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TaskOutput&lt;/td&gt;
&lt;td&gt;Explicit output retrieval (now deprecated)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;&amp;lt;task-notification&amp;gt;&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Completion notification&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Looking at only one of these reveals only a fragment. Together they form one system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Background is an execution mode, not a behavior
&lt;/h2&gt;

&lt;p&gt;Every ordinary tool has a core behavior: Read reads a file, Bash runs a command, and Agent delegates to another Claude. &lt;strong&gt;Background is not another behavior. It is a mode of execution.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The same Bash or Agent action can run in two modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Synchronous (default):&lt;/strong&gt; call → block → receive the result → continue the conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Background (&lt;code&gt;run_in_background: true&lt;/code&gt;):&lt;/strong&gt; call → return a task ID immediately → the main Claude continues → the harness sends a notification when the work completes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Making Background a family of separate tools (&lt;code&gt;BackgroundBash&lt;/code&gt;, &lt;code&gt;BackgroundAgent&lt;/code&gt;, and so on) would double the tool count and the decision burden. Claude Code instead parameterizes the existing behavior. This is the same orthogonal design used by Grep’s &lt;code&gt;output_mode&lt;/code&gt; and Edit’s &lt;code&gt;replace_all&lt;/code&gt;: keep the behavior fixed and switch the mode with a field.&lt;/p&gt;

&lt;p&gt;The analogy to Unix is useful. &lt;code&gt;fork()&lt;/code&gt; creates a child process, the task ID acts like a PID, and &lt;code&gt;&amp;lt;task-notification&amp;gt;&lt;/code&gt; resembles &lt;code&gt;SIGCHLD&lt;/code&gt;. Claude Code has built a process-management system at the harness layer, except a child may be a shell process, another Claude, or a WebSocket connection.&lt;/p&gt;

&lt;h2&gt;
  
  
  One task-ID system
&lt;/h2&gt;

&lt;p&gt;Background Bash, background Agent, Cron jobs, and Monitor instances all receive a handle that the harness can track:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bash(command: "long-training.py", run_in_background: true) → task_id
Agent(prompt: "...", run_in_background: true) → task_id
CronCreate(cron: "*/5 * * * *", recurring: true, prompt: "...") → job_id
Monitor(command: "tail -f log | grep ERROR", persistent: true) → monitor_id
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All can be stopped through the same &lt;code&gt;TaskStop&lt;/code&gt; interface. The notifications differ by type:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Bash completion includes an output-file path.&lt;/li&gt;
&lt;li&gt;Agent completion includes the agent’s result.&lt;/li&gt;
&lt;li&gt;Cron fires a new turn using its prompt.&lt;/li&gt;
&lt;li&gt;Monitor turns each stdout line or WebSocket event into a message.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The interface is unified while the semantics remain specialized. That is good API design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three kinds of background work
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Type A: a one-shot task with a known end
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Examples:&lt;/strong&gt; background Bash and background Agent.&lt;/p&gt;

&lt;p&gt;The task starts now, ends later, and the harness sends one &lt;code&gt;&amp;lt;task-notification&amp;gt;&lt;/code&gt; when it finishes. Typical uses include a long test, build, training run, research subagent, or dependency installation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bash(command: "...", run_in_background: true) → task_id
    (conversation continues)
[task-notification: completed, output: /tmp/.../out.log]
Read("/tmp/.../out.log")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Type B: trigger at a future time
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; CronCreate.&lt;/p&gt;

&lt;p&gt;The job has not started yet. CronCreate records a prompt and a schedule; when the time arrives, the runtime starts a new turn with that prompt. This is useful for a reminder, a check five minutes from now, or tomorrow’s morning self-check.&lt;/p&gt;

&lt;p&gt;That is the key distinction: Type A is an already-running task waiting to finish; Type B is a not-yet-started task waiting for its trigger.&lt;/p&gt;

&lt;h3&gt;
  
  
  Type C: continuous or long-lived listening
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; Monitor, especially &lt;code&gt;persistent: true&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The listener starts now, has no fixed end (or ends at timeout, stream close, or an explicit stop), and sends a notification for every event. Typical uses are following errors in a log, watching filesystem changes, subscribing to a WebSocket, or following a PR until merge.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;One-shot task&lt;/th&gt;
&lt;th&gt;Scheduled trigger&lt;/th&gt;
&lt;th&gt;Continuous listener&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Notifications&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;One&lt;/strong&gt; on completion&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;N&lt;/strong&gt; fires&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Unbounded&lt;/strong&gt;, one per event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lifecycle&lt;/td&gt;
&lt;td&gt;Started, then finishes&lt;/td&gt;
&lt;td&gt;Not started, waiting&lt;/td&gt;
&lt;td&gt;Started, stream continues&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interface&lt;/td&gt;
&lt;td&gt;Bash / Agent&lt;/td&gt;
&lt;td&gt;CronCreate&lt;/td&gt;
&lt;td&gt;Monitor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stop&lt;/td&gt;
&lt;td&gt;Natural completion or TaskStop&lt;/td&gt;
&lt;td&gt;CronDelete&lt;/td&gt;
&lt;td&gt;TaskStop&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The anti-polling principle
&lt;/h2&gt;

&lt;p&gt;Background should create a simple intuition:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When the harness can notify you, do not sleep. When a task is already in the background, do not synchronously wait for it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Do not start a background task, sleep for a minute, and then read its output. The harness will notify you; keep working and read the output after the notification.&lt;/p&gt;

&lt;p&gt;Do not repeatedly &lt;code&gt;curl&lt;/code&gt; a CI endpoint followed by &lt;code&gt;sleep 60&lt;/code&gt;. Choose the primitive that matches the semantics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a command that eventually exits → Bash &lt;code&gt;run_in_background&lt;/code&gt; with an &lt;code&gt;until&lt;/code&gt; loop;&lt;/li&gt;
&lt;li&gt;a stream of events → Monitor;&lt;/li&gt;
&lt;li&gt;one check at a known time → CronCreate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Claude Code also warns that &lt;strong&gt;long leading &lt;code&gt;sleep&lt;/code&gt; commands are blocked&lt;/strong&gt;. This is a hard system constraint, not merely advice. The tool layer actively pushes Claude toward asynchronous thinking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Task ID, Job ID, and Task (the to-do item)
&lt;/h2&gt;

&lt;p&gt;The word “task” is overloaded:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Task (&lt;code&gt;TaskCreate&lt;/code&gt;, &lt;code&gt;TaskList&lt;/code&gt;, …)&lt;/td&gt;
&lt;td&gt;Task family&lt;/td&gt;
&lt;td&gt;A conceptual to-do item&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;task_id&lt;/code&gt; / &lt;code&gt;&amp;lt;task-notification&amp;gt;&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Background&lt;/td&gt;
&lt;td&gt;A running task instance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ID returned by CronCreate&lt;/td&gt;
&lt;td&gt;Cron family&lt;/td&gt;
&lt;td&gt;A scheduled job&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Task-family tasks may not have started. Background tasks are concrete running processes or listeners. &lt;code&gt;TaskStop&lt;/code&gt; and the former &lt;code&gt;TaskOutput&lt;/code&gt; control the latter, not the to-do list. A useful memory aid is that &lt;code&gt;TaskCreate&lt;/code&gt;/&lt;code&gt;TaskList&lt;/code&gt;/&lt;code&gt;TaskGet&lt;/code&gt;/&lt;code&gt;TaskUpdate&lt;/code&gt; feel like CRUD, while &lt;code&gt;TaskStop&lt;/code&gt; is runtime control.&lt;/p&gt;

&lt;h2&gt;
  
  
  A concurrent workflow
&lt;/h2&gt;

&lt;p&gt;Imagine the user wants to start three things at once: run a 15-minute test suite, ask a subagent to research the auth architecture, and start a dev server while watching its log.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bash(command: "pnpm test:all", run_in_background: true) → task_test
Agent(prompt: "research auth architecture", run_in_background: true) → task_auth
Bash(command: "pnpm dev", run_in_background: true) → task_dev
Monitor(
  command: "tail -f /tmp/dev.log | grep -E --line-buffered 'error|warn'",
  description: "dev server errors"
) → monitor_dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Multiple tool calls in one message start concurrently. The main Claude remains available for conversation. The test completion produces a task notification; the research completion produces another; a dev-log error arrives immediately through Monitor; and &lt;code&gt;TaskStop(task_dev)&lt;/code&gt; stops the server. Traditional REPL execution would serialize these tasks. Background turns them into parallel work, with the longest task determining the total elapsed time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Boundaries
&lt;/h2&gt;

&lt;p&gt;Background is not unlimited:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Session-only:&lt;/strong&gt; jobs die when the session ends. Use system cron, launchd, GitHub Actions, or a cloud scheduler for multi-day work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Notifications wait for idle:&lt;/strong&gt; a task notification does not interrupt Claude while it is processing a user turn; it queues until the REPL is idle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate limiting:&lt;/strong&gt; high-volume Monitor streams and accumulated output can be stopped or truncated. Strong filters are essential.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrency limits:&lt;/strong&gt; Bash and Agent jobs are queued beyond the runtime’s host-dependent limit (typically &lt;code&gt;min(16, cpu-2)&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sandboxing remains active:&lt;/strong&gt; background execution does not bypass the sandbox. A separate, explicit dangerous-sandbox option is required to change that boundary.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The four-layer design signal
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Naming
&lt;/h3&gt;

&lt;p&gt;The most important naming decision is the absence of &lt;code&gt;BackgroundBash&lt;/code&gt; or &lt;code&gt;BackgroundAgent&lt;/code&gt;. One Boolean, &lt;code&gt;run_in_background&lt;/code&gt;, declares that Background is an execution mode, not a new behavior.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;TaskStop&lt;/code&gt; is also deliberately generic. Stopping a background Bash task is not &lt;code&gt;BashStop&lt;/code&gt;; stopping an Agent is not &lt;code&gt;AgentStop&lt;/code&gt;. The single verb resembles &lt;code&gt;kill &amp;lt;pid&amp;gt;&lt;/code&gt;: whatever was forked, the same control operation stops it.&lt;/p&gt;

&lt;p&gt;The old &lt;code&gt;TaskOutput&lt;/code&gt; has been de-emphasized in favor of reading the output file with Read. If an existing primitive covers the capability, the API does not add another verb. Cron and Monitor keep their specialized names externally while sharing the task infrastructure internally.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tool-level descriptions
&lt;/h3&gt;

&lt;p&gt;Descriptions carry most of the normative guidance:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;avoid unnecessary &lt;code&gt;sleep&lt;/code&gt; and polling;&lt;/li&gt;
&lt;li&gt;use background Bash for one-shot waiting;&lt;/li&gt;
&lt;li&gt;use Monitor for streaming events;&lt;/li&gt;
&lt;li&gt;strong filters are required because every event costs conversation context;&lt;/li&gt;
&lt;li&gt;Cron jobs are session-only, not system cron jobs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;“Long leading &lt;code&gt;sleep&lt;/code&gt; commands are blocked” upgrades a recommendation into an enforced constraint. Monitor’s warnings about buffering, rate limits, and raw logs do the same for event streams. These descriptions are cross-tool contracts encoded in individual tools.&lt;/p&gt;

&lt;h3&gt;
  
  
  Field-level descriptions
&lt;/h3&gt;

&lt;p&gt;The most visible design signal is that Bash and Agent use the same field with opposite defaults:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;Typical intent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bash&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Most shell commands are short; synchronous is simpler&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;&lt;code&gt;true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Subagents usually take time; keep the main Claude free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Bash’s field description says to use background only when the result is not needed immediately, promises a later notification, and warns not to append &lt;code&gt;&amp;amp;&lt;/code&gt;. Agent’s description explains when to opt out and run in the foreground. The same field is a mirror: the defaults reflect the typical lifecycle of each tool.&lt;/p&gt;

&lt;p&gt;The output path in &lt;code&gt;&amp;lt;task-notification&amp;gt;&lt;/code&gt; naturally directs Claude to Read the result. Monitor’s &lt;code&gt;persistent&lt;/code&gt; Boolean distinguishes a bounded listener from a session-long listener.&lt;/p&gt;

&lt;h3&gt;
  
  
  Schema validation
&lt;/h3&gt;

&lt;p&gt;Background has a thin schema layer because it modifies existing synchronous behavior rather than replacing it. The important safeguards are runtime-level: Monitor’s one-hour timeout, &lt;code&gt;persistent&lt;/code&gt; semantics, cron-expression validation, unified task IDs, concurrency limits, and sandbox enforcement.&lt;/p&gt;

&lt;p&gt;The thinness is evidence for the central design choice. If Background were a new capability, it would need an independent schema. Because it is a mode, the existing &lt;code&gt;command&lt;/code&gt;, &lt;code&gt;prompt&lt;/code&gt;, and &lt;code&gt;cron&lt;/code&gt; schemas are mostly enough; Background adds a flag or a trigger policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Relationship to neighboring tools
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool family&lt;/th&gt;
&lt;th&gt;Background expression&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bash&lt;/td&gt;
&lt;td&gt;&lt;code&gt;run_in_background: false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Foreground&lt;/td&gt;
&lt;td&gt;Long commands without blocking the main loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;&lt;code&gt;run_in_background: true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Background&lt;/td&gt;
&lt;td&gt;Subagents normally have a long lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Task family&lt;/td&gt;
&lt;td&gt;IDs, TaskStop&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;State and a stop interface&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cron family&lt;/td&gt;
&lt;td&gt;Time-driven wakeups&lt;/td&gt;
&lt;td&gt;Scheduled&lt;/td&gt;
&lt;td&gt;Time primitive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monitor&lt;/td&gt;
&lt;td&gt;Event-driven wakeups&lt;/td&gt;
&lt;td&gt;Event stream&lt;/td&gt;
&lt;td&gt;Event primitive&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The anti-polling principle connects Bash, Agent, and Monitor: do not poll a job that will notify, do not schedule checks for an Agent that is already running, and do not use a continuous &lt;code&gt;tail -f&lt;/code&gt; for a one-shot event.&lt;/p&gt;

&lt;p&gt;The first thirteen tools are mostly synchronous: Read, Edit, Write, Grep, Glob, WebFetch, WebSearch, the interaction tools, and so on. Background is the hidden implementation layer that lets those spatial primitives extend into asynchronous collaboration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: the invisible skeleton
&lt;/h2&gt;

&lt;p&gt;Background’s elegance is not simply that it makes AI asynchronous. It is that every layer reinforces &lt;strong&gt;parameterization instead of tool proliferation&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Naming:&lt;/strong&gt; no independent tool, just &lt;code&gt;run_in_background&lt;/code&gt; and a unified TaskStop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool descriptions:&lt;/strong&gt; anti-polling, filter-first, session-only, and notification contracts are distributed across the relevant tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fields:&lt;/strong&gt; Bash and Agent have opposite defaults because their normal lifecycles differ.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema:&lt;/strong&gt; deliberately thin, with timeout, persistence, concurrency, and sandbox safeguards as the backstop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The missing schema is itself evidence. Background is not a new action; it is a mode attached to existing actions. Without it, Claude Code would be a one-action-at-a-time assistant. With it, Claude becomes a multi-threaded collaborator that can run long jobs, keep talking, react to events, and schedule future work.&lt;/p&gt;

&lt;p&gt;The series now closes with a complete map: user alignment → locating → perceiving → executing → fallback → scaling through subagents, time, event streams, and background execution. The goal was never to enumerate tools, but to apply the same four-layer anatomy to each one. The reusable lesson is &lt;strong&gt;restraint&lt;/strong&gt;: each tool does one small thing, and composition creates the collaboration system.&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Code Tools Deep Dive (13): Monitor</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Tue, 01 Sep 2026 07:23:28 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-13-monitor-17n6</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-13-monitor-17n6</guid>
      <description>&lt;p&gt;This is the thirteenth article in my Claude Code tools series. The previous article, &lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-12-the-cron-family-1cnj"&gt;The Cron Family&lt;/a&gt;, explained how Claude can trigger actions &lt;strong&gt;across time&lt;/strong&gt;. Cron is clock-driven: it fires at a scheduled moment, regardless of what is happening outside.&lt;/p&gt;

&lt;p&gt;Engineering has another kind of waiting: &lt;strong&gt;wait until something happens&lt;/strong&gt;, without knowing exactly when. For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“Tell me when &lt;code&gt;ERROR&lt;/code&gt; appears in the logs.”&lt;/li&gt;
&lt;li&gt;“Rebuild when the file changes.”&lt;/li&gt;
&lt;li&gt;“Notify me when the PR status changes.”&lt;/li&gt;
&lt;li&gt;“Report each CI check as it settles.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These cases need an &lt;strong&gt;event-driven asynchronous waiting primitive&lt;/strong&gt;. Claude places a set of sensors, and the runtime reports external events automatically. That is the purpose of &lt;strong&gt;Monitor&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Monitor
&lt;/h2&gt;

&lt;p&gt;Monitor is Claude Code’s built-in &lt;strong&gt;event-stream listener&lt;/strong&gt;. Together with Bash background execution and Cron, it completes the three basic asynchronous waiting primitives:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Trigger&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bash &lt;code&gt;run_in_background&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One task completes&lt;/td&gt;
&lt;td&gt;“Tell me when the build is done.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CronCreate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A clock reaches a time&lt;/td&gt;
&lt;td&gt;“Remind me at 9.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Monitor&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;An event stream, one stdout line per event&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;“Tell me every time X happens.”&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two wait for one point. Monitor waits for a &lt;strong&gt;line&lt;/strong&gt;: the stream may continue indefinitely, until a timeout or Claude stops it.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it solves
&lt;/h3&gt;

&lt;p&gt;Monitor answers the question: &lt;strong&gt;how can Claude continuously perceive changes in the outside world?&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;More than one notification:&lt;/strong&gt; Bash background execution normally reports once; Monitor reports every event.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Event-stream modeling:&lt;/strong&gt; each stdout line is a notification, naturally aligned with Unix conventions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two data sources:&lt;/strong&gt; a shell command &lt;strong&gt;or&lt;/strong&gt; a direct WebSocket connection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forces filter design:&lt;/strong&gt; the prompt makes Claude decide what deserves a notification and what should be ignored.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persistent listening:&lt;/strong&gt; &lt;code&gt;persistent: true&lt;/code&gt; can keep a monitor alive for the whole session, useful for PRs and long-running logs.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Its fundamental difference from Cron is simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cron&lt;/strong&gt; is clock-driven: time is the active party.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitor&lt;/strong&gt; is event-driven: the outside event is the active party.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cron says “I will ask you when the time comes.” Monitor says “call me when something happens.” One is pull; the other is push.&lt;/p&gt;

&lt;h3&gt;
  
  
  A concrete example
&lt;/h3&gt;

&lt;p&gt;Suppose the user says: &lt;strong&gt;“I am starting a 20-minute model training run. Watch the log, tell me about errors immediately, and report progress too.”&lt;/strong&gt; This is a long-running job whose events occur at unknown times.&lt;/p&gt;

&lt;h4&gt;
  
  
  Bad alternative 1: sleep and inspect later
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bash(command: "sleep 1200 &amp;amp;&amp;amp; cat train.log", timeout: 1300000)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the process fails after three minutes, Claude does not know until the end. Early signals are lost.&lt;/p&gt;

&lt;h4&gt;
  
  
  Bad alternative 2: scheduled polling
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CronCreate(cron: "*/2 * * * *", recurring: true, prompt: "check train.log and report ERROR")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Polling every two minutes introduces up to two minutes of latency and repeatedly rereads the file. Each Cron wakeup also consumes conversation context.&lt;/p&gt;

&lt;h4&gt;
  
  
  Bad alternative 3: a background &lt;code&gt;tail&lt;/code&gt;
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bash(command: "tail -f train.log", run_in_background: true)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A background Bash job normally notifies once, when it exits. &lt;code&gt;tail -f&lt;/code&gt; never exits, so it never sends the useful notifications.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Monitor solution
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Monitor(
  command: "tail -f train.log | grep -E --line-buffered 'elapsed_steps=|Traceback|Error|FAILED|Killed|OOM'",
  description: "training log: progress and errors",
  timeout_ms: 1500000
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The runtime starts the shell command and follows the log. &lt;code&gt;grep&lt;/code&gt; lets only matching lines through. &lt;strong&gt;Each stdout line becomes a notification&lt;/strong&gt; delivered to the conversation immediately. Claude can continue talking or do other work, while progress and failures arrive automatically. After 20 minutes the timeout ends the monitor, or the user can stop it earlier.&lt;/p&gt;

&lt;p&gt;The key insight is that Monitor turns Claude from an active poller into a passive receiver. Every relevant external event is known immediately, without repeated polling or a blocked context.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two data sources: command or WebSocket
&lt;/h3&gt;

&lt;p&gt;Monitor has an unusually rich design: the source can be a shell command or a WebSocket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shell command&lt;/strong&gt; is the common mode. Every line written to stdout is an event.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;WebSocket&lt;/strong&gt; can be used directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Monitor(
  ws: { url: "wss://events.example.com/stream", protocols: ["v1"] },
  description: "deployment event stream",
  timeout_ms: 300000
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The runtime opens the connection; every text frame is one event, while a binary frame is represented as &lt;code&gt;[binary frame, N bytes]&lt;/code&gt;. Closing the connection ends the monitor.&lt;/p&gt;

&lt;p&gt;Using &lt;code&gt;websocat&lt;/code&gt; through &lt;code&gt;command&lt;/code&gt; would work, but adds quoting, process, installation, and buffering problems. Built-in WebSocket support removes a process and normalizes the mapping from frames to events. It is a strong example of using a tool to eliminate fragile plumbing, and suggests concrete use cases such as agent-to-agent communication, deployment subscriptions, and long-lived push channels.&lt;/p&gt;

&lt;h3&gt;
  
  
  When to use Monitor
&lt;/h3&gt;

&lt;p&gt;The tool prompt gives a useful selection rule:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use Monitor when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;every occurrence of X should generate a notification;&lt;/li&gt;
&lt;li&gt;every occurrence should be reported until a known terminal condition;&lt;/li&gt;
&lt;li&gt;you need to consume a WebSocket event stream.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Do not use Monitor when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;you only need one completion notification: use Bash &lt;code&gt;run_in_background&lt;/code&gt; with a loop that eventually exits;&lt;/li&gt;
&lt;li&gt;you need a clock trigger: use CronCreate;&lt;/li&gt;
&lt;li&gt;events arrive at a very high rate: tighten the filter, because rate limiting may stop the monitor.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important warning is: &lt;strong&gt;“Don’t use an unbounded command for a single notification.”&lt;/strong&gt; For “tell me once when the build is ready,” use a background Bash command such as &lt;code&gt;until grep -q "Ready" dev.log; do sleep 0.5; done&lt;/code&gt;. Do not use &lt;code&gt;tail -f ... | grep -m 1 "Ready"&lt;/code&gt;: &lt;code&gt;tail -f&lt;/code&gt; may remain alive after the match, leaving Monitor attached until timeout. Monitor is optimized for continuous events, not one-shot completion.&lt;/p&gt;

&lt;h3&gt;
  
  
  Technical design
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Naming
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;Monitor&lt;/code&gt; is a neutral SRE term for continuous observation and alerting. It is broader than &lt;code&gt;Watch&lt;/code&gt;, &lt;code&gt;Tail&lt;/code&gt;, &lt;code&gt;Subscribe&lt;/code&gt;, or &lt;code&gt;Listen&lt;/code&gt;, and steers Claude toward “place a watch and report events,” not “grep the file once.”&lt;/p&gt;

&lt;h4&gt;
  
  
  The tool-level description is a mini operations guide
&lt;/h4&gt;

&lt;p&gt;The prompt covers notification choice, event-stream semantics, output volume, buffering, data-source preference, and observability completeness.&lt;/p&gt;

&lt;p&gt;It explicitly says that each stdout line is an event, then classifies tools by notification count:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One notification:&lt;/strong&gt; Bash with &lt;code&gt;run_in_background&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One per occurrence indefinitely:&lt;/strong&gt; Monitor with an unbounded command.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One per occurrence until a known end:&lt;/strong&gt; Monitor with a command that emits lines and exits.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It also teaches Unix buffering. Every pipeline stage must flush per line: use &lt;code&gt;grep --line-buffered&lt;/code&gt; and &lt;code&gt;awk&lt;/code&gt; with &lt;code&gt;fflush()&lt;/code&gt;. Avoid &lt;code&gt;head&lt;/code&gt;, which may wait for N matches before producing output. This tribal sysadmin knowledge is placed directly in the tool prompt so a naïve command does not make events appear delayed.&lt;/p&gt;

&lt;p&gt;The deepest rule is &lt;strong&gt;“silence is not success.”&lt;/strong&gt; A filter must match every terminal state, not only the happy path. A monitor that watches only &lt;code&gt;elapsed_steps=&lt;/code&gt; stays silent through a crash, hang, or unexpected exit, making failure indistinguishable from “still running.” Before arming a monitor, ask: &lt;em&gt;if the process crashed right now, would my filter emit anything?&lt;/em&gt; If not, widen it to include &lt;code&gt;Traceback&lt;/code&gt;, &lt;code&gt;Error&lt;/code&gt;, &lt;code&gt;FAILED&lt;/code&gt;, &lt;code&gt;Killed&lt;/code&gt;, &lt;code&gt;OOM&lt;/code&gt;, and similar signals.&lt;/p&gt;

&lt;p&gt;“Selective” also does not mean “only good news.” Select the lines you would act on, whether they describe progress or failure. If output becomes excessive, the runtime automatically stops the monitor; Claude should restart it with a tighter filter. Lines arriving within 200 ms are batched into one notification, so a multiline traceback remains readable as one event.&lt;/p&gt;

&lt;p&gt;Finally, the prompt prefers the native &lt;code&gt;ws&lt;/code&gt; source over &lt;code&gt;command: 'websocat wss://…'&lt;/code&gt;, avoiding an extra process and another buffering layer.&lt;/p&gt;

&lt;h4&gt;
  
  
  Fields and runtime rules
&lt;/h4&gt;

&lt;p&gt;Monitor has five fields:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;command&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Shell source; mutually exclusive with &lt;code&gt;ws&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ws&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;WebSocket source with &lt;code&gt;url&lt;/code&gt; and &lt;code&gt;protocols&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;description&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Required label shown with every notification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;timeout_ms&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Defaults to 300,000 ms; maximum 3,600,000 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;persistent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Defaults to false; true keeps it alive for the session&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The schema is moderate, but important constraints live in the runtime:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Exactly one of &lt;code&gt;command&lt;/code&gt; and &lt;code&gt;ws&lt;/code&gt; must be supplied.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;description&lt;/code&gt; is required because it is visible in every notification.&lt;/li&gt;
&lt;li&gt;A non-persistent monitor cannot exceed the one-hour timeout.&lt;/li&gt;
&lt;li&gt;Excessive output triggers rate limiting and an explicit stop.&lt;/li&gt;
&lt;li&gt;With &lt;code&gt;persistent: true&lt;/code&gt;, &lt;code&gt;timeout_ms&lt;/code&gt; is ignored; the monitor ends with the session or an explicit &lt;code&gt;TaskStop&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These are loud failures or loud stops. A bad monitor cannot silently look healthy while doing nothing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Division of responsibility
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Bash background&lt;/th&gt;
&lt;th&gt;CronCreate&lt;/th&gt;
&lt;th&gt;Monitor&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Waits for&lt;/td&gt;
&lt;td&gt;One task to finish&lt;/td&gt;
&lt;td&gt;A clock time&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;An event stream&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Notifications&lt;/td&gt;
&lt;td&gt;One on process exit&lt;/td&gt;
&lt;td&gt;One per schedule match&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;One per event&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wakeup&lt;/td&gt;
&lt;td&gt;Process exit&lt;/td&gt;
&lt;td&gt;Scheduled moment&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;stdout line or WebSocket frame&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sources&lt;/td&gt;
&lt;td&gt;Shell command&lt;/td&gt;
&lt;td&gt;Cron expression&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Command or WebSocket&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical use&lt;/td&gt;
&lt;td&gt;“Wait for CI”&lt;/td&gt;
&lt;td&gt;“Check every five minutes”&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;“Alert on every log error”&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conservative bias&lt;/td&gt;
&lt;td&gt;Notify on exit&lt;/td&gt;
&lt;td&gt;Fire at the time&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Emit only actionable signals&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Monitor completes the waiting model begun by the first twelve tools. Bash waits for a point, Cron waits for a time, and Monitor waits for a line. Compared with Tasks, Task state is pulled by Claude from storage; Monitor state is pushed from the outside world. Compared with Bash polling, Monitor upgrades &lt;code&gt;sleep + poll&lt;/code&gt; into a structured event-stream primitive.&lt;/p&gt;

&lt;h3&gt;
  
  
  Summary
&lt;/h3&gt;

&lt;p&gt;Monitor’s most impressive feature is not merely continuous listening. It embeds an observability methodology in the tool description:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a minimal SRE-oriented name;&lt;/li&gt;
&lt;li&gt;a long prompt explaining notification choice, buffering, rate limiting, batching, and “silence is not success”;&lt;/li&gt;
&lt;li&gt;only five fields, each backed by a meaningful runtime decision;&lt;/li&gt;
&lt;li&gt;hard runtime protection for source selection, timeouts, and output volume.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The schema locks down the basic shape, while the prompt teaches Claude how to build a reliable watch: event-driven, failure-visible, resistant to conversation flooding, and available over both shell and WebSocket sources.&lt;/p&gt;

&lt;p&gt;The next article examines &lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-14-the-background-mechanism-332e"&gt;Background mechanisms&lt;/a&gt;, the final article in the series. It crosses tool boundaries and follows &lt;code&gt;run_in_background&lt;/code&gt; through Bash, Agent, the Task family, and Monitor, showing how asynchronous execution becomes a first-class Claude Code semantic.&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Code Tools Deep Dive (12): The Cron Family</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Wed, 26 Aug 2026 11:26:04 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-12-the-cron-family-1cnj</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-12-the-cron-family-1cnj</guid>
      <description>&lt;p&gt;This is the twelfth article in my series on Claude Code tools. The first eleven explored Claude Code’s &lt;strong&gt;spatial toolkit&lt;/strong&gt;: from the local filesystem to the public web, from one Claude to multiple Claudes, and from immediate actions to persistent task lists. Their temporal model is fundamentally &lt;strong&gt;synchronous&lt;/strong&gt;: Claude calls a tool, it executes, and a result returns immediately.&lt;/p&gt;

&lt;p&gt;Real engineering has another class of requests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“Remind me to check CI in 30 minutes.”&lt;/li&gt;
&lt;li&gt;“Check every five minutes whether the deployment is ready.”&lt;/li&gt;
&lt;li&gt;“Run a morning self-check at 9 tomorrow.”&lt;/li&gt;
&lt;li&gt;“In an hour, review this proposal again.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The common feature is that the action is not “do it now.” It is &lt;strong&gt;“automatically trigger it at a future time.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That requires a &lt;strong&gt;time primitive&lt;/strong&gt;. Claude Code’s answer is the Cron family: three tools—CronCreate, CronDelete, and CronList—that form a scheduled-execution system.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This series begins with a prerequisite article explaining what tools are and how Claude uses them. Like the other articles, this one follows the four-layer framework introduced there.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Cron family: CronCreate / CronDelete / CronList
&lt;/h2&gt;

&lt;p&gt;Like the Task family, these three tools are semantically coupled and share one data model: the session’s list of cron jobs. They are clearer as a family than as isolated operations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Family overview
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CronCreate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Create a future prompt trigger using standard five-field cron syntax&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CronDelete&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cancel a scheduled job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CronList&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;List all jobs scheduled in the current session&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A related tool:&lt;/strong&gt; ScheduleWakeup is a specialized cousin used by the &lt;code&gt;/loop&lt;/code&gt; skill to schedule its next self-wakeup. It is optimized for dynamic loops and is worth mentioning alongside Cron.&lt;/p&gt;

&lt;p&gt;The core division is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CronCreate is the engine.&lt;/strong&gt; Most calls create jobs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CronList and CronDelete are management tools.&lt;/strong&gt; They inspect and clean up.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest difference from Tasks is this: &lt;strong&gt;Task records work that remains to be done; Cron schedules an action for the future.&lt;/strong&gt; Task waits for Claude to choose the next item. Cron triggers automatically when the time arrives.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it does
&lt;/h3&gt;

&lt;p&gt;The Cron family solves how Claude can execute actions &lt;strong&gt;across time&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Break synchronous limits:&lt;/strong&gt; Claude can schedule a future self-wakeup rather than only respond to the current request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schedule precisely:&lt;/strong&gt; standard cron syntax (&lt;code&gt;M H DoM Mon DoW&lt;/code&gt;) supports arbitrary times and periods.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Support one-shot and recurring modes:&lt;/strong&gt; a &lt;code&gt;recurring&lt;/code&gt; boolean selects the lifecycle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provide lightweight reminders:&lt;/strong&gt; “remind me in 30 minutes” does not require a background task.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Proactively observe external state:&lt;/strong&gt; Claude can check CI or a deployment at a scheduled time.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is the first tool family that crosses time. The previous eleven tools are actions at a &lt;strong&gt;point&lt;/strong&gt;: a tool call happens and finishes. Cron is a schedule on a &lt;strong&gt;timeline&lt;/strong&gt;: mark a point, then let the runtime fire automatically.&lt;/p&gt;

&lt;h3&gt;
  
  
  A concrete example
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; The user says, &lt;strong&gt;“I just pushed a deployment. It should finish in about eight minutes. Check the CI status then and tell me if anything is wrong.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is a classic “wait for external state to change” task.&lt;/p&gt;

&lt;h4&gt;
  
  
  Bad alternative 1: sleep
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bash(command: "sleep 480 &amp;amp;&amp;amp; gh run list", timeout: 500000)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The main Claude is blocked for eight minutes and cannot answer another question. Synchronous waiting wastes conversational time.&lt;/p&gt;

&lt;h4&gt;
  
  
  Bad alternative 2: poll every minute
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;while true:
    Bash(command: "gh run list")
    sleep 60
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This consumes context once per minute. Eight minutes means eight calls, and logs fill the main context. The context budget is spent on waiting.&lt;/p&gt;

&lt;h4&gt;
  
  
  How CronCreate solves it
&lt;/h4&gt;

&lt;p&gt;Claude schedules a one-shot wakeup eight minutes in the future:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CronCreate(
  cron: "13 22 29 7 *",           # one trigger at 22:13 on July 29
  recurring: false,
  prompt: "Check CI with gh run list. Tell the user if it failed; otherwise confirm briefly."
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;At runtime:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the job is stored in session memory&lt;/li&gt;
&lt;li&gt;the main Claude immediately returns to the user instead of blocking&lt;/li&gt;
&lt;li&gt;the user can ask other questions or start other work&lt;/li&gt;
&lt;li&gt;at 22:13, the runtime invokes the prompt as a new Claude call&lt;/li&gt;
&lt;li&gt;Claude runs &lt;code&gt;gh run list&lt;/code&gt; and reports the status&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The experience looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[22:05] User: I pushed a deployment; check CI in eight minutes.
[22:05] Claude: Done—I scheduled an automatic check for 22:13.
              You can keep working in the meantime.
[22:05–22:12] User: continues with other work
[22:13] Claude: CI check complete; all three workflows are green.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Key insight:&lt;/strong&gt; CronCreate moves the responsibility for waiting from the main Claude to the runtime. Claude schedules the work and gets out of the way, consuming neither conversation time nor context while waiting.&lt;/p&gt;

&lt;h4&gt;
  
  
  Combining the tools: inspect or cancel
&lt;/h4&gt;

&lt;p&gt;If the user changes their mind—“Never mind, I’ll check CI myself”—Claude can call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CronList()                         # find the job ID
CronDelete(id: "cron_xxx")        # cancel it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the user asks “What did you schedule?”, CronList provides the answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Proactive vs. passive wakeups
&lt;/h3&gt;

&lt;p&gt;Cron supports two modes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One-shot (&lt;code&gt;recurring: false&lt;/code&gt;)&lt;/strong&gt; is for known moments:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;remind me to review this PR tomorrow at 9&lt;/li&gt;
&lt;li&gt;check CI again in 30 minutes&lt;/li&gt;
&lt;li&gt;remind me to eat lunch at noon&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The minute, hour, day-of-month, and month are pinned. The job fires once and disappears.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recurring (&lt;code&gt;recurring: true&lt;/code&gt;)&lt;/strong&gt; is for monitoring with no known end:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;check CI every five minutes until I say stop&lt;/li&gt;
&lt;li&gt;inspect queue length every hour&lt;/li&gt;
&lt;li&gt;run a morning self-check every day&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Typical expressions include &lt;code&gt;*/5 * * * *&lt;/code&gt;, &lt;code&gt;0 * * * *&lt;/code&gt;, and &lt;code&gt;0 9 * * *&lt;/code&gt;. Recurring jobs live for at most seven days, fire one final time, then are deleted. This prevents forgotten jobs from consuming resources indefinitely.&lt;/p&gt;

&lt;h3&gt;
  
  
  When it is triggered
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Use Cron for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;waiting for an external asynchronous event such as CI or deployment&lt;/li&gt;
&lt;li&gt;reminders and actions at a known time&lt;/li&gt;
&lt;li&gt;periodic monitoring every N minutes&lt;/li&gt;
&lt;li&gt;handing a future check back to Claude after the current conversation ends&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Do not use Cron for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;second- or subsecond-level actions; cron has minute resolution&lt;/li&gt;
&lt;li&gt;exact event-driven responses; use Monitor&lt;/li&gt;
&lt;li&gt;waits already covered by harness notifications, such as completion of a background Bash task or subagent&lt;/li&gt;
&lt;li&gt;persistent jobs across sessions; Cron is &lt;strong&gt;session-only&lt;/strong&gt;, stored in memory and gone when Claude exits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Choose among the waiting primitives by meaning:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Need&lt;/th&gt;
&lt;th&gt;Primitive&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One-shot event notification, such as CI completion&lt;/td&gt;
&lt;td&gt;Bash &lt;code&gt;run_in_background&lt;/code&gt; with harness notification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Listening for a change with no fixed time&lt;/td&gt;
&lt;td&gt;Monitor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A reminder or one-time delay&lt;/td&gt;
&lt;td&gt;CronCreate with &lt;code&gt;recurring: false&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Periodic monitoring&lt;/td&gt;
&lt;td&gt;CronCreate with &lt;code&gt;recurring: true&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;/loop&lt;/code&gt; self-wakeup&lt;/td&gt;
&lt;td&gt;ScheduleWakeup&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cron is not the only way to wait. The right primitive depends on whether the trigger is a completion event, a change event, a clock time, or a recurring interval.&lt;/p&gt;

&lt;h3&gt;
  
  
  Technical design
&lt;/h3&gt;

&lt;h4&gt;
  
  
  1. Naming
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;CronCreate&lt;/code&gt; / &lt;code&gt;CronDelete&lt;/code&gt; / &lt;code&gt;CronList&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;This is another complete dual loop: Create attaches a job, Delete removes it, and List observes it. A scheduled job has a lifecycle—created, active, expired or cancelled—so observation and cancellation deserve first-class operations.&lt;/p&gt;

&lt;p&gt;“Cron” reuses forty years of Unix crontab convention rather than inventing a DSL. Anyone who has used &lt;code&gt;crontab -e&lt;/code&gt; on Linux or macOS already understands the idea and much of the syntax. Reusing an industry convention reduces cognitive load. &lt;code&gt;List&lt;/code&gt; is plural rather than &lt;code&gt;Get&lt;/code&gt;, signaling that multiple jobs may be returned.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Tool-level descriptions
&lt;/h4&gt;

&lt;p&gt;Cron’s descriptions cover eight concerns: &lt;strong&gt;session-only lifetime, the seven-day cap, load spreading away from :00 and :30, exceptions where exact half-hours are correct, when to use Monitor instead, language signals for one-shot versus recurring jobs, local-time semantics, and transparent jitter&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;State the session-only lifetime first&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Jobs live only in this Claude session—nothing is written to disk, and the job is gone when Claude exits.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This prevents Claude from promising a weekly job that cannot survive the session. The simplified design avoids the complexity of user authorization, multi-session synchronization, and durable-job failure handling. The tradeoff is that long-lived jobs belong in system cron or a cloud scheduler, not this family.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Explain the seven-day cap proactively&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Recurring tasks auto-expire after seven days—they fire one final time, then are deleted. This bounds session lifetime. Tell the user about the seven-day limit when scheduling recurring jobs.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Claude must tell the user about this limit when creating a recurring job. The cap is also an anti-forgetting mechanism: a forgotten hourly monitor cannot consume resources forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spread load by avoiding :00 and :30&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every user who asks for “9am” gets &lt;code&gt;0 9&lt;/code&gt;, and every user who asks for “hourly” gets &lt;code&gt;0 *&lt;/code&gt;, which means requests from across the planet land on the API at the same instant. When the user’s request is approximate, pick a minute that is NOT 0 or 30.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is a rare case where a tool prompt includes a server-operations concern. If every “9am” becomes 9:00, requests from every timezone create a synchronized load spike. Choosing :03 or :57 spreads the fleet.&lt;/p&gt;

&lt;p&gt;The prompt explains the reason, not just the rule, so Claude can reason about exceptions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Explain when :00 and :30 are correct&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Only use minute 0 or 30 when the user names that exact time and clearly means it, such as “at 9:00 sharp,” “at half past,” or coordinating with a meeting. When in doubt, nudge a few minutes early or late—the user will not notice, and the fleet will.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The exception prevents the load-spreading rule from becoming dogma. Exact meeting coordination remains exact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Point live watching to Monitor&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Not for live watching. CronCreate reruns a prompt at fixed wall-clock intervals. To watch a log file, process, or command output and be notified the moment something changes, use Monitor instead—Monitor streams events as they happen; cron polls on a schedule.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The tool explicitly points to its sibling rather than expecting Claude to compare tools unaided. Cron is scheduled polling; Monitor is event streaming.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Infer one-shot jobs from user language&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For “remind me at X” or “at &lt;time&gt;, do Y” requests—fire once, then auto-delete. Pin minute, hour, day-of-month, and month to specific values.&lt;/time&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The wording itself signals &lt;code&gt;recurring: false&lt;/code&gt;, so Claude need not ask the user whether a reminder should repeat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use local time, not UTC conversion&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Uses standard five-field cron in the user’s local timezone: minute, hour, day-of-month, month, day-of-week. &lt;code&gt;0 9 * * *&lt;/code&gt; means 9am local—no timezone conversion needed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This avoids a classic sysadmin mistake: manually converting a user’s local time to UTC and getting it wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make jitter transparent&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The scheduler adds a small deterministic jitter on top of whatever you pick.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Claude is told that a job chosen for :57 might fire at :58. Recurring tasks may be delayed by up to 10%, capped at 15 minutes; one-shot tasks scheduled exactly at :00 or :30 may fire up to 90 seconds early. The runtime spreads load even when Claude chooses a round time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fire only while the REPL is idle&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Jobs only fire while the REPL is idle, not mid-query.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If a cron job becomes due while Claude is processing a user request, it waits until the current response finishes. A scheduled trigger cannot interrupt the user’s active thought.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Field-level descriptions
&lt;/h4&gt;

&lt;p&gt;CronCreate exposes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;cron&lt;/code&gt;: five fields in the user’s local timezone—&lt;code&gt;minute hour day-of-month month day-of-week&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;prompt&lt;/code&gt;: the prompt to invoke at the scheduled time&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;recurring&lt;/code&gt;: boolean, default &lt;code&gt;true&lt;/code&gt;; &lt;code&gt;false&lt;/code&gt; creates a one-shot job&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;durable&lt;/code&gt;: a legacy field with no effect&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;CronDelete takes the &lt;code&gt;id&lt;/code&gt; returned by CronCreate. CronList takes no arguments. The interesting design work is concentrated in CronCreate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use a standard cron string instead of a new DSL&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;0 9 * * *&lt;/code&gt; means 9am every day, exactly as it does in Unix crontab. A JSON object such as &lt;code&gt;{ minute: "*/5", hour: "*", ... }&lt;/code&gt; might look more structured, but it would make both users and models learn a new language. Industry convention wins.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The &lt;code&gt;recurring: true&lt;/code&gt; default encodes a preference&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Without the field, a job repeats. That aligns with Cron’s typical use cases: CI monitoring, deployment observation, and periodic health checks. One-shot reminders are the less common case and require an explicit &lt;code&gt;recurring: false&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;durable&lt;/code&gt; is a transparent historical artifact&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The description says that &lt;code&gt;durable&lt;/code&gt; has no effect. The field likely remains for compatibility after an earlier persistence design was removed. It is neither hidden nor silently honored; Claude is told not to spend effort configuring it.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. Schema validation
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;Schema constraint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cron&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;none; required&lt;/td&gt;
&lt;td&gt;five-field shape, shallow validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;prompt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;none; required&lt;/td&gt;
&lt;td&gt;no length limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;recurring&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;&lt;code&gt;true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;durable&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;legacy field&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The meaningful constraints are runtime behaviors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;seven-day expiry for recurring jobs&lt;/li&gt;
&lt;li&gt;firing only while the REPL is idle&lt;/li&gt;
&lt;li&gt;automatic jitter and load spreading&lt;/li&gt;
&lt;li&gt;session-only lifetime&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cron’s syntax is too flexible for a schema to classify every schedule as good or bad. Both &lt;code&gt;7 * * * *&lt;/code&gt; and &lt;code&gt;0 * * * *&lt;/code&gt; are legal. Natural-language guidance teaches Claude which valid schedule best matches the user’s intent.&lt;/p&gt;

&lt;h3&gt;
  
  
  ScheduleWakeup: the specialized version
&lt;/h3&gt;

&lt;p&gt;CronCreate is the general scheduler. The &lt;code&gt;/loop&lt;/code&gt; skill uses ScheduleWakeup for dynamic-interval loops:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude itself is the caller rather than an external trigger.&lt;/li&gt;
&lt;li&gt;The previous loop prompt is passed forward automatically.&lt;/li&gt;
&lt;li&gt;The tool description accounts for a five-minute prompt-cache TTL.&lt;/li&gt;
&lt;li&gt;The usual interval is 60–1,200 seconds.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The division is simple: general scheduling uses CronCreate; self-scheduling inside &lt;code&gt;/loop&lt;/code&gt; uses ScheduleWakeup. It is Cron’s loop-specialized relative.&lt;/p&gt;




&lt;h3&gt;
  
  
  Division of responsibility among neighboring tools
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Interaction trio&lt;/th&gt;
&lt;th&gt;Locate + perceive + execute&lt;/th&gt;
&lt;th&gt;Bash&lt;/th&gt;
&lt;th&gt;Agent&lt;/th&gt;
&lt;th&gt;Task family&lt;/th&gt;
&lt;th&gt;Web pair&lt;/th&gt;
&lt;th&gt;Cron family&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Role&lt;/td&gt;
&lt;td&gt;Collaborative alignment&lt;/td&gt;
&lt;td&gt;Modify code&lt;/td&gt;
&lt;td&gt;Execute commands&lt;/td&gt;
&lt;td&gt;Derive Claude&lt;/td&gt;
&lt;td&gt;Externalize memory&lt;/td&gt;
&lt;td&gt;Reach the web&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Trigger the future&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time model&lt;/td&gt;
&lt;td&gt;Present&lt;/td&gt;
&lt;td&gt;Present&lt;/td&gt;
&lt;td&gt;Present&lt;/td&gt;
&lt;td&gt;Present, fork/join&lt;/td&gt;
&lt;td&gt;Across time&lt;/td&gt;
&lt;td&gt;Present&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Future, scheduled&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Disk&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Subagent&lt;/td&gt;
&lt;td&gt;Runtime storage&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Session-only, max seven days&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Naming pattern&lt;/td&gt;
&lt;td&gt;Enter / Exit&lt;/td&gt;
&lt;td&gt;Read / Edit / Write&lt;/td&gt;
&lt;td&gt;Single&lt;/td&gt;
&lt;td&gt;Single&lt;/td&gt;
&lt;td&gt;CRUD family&lt;/td&gt;
&lt;td&gt;Fetch / Search&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Create / Delete / List&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main benefit&lt;/td&gt;
&lt;td&gt;User alignment&lt;/td&gt;
&lt;td&gt;Precise code changes&lt;/td&gt;
&lt;td&gt;Engineering workflow&lt;/td&gt;
&lt;td&gt;Context space&lt;/td&gt;
&lt;td&gt;Against forgetting&lt;/td&gt;
&lt;td&gt;Controlled information&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Wait for the world to change&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Cron vs. Task
&lt;/h3&gt;

&lt;p&gt;Both families create state across time, but in opposite directions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Task:&lt;/strong&gt; leaves unfinished work in the present, recording what should be done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cron:&lt;/strong&gt; schedules a future prompt, recording when something should happen.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Task is a box of notes that Claude checks manually. Cron is an alarm that rings automatically. Task means “Claude chooses to call List”; Cron means “the clock wakes Claude.”&lt;/p&gt;

&lt;h3&gt;
  
  
  Cron vs. Bash background execution
&lt;/h3&gt;

&lt;p&gt;Both are asynchronous, but their triggers differ:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Bash background: the machine waits for a command to finish once.&lt;/li&gt;
&lt;li&gt;Cron: the runtime waits for a time and fires once or repeatedly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bash background handles “wait for CI to finish.” Cron handles “check every five minutes.” One is asynchronous I/O; the other is asynchronous time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cron vs. Agent
&lt;/h3&gt;

&lt;p&gt;Both create parallel work in different dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent creates &lt;strong&gt;spatial parallelism&lt;/strong&gt; by forking a new context.&lt;/li&gt;
&lt;li&gt;Cron creates &lt;strong&gt;temporal parallelism&lt;/strong&gt; by queueing work for the future.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first eleven tools perform actions now. Cron is the only primitive that treats &lt;strong&gt;future time as a first-class input&lt;/strong&gt;. It does not add a new capability so much as provide a trigger moment for every other capability.&lt;/p&gt;




&lt;h3&gt;
  
  
  Summary
&lt;/h3&gt;

&lt;p&gt;The Cron family’s most interesting signal is a single design line visible at every layer: &lt;strong&gt;reuse industry conventions to reduce cognitive load&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Naming:&lt;/strong&gt; reuse forty years of Unix crontab vocabulary rather than inventing concepts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fields:&lt;/strong&gt; represent schedules as standard five-field strings, not a new JSON DSL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Defaults:&lt;/strong&gt; &lt;code&gt;recurring: true&lt;/code&gt; matches monitoring, while one-shot reminders are explicit exceptions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Timezone:&lt;/strong&gt; use local time and avoid UTC conversion.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema:&lt;/strong&gt; remain intentionally thin because valid cron syntax cannot by itself distinguish good intent from bad intent.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Another unusual signal is the &lt;strong&gt;server’s perspective encoded in the tool prompt&lt;/strong&gt;. Avoiding :00 and :30 distributes load across the fleet. Most tools care only about Claude using them correctly; Cron also accounts for the operational consequences of thousands of scheduled requests.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Honest transparency runs throughout the family:&lt;/strong&gt; session-only lifetime is stated at the start, the seven-day cap must be disclosed, &lt;code&gt;durable&lt;/code&gt; is labeled ineffective, and jitter is exposed. Claude should not promise what the runtime cannot deliver.&lt;/p&gt;

&lt;p&gt;The Create / Delete / List trio forms a complete lifecycle, just like the Task family. Create does the substantive work; List observes and Delete manages the active state.&lt;/p&gt;

&lt;p&gt;The next article will examine Monitor: Cron wakes Claude when the clock reaches a point; Monitor wakes Claude when an event occurs. One is proactive scheduled polling; the other is passive event-driven streaming.&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I built a YouTube-to-text tool, and three things turned out much harder than expected</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Wed, 26 Aug 2026 05:59:06 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/i-built-a-youtube-to-text-tool-and-three-things-turned-out-much-harder-than-expected-457o</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/i-built-a-youtube-to-text-tool-and-three-things-turned-out-much-harder-than-expected-457o</guid>
      <description>&lt;p&gt;Someone links a 45-minute conference talk and says "the good part is in the middle somewhere." You want three sentences. You do not want 45 minutes.&lt;/p&gt;

&lt;p&gt;So I built SummarizeVideoToText: paste a video link, get a text workspace — full transcript, AI summary, timestamped chapters, a mind map, and a Q&amp;amp;A panel you can interrogate about the video. No sign-up needed to try it.&lt;/p&gt;

&lt;p&gt;That's the pitch. The interesting part is what broke along the way.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Getting captions is a fallback chain, not an API call
&lt;/h2&gt;

&lt;p&gt;My first version called one endpoint and assumed a transcript came back. In practice, that endpoint fails constantly — YouTube rotates its internals, some videos need a proof-of-origin token, some tracks exist but not in the language you asked for.&lt;/p&gt;

&lt;p&gt;What actually works is a chain of providers where each layer falls back to the next:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChainProvider&lt;/span&gt; &lt;span class="k"&gt;implements&lt;/span&gt; &lt;span class="nx"&gt;TranscriptProvider&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;constructor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;providers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;TranscriptProvider&lt;/span&gt;&lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
  &lt;span class="c1"&gt;// try each in turn; fall through on failure&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The non-obvious part is knowing when &lt;strong&gt;not&lt;/strong&gt; to fall through. Two cases end the chain immediately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;invalid_url&lt;/code&gt; — the link itself is broken. No provider will do better.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;no_transcript&lt;/code&gt; — layer one confirmed the page loads fine and has no caption track at all. A paid provider will confirm the same thing and bill you for it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else falls through. That distinction is the difference between a robust chain and a machine that burns API credits to rediscover the same "nope."&lt;/p&gt;

&lt;p&gt;The honest limitation this leaves: &lt;strong&gt;if a video has no captions in any form, there's nothing to summarize.&lt;/strong&gt; I show that plainly instead of pretending. Audio transcription for YouTube is on the roadmap; TikTok and Instagram already go through AI transcription because they rarely ship captions.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Caching wasn't about speed. It was about money.
&lt;/h2&gt;

&lt;p&gt;I started with Redis and a TTL, like you do. Then I watched the logs: a video would get summarized, sit for a week, the key would expire, someone would open the same URL — and the whole pipeline would run again. New caption fetch, new LLM call, new bill.&lt;/p&gt;

&lt;p&gt;The realization: &lt;strong&gt;a video's content never changes.&lt;/strong&gt; There is no correctness reason to ever evict a summary. TTL made sense for a hot cache, not for the artifact itself.&lt;/p&gt;

&lt;p&gt;So it became two layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Redis&lt;/strong&gt; — hot cache, short TTL, absorbs the repeat traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Postgres&lt;/strong&gt; — permanent store, no TTL. Redis misses land here, not on the model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Storing a few KB of text forever costs orders of magnitude less than regenerating it once. If your pipeline has an expensive deterministic step, "cache expiry" and "delete the result" should not be the same decision.&lt;/p&gt;

&lt;p&gt;The user-visible payoff is that opening a video someone else already summarized is instant and costs nobody anything — which is also why I could leave the free tier usable without an account.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Notion doesn't take Markdown
&lt;/h2&gt;

&lt;p&gt;I wanted "export this whole note to Notion." I assumed I'd POST some Markdown. Notion's API is a &lt;strong&gt;block model&lt;/strong&gt; — every heading, paragraph, and list item is a typed object, and the constraints stack up fast:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Max &lt;strong&gt;100 child blocks&lt;/strong&gt; per request. A transcript is hundreds of lines.&lt;/li&gt;
&lt;li&gt;Max &lt;strong&gt;2000 characters&lt;/strong&gt; per rich-text object. Long paragraphs need chunking.&lt;/li&gt;
&lt;li&gt;Max &lt;strong&gt;2 levels of nesting&lt;/strong&gt; per request. So a collapsible toggle containing a full transcript can't be created in one shot: you create the toggle, read its ID out of the response, then append its children in batches.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the one that cost me an evening of confusion: &lt;strong&gt;an integration without "read content" permission gets partial objects back.&lt;/strong&gt; Creating a page returns an object with an &lt;code&gt;id&lt;/code&gt; and no &lt;code&gt;url&lt;/code&gt;. Searching returns pages with no &lt;code&gt;properties&lt;/code&gt;, so no title. Nothing errors. You just get &lt;code&gt;undefined&lt;/code&gt; where you expected a link, and a page titled &lt;code&gt;""&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Two lessons. First, when an API returns a suspiciously empty field, check the permission scope before you check your code. Second, degrade instead of inventing: my first fix put the string &lt;code&gt;"(Untitled page)"&lt;/code&gt; in the UI, which turned a missing title into a confidently wrong one. The real fix was to reconstruct the URL from the ID (&lt;code&gt;notion.so/&amp;lt;id-without-dashes&amp;gt;&lt;/code&gt; is a valid link) and drop the page name from the message when it isn't known.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bonus: the Obsidian URL that was always too long
&lt;/h2&gt;

&lt;p&gt;Obsidian has a URI scheme: &lt;code&gt;obsidian://new?name=...&amp;amp;content=...&lt;/code&gt;. Clean, one click, works great in the demo.&lt;/p&gt;

&lt;p&gt;It never worked in production. A full transcript blows past the URI length limit every single time, so my code silently fell back to downloading a &lt;code&gt;.md&lt;/code&gt; file. Users clicked "Export to Obsidian" and got a file in their Downloads folder — technically not a failure, so nothing ever showed up in the error logs.&lt;/p&gt;

&lt;p&gt;The fix was one flag I'd missed in the docs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;copyText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;location&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;href&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`obsidian://new?name=&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;encodeURIComponent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;title&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;&amp;amp;clipboard=true`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;clipboard=true&lt;/code&gt; tells Obsidian to pull the body from the clipboard. The URI now carries only the title, and length stops being a factor.&lt;/p&gt;

&lt;h2&gt;
  
  
  The analytics bug that made everything else unmeasurable
&lt;/h2&gt;

&lt;p&gt;One more, because it invalidated a week of numbers.&lt;/p&gt;

&lt;p&gt;My event helper was defensive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;track&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;gtag&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;gtag&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;gtag&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;function&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="c1"&gt;// ← this line&lt;/span&gt;
  &lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sensible: ad blockers exist, analytics should never break a click. But the gtag stub was loading with &lt;code&gt;afterInteractive&lt;/code&gt;, meaning &lt;code&gt;window.gtag&lt;/code&gt; doesn't exist until hydration finishes. On a slow connection that's a multi-second window — and the "Summarize" button is the &lt;em&gt;first&lt;/em&gt; thing anyone clicks. Those events were dropped silently, so my funnel's denominator was quietly too small.&lt;/p&gt;

&lt;p&gt;The fix is to load the tiny stub &lt;code&gt;beforeInteractive&lt;/code&gt; (it only pushes to an array) and let the real script arrive later and replay the queue. That's what Google's own snippet does; I'd split it apart without thinking about ordering.&lt;/p&gt;

&lt;p&gt;Related lesson from the same audit: &lt;strong&gt;instrument outcomes, not intentions.&lt;/strong&gt; I was tracking "user clicked Export" but not whether the export succeeded. 100 clicks could be 97 successes or 3. Every click event that kicks off async work deserves a matching result event with a failure code.&lt;/p&gt;




&lt;h2&gt;
  
  
  What it is now
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;YouTube&lt;/strong&gt; via official captions; &lt;strong&gt;TikTok / Instagram&lt;/strong&gt; via AI transcription&lt;/li&gt;
&lt;li&gt;Summary, timestamped chapters, key insights, an interactive mind map, and Q&amp;amp;A grounded in the actual transcript&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;25 summary templates&lt;/strong&gt; (study notes, meeting minutes, Twitter thread, flashcards, SEO article…) and &lt;strong&gt;14 output languages&lt;/strong&gt;, independent of the video's language&lt;/li&gt;
&lt;li&gt;Export the whole note to &lt;strong&gt;Notion&lt;/strong&gt;, &lt;strong&gt;Obsidian&lt;/strong&gt;, or &lt;strong&gt;Markdown&lt;/strong&gt;, with clickable timestamps that jump back into the video&lt;/li&gt;
&lt;li&gt;Free: 2 videos/day with no account (up to 15 min), 10/day with a free Google sign-in (up to 1 hour)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Try it: &lt;strong&gt;&lt;a href="https://summarizevideototext.com" rel="noopener noreferrer"&gt;summarizevideototext.com&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you've fought with the Notion block API or the YouTube caption endpoints, I'd genuinely like to compare notes in the comments — especially if you found a cleaner answer than a fallback chain.&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>nextjs</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Claude Code Tools Deep Dive (11): WebFetch + WebSearch</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Tue, 25 Aug 2026 05:36:03 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-11-webfetch-websearch-5fen</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-11-webfetch-websearch-5fen</guid>
      <description>&lt;p&gt;This is the eleventh article in my series on Claude Code tools. The first ten covered Claude Code’s mostly &lt;strong&gt;inward-facing toolkit&lt;/strong&gt;: aligning with the user, operating the local filesystem, running commands, spawning subagents, and managing todos. All of those tools are built around the local environment.&lt;/p&gt;

&lt;p&gt;Real engineering work often requires Claude to &lt;strong&gt;leave the local environment&lt;/strong&gt;: read Anthropic’s API documentation, inspect a third-party library’s GitHub README, find a current npm tutorial, or verify an official specification. That information may not be local—or it may postdate the training data.&lt;/p&gt;

&lt;p&gt;Claude Code answers with a pair of internet tools: &lt;strong&gt;WebFetch retrieves content from one known URL; WebSearch finds URLs across the web from a query.&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This series begins with a prerequisite article explaining what tools are and how Claude uses them. Like the other articles, this one follows the four-layer framework introduced there.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  WebFetch + WebSearch
&lt;/h2&gt;

&lt;p&gt;These tools are covered together for the same reason as Grep + Glob: their semantics are tightly coupled. One fetches by URL and the other searches by query, and they are frequently combined—Search finds an entry point, then Fetch extracts the content.&lt;/p&gt;

&lt;h3&gt;
  
  
  Family overview
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Typical use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;WebFetch&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One known URL&lt;/td&gt;
&lt;td&gt;Page content, transformed from HTML to Markdown&lt;/td&gt;
&lt;td&gt;“Read this documentation and extract X”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;WebSearch&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A query&lt;/td&gt;
&lt;td&gt;Search results with titles and URLs&lt;/td&gt;
&lt;td&gt;“Find the latest way to do X”&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The division is simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Know the URL?&lt;/strong&gt; Call WebFetch directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not know the URL?&lt;/strong&gt; Use WebSearch, then WebFetch the useful results.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This mirrors Grep + Glob in the local filesystem. Grep + Glob search and locate locally; WebSearch + WebFetch do the same work on the public internet. &lt;strong&gt;The mental model stays the same; only the domain changes.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What they do
&lt;/h3&gt;

&lt;p&gt;Together, WebFetch and WebSearch solve how Claude can &lt;strong&gt;break through the time and scope limits of training data&lt;/strong&gt; and obtain current, specific external information:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Break through the training cutoff:&lt;/strong&gt; search and fetch can retrieve information published today.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Break through training coverage:&lt;/strong&gt; an obscure library may not appear in training data, but its official docs can be fetched.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify official information:&lt;/strong&gt; when an answer promises citations or official wording, the source must actually be read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compress content:&lt;/strong&gt; WebFetch uses AI to return only what the prompt asks for instead of placing an entire HTML page in context.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This family differs from every earlier tool in one decisive way: it is the only one that &lt;strong&gt;crosses the local boundary&lt;/strong&gt;. The first ten tools operate on the local machine; WebFetch and WebSearch connect Claude to the public web.&lt;/p&gt;

&lt;h3&gt;
  
  
  A concrete example
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; The user says, &lt;strong&gt;“I think Anthropic recently released a Claude 4.5 Sonnet model. Look up how to use its API, especially what changed from Claude 4, and check the pricing.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The information is online, may postdate training, and is too volatile to answer from memory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the model may be newly released&lt;/li&gt;
&lt;li&gt;API parameters may have changed&lt;/li&gt;
&lt;li&gt;pricing numbers are not safe to guess&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Step 1: Find an entry point with WebSearch
&lt;/h4&gt;

&lt;p&gt;Claude does not know the exact URL but knows the information should come from Anthropic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WebSearch(
  query: "Claude 4.5 Sonnet API pricing announcement 2026",
  allowed_domains: ["anthropic.com", "docs.anthropic.com"]
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;allowed_domains&lt;/code&gt; narrows the search to official sites and excludes marketing pages or second-hand summaries.&lt;/p&gt;

&lt;p&gt;The results might be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Claude 4.5 Sonnet — Anthropic
   https://www.anthropic.com/news/claude-4-5-sonnet
2. Models Overview — Anthropic Docs
   https://docs.anthropic.com/en/docs/about-claude/models
3. Pricing — Anthropic
   https://www.anthropic.com/pricing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each result is a title and URL, not the full page. Claude now has three precise entry points.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 2: Extract the details with WebFetch
&lt;/h4&gt;

&lt;p&gt;Claude fetches each URL with a specific prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WebFetch(
  url: "https://www.anthropic.com/news/claude-4-5-sonnet",
  prompt: "Extract the release date, improvements over Claude 4, benchmark numbers, and API model ID."
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second argument does not mean “return the full page.” It means &lt;strong&gt;process the page for this purpose&lt;/strong&gt;. Behind the scenes, the runtime:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fetches the URL&lt;/li&gt;
&lt;li&gt;converts HTML to Markdown&lt;/li&gt;
&lt;li&gt;uses a smaller, faster model to extract what the prompt requests&lt;/li&gt;
&lt;li&gt;returns only the extracted result&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A 5,000-word article may become 200 words in the main context. Like Agent, WebFetch is a context-compression mechanism.&lt;/p&gt;

&lt;p&gt;After three WebFetch calls, Claude has structured summaries sufficient to answer the API, comparison, and pricing questions.&lt;/p&gt;

&lt;h4&gt;
  
  
  Key insight: WebFetch is curl with AI
&lt;/h4&gt;

&lt;p&gt;Traditional &lt;code&gt;curl&lt;/code&gt; means “URL in, raw HTML out.” WebFetch means “URL plus intent in, &lt;strong&gt;processed result&lt;/strong&gt; out.”&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;With curl, Claude must parse HTML, remove CSS and navigation noise, and ignore advertisements.&lt;/li&gt;
&lt;li&gt;With WebFetch, the runtime’s AI does that work and returns content already extracted for the question.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;WebFetch is therefore an &lt;strong&gt;on-demand internet extraction primitive&lt;/strong&gt;, not merely a webpage downloader.&lt;/p&gt;

&lt;h3&gt;
  
  
  When they are triggered
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Use WebSearch when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;you need current information beyond the training cutoff&lt;/li&gt;
&lt;li&gt;you do not know the exact URL&lt;/li&gt;
&lt;li&gt;you want to compare several sources&lt;/li&gt;
&lt;li&gt;you need to search within a particular domain using &lt;code&gt;allowed_domains&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;you want to exclude particular domains using &lt;code&gt;blocked_domains&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use WebFetch when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the URL is known from the user or WebSearch&lt;/li&gt;
&lt;li&gt;you need official docs, specifications, or API references summarized against a specific prompt&lt;/li&gt;
&lt;li&gt;you need to verify a citation&lt;/li&gt;
&lt;li&gt;you need a GitHub README or documentation page, although &lt;code&gt;gh&lt;/code&gt; is usually better for GitHub&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Combine them when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Search finds URLs → choose authoritative results → Fetch details → synthesize the answer&lt;/li&gt;
&lt;li&gt;Search finds three to five sources → Fetch each → cross-check them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Do not use them when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the answer is stable and already in training knowledge, such as basic JavaScript syntax&lt;/li&gt;
&lt;li&gt;the content is GitHub-specific; use &lt;code&gt;gh&lt;/code&gt; through Bash when possible&lt;/li&gt;
&lt;li&gt;the URL requires authentication; WebFetch cannot access Google Docs, Confluence, Jira, or private GitHub repositories&lt;/li&gt;
&lt;li&gt;the information is local; use Grep instead of WebSearch&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The central principle is: &lt;strong&gt;do not go online unnecessarily&lt;/strong&gt;. The web is slower, more expensive, and vulnerable to network failures and page changes. Use it only when local files and training knowledge are insufficient.&lt;/p&gt;

&lt;h3&gt;
  
  
  Technical design
&lt;/h3&gt;

&lt;p&gt;WebFetch and WebSearch are sibling tools with a shared design philosophy. Their roles are distinct but complementary, so we can examine both through the four-layer framework.&lt;/p&gt;




&lt;h2&gt;
  
  
  WebFetch
&lt;/h2&gt;

&lt;h4&gt;
  
  
  1. Naming
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;WebFetch&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The name says exactly what it does: &lt;strong&gt;fetch a web resource&lt;/strong&gt;. “Fetch” is an industry verb from the Fetch API and &lt;code&gt;git fetch&lt;/code&gt;, suggesting retrieval rather than exploration. The &lt;code&gt;url&lt;/code&gt; field is immediately familiar to anyone who has worked with the web.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ReadURL&lt;/code&gt; would be misleading because WebFetch is not a lossless Read operation. It performs AI-assisted, on-demand extraction. &lt;code&gt;HTTPGet&lt;/code&gt; would be too low-level and would lose the promise of processing content against a prompt. &lt;strong&gt;Fetch sits between raw retrieval and AI processing&lt;/strong&gt;, which is exactly the intended semantic boundary.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Tool-level description
&lt;/h4&gt;

&lt;p&gt;WebFetch’s description is heavier than most tools. It begins with an all-caps IMPORTANT and follows with usage notes focused on four concerns: &lt;strong&gt;authentication failures, MCP precedence, GitHub specialization, and redirect safety&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;IMPORTANT: authenticated services are out of scope&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;IMPORTANT: WebFetch WILL FAIL for authenticated or private URLs. Before using this tool, check if the URL points to an authenticated service such as Google Docs, Confluence, Jira, or GitHub. If so, look for a specialized MCP tool that provides authenticated access.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the strongest sentence in the WebFetch description. IMPORTANT plus &lt;strong&gt;WILL FAIL&lt;/strong&gt; prevents a wasted call into a 401 or 403. It also gives a replacement: find an authenticated MCP tool.&lt;/p&gt;

&lt;p&gt;Every “no” comes with a “yes”: not this fetcher, but that specialized tool. Claude learns to inspect the available tool ecosystem before acting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MCP takes precedence&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;IMPORTANT: If an MCP-provided web fetch tool is available, prefer using that tool instead, as it may have fewer restrictions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;WebFetch explicitly acknowledges its limits. If a specialized MCP fetcher exists, use it. This humility is unusual in tool descriptions and reinforces the rule that authentication and capability-specific access belong to a better adapter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub gets a specialized instruction&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For GitHub URLs, prefer using the &lt;code&gt;gh&lt;/code&gt; CLI via Bash instead, such as &lt;code&gt;gh pr view&lt;/code&gt;, &lt;code&gt;gh issue view&lt;/code&gt;, or &lt;code&gt;gh api&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;GitHub is singled out because it is common and because &lt;code&gt;gh&lt;/code&gt; can use the user’s local credentials. It can access private repositories and review comments that an anonymous WebFetch cannot. This is a case where a specific workflow outranks the general-purpose tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cross-host redirects require an explicit protocol&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When a URL redirects to a different host, the tool will inform you and provide the redirect URL in a special format. You should then make a new WebFetch request with the redirect URL.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;WebFetch does not silently follow a cross-host redirect. It reports the new URL and lets Claude decide whether to fetch it. Same-host redirects can remain convenient; cross-host redirects become an explicit security boundary. This protects against pages that quietly send the fetcher somewhere unexpected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Transparent 15-minute caching&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Includes a self-cleaning 15-minute cache for faster responses when repeatedly accessing the same URL.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Claude is told that repeated calls to the same URL may be faster. The disclosure encourages useful re-fetching in one session without making Claude worry that every retry is wasteful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Automatic HTTP-to-HTTPS upgrade&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;HTTP URLs will be automatically upgraded to HTTPS.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This makes a hidden behavior explicit. Claude can provide &lt;code&gt;http://&lt;/code&gt; and the runtime upgrades it without requiring a manual edit—lowering error rates without relying on silent magic.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Field-level descriptions
&lt;/h4&gt;

&lt;p&gt;WebFetch has only two fields, but both are required:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;url&lt;/code&gt;: the complete URL&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;prompt&lt;/code&gt;: what to extract from the page&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Why is &lt;code&gt;prompt&lt;/code&gt; required? Because WebFetch does &lt;strong&gt;not&lt;/strong&gt; return the full page. It returns content processed according to the prompt. Without one, the small runtime model would not know what to extract or how much to return.&lt;/p&gt;

&lt;p&gt;Compare the two mental models:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;curl https://example.com
→ raw HTML, potentially tens of thousands of words

WebFetch(url, prompt: "Extract the three key points")
→ a focused summary
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Writing the prompt is like briefing a new colleague. “Read this page” is shallow; “Find every rate-limit number, list each one, and say explicitly if none is present” is precise.&lt;/p&gt;

&lt;p&gt;Making &lt;code&gt;prompt&lt;/code&gt; required is WebFetch’s most elegant field decision. It forces Claude to decide what it needs &lt;strong&gt;before fetching&lt;/strong&gt;, protecting the context budget instead of fetching first and filtering later.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. Schema validation
&lt;/h4&gt;

&lt;p&gt;WebFetch’s schema has almost no hard constraints:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Constraint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;url&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;URI format validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;prompt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;required; no length constraint&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The one physical barrier is &lt;code&gt;url&lt;/code&gt; with a URI format. A value such as &lt;code&gt;foo&lt;/code&gt; is rejected before the tool call is sent. Claude must provide a complete URL.&lt;/p&gt;

&lt;p&gt;Everything else—authentication, GitHub, MCP, redirects, and when to use the tool—lives in natural-language guidance. WebFetch’s complexity is not parameter validation; it is judging when the tool should not be used.&lt;/p&gt;




&lt;h2&gt;
  
  
  WebSearch
&lt;/h2&gt;

&lt;h4&gt;
  
  
  1. Naming
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;WebSearch&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The tool uses Search rather than &lt;code&gt;WebQuery&lt;/code&gt; or &lt;code&gt;GoogleSearch&lt;/code&gt;, preserving generality and avoiding a search-engine brand. Its behavior is “query in, result set out,” exactly what Search means.&lt;/p&gt;

&lt;p&gt;Together with WebFetch, the names form a clean pair: &lt;strong&gt;Fetch retrieves a known URL; Search discovers URLs from terms&lt;/strong&gt;.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Tool-level description
&lt;/h4&gt;

&lt;p&gt;WebSearch contains two unusually strong constraints: &lt;strong&gt;citation obligations and explicit time awareness&lt;/strong&gt;. Its description focuses on four ideas: capability, mandatory Sources, domain filtering, and the correct year.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Basic capability&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Allows Claude to search the web and use the results to inform responses. Provides up-to-date information for current events and recent data.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These two sentences state why WebSearch exists: to overcome the freshness limits of training data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mandatory citations at CRITICAL strength&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;CRITICAL REQUIREMENT - You MUST follow this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;After answering the user's question, you MUST include a “Sources:” section at the end of your response.&lt;/li&gt;
&lt;li&gt;In that section, list all relevant URLs from the search results as Markdown hyperlinks: &lt;a href="https://dev.toURL"&gt;Title&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;This is MANDATORY—never skip the Sources section.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the heaviest paragraph in the WebSearch description. CRITICAL, MUST repeated three times, and MANDATORY elevate citation from advice to law.&lt;/p&gt;

&lt;p&gt;Why require Sources? Search results come from uncontrolled sources: some are biased, stale, or optimized for SEO. A Sources section provides traceability so the user can inspect whether the answer rests on reliable material.&lt;/p&gt;

&lt;p&gt;It also enforces the fact-checking discipline established at the beginning of the series. If Claude promises official or cited information, it must actually retrieve the source, and WebSearch must expose the URLs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain filtering&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Domain filtering is supported to include or block specific websites.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This reminds Claude that &lt;code&gt;allowed_domains&lt;/code&gt; can create an official-source whitelist, while &lt;code&gt;blocked_domains&lt;/code&gt; can exclude low-quality or unwanted sites. To verify Anthropic’s official wording, for example, search only &lt;code&gt;anthropic.com&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The current-year rule&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;IMPORTANT - Use the correct year in search queries. The current month is July 2026. You MUST use this year when searching for recent information, documentation, or current events. For “latest React docs,” search with the current year, not last year.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hardcoding the current time in a tool description is unusual, but the reason is profound: Claude may not know which month it is after its training cutoff. Search freshness depends on dates. Adding the year to the query helps distinguish current documentation from old results.&lt;/p&gt;

&lt;p&gt;The React example turns an abstract rule into a concrete comparison: correct current year versus stale year.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;US-only availability&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Domain filtering is supported to include or block specific websites. Web search is only available in the US.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This small boundary statement prevents failed calls in unsupported regions.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Field-level descriptions
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;query&lt;/code&gt;: required search terms&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;allowed_domains&lt;/code&gt;: optional whitelist array&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;blocked_domains&lt;/code&gt;: optional blacklist array&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The two domain filters expose two independent information postures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;allowed_domains&lt;/code&gt;: search only sources Claude trusts, such as &lt;code&gt;anthropic.com&lt;/code&gt; or &lt;code&gt;docs.python.org&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;blocked_domains&lt;/code&gt;: exclude sources Claude does not want, such as outdated or low-quality sites.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same domain should not logically appear in both lists, but the tool leaves that judgment to Claude rather than forbidding both arrays at the schema level.&lt;/p&gt;

&lt;p&gt;The only explicit field-level minimum is that &lt;code&gt;query&lt;/code&gt; must contain at least two characters, preventing meaningless one-character searches.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. Schema validation
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Constraint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;query&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;&lt;code&gt;minLength: 2&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;allowed_domains&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;array of strings&lt;/td&gt;
&lt;td&gt;optional&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;blocked_domains&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;array of strings&lt;/td&gt;
&lt;td&gt;optional&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;minLength: 2&lt;/code&gt; is a physical schema barrier. A one-character query is rejected directly, even though a single Chinese character might sometimes be meaningful. The tool chooses a simple universal rule.&lt;/p&gt;

&lt;p&gt;The domain arrays are intentionally permissive. The schema allows both to be populated; prompt guidance handles the logical conflict. Capability remains open while judgment stays with Claude.&lt;/p&gt;




&lt;h3&gt;
  
  
  Why dedicated WebFetch and WebSearch instead of Bash + curl/search APIs?
&lt;/h3&gt;

&lt;p&gt;Bash could theoretically combine &lt;code&gt;curl&lt;/code&gt; with a search API, but that creates several problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;HTML parsing:&lt;/strong&gt; raw curl output includes CSS, navigation, and advertisements.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credential leakage:&lt;/strong&gt; a local curl might accidentally use &lt;code&gt;.netrc&lt;/code&gt; or cookies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search API key management:&lt;/strong&gt; someone must provide and protect Google or Bing API credentials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No citation obligation:&lt;/strong&gt; Claude could summarize curl output without reporting sources.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No content compression:&lt;/strong&gt; a 50,000-word page could flood the context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The dedicated tools solve these problems: HTML becomes Markdown automatically, fetching is anonymous, API credentials are managed by the runtime, WebSearch requires Sources, and WebFetch extracts against a prompt. This is another example of the division: &lt;strong&gt;Bash is the catch-all; dedicated tools are the refined interface.&lt;/strong&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Division of responsibility among neighboring tools
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Interaction trio&lt;/th&gt;
&lt;th&gt;Locate + perceive + execute&lt;/th&gt;
&lt;th&gt;Bash&lt;/th&gt;
&lt;th&gt;Agent&lt;/th&gt;
&lt;th&gt;Task family&lt;/th&gt;
&lt;th&gt;WebFetch / WebSearch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Role&lt;/td&gt;
&lt;td&gt;Collaborative alignment&lt;/td&gt;
&lt;td&gt;Modify code&lt;/td&gt;
&lt;td&gt;Execute commands&lt;/td&gt;
&lt;td&gt;Derive Claude&lt;/td&gt;
&lt;td&gt;Externalize working memory&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Reach the public web&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input source&lt;/td&gt;
&lt;td&gt;User&lt;/td&gt;
&lt;td&gt;Disk&lt;/td&gt;
&lt;td&gt;Command&lt;/td&gt;
&lt;td&gt;Prompt&lt;/td&gt;
&lt;td&gt;User / AI&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;URL / query&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;Structured&lt;/td&gt;
&lt;td&gt;Text / diff&lt;/td&gt;
&lt;td&gt;Raw text&lt;/td&gt;
&lt;td&gt;Subagent result&lt;/td&gt;
&lt;td&gt;State&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;HTML → Markdown / summaries&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authentication&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Existing user session&lt;/td&gt;
&lt;td&gt;User credentials&lt;/td&gt;
&lt;td&gt;Forked session&lt;/td&gt;
&lt;td&gt;User session&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Anonymous; no credentials&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main benefit&lt;/td&gt;
&lt;td&gt;User alignment&lt;/td&gt;
&lt;td&gt;Precise code changes&lt;/td&gt;
&lt;td&gt;Engineering workflow&lt;/td&gt;
&lt;td&gt;Context space&lt;/td&gt;
&lt;td&gt;Against forgetting&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Controllable information access&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;WebFetch + WebSearch mirror Grep + Glob:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Grep + Glob: find by content or path &lt;strong&gt;inside a project&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;WebFetch + WebSearch: fetch by URL or search by terms &lt;strong&gt;on the public web&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both pairs implement “one precise target plus one exploratory search,” but the web pair faces an untrusted external world. That is why it adds MCP precedence, explicit cross-host redirects, and mandatory Sources.&lt;/p&gt;

&lt;p&gt;The boundary with Bash is equally clear. Curl and search APIs could do the work, but they expose parsing, credentials, keys, citations, and context to risk. Dedicated tools package those concerns into a controlled interface.&lt;/p&gt;




&lt;h3&gt;
  
  
  Summary
&lt;/h3&gt;

&lt;p&gt;WebFetch + WebSearch split the broad request “let the AI browse the web” into two specialized tools and use all four design layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Naming:&lt;/strong&gt; Fetch and Search borrow industry conventions. Fetch implies a known target; Search implies exploratory discovery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool-level descriptions:&lt;/strong&gt; WebFetch emphasizes authentication failures, MCP precedence, GitHub specialization, and redirect safety. WebSearch emphasizes mandatory Sources and the current-year rule—one protects traceability, the other compensates for Claude’s weak time awareness.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Field design:&lt;/strong&gt; WebFetch has only two fields, but making &lt;code&gt;prompt&lt;/code&gt; required elevates on-demand extraction and protects context. WebSearch exposes allowed and blocked domains as independent trust controls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema validation:&lt;/strong&gt; WebFetch rejects non-URI strings; WebSearch rejects one-character queries. Basic mistakes are stopped physically at the schema boundary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Several signals are unique across the tool ecosystem:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Required prompt:&lt;/strong&gt; WebFetch becomes an extraction primitive rather than a page downloader.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mandatory Sources:&lt;/strong&gt; WebSearch is the only tool whose description uses CRITICAL and MANDATORY to enforce response formatting and citation transparency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hardcoded current time:&lt;/strong&gt; a dynamic date is inserted into a static prompt to compensate for Claude’s missing sense of the current month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Humility toward MCP:&lt;/strong&gt; the tool explicitly says “not me—use the authenticated or more capable MCP tool.”&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No automatic cross-host redirects:&lt;/strong&gt; the security decision is handed to Claude through an explicit protocol rather than silent behavior.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Together these signals turn the ability to reach the public web into an external information interface that is &lt;strong&gt;controlled, traceable, and willing to defer to better tools&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The next article will examine the Cron family, shifting from the spatial dimension—project and public web—to the time dimension: scheduled and future execution.&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Code Tools Deep Dive (10): The Task Family</title>
      <dc:creator>Zhengxin</dc:creator>
      <pubDate>Mon, 24 Aug 2026 14:34:38 +0000</pubDate>
      <link>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-10-the-task-family-569p</link>
      <guid>https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-10-the-task-family-569p</guid>
      <description>&lt;p&gt;This is the tenth article in my series on Claude Code tools. The first nine covered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The interaction primitive trio&lt;/strong&gt;—AskUserQuestion, EnterPlanMode, and ExitPlanMode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The execution primitive chain&lt;/strong&gt;—&lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-4-grep-glob-3fh4"&gt;Grep + Glob&lt;/a&gt; → &lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-5-read-2ba6"&gt;Read&lt;/a&gt; → &lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-6-edit-243f"&gt;Edit&lt;/a&gt; / &lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-7-write-2mhl"&gt;Write&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The general-purpose fallback&lt;/strong&gt;, &lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-8-bash-5f21"&gt;Bash&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The meta-tool&lt;/strong&gt;, &lt;a href="https://dev.to/_94be737e156beb4d74df2/claude-code-tools-deep-dive-9-agent-1k02"&gt;Agent&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first nine tools are about Claude doing &lt;strong&gt;the thing happening now&lt;/strong&gt;. Each tool call performs one immediate action. Real projects also require Claude to &lt;strong&gt;remember what needs to happen, track progress, decompose a large task, and share one checklist across multiple Claudes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That requires a task-management system. Claude Code’s answer is the &lt;strong&gt;Task family&lt;/strong&gt;—six tools that form a todo system.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This series begins with a prerequisite article explaining what tools are and how Claude uses them. Like the other articles, this one follows the four-layer framework introduced there.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Task family: TaskCreate / TaskList / TaskGet / TaskUpdate / TaskStop / TaskOutput
&lt;/h2&gt;

&lt;p&gt;This is the first time the series examines six tools in one article. Why group them? Because they share one data model—a task list—and are semantically coupled. Discussing one in isolation would shift attention from the system to a single operation. It would be like explaining how to create one Jira ticket without explaining the Jira system around it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Family overview
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;th&gt;Typical moment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TaskCreate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Create a task&lt;/td&gt;
&lt;td&gt;Decomposing a complex request or receiving multiple requirements&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TaskList&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;List all tasks&lt;/td&gt;
&lt;td&gt;Finding the next available task or reporting progress&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TaskGet&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Retrieve one task’s details&lt;/td&gt;
&lt;td&gt;Before starting work or inspecting dependencies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TaskUpdate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Change task status or metadata&lt;/td&gt;
&lt;td&gt;Starting, completing, or linking tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TaskStop&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stop a background task&lt;/td&gt;
&lt;td&gt;Aborting a background Bash process or subagent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TaskOutput&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Retrieve background output&lt;/td&gt;
&lt;td&gt;&lt;em&gt;Deprecated; use Read on the output file instead&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The family actually contains two groups:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The first four are CRUD for &lt;strong&gt;conceptual todo tasks&lt;/strong&gt;—something Claude remembers needs to happen.&lt;/li&gt;
&lt;li&gt;The last two control &lt;strong&gt;runtime tasks&lt;/strong&gt;—a real Bash process or subagent that is currently running.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They all say Task, but they operate on different things. This is the family’s most confusing design choice and will matter throughout the article.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it does
&lt;/h3&gt;

&lt;p&gt;The Task family—especially the four todo tools—solves how Claude can manage multi-step work &lt;strong&gt;across tool calls and across time&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Make decomposition visible:&lt;/strong&gt; complex requirements become entries whose progress the user can see.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Track progress:&lt;/strong&gt; every task has a &lt;code&gt;pending&lt;/code&gt;, &lt;code&gt;in_progress&lt;/code&gt;, or &lt;code&gt;completed&lt;/code&gt; state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model dependencies:&lt;/strong&gt; “A blocks B” becomes explicit and enforces order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coordinate multiple Claudes:&lt;/strong&gt; the main Claude decomposes work, subagents claim owners, and everyone shares one list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compress context:&lt;/strong&gt; a short subject can stand for an entire chunk of work, reducing what the main Claude must keep in mind.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The crucial difference from earlier tools is that &lt;strong&gt;Task is the only family with persistent state&lt;/strong&gt;. Read, Edit, and Bash return a result once in a tool call. A TaskCreate entry remains in the runtime and appears in future TaskList calls until it is completed or deleted.&lt;/p&gt;

&lt;h3&gt;
  
  
  A concrete example
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; The user says, &lt;strong&gt;“Add a user-profile page. It needs a backend API, a frontend component, a database schema, tests, and permission checks.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is a multi-task requirement.&lt;/p&gt;

&lt;h4&gt;
  
  
  The bad alternative: no Task family
&lt;/h4&gt;

&lt;p&gt;Without tasks, Claude can only:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Keep the plan in short-term memory while working.&lt;/li&gt;
&lt;li&gt;Announce every next step in chat as a textual progress log.&lt;/li&gt;
&lt;li&gt;Eventually lose a requirement—perhaps the tests—when the conversation becomes long.&lt;/li&gt;
&lt;li&gt;Reconstruct the entire conversation whenever the user asks, “How far along are you?”&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The core problem is that the checklist exists only in Claude’s short-term context. Context compression, a subagent handoff, or session recovery can make it disappear.&lt;/p&gt;

&lt;h4&gt;
  
  
  How the Task family solves it
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Create tasks immediately&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TaskCreate(subject: "Design database schema", description: "Add profile fields to users or create a profiles table")
TaskCreate(subject: "Write migration", description: "Generate the Knex migration")
TaskCreate(subject: "Implement backend API", description: "GET/PATCH /api/profile through the auth middleware")
TaskCreate(subject: "Build ProfilePage", description: "Add the /profile route, form, and API submission")
TaskCreate(subject: "Add tests", description: "API tests plus frontend component tests")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each call returns an ID, such as &lt;code&gt;task_001&lt;/code&gt; through &lt;code&gt;task_005&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2: Add dependencies&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The schema must exist before the API, and the API must exist before the frontend. TaskUpdate expresses those relationships through &lt;code&gt;blockedBy&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TaskUpdate(taskId: "task_002", addBlockedBy: ["task_001"])
TaskUpdate(taskId: "task_003", addBlockedBy: ["task_002"])
TaskUpdate(taskId: "task_004", addBlockedBy: ["task_003"])
TaskUpdate(taskId: "task_005", addBlockedBy: ["task_003"])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The list now forms a dependency graph: schema → migration → API → (frontend + tests).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3: Find the next available task&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;TaskList returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;task_001 · pending · Design database schema · blockedBy: []
task_002 · pending · Write migration         · blockedBy: [task_001]
task_003 · pending · Implement backend API    · blockedBy: [task_002]
task_004 · pending · Build ProfilePage        · blockedBy: [task_003]
task_005 · pending · Add tests                · blockedBy: [task_003]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only &lt;code&gt;task_001&lt;/code&gt; is pending and unblocked, so it is the next task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4: Claim, execute, and complete&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TaskUpdate(taskId: "task_001", status: "in_progress")
# Claude designs the schema and records the decision
TaskUpdate(taskId: "task_001", status: "completed")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once &lt;code&gt;task_001&lt;/code&gt; is complete, the migration task is unblocked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5: Delegate a task to a subagent&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The frontend task can be delegated:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent(
  description: "Build ProfilePage",
  prompt: "Task task_004: build the ProfilePage at /profile. Use TaskGet for the full details."
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The subagent can use the task ID to call TaskGet, claim the task with TaskUpdate, and mark it completed. The main Claude and the subagent coordinate through the shared task system rather than sending ad hoc messages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 6: Report progress&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Whenever the user asks for an update, TaskList is enough:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✅ task_001 · completed  · Design database schema
✅ task_002 · completed  · Write migration
🔄 task_003 · in_progress · Implement backend API (Claude)
⏸️ task_004 · pending    · Build ProfilePage (blocked by 003)
⏸️ task_005 · pending    · Add tests (blocked by 003)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One list makes the state immediately legible.&lt;/p&gt;

&lt;h4&gt;
  
  
  The key insight: Task externalizes Claude’s working memory
&lt;/h4&gt;

&lt;p&gt;Earlier tools are about &lt;strong&gt;doing&lt;/strong&gt;. The Task family is about &lt;strong&gt;remembering&lt;/strong&gt;. It moves Claude’s short-term plan into runtime storage.&lt;/p&gt;

&lt;p&gt;That creates two major effects:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Persistence across contexts:&lt;/strong&gt; tasks survive context compression, switching, and recovery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sharing across Claudes:&lt;/strong&gt; the main Claude and subagents synchronize through the Task system instead of messaging each other manually.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This resembles a human engineering team writing work into Jira. It is not a rejection of personal memory; &lt;strong&gt;memory is individual, while tasks are shared&lt;/strong&gt;. Writing them down enables collaboration, tracking, and completeness.&lt;/p&gt;

&lt;h3&gt;
  
  
  When it is triggered
&lt;/h3&gt;

&lt;p&gt;The family’s prompts provide fairly strict guidance:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use Task tools for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Complex work with three or more steps.&lt;/strong&gt; A one-step task does not need a task entry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nontrivial multi-operation work.&lt;/strong&gt; Planning and tracking have value here.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;An explicit user request for a todo list.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multiple requirements in one instruction.&lt;/strong&gt; Create them together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan mode.&lt;/strong&gt; Track the plan’s steps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Starting work.&lt;/strong&gt; Claim the task and mark it &lt;code&gt;in_progress&lt;/code&gt; before acting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Finishing work.&lt;/strong&gt; Mark it &lt;code&gt;completed&lt;/code&gt; immediately and inspect newly unblocked tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Do not use Task tools for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one direct operation&lt;/li&gt;
&lt;li&gt;a trivial task where tracking creates more noise than value&lt;/li&gt;
&lt;li&gt;a simple job with fewer than three steps&lt;/li&gt;
&lt;li&gt;pure conversation or an informational answer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The core judgment is: &lt;strong&gt;the Task family is for work with meaningful scale&lt;/strong&gt;. If a single tool call completes the task, a Task entry is noise. If the work has decomposition, dependencies, or progress worth tracking, failing to create one is a process failure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Technical design
&lt;/h3&gt;

&lt;h4&gt;
  
  
  1. Naming
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;TaskCreate&lt;/code&gt; / &lt;code&gt;TaskList&lt;/code&gt; / &lt;code&gt;TaskGet&lt;/code&gt; / &lt;code&gt;TaskUpdate&lt;/code&gt; / &lt;code&gt;TaskStop&lt;/code&gt; / &lt;code&gt;TaskOutput&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The shared &lt;code&gt;Task&lt;/code&gt; prefix replaces alternatives such as &lt;code&gt;Todo&lt;/code&gt;, &lt;code&gt;Ticket&lt;/code&gt;, or &lt;code&gt;Job&lt;/code&gt;. “Task” implies a clear execution owner; a todo can merely mean “look at this someday.” The name itself hints that an &lt;code&gt;owner&lt;/code&gt; field exists.&lt;/p&gt;

&lt;p&gt;The CRUD suffixes—Create, List, Get, Update—are standard database-like verbs: create one, list all, retrieve one, update one. The four names immediately establish a mental model of an enumerable, addressable, mutable entity collection.&lt;/p&gt;

&lt;p&gt;There is deliberately no &lt;code&gt;TaskDelete&lt;/code&gt;. Hard deletion is represented by &lt;code&gt;TaskUpdate(status: "deleted")&lt;/code&gt;. Deletion is treated as a terminal state in the state machine rather than a separate operation, concentrating all status transitions in TaskUpdate and reducing decision overhead.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;TaskStop&lt;/code&gt; and &lt;code&gt;TaskOutput&lt;/code&gt; introduce semantic drift. They reuse the Task namespace but operate on running background processes—Bash or subagents—rather than conceptual todos. The designers chose one namespace over a separate runtime-task family, but this is also the family’s most obvious source of confusion.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;activeForm&lt;/code&gt; is the most ambitious field name in the family. It is not called &lt;code&gt;presentContinuous&lt;/code&gt;, &lt;code&gt;verbForm&lt;/code&gt;, or &lt;code&gt;spinnerLabel&lt;/code&gt;; &lt;code&gt;activeForm&lt;/code&gt; sounds grammatical. When Claude writes it, the name nudges Claude to convert the action into the present progressive rather than enter a generic UI label.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Tool-level descriptions
&lt;/h4&gt;

&lt;p&gt;Each Task tool has its own description, but they share a positioning: &lt;strong&gt;these are members of one collaboration contract&lt;/strong&gt;, not isolated utilities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TaskCreate: a quantitative threshold&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Use this tool proactively in these scenarios: Complex multi-step tasks—when a task requires 3 or more distinct steps or actions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;“Three or more” is an explicit threshold. It trains Claude not to create tasks for every small operation and replaces the vague word “complex” with a measurable rule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A three-stage timing protocol&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;After receiving new instructions—immediately capture requirements as tasks.&lt;br&gt;
When you start working—mark the task &lt;code&gt;in_progress&lt;/code&gt; before beginning.&lt;br&gt;
After completing—mark it &lt;code&gt;completed&lt;/code&gt; and add follow-up tasks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The rhythm is exact: &lt;strong&gt;receive → create; start → in progress; finish → completed&lt;/strong&gt;. It wraps every work segment and prevents work from silently beginning or ending.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TaskUpdate: a strict completion standard&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Only mark a task completed when it is fully accomplished. If there are errors, blockers, or unfinished work, keep it &lt;code&gt;in_progress&lt;/code&gt;. Never mark it completed when tests fail, implementation is partial, or unresolved errors remain.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This blocks “fake completion”—the tendency to mark something done because the broad direction looks right while leaving half-finished work behind.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TaskList: a default scheduling intuition&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Prefer working on tasks in ID order (lowest ID first) when multiple tasks are available.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Earlier tasks are often prerequisites for later tasks, so ID order makes the default schedule match creation order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TaskGet before TaskUpdate: staleness awareness&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Make sure to read a task’s latest state using &lt;code&gt;TaskGet&lt;/code&gt; before updating it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another agent may have changed the task, especially in a multi-Claude workflow. Fetching the latest state before writing is a simple form of optimistic concurrency control: &lt;strong&gt;read before write, never overwrite stale state blindly&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TaskOutput: transparent deprecation&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;DEPRECATED: Background tasks return their output file path in the tool result and in the completion notification. For Bash tasks, prefer Read on that output path.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The tool description directly says not to use it and gives the replacement. This reflects a broader design principle: if an existing primitive can cover a capability, do not maintain a separate tool for it. Fewer tools mean less API surface and less decision burden.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A family-specific reminder hook&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If Claude goes a long time without using task tools, the harness can insert a system reminder:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The task tools haven’t been used recently. If your work would benefit from tracking progress, consider using TaskCreate and TaskUpdate.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This nudges Claude toward progress tracking without making it mandatory. The reminder ends with the equivalent of “only use these if relevant.” Earlier tools do not need this hook because their value is immediate; Task tools need a cross-time nudge.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Field-level descriptions
&lt;/h4&gt;

&lt;p&gt;A complete Task object includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;id&lt;/code&gt;: system-generated unique identifier&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;subject&lt;/code&gt;: short imperative title, such as “Run tests”&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;description&lt;/code&gt;: detailed explanation&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;activeForm&lt;/code&gt;: present-progressive form, such as “Running tests,” for a spinner&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;status&lt;/code&gt;: &lt;code&gt;pending&lt;/code&gt;, &lt;code&gt;in_progress&lt;/code&gt;, &lt;code&gt;completed&lt;/code&gt;, or &lt;code&gt;deleted&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;owner&lt;/code&gt;: the agent doing the work; empty means unclaimed&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;blocks&lt;/code&gt;: tasks blocked by this task&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;blockedBy&lt;/code&gt;: tasks blocking this task&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;metadata&lt;/code&gt;: arbitrary key-value data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Four design choices stand out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three representations: subject, description, activeForm&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The same task appears in three forms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;subject&lt;/code&gt;: a short imperative, “Run tests”&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;description&lt;/code&gt;: “Run the unit tests and confirm all four auth tests pass”&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;activeForm&lt;/code&gt;: “Running tests”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They map to different UI locations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;list view shows the short subject&lt;/li&gt;
&lt;li&gt;detail view shows the full description&lt;/li&gt;
&lt;li&gt;a spinner uses the present-progressive active form&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The forced progressive form is more than cosmetic. Claude must provide both “what to do” and “what is happening now,” encoding the distinction between intending to start and having started.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;blocks&lt;/code&gt; and &lt;code&gt;blockedBy&lt;/code&gt;: bidirectional dependencies&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;They are two views of the same relationship:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A blocks B  ⇔  B is blockedBy A
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The runtime keeps both directions consistent. Claude can add one side with &lt;code&gt;addBlocks&lt;/code&gt; or &lt;code&gt;addBlockedBy&lt;/code&gt;, and the other side synchronizes automatically.&lt;/p&gt;

&lt;p&gt;This is redundancy in favor of readable scheduling semantics: “what do I block?” and “what blocks me?” are different questions for Claude even though they describe one edge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Incremental merge semantics&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;TaskUpdate accepts &lt;code&gt;addBlocks&lt;/code&gt; and &lt;code&gt;addBlockedBy&lt;/code&gt;, not a replacement-style &lt;code&gt;blocks: [...]&lt;/code&gt;. Adding one dependency therefore cannot accidentally erase existing dependencies. Incremental updates are safer and naturally idempotent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Status: a linear state machine with a deleted escape hatch&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The normal path is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pending → in_progress → completed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Completed tasks cannot move backward to &lt;code&gt;in_progress&lt;/code&gt;; if the work must be redone, create a new task. This prevents unpredictable state oscillation.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;deleted&lt;/code&gt; is a terminal cleanup state for mistaken tasks. Deleted tasks disappear from normal lists while their IDs remain reserved, preventing reuse. Deletion is therefore part of the state machine rather than disappearance from the database.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;blockedBy&lt;/code&gt; also constrains status transitions. A task with incomplete dependencies cannot be claimed as &lt;code&gt;in_progress&lt;/code&gt;. Status is not an isolated field; it is a multi-field transition governed by the current dependency graph.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;owner&lt;/code&gt; and &lt;code&gt;metadata&lt;/code&gt;: two switches for multi-Claude work&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;owner&lt;/code&gt; identifies which agent currently owns a task:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the main Claude creates it with no owner&lt;/li&gt;
&lt;li&gt;a subagent claims it and records its agent name&lt;/li&gt;
&lt;li&gt;TaskList shows which work is claimed and which is available&lt;/li&gt;
&lt;li&gt;after completion, another agent can take the next task&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the basic pattern of a distributed work queue, with Claude instances as consumers.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;metadata&lt;/code&gt; is a free-form key-value escape hatch for file paths, reference links, subagent context, or temporary notes. &lt;code&gt;owner&lt;/code&gt; is a core contract; &lt;code&gt;metadata&lt;/code&gt; leaves room for extension.&lt;/p&gt;

&lt;h4&gt;
  
  
  4. Schema validation
&lt;/h4&gt;

&lt;p&gt;The family combines schema-level static checks with runtime state-machine checks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Constraint&lt;/th&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;activeForm&lt;/code&gt; required&lt;/td&gt;
&lt;td&gt;Schema&lt;/td&gt;
&lt;td&gt;TaskCreate requires the progressive form&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;status&lt;/code&gt; enum&lt;/td&gt;
&lt;td&gt;Schema&lt;/td&gt;
&lt;td&gt;Only &lt;code&gt;pending&lt;/code&gt;, &lt;code&gt;in_progress&lt;/code&gt;, &lt;code&gt;completed&lt;/code&gt;, or &lt;code&gt;deleted&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;subject&lt;/code&gt; length&lt;/td&gt;
&lt;td&gt;Schema&lt;/td&gt;
&lt;td&gt;Short title has a version-dependent maximum&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backward status transition&lt;/td&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;completed → in_progress&lt;/code&gt; is rejected&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unresolved &lt;code&gt;blockedBy&lt;/code&gt; → &lt;code&gt;in_progress&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;A locked task cannot be claimed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TaskGet before TaskUpdate&lt;/td&gt;
&lt;td&gt;Runtime guidance&lt;/td&gt;
&lt;td&gt;Strongly recommended, but primarily prompt-enforced&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Static constraints belong in the schema; dynamic constraints such as dependencies, concurrency, and legal state transitions belong in runtime. The Task family is more balanced than Read or Edit: schemas protect inputs while runtime protects transitions.&lt;/p&gt;

&lt;p&gt;TaskOutput’s deprecation also illustrates transparent fallback. Output retrieval is not replaced by a new specialized tool; it is reduced to Read on the existing output path. &lt;strong&gt;Capabilities that existing primitives can cover do not need another tool.&lt;/strong&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Division of responsibility among neighboring tools
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Interaction trio&lt;/th&gt;
&lt;th&gt;Locate + perceive + execute&lt;/th&gt;
&lt;th&gt;Bash&lt;/th&gt;
&lt;th&gt;Agent&lt;/th&gt;
&lt;th&gt;Task family&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Role&lt;/td&gt;
&lt;td&gt;Collaborative alignment&lt;/td&gt;
&lt;td&gt;Modify code&lt;/td&gt;
&lt;td&gt;Execute commands&lt;/td&gt;
&lt;td&gt;Derive Claude&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Externalize working memory&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time model&lt;/td&gt;
&lt;td&gt;Present, one interaction&lt;/td&gt;
&lt;td&gt;Present, one operation&lt;/td&gt;
&lt;td&gt;Present, command lifecycle&lt;/td&gt;
&lt;td&gt;Present, fork/join&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Persistent across time&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State location&lt;/td&gt;
&lt;td&gt;None, conversation-driven&lt;/td&gt;
&lt;td&gt;Disk + harness&lt;/td&gt;
&lt;td&gt;Gone after command&lt;/td&gt;
&lt;td&gt;Inside the subagent&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Runtime storage&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main benefit&lt;/td&gt;
&lt;td&gt;User alignment&lt;/td&gt;
&lt;td&gt;Precise code changes&lt;/td&gt;
&lt;td&gt;Engineering workflow&lt;/td&gt;
&lt;td&gt;Context space&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Against forgetting; visible collaboration&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Naming pattern&lt;/td&gt;
&lt;td&gt;Enter / Exit pair&lt;/td&gt;
&lt;td&gt;Read / Edit / Write family&lt;/td&gt;
&lt;td&gt;Single tool&lt;/td&gt;
&lt;td&gt;Single tool&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;CRUD + Stop / Output&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Task family is most tightly coupled with Agent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent delegates work whose outcome may fail, hang, or need stopping.&lt;/li&gt;
&lt;li&gt;Tasks provide the work-item container that makes subagent work trackable.&lt;/li&gt;
&lt;li&gt;TaskStop accepts a subagent ID or task ID, creating a unified stop entry point.&lt;/li&gt;
&lt;li&gt;TaskOutput used to retrieve subagent results directly; now Read handles the output file.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Task and Bash also have a useful analogy. Bash’s &lt;code&gt;run_in_background&lt;/code&gt; puts a command in the background; TaskCreate puts a todo in persistent storage. Both prevent the main loop from blocking, but they solve different problems: Bash is asynchronous machine I/O, while Task is asynchronous human-AI coordination.&lt;/p&gt;

&lt;p&gt;The family’s position in the ecosystem is therefore unique. The first nine tools perform one immediate operation per call. Task is a &lt;strong&gt;meta-primitive that stores what should happen across time&lt;/strong&gt;.&lt;/p&gt;




&lt;h3&gt;
  
  
  Summary
&lt;/h3&gt;

&lt;p&gt;The Task family’s elegance is not the existence of a todo list. It is the way its signals span four layers and form a complete dual system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Naming:&lt;/strong&gt; six tools—four CRUD operations plus Stop and Output. &lt;code&gt;activeForm&lt;/code&gt; encodes grammar in a field name; &lt;code&gt;TaskDelete&lt;/code&gt; is omitted in favor of &lt;code&gt;status: "deleted"&lt;/code&gt;; output retrieval falls back to Read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool-level descriptions:&lt;/strong&gt; each tool has an independent prompt, but they reference one another and encode the collaboration contract—three-step threshold, timing protocol, false-completion prohibition, staleness warning, deprecation notice, and reminder hook.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Field-level descriptions:&lt;/strong&gt; subject, description, and activeForm map to three UI contexts; blocks and blockedBy expose both directions of a dependency; add-prefixed fields prevent destructive replacement; owner is a hard collaboration field while metadata is a flexible escape hatch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema validation:&lt;/strong&gt; static constraints such as required activeForm and status enums live in the schema; dynamic constraints such as state transitions, dependency locks, and stale updates live in runtime.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Task extends Claude Code from the &lt;strong&gt;present tense to the future tense&lt;/strong&gt;. The first nine tools do something now; Task stores what must happen in runtime storage across tool calls, time, and Claude instances. The result is a shift from relying on mental effort against forgetting to using a system against forgetting. Forgetting is no longer catastrophic because the list remains.&lt;/p&gt;

&lt;p&gt;The deeper insight is that a dual tool family needs both a complete lifecycle and an exit path. CRUD is not “create without delete”: completed ends the normal lifecycle, deleted cleans up mistaken tasks, and dependency updates release blocked work. Every task has a defined way to end.&lt;/p&gt;

&lt;p&gt;The forced progressive &lt;code&gt;activeForm&lt;/code&gt; is the family’s boldest field design. It turns “fill in a UI label” into a grammar transformation and trains Claude to see work as &lt;strong&gt;currently happening&lt;/strong&gt;, not merely intended. That distinction is the difference between having started and planning to start—and the difference between a static todo list and a live work rhythm.&lt;/p&gt;

&lt;p&gt;The next article will examine WebFetch + WebSearch, the sister tools that take Claude beyond the filesystem and asynchronous tasks into the external web: one is “curl with AI,” the other “search with filters.”&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
