<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Umesh Malik</title>
    <description>The latest articles on DEV Community by Umesh Malik (@umesh_malik).</description>
    <link>https://dev.to/umesh_malik</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3777486%2F9bb4f37b-acd0-4752-9675-5e1cf9dd0b78.jpg</url>
      <title>DEV Community: Umesh Malik</title>
      <link>https://dev.to/umesh_malik</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/umesh_malik"/>
    <language>en</language>
    <item>
      <title>How to Increase SEO Traffic in the AI Era: 10 Techniques (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Wed, 29 Jul 2026 17:49:55 +0000</pubDate>
      <link>https://dev.to/umesh_malik/how-to-increase-seo-traffic-in-the-ai-era-10-techniques-2026-3dca</link>
      <guid>https://dev.to/umesh_malik/how-to-increase-seo-traffic-in-the-ai-era-10-techniques-2026-3dca</guid>
      <description>&lt;p&gt;&lt;strong&gt;How to increase SEO traffic in the AI era&lt;/strong&gt; comes down to one uncomfortable truth: you no longer earn traffic by ranking — you earn it by being the source an AI reaches for, and then by giving the reader a reason to click through anyway. Put plainly, &lt;strong&gt;AI-era SEO&lt;/strong&gt; is the practice of structuring your content and technical surfaces so AI engines retrieve it, quote it, and attribute it to you — while the occasional click still lands. This is the hands-on companion to &lt;a href="https://umesh-malik.com/blog/seo-in-the-ai-era-geo-playbook" rel="noopener noreferrer"&gt;SEO in the AI Era: The 2026 GEO Playbook&lt;/a&gt;. That post explained &lt;em&gt;what changed and why&lt;/em&gt;. This one is the field guide: ten genuine, white-hat techniques you can start running this week, no budget and no ten-person content team required.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Growth now has two dials, not one.&lt;/strong&gt; Citation share (do AI answers quote you?) and residual click-through (does the reader still visit?). You have to move both — optimizing only for rank moves neither.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The highest-leverage technique costs nothing:&lt;/strong&gt; rewrite the first 60 words of your top pages into self-contained, quotable answers. It's the single change that most reliably turns a ranking page into a cited one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;First-party data is the only durable moat.&lt;/strong&gt; Models synthesize summaries for free; they cannot synthesize your benchmark, your invoice, your incident. One original number beats a thousand rephrased paragraphs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Genuine techniques only.&lt;/strong&gt; Everything here is white-hat: no link schemes, no astroturfing, no AI-filler at volume. Those tactics violate search-engine spam policies and AI engines punish them too. Growth that survives is growth you'd be happy to explain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distribution is half the job.&lt;/strong&gt; Being on Reddit, YouTube and the two forums that own your niche puts you inside the sources the models already trust — often faster than ranking your own domain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure citations, not average CTR.&lt;/strong&gt; In a world where ~68% of searches end without a click, average CTR falls even as your real traffic and revenue grow. Track the funnel that actually pays.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to Increase SEO Traffic in the AI Era, Concretely
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;In the AI era, traffic grows when your content gets retrieved and quoted inside AI answers, when those answers earn the residual click, and when classic organic still sends the qualified visitors AI engines can't intercept.&lt;/strong&gt; Three overlapping funnels, not one. The mistake almost everyone makes is pouring effort into the middle of the old funnel — more keywords, more backlinks, more word count — while the new one quietly decides who wins.&lt;/p&gt;

&lt;p&gt;Here's the mental model I use. Ranking gets you into the &lt;em&gt;candidate set&lt;/em&gt; an AI engine draws from. Structure and specificity decide whether you get &lt;em&gt;extracted&lt;/em&gt; from that set. Entity authority decides whether you get &lt;em&gt;attributed&lt;/em&gt;. And genuine value decides whether the human bothers to &lt;em&gt;click&lt;/em&gt;. Every technique below moves exactly one of those four gates — I'll tell you which.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: You can't "increase SEO traffic" as a single number anymore. You increase citation share and residual clicks separately, and they respond to different levers. Confusing the two is why so much effort produces so little movement.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Ten Techniques, Ranked by Effort-to-Payoff
&lt;/h2&gt;

&lt;p&gt;I ordered these by payoff per hour, not by importance. The first three you can do this week with content you already have. The last three are where you leave everyone else behind — and where almost nobody is competing yet.&lt;/p&gt;

&lt;p&gt;The technical half — techniques 09 and 10 — is a weekend of work if you're already on Cloudflare. I documented the exact implementation in &lt;a href="https://umesh-malik.com/blog/agentic-browsing-pagespeed-ai-ready" rel="noopener noreferrer"&gt;Agentic Browsing in PageSpeed Insights&lt;/a&gt;, the performance side in &lt;a href="https://umesh-malik.com/blog/core-web-vitals-optimization-guide" rel="noopener noreferrer"&gt;How to Fix Core Web Vitals&lt;/a&gt; (a stable, fast page is one screenshotting agents can actually read), and the callable endpoint end to end in &lt;a href="https://umesh-malik.com/blog/deploy-mcp-server-cloudflare-workers" rel="noopener noreferrer"&gt;Deploy an MCP Server on Cloudflare Workers&lt;/a&gt; and &lt;a href="https://umesh-malik.com/blog/how-to-build-mcp-server" rel="noopener noreferrer"&gt;How to Build an MCP Server&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Rewrite That Turns a Ranking Page Into a Cited One
&lt;/h2&gt;

&lt;p&gt;Technique 01 sounds abstract until you see it side by side. Same facts, same page — only one version can be lifted into an AI answer without a human editing it first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Genuine vs. Spam: The Line You Don't Cross
&lt;/h2&gt;

&lt;p&gt;Every "grow your traffic fast" thread eventually recommends something that works for a month and then torches your domain. In the AI era the blast radius is bigger, because both search engines &lt;em&gt;and&lt;/em&gt; the models learn to distrust the pattern. Here's the honest split.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why the spam column is worse than useless now&lt;/strong&gt;&lt;br&gt;
Classic spam tactics degraded slowly — a penalty here, a de-index there. AI engines add a second, harsher feedback loop: once a model learns your content is low-trust or synthetic, it stops reaching for you across every query, and there's no "reconsideration request" for a model's retrieval prior. The genuine techniques are slower, but they're the only ones that compound instead of detonate.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Does AI Search Traffic Actually Convert?
&lt;/h2&gt;

&lt;p&gt;If you only looked at CTR you'd conclude the effort isn't worth it. Look at composition instead.&lt;/p&gt;

&lt;p&gt;The mechanism is &lt;strong&gt;intent compression&lt;/strong&gt;. Someone who spent four turns in ChatGPT refining "I need a vector DB for a 50M-embedding hybrid-search workload" has already done the comparison shopping. When they land on you, they're at the &lt;em&gt;end&lt;/em&gt; of the funnel. That's why techniques 04 and 05 — original data and comparison tables — pay off twice: they're what gets you cited, and they're what closes the visitor who arrives pre-qualified.&lt;/p&gt;

&lt;h2&gt;
  
  
  The One-Page Weekly Routine
&lt;/h2&gt;

&lt;p&gt;Techniques don't move traffic; &lt;em&gt;repeated&lt;/em&gt; techniques do. This is the entire routine, small enough to actually keep.&lt;/p&gt;

&lt;p&gt;Once a quarter, layer in the bigger swings: publish one piece of genuinely first-party data (technique 04), refresh your top 20 pages (08), and — if you haven't yet — ship &lt;code&gt;llms.txt&lt;/code&gt;, schema, a Markdown mirror and an MCP endpoint (09–10). The weekly routine moves extraction and attribution; the quarterly swings move the candidate set.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Mistakes That Cap Your Growth
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Optimizing for a rank you already have.&lt;/strong&gt; Position #3 and not being quoted? More backlinks won't fix it — rewrite the passage so it's extractable. Ranking is the qualifier; structure is the prize.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chasing volume over evidence.&lt;/strong&gt; Ten AI-written pages restating the docs will get you cited zero times. One page with a real benchmark gets cited repeatedly. Publish less, prove more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treating &lt;code&gt;llms.txt&lt;/code&gt; as a keyword dump.&lt;/strong&gt; It's a map — an H1, a real description, links to your best content. Stuffing it is the 2007 meta-keywords mistake in a new file. Ship it because it's cheap and now a Lighthouse audit, not because it's a proven ranking lever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blocking every AI crawler in a panic.&lt;/strong&gt; Blocking retrieval bots (&lt;code&gt;OAI-SearchBot&lt;/code&gt;, &lt;code&gt;PerplexityBot&lt;/code&gt;, &lt;code&gt;ChatGPT-User&lt;/code&gt;) cuts you out of the only path to being cited. Decide per bot — it's reasonable to allow retrieval crawlers while restricting pure-training ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reporting average CTR as your headline metric.&lt;/strong&gt; With ~68% of searches ending click-free, average CTR falls while total clicks and revenue climb. Report total clicks, AI-channel conversions and citation share instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Abandoning classic SEO.&lt;/strong&gt; AI Overviews are largely assembled from pages that already rank, and cited brands earn meaningfully more organic clicks than uncited ones. GEO is a layer on top of technical SEO — an unindexed page is uncitable.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;Increasing SEO traffic in the AI era isn't a new trick bolted onto the old playbook — it's a different objective that happens to share a name. You're no longer buying a rank and collecting the clicks it pays out. You're earning a place in the answer, and hoping the reader still wants the source.&lt;/p&gt;

&lt;p&gt;The good news is that the techniques that work are the ones you'd want to do anyway: write clearer answers, prove your claims with real numbers, show up genuinely where your audience already is, and make your site trivially easy for a machine to read. None of it requires a budget or a team. All of it compounds. And unlike the shortcuts, none of it blows up in your face a month later.&lt;/p&gt;

&lt;p&gt;Start with technique 01 on your single best page this week. Then read part one — &lt;a href="https://umesh-malik.com/blog/seo-in-the-ai-era-geo-playbook" rel="noopener noreferrer"&gt;SEO in the AI Era: The 2026 GEO Playbook&lt;/a&gt; — for the full strategic picture behind why these ten moves are the ones that matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2311.09735" rel="noopener noreferrer"&gt;GEO: Generative Engine Optimization — the original research paper&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://searchengineland.com/google-zero-click-searches-2026-study-479717" rel="noopener noreferrer"&gt;Google zero-click searches reach 68% in early 2026 — Search Engine Land&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/" rel="noopener noreferrer"&gt;The crawl before the fall of referrals — Cloudflare Radar&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.similarweb.com/blog/marketing/geo/gen-ai-stats/" rel="noopener noreferrer"&gt;Gen AI stats 2026: AI visibility trends — Similarweb&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.semrush.com/blog/most-cited-domains-ai/" rel="noopener noreferrer"&gt;The most-cited domains in AI: a 3-month study — Semrush&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering. This is part two of the AI-era SEO series — start with &lt;a href="https://umesh-malik.com/blog/seo-in-the-ai-era-geo-playbook" rel="noopener noreferrer"&gt;SEO in the AI Era: The 2026 GEO Playbook&lt;/a&gt; for the strategy behind these techniques.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/increase-seo-traffic-ai-era-techniques" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/seo-in-the-ai-era-geo-playbook" rel="noopener noreferrer"&gt;SEO in the AI Era: The 2026 GEO Playbook for Winning AI Search Traffic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/agentic-browsing-pagespeed-ai-ready" rel="noopener noreferrer"&gt;Agentic Browsing in PageSpeed Insights: How to Make Your Website AI-Ready (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/react-server-components-guide" rel="noopener noreferrer"&gt;React Server Components in 2026: The Mental Model, the use client Boundary &amp;amp; When Not to Use Them&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>seo</category>
      <category>geo</category>
      <category>aisearch</category>
      <category>contentstrategy</category>
    </item>
    <item>
      <title>SEO in the AI Era: The 2026 GEO Playbook for Winning AI Search Traffic</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Mon, 27 Jul 2026 15:44:19 +0000</pubDate>
      <link>https://dev.to/umesh_malik/seo-in-the-ai-era-the-2026-geo-playbook-for-winning-ai-search-traffic-4iop</link>
      <guid>https://dev.to/umesh_malik/seo-in-the-ai-era-the-2026-geo-playbook-for-winning-ai-search-traffic-4iop</guid>
      <description>&lt;p&gt;&lt;strong&gt;SEO in the AI era&lt;/strong&gt; is no longer a competition for ten blue links. It's a competition to be the source an AI model reaches for when it writes the answer — and then, occasionally, to earn the click that follows. The ranking is still real. The click is now optional.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The click collapsed, not the search.&lt;/strong&gt; Zero-click searches hit roughly &lt;strong&gt;68%&lt;/strong&gt; in early 2026, up from about 45% a decade ago. Pew Research measured click-through at &lt;strong&gt;8% when an AI Overview is present vs 15% when it isn't&lt;/strong&gt;, and Seer Interactive clocked organic CTR falling &lt;strong&gt;61%&lt;/strong&gt; (1.76% → 0.61%) across 3,119 informational queries. Google's AI Mode is worse still: a &lt;strong&gt;~93% zero-click rate&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The remaining traffic is dramatically better.&lt;/strong&gt; Similarweb clickstream data (Apr–May 2026) puts ChatGPT referral conversion at &lt;strong&gt;7.1%&lt;/strong&gt; — second only to paid search. Ahrefs found AI search was &lt;strong&gt;0.5% of traffic but 12.1% of signups&lt;/strong&gt;. Fewer visitors, far higher intent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GEO is not "SEO but with AI in the title."&lt;/strong&gt; Generative Engine Optimization optimizes for &lt;em&gt;extraction and citation&lt;/em&gt;, not position. The peer-reviewed GEO research found &lt;strong&gt;quotations lift AI visibility 41%, statistics 32%, cited sources 30%, and fluency 28%&lt;/strong&gt; — levers that have no equivalent in classic SEO.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;There is no single algorithm to game.&lt;/strong&gt; An analysis of 680M citations found only &lt;strong&gt;11% of domains are cited by both ChatGPT and Perplexity&lt;/strong&gt;. You are optimizing for a fragmented committee, not one crawler.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The technical bar moved from "crawlable" to "callable."&lt;/strong&gt; Clean semantic HTML, &lt;code&gt;llms.txt&lt;/code&gt;, structured data, a Markdown mirror of every page, and increasingly an MCP endpoint. Machines are now a first-class audience with their own read path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Track it or you'll conclude the wrong thing.&lt;/strong&gt; Only ~14% of marketers separate AI search as a channel; most of it is silently misfiled as "direct" in GA4. You cannot manage a channel you can't see.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Your Traffic Didn't Drop Because You Got Worse at SEO
&lt;/h2&gt;

&lt;p&gt;Here's the conversation I keep having. Someone shows me a Search Console chart where impressions are flat or up, average position is stable or improving — and clicks have fallen off a cliff since late 2025. Then they ask what they broke.&lt;/p&gt;

&lt;p&gt;They broke nothing. &lt;strong&gt;The search engine stopped being a referral machine and started being an answer machine.&lt;/strong&gt; Your content still ranked. It just got read, summarized, and delivered to the user inside a chat box with a citation chip they didn't click.&lt;/p&gt;

&lt;p&gt;This is the single most important mental shift of the last two years, and most SEO advice hasn't caught up. The industry is still optimizing for a rank position that increasingly determines &lt;em&gt;whether you get quoted&lt;/em&gt;, not &lt;em&gt;whether you get visited&lt;/em&gt;. Those are different objectives with different tactics, and confusing them is why so many teams are working hard and losing ground.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: In the AI era, ranking is the qualifier and citation is the prize. Position #1 that never gets extracted is worth less than position #6 that gets quoted in every answer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What Is GEO (Generative Engine Optimization)?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Generative Engine Optimization (GEO) is the practice of structuring content, entities and technical surfaces so that AI systems — ChatGPT, Claude, Perplexity, Google AI Overviews and AI Mode, Copilot — retrieve your page, extract a specific claim from it, and attribute that claim to you in a generated answer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The distinction that matters: classic SEO optimizes a &lt;em&gt;document&lt;/em&gt; for a &lt;em&gt;position&lt;/em&gt;. GEO optimizes a &lt;em&gt;passage&lt;/em&gt; for &lt;em&gt;retrieval and quotation&lt;/em&gt;. An AI engine doesn't rank your page — it chunks it, embeds the chunks, retrieves the two or three most relevant, and synthesizes. Your unit of competition shrank from "the article" to "the paragraph."&lt;/p&gt;

&lt;p&gt;That single fact drives almost every practical tactic below. If your best insight is buried in paragraph nine of a section that requires the previous eight paragraphs for context, it cannot be extracted cleanly, so it will not be cited. A self-contained, factual, quotable paragraph beats a beautifully-argued essay that only works as a whole.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;GEO, AEO, AI SEO — the naming is a mess&lt;/strong&gt;&lt;br&gt;
You'll see Generative Engine Optimization, Answer Engine Optimization, LLM SEO and AI Search Optimization used interchangeably. They describe the same job: earning visibility inside generated answers rather than inside a list of links. Pick one term and move on — arguing about the acronym is the least valuable thing you can do this quarter.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What Actually Changed — Nine Shifts, One Table
&lt;/h2&gt;

&lt;p&gt;Most "AI changed SEO" posts wave at the vibe. Here is the concrete delta, item by item.&lt;/p&gt;

&lt;p&gt;Three of those deserve unpacking, because they're the ones people get wrong.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Query fan-out means you can't see the query anymore
&lt;/h3&gt;

&lt;p&gt;Google's AI Mode doesn't run your keyword. It decomposes the user's task into a set of synthetic sub-queries, runs them in parallel, and assembles the result. You will never see those sub-queries in Search Console. This is why keyword-level reporting is quietly becoming fiction, and why &lt;strong&gt;topical coverage beats keyword targeting&lt;/strong&gt;: you want to be a plausible answer to a whole neighborhood of questions you cannot enumerate.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The citation market is fragmented, and that's good news
&lt;/h3&gt;

&lt;p&gt;An analysis of 680 million citations across ChatGPT, Google AI Overviews and Perplexity found that &lt;strong&gt;only 11% of domains are cited by both ChatGPT and Perplexity&lt;/strong&gt;. Each engine has different retrieval logic, different index freshness, different trust priors. And critically, even the most-cited domain on any platform rarely exceeds ~5% of total citations — versus classic SEO, where the top 10 results eat roughly two-thirds of clicks.&lt;/p&gt;

&lt;p&gt;Translation: the AI citation market is &lt;em&gt;less&lt;/em&gt; winner-take-all than the blue-link market ever was. A small, sharp, well-structured site can get cited alongside a Fortune 500 in a way it could never outrank one.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Crawling exploded; referrals didn't follow
&lt;/h3&gt;

&lt;p&gt;This is the uncomfortable part. Cloudflare Radar's crawl-to-refer ratios (May 2026) show what the exchange actually looks like:&lt;/p&gt;

&lt;p&gt;Cloudflare attributed &lt;strong&gt;51.8% of AI crawler requests to training&lt;/strong&gt;, 35.7% to mixed training-plus-retrieval, and only &lt;strong&gt;9.3% to search-only&lt;/strong&gt; purposes. So most of the machine attention on your site is not shopping for a link to send you. Deciding what to do about that — block, allow, or monetize — is now a real strategic call, not a robots.txt afterthought.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Bother? Because the Traffic That Survives Is Better
&lt;/h2&gt;

&lt;p&gt;If you stopped at the CTR numbers you'd conclude the web is over. It isn't. The composition changed.&lt;/p&gt;

&lt;p&gt;Three independent datasets point the same direction. Similarweb's clickstream panel puts ChatGPT referral conversion at &lt;strong&gt;7.1%&lt;/strong&gt;, behind only paid search (7.8%) and ahead of direct, organic, social and email. A 12-month GA4 study by Visibility Labs found ChatGPT traffic converting at &lt;strong&gt;1.81% vs 1.39%&lt;/strong&gt; for non-branded organic — a more modest but real 31% edge. Semrush reports AI-driven visitors converting at roughly 4.4x standard organic.&lt;/p&gt;

&lt;p&gt;The mechanism is &lt;strong&gt;intent compression&lt;/strong&gt;. Someone who spent four turns in ChatGPT narrowing "I need a vector database for a 50M-embedding workload with hybrid search" has already done the comparison shopping. When they land on you, they're at the end of the funnel, not the start. Traditional organic sends you people still browsing; AI sends you people already decided.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The honest caveat&lt;/strong&gt;&lt;br&gt;
Volume is small. AI referrals are typically low single-digit percentages of total sessions for most sites right now, and every one of these conversion studies uses a different panel, attribution window and definition of "AI traffic." Treat the direction as solid and the exact multiplier as marketing. The right posture is "this is a high-value emerging channel worth engineering for," not "abandon organic search."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The GEO Playbook — Nine Moves That Actually Move Citations
&lt;/h2&gt;

&lt;p&gt;The peer-reviewed GEO research gives us the closest thing to a measured baseline: &lt;strong&gt;quotations increased AI visibility by 41%, statistics by 32%, cited sources by 30%, and fluency optimization by 28%.&lt;/strong&gt; Notice what's absent from that list — keyword density, word count, exact-match headings. The levers changed.&lt;/p&gt;

&lt;p&gt;Here's the playbook I actually run, in priority order.&lt;/p&gt;

&lt;p&gt;For the technical half of moves 07 and 09, I documented the exact implementation — Worker, discovery files, MCP endpoint and all — in &lt;a href="https://umesh-malik.com/blog/agentic-browsing-pagespeed-ai-ready" rel="noopener noreferrer"&gt;Agentic Browsing in PageSpeed Insights&lt;/a&gt;, and the endpoint itself in &lt;a href="https://umesh-malik.com/blog/how-to-build-mcp-server" rel="noopener noreferrer"&gt;How to Build a Production MCP Server&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Rewrite That Doubles Your Citation Odds
&lt;/h2&gt;

&lt;p&gt;Abstract advice is easy to nod at and hard to apply. Here is the same content, written both ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Measure AI Traffic When Analytics Lies to You
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Most AI traffic is invisible by default because AI clients strip or mangle the referrer, so GA4 files the session under "direct."&lt;/strong&gt; Only about 14% of marketers currently track AI search as a distinct channel — which means most teams are either underestimating a growing channel or crediting it to the wrong one.&lt;/p&gt;

&lt;p&gt;Fix it in three layers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1 — a referral channel group.&lt;/strong&gt; Create a custom channel in GA4 matching AI hostnames on the session source. A regex that covers today's field:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chatgpt\.com|chat\.openai\.com|perplexity\.ai|claude\.ai|copilot\.microsoft\.com|gemini\.google\.com|you\.com|phind\.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note that &lt;code&gt;google.com&lt;/code&gt; referrals from AI Overviews and AI Mode look identical to classic organic — you cannot cleanly separate them in GA4 today. Don't pretend otherwise in your reporting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2 — server logs for the crawl side.&lt;/strong&gt; Referrals only show you the ~1% that clicked. Your access logs show the other side of the trade: who's reading you. Filter by user agent for &lt;code&gt;GPTBot&lt;/code&gt;, &lt;code&gt;OAI-SearchBot&lt;/code&gt;, &lt;code&gt;ChatGPT-User&lt;/code&gt;, &lt;code&gt;ClaudeBot&lt;/code&gt;, &lt;code&gt;Claude-User&lt;/code&gt;, &lt;code&gt;PerplexityBot&lt;/code&gt;, &lt;code&gt;Google-Extended&lt;/code&gt;, &lt;code&gt;Applebot-Extended&lt;/code&gt;, &lt;code&gt;Bytespider&lt;/code&gt;, &lt;code&gt;CCBot&lt;/code&gt;. Compute your own crawl-to-refer ratio per bot. That ratio is the honest scoreboard for whether the exchange is working for you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 — citation tracking.&lt;/strong&gt; Run your 20 highest-value questions through ChatGPT, Perplexity, AI Mode and Claude on a fixed monthly cadence and record whether you're cited. It's manual and it's noisy — answers vary run to run — but it's the only direct read on the metric that matters. Tools exist for this; a spreadsheet and a recurring calendar block works fine to start.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The metric to actually report&lt;/strong&gt;&lt;br&gt;
Stop leading with average CTR. In a world where 68% of searches end without a click, a falling average CTR can coexist with growing total clicks and growing revenue. Report total clicks, AI-channel conversions, and citation share. Average CTR is now a ratio whose denominator is being inflated by impressions you were never going to convert.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Common Mistakes That Are Costing You Citations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Optimizing for a rank you already have.&lt;/strong&gt; If you're position #3 and not being quoted, more backlinks won't fix it. Rewrite the passage so it's extractable. Ranking gets you into the retrieval candidate set; structure gets you into the answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treating &lt;code&gt;llms.txt&lt;/code&gt; as a keyword dump.&lt;/strong&gt; It's a map: an H1, a real description, and links to your genuinely best content. Stuffing it is the 2007 meta-keywords mistake in a new file. Adoption is still early and even Google has been lukewarm — ship it because it's cheap and now an explicit Lighthouse audit, not because it's a proven ranking lever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publishing AI-written filler at volume.&lt;/strong&gt; The engines are retrieving from a corpus increasingly full of generated text. Undifferentiated content has no reason to be picked. First-party data, original benchmarks and named opinions are the only durable moat, and they're the one thing a model can't synthesize from the existing web.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blocking every AI crawler in a panic.&lt;/strong&gt; Blocking &lt;code&gt;ClaudeBot&lt;/code&gt; and &lt;code&gt;GPTBot&lt;/code&gt; cuts training use — and also cuts you out of the retrieval paths that share user agents or infrastructure. Decide deliberately per bot: it's reasonable to allow retrieval bots (&lt;code&gt;OAI-SearchBot&lt;/code&gt;, &lt;code&gt;PerplexityBot&lt;/code&gt;, &lt;code&gt;ChatGPT-User&lt;/code&gt;) while restricting pure training crawlers. Just know which you're doing and why.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chasing an "AI visibility score" from a tool.&lt;/strong&gt; Every vendor has an index and none of them can see inside the models. Use them for directional trend, never as a KPI. Your own citation spot-checks and server logs are more honest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Abandoning classic SEO.&lt;/strong&gt; AI Overviews are largely assembled from pages that already rank. Brands cited inside AI Overviews earn roughly 35% more organic clicks than uncited competitors. GEO is a layer on top of technical SEO, not a replacement for it — an unindexed page is uncitable.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Being PRO at SEO in the AI Era
&lt;/h2&gt;

&lt;p&gt;Everything above gets you to competent. Four things separate the professionals.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. You publish things that cannot be synthesized.&lt;/strong&gt; The web is filling with plausible restatements of existing knowledge, and models are excellent at producing those for free. The only content with a structural advantage is content the model cannot generate: your benchmark run, your production incident, your pricing comparison with real invoices, your opinion with your name on it. Every hour spent producing first-party evidence is worth ten spent rephrasing documentation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. You own an entity, not a keyword list.&lt;/strong&gt; Pick a narrow territory and be visibly, consistently the person or brand associated with it across every surface a model reads — your site, GitHub, YouTube, conference decks, forum answers, other people's posts. Entity strength is what breaks ties, and ties are most of the game now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. You treat machines as a first-class audience with their own read path.&lt;/strong&gt; Humans get the designed page. Agents get semantic HTML, a Markdown mirror, &lt;code&gt;llms.txt&lt;/code&gt;, structured data, and a callable endpoint. Same content, three formats, one source of truth. Most sites are still serving agents a JavaScript-heavy page and hoping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. You measure the funnel, not the vanity metric.&lt;/strong&gt; Impressions → citations → clicks → conversions, with the AI channel broken out. You'll often find the channel that looks worst on CTR is the best on revenue per session. Teams that report on the wrong end of that funnel keep optimizing away their most valuable traffic.&lt;/p&gt;

&lt;p&gt;Two of those overlap with work you may already have done: the CLS and performance items are covered in my &lt;a href="https://umesh-malik.com/blog/core-web-vitals-optimization-guide" rel="noopener noreferrer"&gt;Core Web Vitals optimization guide&lt;/a&gt;, and the callable layer is a weekend's work if you're already on Cloudflare — see &lt;a href="https://umesh-malik.com/blog/deploy-mcp-server-cloudflare-workers" rel="noopener noreferrer"&gt;Deploy an MCP Server on Cloudflare Workers&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;SEO isn't dying. It's being demoted from "the traffic channel" to "the qualification round." Ranking still decides whether you're in the retrieval candidate set. Everything after that — whether you get extracted, quoted, attributed, and occasionally clicked — is a different discipline with different levers, and almost nobody is running it deliberately yet.&lt;/p&gt;

&lt;p&gt;That's the opportunity. The citation market is fragmented enough that a small site with sharp, specific, well-structured, genuinely original content can sit in the same generated answer as a company with a hundred-person content team. That was never true of the blue links.&lt;/p&gt;

&lt;p&gt;So: write answers, not articles. Publish evidence, not summaries. Serve machines a format they can read without guessing. Measure citations, not average CTR. And stop optimizing for a click-through rate that the interface itself decided to take away from you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://searchengineland.com/google-zero-click-searches-2026-study-479717" rel="noopener noreferrer"&gt;Google zero-click searches reach 68% in early 2026 — Search Engine Land&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/" rel="noopener noreferrer"&gt;The crawl before the fall of referrals — Cloudflare Radar&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://searchengineland.com/chatgpt-vs-non-branded-organic-search-conversions-470321" rel="noopener noreferrer"&gt;ChatGPT traffic converts 31% higher than non-branded organic search — Search Engine Land&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.similarweb.com/blog/marketing/geo/gen-ai-stats/" rel="noopener noreferrer"&gt;Gen AI stats 2026: AI visibility trends — Similarweb&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.semrush.com/blog/most-cited-domains-ai/" rel="noopener noreferrer"&gt;The most-cited domains in AI: a 3-month study — Semrush&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2311.09735" rel="noopener noreferrer"&gt;GEO: Generative Engine Optimization — the original research paper&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering. Want the technical half in depth? Read &lt;a href="https://umesh-malik.com/blog/agentic-browsing-pagespeed-ai-ready" rel="noopener noreferrer"&gt;Agentic Browsing in PageSpeed Insights: How to Make Your Website AI-Ready&lt;/a&gt; next.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/seo-in-the-ai-era-geo-playbook" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/agentic-browsing-pagespeed-ai-ready" rel="noopener noreferrer"&gt;Agentic Browsing in PageSpeed Insights: How to Make Your Website AI-Ready (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/react-server-components-guide" rel="noopener noreferrer"&gt;React Server Components in 2026: The Mental Model, the use client Boundary &amp;amp; When Not to Use Them&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/core-web-vitals-optimization-guide" rel="noopener noreferrer"&gt;How to Fix Core Web Vitals: LCP, INP &amp;amp; CLS (2026)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>seo</category>
      <category>geo</category>
      <category>aisearch</category>
      <category>contentstrategy</category>
    </item>
    <item>
      <title>Claude Opus 5: Benchmarks, Pricing &amp; Two Breaking Changes</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Fri, 24 Jul 2026 18:55:22 +0000</pubDate>
      <link>https://dev.to/umesh_malik/claude-opus-5-benchmarks-pricing-two-breaking-changes-4hnj</link>
      <guid>https://dev.to/umesh_malik/claude-opus-5-benchmarks-pricing-two-breaking-changes-4hnj</guid>
      <description>&lt;p&gt;&lt;strong&gt;Claude Opus 5&lt;/strong&gt; landed on July 24, 2026, and it quietly did the most disruptive thing a model launch can do: it kept the old price. Same &lt;code&gt;$5 / $25&lt;/code&gt; per million tokens as Opus 4.8, and it more than doubles Opus 4.8's score on Anthropic's hardest agentic coding benchmark. In the API it's &lt;code&gt;claude-opus-5&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That combination breaks the mental model most teams have been running on. For the last two months the ladder was simple — default to Sonnet, escalate to Opus, and pay double for &lt;a href="https://umesh-malik.com/blog/claude-fable-5-guide" rel="noopener noreferrer"&gt;Fable 5&lt;/a&gt; when the task was genuinely brutal. Opus 5 collapses the top two rungs: it lands within half a percentage point of Fable 5 on real-world coding evals at roughly half the cost per task. The interesting question is no longer "when do I escalate to Fable?" It's "is there still a reason to?"&lt;/p&gt;

&lt;p&gt;There's a catch, and it's the part the launch post doesn't lead with. Two API behaviors changed, and if you migrate by swapping the model string, one of them will silently truncate your responses and the other will 400 your requests.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Opus 5 is a drop-in upgrade at Opus 4.8's exact price&lt;/strong&gt; — &lt;code&gt;$5&lt;/code&gt; per million input tokens, &lt;code&gt;$25&lt;/code&gt; output, 1M context, 128K max output. No price increase, no new tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It gets near-Fable 5 capability at half Fable's cost.&lt;/strong&gt; Within 0.5% of Fable 5 on CursorBench 3.2, past Fable 5's peak OSWorld 2.0 score at about a third of the budget, and 43.3% on Frontier-Bench v0.1 versus Fable 5's 33.7%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two breaking changes will bite a string-swap migration.&lt;/strong&gt; Thinking is now &lt;strong&gt;on by default&lt;/strong&gt; (so a tight &lt;code&gt;max_tokens&lt;/code&gt; can truncate mid-answer), and &lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; returns a &lt;code&gt;400&lt;/code&gt; at &lt;code&gt;xhigh&lt;/code&gt; or &lt;code&gt;max&lt;/code&gt; effort.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-sweep your effort levels.&lt;/strong&gt; &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; are meaningfully stronger on Opus 5 than on any earlier Opus — they're now the primary cost lever, not a compromise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The play:&lt;/strong&gt; make Opus 5 your default ceiling, keep &lt;a href="https://umesh-malik.com/blog/claude-sonnet-5-guide" rel="noopener noreferrer"&gt;Sonnet 5&lt;/a&gt; for high-volume work, and reserve Fable 5 for the narrow band where you've &lt;em&gt;measured&lt;/em&gt; that Opus 5 falls short.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Is Claude Opus 5?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Claude Opus 5 is Anthropic's model for complex agentic coding and enterprise work — a frontier-class model priced at the mid-flagship tier.&lt;/strong&gt; It has a 1-million-token context window, up to 128K output tokens per request, adaptive thinking that's on by default, high-resolution vision, and a May 2026 knowledge cutoff. You call it as &lt;code&gt;claude-opus-5&lt;/code&gt; on the Claude API, &lt;code&gt;anthropic.claude-opus-5&lt;/code&gt; on Amazon Bedrock, and &lt;code&gt;claude-opus-5&lt;/code&gt; on Google Cloud and Microsoft Foundry.&lt;/p&gt;

&lt;p&gt;The one-sentence version: &lt;strong&gt;it's the model that made "escalate to the premium tier" a much rarer decision.&lt;/strong&gt; Anthropic positions Fable 5 as the answer when you need the highest available capability, and Opus 5 as the answer for everything else demanding — which, after you look at the benchmarks, turns out to be nearly everything.&lt;/p&gt;

&lt;p&gt;Three things define it in practice:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It verifies its own work.&lt;/strong&gt; This is the headline behavioral change over Opus 4.8. It opens pages in a browser at desktop and phone widths, checks its output against reality, and iterates before handing back.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It finishes in fewer turns.&lt;/strong&gt; Across effort levels, Opus 5 averaged roughly 9 percentage points higher accuracy while using about a third fewer turns and tool calls, and 60% less wall-clock time. Fewer turns is not a vanity metric — it's the token bill.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It's the most aligned Opus Anthropic has shipped.&lt;/strong&gt; It scored 2.3 on their automated overall-misalignment audit, the lowest of any recent model, with the lowest rates of deceptive behavior and the least susceptibility to being tricked into misuse.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: The story of Opus 5 is not a capability jump — those happen every few months now. It's a &lt;em&gt;price-tier&lt;/em&gt; jump. Frontier-class results arrived on the shelf below the frontier-class price.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Claude Opus 5 Benchmarks
&lt;/h2&gt;

&lt;p&gt;Anthropic leaned on a newer, harder set of evals for &lt;a href="https://www.anthropic.com/news/claude-opus-5" rel="noopener noreferrer"&gt;this launch&lt;/a&gt; — the older saturated ones stopped separating models. Here's what the numbers actually say.&lt;/p&gt;

&lt;p&gt;Two results deserve a second look.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frontier-Bench v0.1 is the one that matters for engineering teams.&lt;/strong&gt; Opus 5 more than doubles Opus 4.8 (43.3% vs 18.7%) &lt;em&gt;and&lt;/em&gt; clears Fable 5 (33.7%) — at half Fable's price. When a mid-tier model beats the premium tier on the premium tier's home turf, the ladder has changed shape.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ARC-AGI 3 is the one that matters for everyone else.&lt;/strong&gt; 30.2% roughly triples the previous best any model had posted, and sits about four times GPT-5.6 Sol's 7.8%. ARC-AGI is deliberately built to resist memorization — it tests reasoning on problems the model hasn't seen a template for. A jump that size on that benchmark is not benchmark-tuning.&lt;/p&gt;

&lt;p&gt;The customer quotes line up with the numbers rather than contradicting them, which is rarer than it should be. Devin called it "approaches Fable-level performance at half the cost." Cursor: "near Fable 5 intelligence at Opus speed and cost." Lovable called it the "biggest leap in the Opus family since 4.5," with "far less variance run to run" — and run-to-run variance is the thing that actually decides whether you can put a model in an unattended pipeline.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where Opus 5 is still behind&lt;/strong&gt;&lt;br&gt;
Opus 5 deliberately does not advance the frontier on dual-use capability. It stays behind Mythos 5 on biology research and offensive cybersecurity: on OSS-Fuzz it identifies vulnerabilities at a similar rate but is considerably less successful at developing exploits. Its cyber classifiers are also less restrictive than Fable 5's — expected to intervene around 85% less often — with flagged requests falling back to Opus 4.8 by default.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Real Cost Math
&lt;/h2&gt;

&lt;p&gt;This is the section that changes what you do on Monday. Opus 5 did not get a price increase. It's &lt;code&gt;$5&lt;/code&gt; per million input tokens and &lt;code&gt;$25&lt;/code&gt; per million output — the same numbers Opus 4.8, 4.7, 4.6, and 4.5 all carried.&lt;/p&gt;

&lt;p&gt;Run the arithmetic on a real workload. Take an agentic run that consumes 500K input tokens (large context, re-sent across turns) and 100K output tokens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Opus 5:&lt;/strong&gt; &lt;code&gt;$2.50 + $2.50 = $5.00&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fable 5:&lt;/strong&gt; &lt;code&gt;$5.00 + $5.00 = $10.00&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sonnet 5:&lt;/strong&gt; &lt;code&gt;$1.50 + $1.50 = $3.00&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But raw per-token price understates the gap, because Opus 5 needs fewer tokens to finish. It hit comparable quality while generating about &lt;strong&gt;26% fewer tokens&lt;/strong&gt; than Opus 4.8 at max reasoning, with a third fewer tool calls. Applied to the run above, that's closer to &lt;code&gt;$3.70&lt;/code&gt; of real spend — against &lt;code&gt;$10.00&lt;/code&gt; on Fable 5 for a result within half a percentage point on CursorBench. &lt;strong&gt;That's not a 2x saving. It's closer to 2.7x.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Two pricing details worth knowing before you budget:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fast mode is a separate line item.&lt;/strong&gt; &lt;code&gt;speed: "fast"&lt;/code&gt; gives you up to 2.5x output throughput, priced at &lt;code&gt;$10 / $50&lt;/code&gt; — double the base rate, and Claude API only (not Bedrock, Google Cloud, or Foundry). It also draws on its own rate-limit pool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Opus 5 has its own rate-limit bucket.&lt;/strong&gt; Opus 4.8, 4.7, 4.6, and 4.5 share one combined Opus limit. Opus 5 does not draw from it. Shifting traffic over neither frees headroom on the old bucket nor inherits it — check your tier's Opus 5 limits before you move volume, or you'll rate-limit yourself on launch day.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Two Breaking Changes
&lt;/h2&gt;

&lt;p&gt;Here's the part that costs people an afternoon. Everything from the Opus 4.7 API surface still holds — &lt;code&gt;budget_tokens&lt;/code&gt; is gone, sampling parameters (&lt;code&gt;temperature&lt;/code&gt;, &lt;code&gt;top_p&lt;/code&gt;, &lt;code&gt;top_k&lt;/code&gt;) are rejected, and last-assistant-turn prefills 400. If you're already on 4.8, those are clean. Two things are genuinely new.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Thinking is on by default — and it eats your &lt;code&gt;max_tokens&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;On Opus 4.8 and 4.7, omitting the &lt;code&gt;thinking&lt;/code&gt; parameter meant &lt;strong&gt;no thinking&lt;/strong&gt;. On Opus 5, omitting it runs &lt;strong&gt;adaptive thinking&lt;/strong&gt;. The wire value didn't change; the default did.&lt;/p&gt;

&lt;p&gt;This is a silent cost and truncation change, not just a behavior one. &lt;code&gt;max_tokens&lt;/code&gt; is a hard cap on thinking &lt;em&gt;plus&lt;/em&gt; response text. A route that ran thinking-off on Opus 4.8 and sized &lt;code&gt;max_tokens&lt;/code&gt; tightly around its expected answer can now truncate mid-response — with no error, just a shorter reply and &lt;code&gt;stop_reason: "max_tokens"&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# On Opus 4.8 this ran with NO thinking. On Opus 5 it thinks —
# and thinking tokens come out of the same 2000-token budget.
&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# ← now shared between thinking and the answer
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fix is one of two lines: raise &lt;code&gt;max_tokens&lt;/code&gt; to leave room, or opt out explicitly with &lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; — subject to the next change.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Disabling thinking 400s at &lt;code&gt;xhigh&lt;/code&gt; and &lt;code&gt;max&lt;/code&gt; effort
&lt;/h3&gt;

&lt;p&gt;You can only disable thinking at effort &lt;code&gt;high&lt;/code&gt; or lower. Pair &lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; with &lt;code&gt;xhigh&lt;/code&gt; or &lt;code&gt;max&lt;/code&gt; and the request returns a &lt;code&gt;400&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# 400 on Claude Opus 5 — disabled thinking above "high" effort
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;thinking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;disabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;output_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;effort&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;xhigh&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;   &lt;span class="c1"&gt;# ← invalid combination
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The trap is that &lt;strong&gt;this is validated per request&lt;/strong&gt;, not per conversation. A later call that bumps effort to &lt;code&gt;xhigh&lt;/code&gt; while thinking is still disabled gets rejected even though every earlier call in the same session succeeded. Audit every call site, not just the first one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The honest recommendation on disabled thinking&lt;/strong&gt;&lt;br&gt;
Don't disable it. On Opus 5 the thinking-off path has two documented failure modes: the model occasionally writes a tool call into its visible response text instead of emitting a structured tool_use block — the turn succeeds, the call never runs, no error is raised — and it can leak internal &amp;lt;thinking&amp;gt; tags into the response. Running thinking-on at low or medium effort is cheaper and avoids both. If you truly must stay thinking-off, add "You may say a brief sentence before using a tool," delete any don't-reason instruction (it makes tag leakage worse), and ask generically for no internal XML tags.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Effort: Re-Sweep It, Don't Reuse It
&lt;/h2&gt;

&lt;p&gt;Opus 5 supports all five effort levels — &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;xhigh&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt; — with &lt;code&gt;high&lt;/code&gt; as the API default. The guidance genuinely changed from 4.8, and copying your old settings over is the most common way to leave money on the table.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Effort&lt;/th&gt;
&lt;th&gt;When to use it on Opus 5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Genuinely frontier problems where correctness outranks cost. Can overthink simpler tasks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;xhigh&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;The starting point for coding and agentic work.&lt;/strong&gt; Set &lt;code&gt;max_tokens&lt;/code&gt; to at least 64K here.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;high&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The default. Most intelligence-sensitive workloads that aren't long-horizon agentic.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;medium&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Now a real option, not a compromise. Sweep it — quality often holds.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;low&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Subagents, scoped tasks, latency-sensitive paths. Stronger here than on any earlier Opus.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The line from Anthropic's &lt;a href="https://platform.claude.com/docs/en/build-with-claude/effort" rel="noopener noreferrer"&gt;own effort docs&lt;/a&gt; is unusually direct: &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; are stronger on Opus 5 than on earlier Opus models, and you should "use them liberally as your primary control for token cost and response time wherever your evals show quality holds." That is the opposite of the reflex most teams built up over the 4.x series, where dropping below &lt;code&gt;high&lt;/code&gt; meant visibly worse output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One trap:&lt;/strong&gt; effort controls &lt;em&gt;thinking volume&lt;/em&gt;, not visible response length. If Opus 5's answers are too long, lowering effort won't reliably fix it — you have to prompt for brevity. A short conciseness instruction cut user-facing response length by about 20% in testing; changing effort didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Else Is New
&lt;/h2&gt;

&lt;p&gt;The refusal path deserves a code snippet, because it's the one that crashes production:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;beta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;16000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;betas&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;server-side-fallback-2026-07-01&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;fallbacks&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="c1"&gt;# routes by refusal category, no model list to maintain
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stop_reason&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;handle_refusal&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A refused request is a &lt;strong&gt;successful HTTP 200&lt;/strong&gt;, not an exception. Code that reads &lt;code&gt;response.content[0]&lt;/code&gt; unconditionally breaks on it. Check &lt;code&gt;stop_reason&lt;/code&gt; first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opus 5 vs Fable 5 vs Sonnet 5
&lt;/h2&gt;

&lt;p&gt;{#snippet newContent()}&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Opus 5 is now the right answer for almost all of it.&lt;/strong&gt; It's within half a percentage point of Fable 5 on real-world coding evals at half the cost per task, beats Fable 5 outright on Frontier-Bench and OSWorld 2.0, finishes in fewer turns, and carries a four-months-fresher knowledge cutoff. It's also the most aligned Opus shipped, with less restrictive cyber classifiers than Fable 5 — so fewer false-positive refusals on benign security work.&lt;/p&gt;

&lt;p&gt;{/snippet}&lt;br&gt;
  {#snippet oldContent()}&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fable 5 still wins a narrow band.&lt;/strong&gt; The very longest unattended runs, where Anthropic still recommends it for extended autonomous operation, and cyber/bio-adjacent research where its higher dual-use ceiling is the point. If you're paying double, you should be able to name the specific task where you &lt;em&gt;measured&lt;/em&gt; Opus 5 falling short — not "it's the flagship."&lt;/p&gt;

&lt;p&gt;{/snippet}&lt;/p&gt;

&lt;p&gt;And Sonnet 5 hasn't been displaced — it's just been re-scoped. It remains the right default for high-volume backends, chat, RAG, and anything latency-sensitive, where &lt;code&gt;$3 / $15&lt;/code&gt; and fast turnaround beat a capability ceiling you never touch. Just remember its &lt;a href="https://umesh-malik.com/blog/claude-sonnet-5-guide" rel="noopener noreferrer"&gt;tokenizer inflates counts by roughly 30%&lt;/a&gt;, which eats into the apparent saving versus Opus 5.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The heuristic I'd use in July 2026:&lt;/strong&gt; Sonnet 5 for volume, Opus 5 as the ceiling for everything hard, Fable 5 only where a measured gap justifies double. That's a simpler ladder than we had last month, and it's cheaper at the top.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Honest Ledger
&lt;/h2&gt;

&lt;h2&gt;
  
  
  The Migration Checklist
&lt;/h2&gt;

&lt;p&gt;That seventh item is the counterintuitive one and it's worth repeating: &lt;strong&gt;delete your verification scaffolding.&lt;/strong&gt; Instructions like "include a final verification step" or "double-check your answer" were good practice on earlier models. On Opus 5 they cause over-verification, and removing them reduces it with no capability regression. It's a delete, not a rewrite — which inverts a prompting rule most teams have internalized. If you maintain a shared prompt library or a house &lt;a href="https://umesh-malik.com/blog/how-to-write-claude-md" rel="noopener noreferrer"&gt;CLAUDE.md&lt;/a&gt;, that's the line to carve out.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;Claude Opus 5 is the rare launch where the interesting number is the one that &lt;em&gt;didn't&lt;/em&gt; change. Same &lt;code&gt;$5 / $25&lt;/code&gt; as the model it replaces, near-Fable capability, fewer turns to get there, and a knowledge cutoff four months fresher than the flagship. If you're running Opus 4.8 today, this is a migration where the ceiling goes up and the bill goes down — those don't come along often.&lt;/p&gt;

&lt;p&gt;Just don't swap the string and walk away. Thinking is on by default now, which means the &lt;code&gt;max_tokens&lt;/code&gt; you tuned last quarter is quietly a truncation risk, and disabled thinking plus &lt;code&gt;xhigh&lt;/code&gt; effort is a &lt;code&gt;400&lt;/code&gt; waiting for whichever call site raises effort first. Fifteen minutes of auditing buys you the whole upgrade.&lt;/p&gt;

&lt;p&gt;Then do the thing most teams skip: &lt;strong&gt;re-sweep your effort levels.&lt;/strong&gt; The single biggest cost lever in this release isn't the price tag — it's that &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; finally hold quality on hard work. Measure it on your own evals rather than trusting your Opus 4.8 instincts, and the same workload can get better &lt;em&gt;and&lt;/em&gt; cheaper at the same time.&lt;/p&gt;

&lt;p&gt;Next: if you're deciding where Opus 5 fits alongside your tooling, read the &lt;a href="https://umesh-malik.com/blog/claude-code-vs-cursor-production-work-2026" rel="noopener noreferrer"&gt;Claude Code vs Cursor production breakdown&lt;/a&gt;, and the &lt;a href="https://umesh-malik.com/blog/claude-sonnet-5-guide" rel="noopener noreferrer"&gt;Sonnet 5 tokenizer gotcha&lt;/a&gt; before you assume Sonnet is always the cheaper option.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.anthropic.com/news/claude-opus-5" rel="noopener noreferrer"&gt;Introducing Claude Opus 5&lt;/a&gt; — Anthropic's launch announcement (benchmarks, alignment audit, customer results)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://platform.claude.com/docs/en/about-claude/models/overview" rel="noopener noreferrer"&gt;Models overview&lt;/a&gt; — Anthropic docs: pricing, context window, knowledge cutoff, platform IDs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://platform.claude.com/docs/en/build-with-claude/effort" rel="noopener noreferrer"&gt;Effort&lt;/a&gt; — Anthropic docs: effort levels and the Opus 5 recommendations&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://platform.claude.com/docs/en/about-claude/models/migration-guide" rel="noopener noreferrer"&gt;Model migration guide&lt;/a&gt; — Anthropic docs: the Opus 4.8 → Opus 5 breaking changes&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows" rel="noopener noreferrer"&gt;Anthropic launches Claude Opus 5&lt;/a&gt; — VentureBeat's launch coverage&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/claude-opus-5-guide" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/claude-sonnet-5-guide" rel="noopener noreferrer"&gt;Claude Sonnet 5: Pricing, Benchmarks, Tokenizer Gotcha&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/claude-fable-5-guide" rel="noopener noreferrer"&gt;Claude Fable 5: Capabilities, Cost &amp;amp; When to Use It (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/how-to-write-claude-md" rel="noopener noreferrer"&gt;How to Write a CLAUDE.md That Actually Helps&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudeopus5</category>
      <category>anthropic</category>
      <category>llm</category>
      <category>aicodingagents</category>
    </item>
    <item>
      <title>Vercel AI SDK in Production: Streaming, Tool-Calling &amp; the Gotchas Nobody Tells You (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Mon, 20 Jul 2026 21:43:47 +0000</pubDate>
      <link>https://dev.to/umesh_malik/vercel-ai-sdk-in-production-streaming-tool-calling-the-gotchas-nobody-tells-you-2026-17co</link>
      <guid>https://dev.to/umesh_malik/vercel-ai-sdk-in-production-streaming-tool-calling-the-gotchas-nobody-tells-you-2026-17co</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Vercel AI SDK in production&lt;/strong&gt; is easy to start and easy to get wrong — it collapses a week of streaming plumbing into an afternoon, then quietly leaves the hard parts to you.&lt;/li&gt;
&lt;li&gt;The demo is &lt;code&gt;streamText&lt;/code&gt; + &lt;code&gt;useChat&lt;/code&gt;. &lt;strong&gt;Production is everything around it&lt;/strong&gt;: aborting generations, tool-calling loops that don't run forever, error and retry UX, rate limiting, and cost you can attribute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wire the abort signal all the way through.&lt;/strong&gt; Stopping only on the client keeps the model generating — and billing — in the background.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate limiting and cost control are your job, not the SDK's.&lt;/strong&gt; Add a &lt;code&gt;429&lt;/code&gt; gate before &lt;code&gt;streamText&lt;/code&gt;, and log &lt;code&gt;usage&lt;/code&gt; from &lt;code&gt;onFinish&lt;/code&gt; per user.&lt;/li&gt;
&lt;li&gt;Default your routes to the &lt;strong&gt;Node runtime&lt;/strong&gt;; reach for Edge only when global cold-start latency is a real, measured problem.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Most Vercel AI SDK tutorials lie by omission
&lt;/h2&gt;

&lt;p&gt;You've seen the tutorial. Install &lt;code&gt;ai&lt;/code&gt; and &lt;code&gt;@ai-sdk/react&lt;/code&gt;, drop a &lt;code&gt;streamText&lt;/code&gt; call in a route handler, wire up &lt;code&gt;useChat&lt;/code&gt;, and — magic — tokens stream into a chat bubble. Ship it Friday.&lt;/p&gt;

&lt;p&gt;Then real users show up. Someone fires ten prompts and cancels each one; your model keeps generating all ten in the background because "stop" only stopped the UI. A tool call loops on itself and burns 40 model calls for one question. A provider hiccups and your chat just... freezes, with no error, no retry, nothing. At the end of the month you get a bill and no idea which feature caused it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Vercel AI SDK is genuinely excellent — it's the fastest way to build AI features in a TypeScript app.&lt;/strong&gt; But the gap between a working demo and a production feature is real, and it's exactly the part nobody writes about. This post is that part.&lt;/p&gt;

&lt;p&gt;Everything below targets &lt;strong&gt;AI SDK 5&lt;/strong&gt; (the transport-based &lt;code&gt;useChat&lt;/code&gt;) on &lt;strong&gt;Next.js 15&lt;/strong&gt; with the App Router. The patterns port to Vue and Svelte too — the core is framework-agnostic.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Vercel AI SDK actually is
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Vercel AI SDK is a TypeScript toolkit that gives you one API to talk to any LLM provider, plus framework hooks that stream model output into your UI.&lt;/strong&gt; It has two halves you should keep straight in your head:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI SDK Core&lt;/strong&gt; — server-side functions like &lt;code&gt;streamText&lt;/code&gt;, &lt;code&gt;generateText&lt;/code&gt;, and &lt;code&gt;tool&lt;/code&gt;. Provider-agnostic: swap &lt;code&gt;@ai-sdk/openai&lt;/code&gt; for &lt;code&gt;@ai-sdk/anthropic&lt;/code&gt; and the rest of your code doesn't change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI SDK UI&lt;/strong&gt; — framework hooks like &lt;code&gt;useChat&lt;/code&gt; and &lt;code&gt;useCompletion&lt;/code&gt; that consume a stream and manage message state for you.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The whole value proposition is that you write a Route Handler with &lt;code&gt;streamText&lt;/code&gt; and drop &lt;code&gt;useChat&lt;/code&gt; into a client component, and the SSE-style streaming, message parsing, and state updates are handled. You don't hand-roll an event stream parser. That's real leverage — and it's why the SDK is worth using over raw &lt;code&gt;fetch&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: The SDK owns the &lt;em&gt;transport&lt;/em&gt; — getting tokens from a model to a React state update. It owns nothing about &lt;em&gt;policy&lt;/em&gt; — who's allowed to call, when to stop, what a failure looks like. Production is all policy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why the Vercel AI SDK in production isn't the demo
&lt;/h2&gt;

&lt;p&gt;Here's the demo, honestly the smallest useful version. A route handler:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// app/api/chat/route.ts&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@ai-sdk/openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;UIMessage&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;}:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;UIMessage&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toUIMessageStreamResponse&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;use client&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useChat&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@ai-sdk/react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;Chat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;sendMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useChat&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setInput&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
          &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;strong&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;role&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;:&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;strong&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
          &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;part&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
            &lt;span class="nx"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;form&lt;/span&gt;
        &lt;span class="na"&gt;onSubmit&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;preventDefault&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
          &lt;span class="nf"&gt;sendMessage&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
          &lt;span class="nf"&gt;setInput&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;input&lt;/span&gt; &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt; &lt;span class="na"&gt;onChange&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;setInput&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;form&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. It streams. Note two things AI SDK 5 changed that trip people up: &lt;code&gt;useChat&lt;/code&gt; &lt;strong&gt;no longer manages your input state&lt;/strong&gt; — you own it with &lt;code&gt;useState&lt;/code&gt; — and messages are made of &lt;strong&gt;parts&lt;/strong&gt;, not a single content string, because a message can interleave text, tool calls, and tool results.&lt;/p&gt;

&lt;p&gt;This works. It's also naive in five specific ways. Let's fix each one.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Tool-calling loops that actually terminate
&lt;/h2&gt;

&lt;p&gt;Tool calling is where the SDK earns its keep — and where a naive setup quietly runs up your bill. A tool is a function the model can decide to call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;stepCountIs&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;zod&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;getWeather&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Get the current weather for a city.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;inputSchema&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;object&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="na"&gt;city&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;City name, e.g. "Berlin"&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
      &lt;span class="p"&gt;}),&lt;/span&gt;
      &lt;span class="na"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;city&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetchWeather&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;city&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;tempC&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tempC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;condition&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;condition&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;stopWhen&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;stepCountIs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two production details. First, &lt;code&gt;inputSchema&lt;/code&gt; is a &lt;strong&gt;Zod schema&lt;/strong&gt; — it's fed to the model &lt;em&gt;and&lt;/em&gt; validates the model's arguments before your &lt;code&gt;execute&lt;/code&gt; runs, so malformed tool calls never reach your code. Use &lt;code&gt;.describe()&lt;/code&gt; liberally; the descriptions are prompt engineering.&lt;/p&gt;

&lt;p&gt;Second — and this is the one that bites — &lt;code&gt;stopWhen: stepCountIs(5)&lt;/code&gt;. Without a stop condition, a multi-step tool loop (model calls a tool, gets a result, decides to call another) can iterate far more than you expect when the model gets confused. &lt;code&gt;stepCountIs(5)&lt;/code&gt; caps the loop at five steps. Set it deliberately based on how many tool hops your feature legitimately needs. An uncapped loop is an uncapped bill.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: A tool-calling agent without a step cap is the AI equivalent of an infinite loop with a network call inside. Always bound the loop.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Aborting a generation without leaking cost
&lt;/h2&gt;

&lt;p&gt;Users cancel. They rephrase mid-stream, hit stop, or navigate away. &lt;code&gt;useChat&lt;/code&gt; gives you this for free — almost:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;sendMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;stop&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useChat&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// status is 'submitted' | 'streaming' | 'ready' | 'error'&lt;/span&gt;
&lt;span class="p"&gt;{(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;submitted&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;streaming&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"button"&lt;/span&gt; &lt;span class="na"&gt;onClick&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;stop&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;Stop&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;)}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Calling &lt;code&gt;stop()&lt;/code&gt; aborts the in-flight &lt;code&gt;fetch&lt;/code&gt;. Here's the trap: &lt;strong&gt;aborting the fetch does not, by itself, stop the model.&lt;/strong&gt; The provider request keeps running server-side unless you propagate the abort. Wire the request's signal into &lt;code&gt;streamText&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;abortSignal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// &amp;lt;-- the line everyone forgets&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toUIMessageStreamResponse&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;req.signal&lt;/code&gt; aborts when the client disconnects. Pass it to &lt;code&gt;streamText&lt;/code&gt; and the provider call is genuinely cancelled — you stop paying for tokens the user will never see. Skip this line and every "stop" is cosmetic: the UI freezes the output while the model finishes generating, invisibly, on your dime.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Error handling and retry UX
&lt;/h2&gt;

&lt;p&gt;Providers fail. Rate limits, timeouts, transient 500s — at any real volume you'll hit all of them. The demo has no error path; the freeze &lt;em&gt;is&lt;/em&gt; the error handling. Do better.&lt;/p&gt;

&lt;p&gt;On the client, &lt;code&gt;useChat&lt;/code&gt; surfaces an &lt;code&gt;error&lt;/code&gt; object and a &lt;code&gt;regenerate&lt;/code&gt; function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;regenerate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useChat&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"alert"&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
    Something went wrong.
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"button"&lt;/span&gt; &lt;span class="na"&gt;onClick&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;regenerate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;Retry&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;)}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the server, control what leaks to the client. By default the SDK masks error details in the stream (good — don't leak provider internals or keys). When you need a real message, pass an &lt;code&gt;onError&lt;/code&gt; mapper to your response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toUIMessageStreamResponse&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;onError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// log the real error server-side; return a safe message to the client&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;[chat] stream error&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;The model is temporarily unavailable. Please retry.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule: &lt;strong&gt;log the truth on the server, show something safe and actionable on the client.&lt;/strong&gt; A user staring at a frozen cursor churns. A user who sees "temporarily unavailable — retry" clicks retry.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Rate limiting — the SDK won't do it for you
&lt;/h2&gt;

&lt;p&gt;Nothing in the AI SDK stops a single user from hammering your route. LLM calls cost real money per request, so an unthrottled chat endpoint is a standing invitation to run up your bill — accidentally or not. Add a gate &lt;strong&gt;before&lt;/strong&gt; you call &lt;code&gt;streamText&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Ratelimit&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@upstash/ratelimit&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Redis&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@upstash/redis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ratelimit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Ratelimit&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromEnv&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="na"&gt;limiter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Ratelimit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slidingWindow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;1 m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="c1"&gt;// 20 requests/min/user&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getUserId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// your auth&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;success&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ratelimit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;success&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Rate limit exceeded. Slow down.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;abortSignal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toUIMessageStreamResponse&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Key it on a &lt;strong&gt;user ID&lt;/strong&gt; where you can, IP as a fallback. Per-user limits survive shared networks and NAT; pure IP limits punish everyone behind one office router. Return &lt;code&gt;429&lt;/code&gt; early so you never pay for the model call you were about to reject.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Cost you can actually attribute
&lt;/h2&gt;

&lt;p&gt;"Our AI bill went up" is a useless sentence if you can't say &lt;em&gt;which feature&lt;/em&gt; or &lt;em&gt;which user&lt;/em&gt;. The provider dashboard gives you a monthly total. &lt;code&gt;onFinish&lt;/code&gt; gives you per-request truth:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;abortSignal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;onFinish&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;usage&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// usage carries inputTokens, outputTokens, totalTokens&lt;/span&gt;
    &lt;span class="nf"&gt;logUsage&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;feature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;chat&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;usage&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Log token usage per user and per feature, then multiply by your model's price. Now "the bill went up" becomes "the summarize feature's output tokens tripled after we widened the context window" — a sentence you can act on. This is also how you decide when to route cheap requests to a smaller model: you can't optimize a cost you don't measure. If you're weighing models, the &lt;a href="https://umesh-malik.com/blog/claude-sonnet-5-guide" rel="noopener noreferrer"&gt;Claude Sonnet 5 guide&lt;/a&gt; walks through the real cost math on exactly this tradeoff.&lt;/p&gt;

&lt;h2&gt;
  
  
  Edge vs Node: pick Node by default
&lt;/h2&gt;

&lt;p&gt;The SDK streams fine on both the Edge and Node runtimes, and a lot of tutorials reach for &lt;code&gt;export const runtime = 'edge'&lt;/code&gt; reflexively. Don't, unless you've measured a reason to.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;Edge runtime&lt;/th&gt;
&lt;th&gt;Node runtime&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cold start&lt;/td&gt;
&lt;td&gt;Very fast&lt;/td&gt;
&lt;td&gt;Slower&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Streaming&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Node APIs / many SDKs&lt;/td&gt;
&lt;td&gt;Partial / unavailable&lt;/td&gt;
&lt;td&gt;Full&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long generations&lt;/td&gt;
&lt;td&gt;Tighter platform limits&lt;/td&gt;
&lt;td&gt;Generous timeouts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vector DB / ORM clients&lt;/td&gt;
&lt;td&gt;Often unsupported&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Edge wins on global cold-start latency. But the moment you need a database client, a Node-only SDK, or a long generation, Edge fights you. &lt;strong&gt;Default to Node, and move a route to Edge only when startup latency is a measured problem for that specific route.&lt;/strong&gt; Premature Edge adoption is a top source of "works locally, breaks in prod."&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stopping on the client only.&lt;/strong&gt; No &lt;code&gt;abortSignal: req.signal&lt;/code&gt; means "stop" is cosmetic and you keep paying. The single most common production bug.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No &lt;code&gt;stopWhen&lt;/code&gt; on tool calls.&lt;/strong&gt; An uncapped multi-step loop is an uncapped bill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treating &lt;code&gt;useChat&lt;/code&gt; like v4.&lt;/strong&gt; In AI SDK 5 you own the input state and read message parts, not a content string. Copy-pasting old tutorials breaks in confusing ways.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No rate limit.&lt;/strong&gt; One script, or one frustrated user mashing send, and your endpoint is a money leak.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leaking raw provider errors&lt;/strong&gt; to the client instead of mapping them with &lt;code&gt;onError&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reflexive Edge runtime&lt;/strong&gt; that breaks the moment you add a DB client.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Best practices, in order
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Always pass &lt;code&gt;abortSignal: req.signal&lt;/code&gt;&lt;/strong&gt; into &lt;code&gt;streamText&lt;/code&gt;. Non-negotiable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Always set &lt;code&gt;stopWhen&lt;/code&gt;&lt;/strong&gt; when tools are involved. Pick the number deliberately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gate the route with a &lt;code&gt;429&lt;/code&gt;&lt;/strong&gt; before calling the model, keyed on user ID.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log &lt;code&gt;usage&lt;/code&gt; in &lt;code&gt;onFinish&lt;/code&gt;&lt;/strong&gt; per user and feature from day one — retrofitting cost attribution is painful.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Map errors with &lt;code&gt;onError&lt;/code&gt;&lt;/strong&gt;; log the real one, return a safe, actionable message.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Start on Node.&lt;/strong&gt; Promote to Edge per-route only when you've measured a latency win.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read message parts&lt;/strong&gt; and render text, tool calls, and tool results distinctly — it's how you build transparent, debuggable agent UIs.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Vercel AI SDK vs raw fetch vs LangChain.js
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Vercel AI SDK&lt;/th&gt;
&lt;th&gt;Raw &lt;code&gt;fetch&lt;/code&gt; + SSE&lt;/th&gt;
&lt;th&gt;LangChain.js&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Streaming to React&lt;/td&gt;
&lt;td&gt;Built-in (&lt;code&gt;useChat&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Hand-rolled parser&lt;/td&gt;
&lt;td&gt;Wrapper, less UI-native&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider switching&lt;/td&gt;
&lt;td&gt;One line&lt;/td&gt;
&lt;td&gt;Rewrite per provider&lt;/td&gt;
&lt;td&gt;Abstracted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool calling + loop control&lt;/td&gt;
&lt;td&gt;First-class (&lt;code&gt;tool&lt;/code&gt;, &lt;code&gt;stopWhen&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;DIY&lt;/td&gt;
&lt;td&gt;First-class, heavier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bundle / footprint&lt;/td&gt;
&lt;td&gt;Small&lt;/td&gt;
&lt;td&gt;Smallest&lt;/td&gt;
&lt;td&gt;Largest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Web apps with streaming UI&lt;/td&gt;
&lt;td&gt;Total control, minimal deps&lt;/td&gt;
&lt;td&gt;Complex chains/orchestration&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My take: for a &lt;strong&gt;web app that streams AI output to a UI, the Vercel AI SDK is the right default&lt;/strong&gt; — you get the transport for free and keep control of policy. Reach for raw &lt;code&gt;fetch&lt;/code&gt; only when you need absolute control and minimal dependencies, and for LangChain.js when your orchestration is genuinely complex (multi-agent graphs, elaborate retrieval chains). For most product teams, that's not day one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The runnable example
&lt;/h2&gt;

&lt;p&gt;The complete, runnable version of everything above — route handler with abort, rate limit, and usage logging, a tool-calling loop with a step cap, and a client with stop and retry — lives in a self-contained project you can clone and run: &lt;strong&gt;&lt;a href="https://github.com/Umeshmalik/examples/tree/main/vercel-ai-sdk-production" rel="noopener noreferrer"&gt;&lt;code&gt;vercel-ai-sdk-production&lt;/code&gt; on GitHub&lt;/a&gt;&lt;/strong&gt;. Copy the patterns, not just the happy path.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The Vercel AI SDK isn't the hard part of shipping AI features — it's the easy part, and it's very good at being easy. &lt;strong&gt;The hard part is the production layer the SDK deliberately leaves to you: when to stop, who's allowed to call, what failure looks like, and what it all costs.&lt;/strong&gt; Wire the abort signal through, cap your tool loops, gate the route, and measure your tokens, and you've closed the gap between a Friday demo and something you can actually put in front of users.&lt;/p&gt;

&lt;p&gt;Next, put this to work on something real: building a &lt;strong&gt;retrieval-augmented chatbot&lt;/strong&gt; with the same SDK. Until that follow-up lands, &lt;a href="https://umesh-malik.com/blog/build-rag-pipeline-from-scratch" rel="noopener noreferrer"&gt;build the RAG pipeline it sits on&lt;/a&gt; and get clear on &lt;a href="https://umesh-malik.com/blog/rag-vs-fine-tuning-llms-2026" rel="noopener noreferrer"&gt;RAG vs fine-tuning&lt;/a&gt; so you're grounding on the right foundation. If you're deploying the model-facing infra too, here's &lt;a href="https://umesh-malik.com/blog/deploy-mcp-server-cloudflare-workers" rel="noopener noreferrer"&gt;how to deploy an MCP server on Cloudflare Workers&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ai-sdk.dev/docs/introduction" rel="noopener noreferrer"&gt;AI SDK by Vercel — documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai-sdk.dev/docs/reference/ai-sdk-core/stream-text" rel="noopener noreferrer"&gt;AI SDK Core: streamText&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai-sdk.dev/docs/ai-sdk-core/tools-and-tool-calling" rel="noopener noreferrer"&gt;AI SDK Core: Tool Calling &amp;amp; stopWhen / stepCountIs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai-sdk.dev/docs/reference/ai-sdk-ui/use-chat" rel="noopener noreferrer"&gt;AI SDK UI: useChat&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vercel.com/blog/ai-sdk-5" rel="noopener noreferrer"&gt;Vercel: AI SDK 5 announcement&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/upstash/ratelimit-js" rel="noopener noreferrer"&gt;Upstash Ratelimit&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/vercel-ai-sdk-production-guide" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/rag-chatbot-nextjs-guide" rel="noopener noreferrer"&gt;Build a RAG Chatbot in Next.js: Retrieval, Streaming &amp;amp; Citations (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/cloudflare-vinext-next-js-vite-revolution" rel="noopener noreferrer"&gt;Cloudflare viNext: The $1,100 Next.js-on-Vite Rebuild&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/react-server-components-guide" rel="noopener noreferrer"&gt;React Server Components in 2026: The Mental Model, the use client Boundary &amp;amp; When Not to Use Them&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>vercelaisdk</category>
      <category>nextjs</category>
      <category>react</category>
      <category>streaming</category>
    </item>
    <item>
      <title>React Server Components in 2026: The Mental Model, the use client Boundary &amp; When Not to Use Them</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Mon, 20 Jul 2026 21:43:15 +0000</pubDate>
      <link>https://dev.to/umesh_malik/react-server-components-in-2026-the-mental-model-the-use-client-boundary-when-not-to-use-them-1l8j</link>
      <guid>https://dev.to/umesh_malik/react-server-components-in-2026-the-mental-model-the-use-client-boundary-when-not-to-use-them-1l8j</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;React Server Components are components that render only on the server and ship zero JavaScript to the browser&lt;/strong&gt; — the default in Next.js 15's App Router.&lt;/li&gt;
&lt;li&gt;The mental model: your app is &lt;strong&gt;two graphs&lt;/strong&gt; — a server graph (default) and a client graph (opt-in with &lt;code&gt;use client&lt;/code&gt;). The whole skill is knowing where to draw the line.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;use client&lt;/code&gt; marks a &lt;strong&gt;boundary, not a file&lt;/strong&gt; — everything imported into a client module joins the client bundle.&lt;/li&gt;
&lt;li&gt;The payoff is &lt;strong&gt;Core Web Vitals&lt;/strong&gt;: teams report 60–70% less client JavaScript, which directly improves LCP and INP.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server Components are not always the answer.&lt;/strong&gt; Anything interactive — state, effects, event handlers, browser APIs — must be a Client Component. Don't fight that.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Server Components aren't experimental anymore — they're the default
&lt;/h2&gt;

&lt;p&gt;For two years React Server Components were the thing everyone had opinions about and nobody fully understood. That era is over. In 2026, &lt;strong&gt;Next.js 15 ships RSC as the default&lt;/strong&gt;, Remix adopted them, and the rest of the ecosystem is falling in line. If you're writing React and still treating every component as a client component, you're shipping JavaScript you don't need to.&lt;/p&gt;

&lt;p&gt;Here's the part that trips people up: RSC isn't a feature you turn on. It's a &lt;strong&gt;change in the default&lt;/strong&gt;. In the App Router, every component is a Server Component &lt;em&gt;unless you say otherwise&lt;/em&gt;. The question flipped from "should this be a server component?" to "does this &lt;em&gt;really&lt;/em&gt; need to be a client component?"&lt;/p&gt;

&lt;p&gt;This post gives you the mental model that makes RSC click, the exact rules for the &lt;code&gt;use client&lt;/code&gt; boundary, and — the part tutorials skip — &lt;strong&gt;when Server Components are the wrong tool.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What React Server Components actually are
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;React Server Components are components that render exclusively on the server and send only their resulting UI to the browser, shipping zero JavaScript for themselves.&lt;/strong&gt; That last clause is the whole point. A Server Component's code — its imports, its data-fetching, its formatting libraries — never reaches the client bundle.&lt;/p&gt;

&lt;p&gt;This is different from SSR, and the confusion between the two is the root of most RSC misunderstanding:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Server-Side Rendering (SSR)&lt;/strong&gt; renders your components to HTML on the server, then ships the JavaScript &lt;em&gt;and&lt;/em&gt; re-runs (hydrates) it on the client. The JS still goes over the wire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;React Server Components&lt;/strong&gt; render on the server and ship &lt;em&gt;no&lt;/em&gt; JavaScript for those components at all. There's nothing to hydrate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So a Server Component can &lt;code&gt;await&lt;/code&gt; your database directly, import a heavy markdown parser, or read the filesystem — and none of that weight lands on the user's phone.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="c1"&gt;// app/page.tsx — a Server Component (no 'use client')&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@/lib/db&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;Page&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;posts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;post&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findMany&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// runs on the server, ships no JS&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;ul&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;posts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;li&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;title&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;li&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;ul&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No &lt;code&gt;useEffect&lt;/code&gt; to fetch, no loading spinner, no client-side data library. The data is there before the component renders.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: SSR is about &lt;em&gt;when&lt;/em&gt; you render (on the server, first). RSC is about &lt;em&gt;where the code lives&lt;/em&gt; (on the server, permanently). One ships JS and hydrates; the other ships none.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Think in two graphs
&lt;/h2&gt;

&lt;p&gt;The mental model that makes everything else obvious: &lt;strong&gt;your component tree is split into two module graphs.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;server graph&lt;/strong&gt; is the default. Components here render on the server, can be &lt;code&gt;async&lt;/code&gt;, can touch server-only resources, and ship no JS.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;client graph&lt;/strong&gt; is opt-in. Components here are your familiar React — &lt;code&gt;useState&lt;/code&gt;, &lt;code&gt;useEffect&lt;/code&gt;, event handlers, browser APIs — and they ship JavaScript.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;use client&lt;/code&gt; is the doorway between them. Everything on the far side of that door is client code. The entire practice of RSC is deciding &lt;em&gt;where to put the door&lt;/em&gt; — as far down the tree as possible, so the interactive leaves are client components and everything above them stays on the server.&lt;/p&gt;

&lt;h2&gt;
  
  
  The &lt;code&gt;use client&lt;/code&gt; boundary
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;use client&lt;/code&gt; at the top of a file marks a &lt;strong&gt;boundary&lt;/strong&gt;, not just that one file. Once a module is a client module, &lt;strong&gt;every module it imports is pulled into the client bundle too.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;use client&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// this file and everything it imports is now client code&lt;/span&gt;

&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;Counter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setN&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"button"&lt;/span&gt; &lt;span class="na"&gt;onClick&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;setN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;button&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mistake this causes: put &lt;code&gt;use client&lt;/code&gt; too high in the tree and you drag half your app into the client bundle. A single &lt;code&gt;use client&lt;/code&gt; at the top of a layout can turn every child into client code, silently erasing the RSC benefit.&lt;/p&gt;

&lt;p&gt;The rule: &lt;strong&gt;push &lt;code&gt;use client&lt;/code&gt; to the leaves.&lt;/strong&gt; Keep the interactive `` a client component; keep the page that renders it a server component.&lt;/p&gt;

&lt;h2&gt;
  
  
  Server Components as children — the pattern that unlocks it
&lt;/h2&gt;

&lt;p&gt;"But my interactive component needs to wrap server content" is the objection everyone hits. The answer is the composition pattern that makes RSC actually usable: &lt;strong&gt;a Client Component can render Server Components passed to it as &lt;code&gt;children&lt;/code&gt; (or any prop).&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;`&lt;code&gt;&lt;/code&gt;tsx&lt;br&gt;
// ClientShell.tsx&lt;br&gt;
'use client';&lt;br&gt;
import { useState } from 'react';&lt;/p&gt;

&lt;p&gt;export function ClientShell({ children }: { children: React.ReactNode }) {&lt;br&gt;
  const [open, setOpen] = useState(true);&lt;br&gt;
  return (&lt;br&gt;
    &lt;/p&gt;
&lt;br&gt;
       setOpen(!open)}&amp;gt;Toggle&lt;br&gt;
      {open &amp;amp;&amp;amp; children}&lt;br&gt;
    &lt;br&gt;
  );&lt;br&gt;
}&lt;br&gt;
&lt;code&gt;&lt;/code&gt;`

&lt;p&gt;`&lt;code&gt;&lt;/code&gt;tsx&lt;br&gt;
// page.tsx — a Server Component&lt;br&gt;
import { ClientShell } from './ClientShell';&lt;br&gt;
import { ServerData } from './ServerData'; // stays on the server!&lt;/p&gt;

&lt;p&gt;export default function Page() {&lt;br&gt;
  return (&lt;/p&gt;

&lt;p&gt;);&lt;br&gt;
}&lt;br&gt;
&lt;code&gt;&lt;/code&gt;`&lt;/p&gt;

&lt;p&gt;`` renders on the server and stays there, even though it's displayed &lt;em&gt;inside&lt;/em&gt; a client component. The client component receives it as already-rendered UI — it never imports it, so it never pulls it into the client bundle. &lt;strong&gt;Interactivity on the outside, server rendering on the inside.&lt;/strong&gt; This is the pattern that lets you keep the boundary low.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming with Suspense
&lt;/h2&gt;

&lt;p&gt;Server Components pair with `` to stream UI progressively. Wrap a slow server component and the rest of the page ships immediately while the slow part streams in when ready:&lt;/p&gt;

&lt;p&gt;`&lt;code&gt;&lt;/code&gt;tsx&lt;br&gt;
import { Suspense } from 'react';&lt;/p&gt;

&lt;p&gt;export default function Page() {&lt;br&gt;
  return (&lt;br&gt;
    &amp;lt;&amp;gt;&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  }&amp;gt;
     {/* streams in when its data resolves */}

&amp;lt;/&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;);&lt;br&gt;
}&lt;br&gt;
&lt;code&gt;&lt;/code&gt;`&lt;/p&gt;

&lt;p&gt;The user sees the shell instantly instead of waiting for the slowest query. If you want the deeper story on out-of-order streaming, I wrote about &lt;a href="https://umesh-malik.com/blog/streaming-html-out-of-order-without-javascript" rel="noopener noreferrer"&gt;streaming HTML out of order without JavaScript&lt;/a&gt; — RSC streaming is the React-flavored version of the same idea.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Web Vitals payoff
&lt;/h2&gt;

&lt;p&gt;This is why I care about RSC, and why you should. &lt;strong&gt;Less client JavaScript is the most direct lever on Core Web Vitals there is.&lt;/strong&gt; Teams moving to RSC report 60–70% reductions in client bundle size, and that shows up where it counts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LCP&lt;/strong&gt; improves because the browser parses and executes less JS before it can paint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;INP&lt;/strong&gt; improves because the main thread isn't clogged hydrating components that never needed to be interactive.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you've been fighting Web Vitals by code-splitting and lazy-loading your way around a giant bundle, RSC attacks the problem at the source: it never sends the bundle. Pair it with the fundamentals in my &lt;a href="https://umesh-malik.com/blog/core-web-vitals-optimization-guide" rel="noopener noreferrer"&gt;Core Web Vitals guide&lt;/a&gt; and the &lt;a href="https://umesh-malik.com/blog/react-performance-optimization-techniques" rel="noopener noreferrer"&gt;React performance techniques&lt;/a&gt; post, and you're optimizing the cause, not the symptom.&lt;/p&gt;

&lt;h2&gt;
  
  
  When NOT to use Server Components
&lt;/h2&gt;

&lt;p&gt;Bold take, honestly held: &lt;strong&gt;Server Components are the right default, and the wrong tool for anything interactive.&lt;/strong&gt; Reach for a Client Component — without guilt — when you need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;State or effects&lt;/strong&gt; — &lt;code&gt;useState&lt;/code&gt;, &lt;code&gt;useReducer&lt;/code&gt;, &lt;code&gt;useEffect&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Event handlers&lt;/strong&gt; — &lt;code&gt;onClick&lt;/code&gt;, &lt;code&gt;onChange&lt;/code&gt;, anything user-driven.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Browser APIs&lt;/strong&gt; — &lt;code&gt;window&lt;/code&gt;, &lt;code&gt;localStorage&lt;/code&gt;, &lt;code&gt;IntersectionObserver&lt;/code&gt;, geolocation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client-only libraries&lt;/strong&gt; — most animation, charting, and map libraries assume the DOM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context that changes on interaction&lt;/strong&gt; — live theme toggles, open/close state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The failure mode isn't "used a client component." It's using client components &lt;em&gt;by default&lt;/em&gt; out of habit, or contorting a genuinely interactive feature to stay on the server. Draw the boundary deliberately: server by default, client at the interactive leaves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;use client&lt;/code&gt; at the top of the tree&lt;/strong&gt; — drags everything below it into the client bundle and erases the benefit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confusing RSC with SSR&lt;/strong&gt; — expecting hydration where there is none, or thinking SSR already gave you the JS savings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;useState&lt;/code&gt;/&lt;code&gt;useEffect&lt;/code&gt; in a Server Component&lt;/strong&gt; — a hard error; those need the client graph.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Importing a Server Component into a Client Component&lt;/strong&gt; — pulls it client-side. Pass it as &lt;code&gt;children&lt;/code&gt; instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fetching in &lt;code&gt;useEffect&lt;/code&gt; out of habit&lt;/strong&gt; — in a Server Component you just &lt;code&gt;await&lt;/code&gt; the data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accessing &lt;code&gt;window&lt;/code&gt; in a Server Component&lt;/strong&gt; — it doesn't exist there; guard or move to the client.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Best practices, in order
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Server by default.&lt;/strong&gt; Only add &lt;code&gt;use client&lt;/code&gt; when a component genuinely needs the client graph.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Push the boundary to the leaves&lt;/strong&gt; — small, interactive client components; server everything above.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compose with &lt;code&gt;children&lt;/code&gt;&lt;/strong&gt; to keep server content inside client shells without importing it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;await&lt;/code&gt; data in Server Components&lt;/strong&gt; instead of client-side fetching where you can.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wrap slow server components in ``&lt;/strong&gt; to stream and protect perceived performance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat client JS as a budget&lt;/strong&gt; — every &lt;code&gt;use client&lt;/code&gt; spends it. Measure your bundle.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Common questions, answered
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Short answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Is RSC the same as SSR?&lt;/td&gt;
&lt;td&gt;No — SSR ships and hydrates JS; RSC ships none for server components.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Do Server Components replace Client Components?&lt;/td&gt;
&lt;td&gt;No — they're the default; client components handle interactivity.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can a Client Component render a Server Component?&lt;/td&gt;
&lt;td&gt;Yes, if passed as &lt;code&gt;children&lt;/code&gt;/props — not if imported.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Where should &lt;code&gt;use client&lt;/code&gt; go?&lt;/td&gt;
&lt;td&gt;At the interactive leaves, as low in the tree as possible.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;React Server Components aren't a new API to memorize — they're a new default to internalize. &lt;strong&gt;Think in two graphs, keep the &lt;code&gt;use client&lt;/code&gt; boundary at the interactive leaves, compose server content into client shells with &lt;code&gt;children&lt;/code&gt;, and you get dramatically less JavaScript for free.&lt;/strong&gt; That's not a micro-optimization; it's the most direct Core Web Vitals win available to a React app in 2026.&lt;/p&gt;

&lt;p&gt;Put it to work: attack your bundle at the source with RSC, then close out the fundamentals with the &lt;a href="https://umesh-malik.com/blog/core-web-vitals-optimization-guide" rel="noopener noreferrer"&gt;Core Web Vitals guide&lt;/a&gt; and &lt;a href="https://umesh-malik.com/blog/react-performance-optimization-techniques" rel="noopener noreferrer"&gt;React performance techniques&lt;/a&gt;. And if you're weighing where React and the framework layer are headed, &lt;a href="https://umesh-malik.com/blog/cloudflare-vinext-next-js-vite-revolution" rel="noopener noreferrer"&gt;Cloudflare viNext and Next.js on Vite&lt;/a&gt; is the wider context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://nextjs.org/docs/app/getting-started/server-and-client-components" rel="noopener noreferrer"&gt;Next.js — Server and Client Components&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://react.dev/reference/rsc/use-client" rel="noopener noreferrer"&gt;React — &lt;code&gt;use client&lt;/code&gt; directive&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://react.dev/reference/rsc/server-components" rel="noopener noreferrer"&gt;React — Server Components&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nextjs.org/learn/react-foundations/server-and-client-components" rel="noopener noreferrer"&gt;Next.js Learn — Server and Client Components&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/react-server-components-guide" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/cloudflare-vinext-next-js-vite-revolution" rel="noopener noreferrer"&gt;Cloudflare viNext: The $1,100 Next.js-on-Vite Rebuild&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/sveltekit-vs-nextjs-comparison" rel="noopener noreferrer"&gt;SvelteKit vs Next.js 2026: Which Should You Choose?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/seo-in-the-ai-era-geo-playbook" rel="noopener noreferrer"&gt;SEO in the AI Era: The 2026 GEO Playbook for Winning AI Search Traffic&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>react</category>
      <category>reactservercomponents</category>
      <category>nextjs</category>
      <category>webperf</category>
    </item>
    <item>
      <title>Build a RAG Chatbot in Next.js: Retrieval, Streaming &amp; Citations (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Mon, 20 Jul 2026 21:43:14 +0000</pubDate>
      <link>https://dev.to/umesh_malik/build-a-rag-chatbot-in-nextjs-retrieval-streaming-citations-2026-20fk</link>
      <guid>https://dev.to/umesh_malik/build-a-rag-chatbot-in-nextjs-retrieval-streaming-citations-2026-20fk</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A RAG chatbot in Next.js&lt;/strong&gt; is a chat UI that retrieves your own data before it answers — so the model responds from your documents, not just its training set.&lt;/li&gt;
&lt;li&gt;The loop is three steps: &lt;strong&gt;embed the question → search a vector store → inject the hits into the prompt → stream a grounded answer.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Use &lt;strong&gt;Supabase + pgvector&lt;/strong&gt; for storage and the &lt;strong&gt;AI SDK&lt;/strong&gt; for embeddings and streaming. No separate vector database required.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Render citations in the UI.&lt;/strong&gt; A RAG answer with no visible sources is just a chatbot with extra steps — citations are what make it trustworthy.&lt;/li&gt;
&lt;li&gt;Most RAG chatbots fail at &lt;strong&gt;retrieval&lt;/strong&gt;, not generation. If the right chunks don't surface, no prompt saves the answer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A chatbot that makes things up is a liability
&lt;/h2&gt;

&lt;p&gt;A plain LLM chatbot is confidently wrong about your business. Ask it about your refund policy, your API, or last quarter's numbers and it will invent something plausible, because it has never seen your data. That's not a bug you can prompt your way out of — the information simply isn't in the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval-augmented generation fixes this by fetching the relevant facts at query time and handing them to the model as context.&lt;/strong&gt; The model stops guessing and starts summarizing what you gave it. Done right, it also &lt;em&gt;cites&lt;/em&gt; where each claim came from, so a human can verify it.&lt;/p&gt;

&lt;p&gt;This post builds the whole thing in &lt;strong&gt;Next.js 15&lt;/strong&gt; with &lt;strong&gt;AI SDK 5&lt;/strong&gt;: the retrieval step, the streaming chat UI, and — the part most tutorials skip — &lt;strong&gt;citations rendered in the interface&lt;/strong&gt; plus the guardrails that keep the model honest.&lt;/p&gt;

&lt;p&gt;One boundary up front: this is the &lt;em&gt;product&lt;/em&gt; layer. If you want how the underlying pipeline works — chunking strategy, embedding choice, hybrid search, reranking — &lt;a href="https://umesh-malik.com/blog/build-rag-pipeline-from-scratch" rel="noopener noreferrer"&gt;build the RAG pipeline from scratch first&lt;/a&gt;. This post assumes you have documents in a vector store and focuses on wiring them into a chatbot users actually trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a RAG chatbot in Next.js actually is
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A RAG chatbot in Next.js is a chat interface where every user question triggers a retrieval step — the app embeds the question, finds the most relevant chunks of your own content, and passes them to the LLM as grounding context before generating a streamed answer.&lt;/strong&gt; Strip away the framing and it's a three-stage request:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Retrieve&lt;/strong&gt; — turn the question into a vector, find the nearest chunks in your store.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Augment&lt;/strong&gt; — build a prompt that says "answer using only this context," with the chunks inlined.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generate&lt;/strong&gt; — stream the model's answer back to the UI, with citations.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model never sees your whole knowledge base. It sees the handful of chunks retrieval decided were relevant. That's the entire game — and it's why retrieval quality, not prompt cleverness, decides whether your chatbot is useful.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: RAG doesn't make the model smarter. It makes the model &lt;em&gt;informed&lt;/em&gt;. Everything good about the answer traces back to what retrieval put in front of it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The vector store: Supabase + pgvector
&lt;/h2&gt;

&lt;p&gt;You don't need a dedicated vector database. &lt;code&gt;pgvector&lt;/code&gt; runs inside the Postgres you probably already have, which means one less service to operate. Enable the extension and create a table:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="n"&gt;extension&lt;/span&gt; &lt;span class="n"&gt;if&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;exists&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="k"&gt;table&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;bigserial&lt;/span&gt; &lt;span class="k"&gt;primary&lt;/span&gt; &lt;span class="k"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="k"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;source&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;               &lt;span class="c1"&gt;-- url or title, for citations&lt;/span&gt;
  &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1536&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;-- text-embedding-3-small = 1536 dims&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;-- cosine-distance similarity search&lt;/span&gt;
&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="k"&gt;replace&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;match_documents&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;query_embedding&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1536&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="n"&gt;match_count&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;returns&lt;/span&gt; &lt;span class="k"&gt;table&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;bigint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;similarity&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;language&lt;/span&gt; &lt;span class="k"&gt;sql&lt;/span&gt; &lt;span class="k"&gt;stable&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;
  &lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;similarity&lt;/span&gt;
  &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;
  &lt;span class="k"&gt;order&lt;/span&gt; &lt;span class="k"&gt;by&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;query_embedding&lt;/span&gt;
  &lt;span class="k"&gt;limit&lt;/span&gt; &lt;span class="n"&gt;match_count&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="err"&gt;$$&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;&amp;lt;=&amp;gt;&lt;/code&gt; operator is cosine distance; &lt;code&gt;1 - distance&lt;/code&gt; gives you a similarity score you can threshold on. Add an index (&lt;code&gt;ivfflat&lt;/code&gt; or &lt;code&gt;hnsw&lt;/code&gt;) once you have real volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  Embedding and retrieval with the AI SDK
&lt;/h2&gt;

&lt;p&gt;The AI SDK gives you &lt;code&gt;embed&lt;/code&gt; for a single value and &lt;code&gt;embedMany&lt;/code&gt; for batches. Embed the user's question, then hand the vector to your &lt;code&gt;match_documents&lt;/code&gt; function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// lib/retrieve.ts&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;embed&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@ai-sdk/openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@supabase/supabase-js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;supabase&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;SUPABASE_URL&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;SUPABASE_KEY&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;embedding&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;embed&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;textEmbeddingModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;text-embedding-3-small&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;supabase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rpc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;match_documents&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;match_count&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Drop weak matches — a bad hit is worse than no hit.&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="p"&gt;[]).&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;similarity&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.75&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That similarity filter matters more than it looks. &lt;strong&gt;A low-relevance chunk doesn't just waste tokens — it actively misleads the model&lt;/strong&gt;, which will dutifully summarize garbage if you feed it garbage. Threshold, don't just take the top 5.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wiring retrieval into the streaming route
&lt;/h2&gt;

&lt;p&gt;Now connect retrieval to generation. The route embeds the latest question, retrieves context, builds a grounding system prompt, and streams the answer. It reuses every production pattern from &lt;a href="https://umesh-malik.com/blog/vercel-ai-sdk-production-guide" rel="noopener noreferrer"&gt;the Vercel AI SDK production guide&lt;/a&gt; — abort signal, error mapping:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// app/api/chat/route.ts&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@ai-sdk/openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;UIMessage&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;retrieve&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@/lib/retrieve&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;runtime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;nodejs&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;}:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;UIMessage&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="c1"&gt;// The question is the last user message's text.&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;last&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;at&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;question&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;last&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)?.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;question&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;`[&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;] (&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;source&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;)\n&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;streamText&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;abortSignal&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;system&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
      &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;You answer using ONLY the context below.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;If the context does not contain the answer, say you do not know.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Cite sources inline with their bracket numbers, e.g. [1].&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="nx"&gt;context&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;(no relevant context found)&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;convertToModelMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toUIMessageStreamResponse&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="c1"&gt;// Attach the retrieved sources so the client can render citations.&lt;/span&gt;
    &lt;span class="na"&gt;messageMetadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;source&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two decisions that make or break trust. First, the system prompt says &lt;strong&gt;"answer using ONLY the context"&lt;/strong&gt; and &lt;strong&gt;"say you do not know"&lt;/strong&gt; — this is your primary hallucination guardrail. Second, &lt;code&gt;messageMetadata&lt;/code&gt; ships the retrieved sources alongside the stream so the UI can show them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rendering citations in the chat UI
&lt;/h2&gt;

&lt;p&gt;A grounded answer nobody can verify isn't grounded in practice. Surface the sources with the message:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;use client&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useChat&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@ai-sdk/react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;Chat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;sendMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useChat&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setInput&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
          &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;strong&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;role&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;: &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;strong&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
          &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;

          &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="cm"&gt;/* Citations from messageMetadata */&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
          &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;role&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;assistant&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;ul&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
              &lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;src&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;li&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
                  [&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;] &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;a&lt;/span&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;src&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;src&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;a&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
                &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;li&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
              &lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;ul&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
          &lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;

      &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;form&lt;/span&gt;
        &lt;span class="na"&gt;onSubmit&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;preventDefault&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
          &lt;span class="nf"&gt;sendMessage&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
          &lt;span class="nf"&gt;setInput&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
        &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;input&lt;/span&gt; &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt; &lt;span class="na"&gt;onChange&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;setInput&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
      &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;form&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;div&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now every answer arrives with a "Sources" list. Users can click through and check. That single UI affordance is the difference between a demo and something a team will actually rely on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Grounding guardrails that stop hallucinations
&lt;/h2&gt;

&lt;p&gt;The system prompt is necessary but not sufficient. Layer these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Threshold retrieval&lt;/strong&gt; (shown above). No chunk clears the bar? Return "I don't have information on that" instead of an empty-context guess.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instruct the "I don't know" path explicitly.&lt;/strong&gt; Models default to helpfulness; you have to make refusal an allowed, expected outcome.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep context tight.&lt;/strong&gt; Five focused chunks beat twenty loose ones — more context dilutes attention and raises cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cite by construction.&lt;/strong&gt; Number the chunks in the prompt and require bracket citations. If the model can't point to a chunk, it shouldn't make the claim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log the retrieved chunks&lt;/strong&gt; per answer. When the bot is wrong, you'll almost always find retrieval, not generation, was the culprit.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: You cannot debug a RAG chatbot by reading its answers. You debug it by reading what retrieval fed it. Log the chunks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Common mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Skipping the similarity threshold&lt;/strong&gt; — feeding low-relevance chunks that mislead the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No "I don't know" path&lt;/strong&gt; — the model invents an answer when context is empty.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No citations in the UI&lt;/strong&gt; — users can't verify, so they either over-trust or don't trust at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dumping the whole knowledge base&lt;/strong&gt; into context — expensive, slower, and &lt;em&gt;worse&lt;/em&gt; answers from diluted attention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blaming the model for retrieval failures&lt;/strong&gt; — most wrong answers are wrong chunks. Fix retrieval first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embedding the raw multi-turn history&lt;/strong&gt; instead of the actual question — retrieval quality craters.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Best practices, in order
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Threshold every retrieval&lt;/strong&gt; and return a graceful "I don't know" when nothing clears the bar.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ground hard in the system prompt&lt;/strong&gt; — "only this context," and refusal is allowed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Render citations&lt;/strong&gt; with every assistant message; make them clickable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep context small and focused&lt;/strong&gt; — quality of chunks over quantity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log retrieved chunks&lt;/strong&gt; per query for debugging and evals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reuse production patterns&lt;/strong&gt; from the AI SDK — abort signal, rate limit, cost logging — a RAG route is still a model route.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure retrieval with an eval set&lt;/strong&gt; before you tune prompts. If recall is bad, prompting is rearranging deck chairs.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  RAG vs the alternatives
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;Cost / effort&lt;/th&gt;
&lt;th&gt;Freshness&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;RAG chatbot&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Answering from your own, changing data&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;Live — update the store&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fine-tuning&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Style, format, narrow domain behavior&lt;/td&gt;
&lt;td&gt;High, retrain to update&lt;/td&gt;
&lt;td&gt;Frozen at training time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Long-context stuffing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Small, static doc sets&lt;/td&gt;
&lt;td&gt;Low setup, high per-call cost&lt;/td&gt;
&lt;td&gt;Manual&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a chatbot over documents that change, &lt;strong&gt;RAG is the default&lt;/strong&gt; — it's cheaper to keep fresh and it can cite. Fine-tuning teaches behavior, not facts; the two are complementary, not competing. If you're deciding, &lt;a href="https://umesh-malik.com/blog/rag-vs-fine-tuning-llms-2026" rel="noopener noreferrer"&gt;RAG vs fine-tuning&lt;/a&gt; breaks down exactly when each wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  The runnable example
&lt;/h2&gt;

&lt;p&gt;A complete, runnable RAG chatbot — Supabase schema and &lt;code&gt;match_documents&lt;/code&gt; function, an ingestion script that chunks and embeds with &lt;code&gt;embedMany&lt;/code&gt;, the streaming route with grounding, and the citations UI — is in a self-contained project: &lt;strong&gt;&lt;a href="https://github.com/Umeshmalik/examples/tree/main/rag-chatbot-nextjs" rel="noopener noreferrer"&gt;&lt;code&gt;rag-chatbot-nextjs&lt;/code&gt; on GitHub&lt;/a&gt;&lt;/strong&gt;. Clone it, point it at your own docs, and you have a grounded chatbot in an afternoon.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;A RAG chatbot is only as good as what retrieval feeds it — and only as trustworthy as the citations it shows. &lt;strong&gt;Get retrieval right, ground the model hard, and render sources users can click, and you've built a chatbot a team will actually rely on instead of quietly distrust.&lt;/strong&gt; The Next.js and AI SDK plumbing is the easy 20%; the retrieval quality and the trust affordances are the 80% that decides whether it ships.&lt;/p&gt;

&lt;p&gt;From here: harden it with the &lt;a href="https://umesh-malik.com/blog/vercel-ai-sdk-production-guide" rel="noopener noreferrer"&gt;production patterns from the AI SDK guide&lt;/a&gt;, deepen retrieval with the &lt;a href="https://umesh-malik.com/blog/build-rag-pipeline-from-scratch" rel="noopener noreferrer"&gt;from-scratch pipeline&lt;/a&gt;, and settle the &lt;a href="https://umesh-malik.com/blog/rag-vs-fine-tuning-llms-2026" rel="noopener noreferrer"&gt;RAG vs fine-tuning&lt;/a&gt; question for your use case.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ai-sdk.dev/docs/ai-sdk-core/embeddings" rel="noopener noreferrer"&gt;AI SDK — Embeddings (&lt;code&gt;embed&lt;/code&gt; / &lt;code&gt;embedMany&lt;/code&gt;)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai-sdk.dev/cookbook/guides/rag-chatbot" rel="noopener noreferrer"&gt;AI SDK RAG guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://supabase.com/blog/openai-embeddings-postgres-vector" rel="noopener noreferrer"&gt;Supabase — Storing OpenAI embeddings in Postgres with pgvector&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://supabase.com/docs/guides/ai" rel="noopener noreferrer"&gt;Supabase — AI &amp;amp; Vectors docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/pgvector/pgvector" rel="noopener noreferrer"&gt;pgvector&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/rag-chatbot-nextjs-guide" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/vercel-ai-sdk-production-guide" rel="noopener noreferrer"&gt;Vercel AI SDK in Production: Streaming, Tool-Calling &amp;amp; the Gotchas Nobody Tells You (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/build-rag-pipeline-from-scratch" rel="noopener noreferrer"&gt;Build a RAG Pipeline From Scratch (Production Patterns That Actually Matter)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/rag-vs-fine-tuning-llms-2026" rel="noopener noreferrer"&gt;RAG vs Fine-Tuning for LLMs in 2026: A Production Decision Framework With Real Tradeoffs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>rag</category>
      <category>nextjs</category>
      <category>vectordatabases</category>
      <category>llmengineering</category>
    </item>
    <item>
      <title>Kimi K3 vs Claude Fable 5: The Full Head-to-Head Benchmarks, Pricing, and Where the Open Model Wins (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Sun, 19 Jul 2026 10:25:14 +0000</pubDate>
      <link>https://dev.to/umesh_malik/kimi-k3-vs-claude-fable-5-the-full-head-to-head-benchmarks-pricing-and-where-the-open-model-wins-47gl</link>
      <guid>https://dev.to/umesh_malik/kimi-k3-vs-claude-fable-5-the-full-head-to-head-benchmarks-pricing-and-where-the-open-model-wins-47gl</guid>
      <description>&lt;p&gt;&lt;strong&gt;Kimi K3 vs Claude Fable 5&lt;/strong&gt; is the matchup nobody in the West saw coming this fast. On &lt;strong&gt;July 16, 2026&lt;/strong&gt;, Moonshot AI shipped &lt;strong&gt;Kimi K3&lt;/strong&gt; — a &lt;strong&gt;2.8-trillion-parameter&lt;/strong&gt; open-weight model — and it did something no Chinese lab had done before: it beat Anthropic's flagship &lt;strong&gt;Claude Fable 5&lt;/strong&gt; outright on multiple frontier benchmarks, including topping the &lt;strong&gt;Frontend Code Arena&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Let me be honest up front, because the headlines are not: &lt;strong&gt;K3 does not dethrone Fable 5 overall.&lt;/strong&gt; Fable 5 still wins the broader set of evaluations and leads the composite intelligence index. Moonshot's own launch post admits K3 trails Fable 5 and GPT-5.6 Sol on aggregate.&lt;/p&gt;

&lt;p&gt;But that framing misses the real story. &lt;strong&gt;Where Kimi K3 wins, it wins the lanes that matter most for agent builders — long-horizon coding, terminal automation, agentic browsing, and frontend generation — and it does it at roughly one-third the price.&lt;/strong&gt; For a huge class of production workloads, "slightly behind the absolute frontier, at 70% lower cost, with open weights" is not second place. It is the new default.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kimi K3 launched July 16, 2026&lt;/strong&gt; from Moonshot AI: a &lt;strong&gt;2.8T-parameter&lt;/strong&gt; mixture-of-experts model (&lt;strong&gt;16 of 896 experts&lt;/strong&gt; active per token), &lt;strong&gt;1M context&lt;/strong&gt;, native multimodal. Open weights followed within weeks — the &lt;strong&gt;largest open model ever released&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It genuinely beats Claude Fable 5&lt;/strong&gt; on &lt;strong&gt;Terminal-Bench 2.1 (88.3 vs 84.6)&lt;/strong&gt;, &lt;strong&gt;SWE-Marathon (42.0 vs 35.0)&lt;/strong&gt;, &lt;strong&gt;BrowseComp (91.2 vs 88.0)&lt;/strong&gt;, and the &lt;strong&gt;Frontend Code Arena (#1, 1,679 Elo)&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fable 5 still wins overall&lt;/strong&gt;: it leads the &lt;strong&gt;Artificial Analysis Intelligence Index (59.9 vs 57.1)&lt;/strong&gt; and takes &lt;strong&gt;FrontierSWE (86.6 vs 81.2)&lt;/strong&gt;, &lt;strong&gt;DeepSWE (70.0 vs 67.5)&lt;/strong&gt;, &lt;strong&gt;GDPval&lt;/strong&gt;, &lt;strong&gt;JobBench&lt;/strong&gt;, and vision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;K3 costs about 70% less&lt;/strong&gt;: &lt;strong&gt;$3 / $15&lt;/strong&gt; per 1M tokens vs Fable 5's &lt;strong&gt;$10 / $50&lt;/strong&gt; — roughly &lt;strong&gt;3.3x cheaper&lt;/strong&gt;, flat across the full 1M window.&lt;/li&gt;
&lt;li&gt;The right move is &lt;strong&gt;workload routing&lt;/strong&gt;, not wholesale switching: &lt;strong&gt;K3 for agentic/terminal/browsing/high-volume&lt;/strong&gt;, &lt;strong&gt;Fable 5 for frontier-difficulty engineering, vision, and knowledge work&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Big caveat&lt;/strong&gt;: the two ran through &lt;strong&gt;different harnesses&lt;/strong&gt; (KimiCode vs Claude Code/Codex), which limits clean model-to-model conclusions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdj9cle6y7cc6gs5xd6kx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdj9cle6y7cc6gs5xd6kx.png" alt="Kimi K3 vs Claude Fable 5 animated benchmark scorecard showing K3 winning terminal, SWE-Marathon and browsing while Fable 5 wins FrontierSWE, DeepSWE and JobBench" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Kimi K3 vs Claude Fable 5: The Benchmark Scorecard
&lt;/h2&gt;

&lt;p&gt;Here is the full head-to-head on vendor-reported, max-effort numbers. I've marked the winner of each row honestly — including the ones where the open model loses.&lt;/p&gt;

&lt;p&gt;Count it up and the pattern is clear: &lt;strong&gt;Fable 5 wins the broader board, but every single one of K3's wins is in the agentic / long-horizon / automation lane.&lt;/strong&gt; That is not a coincidence — it is what Moonshot optimized for.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: K3 and Fable 5 are not fighting for the same crown. Fable 5 is the frontier-intelligence ceiling; K3 is the agentic-throughput value leader. Which one is "better" depends entirely on whether your workload looks like a hard exam or a long, cheap, repetitive grind.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Watch the Agentic Gap Animate
&lt;/h2&gt;

&lt;p&gt;The wins that matter for agent builders are the long-horizon ones — the tasks where the model has to keep going for dozens or hundreds of steps without losing the thread. Here is how K3 stacks up against Fable 5 on exactly those, as an animated reveal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Kimi K3 Actually Wins
&lt;/h2&gt;

&lt;p&gt;Strip away the leaderboard noise and K3's edge is specific and real.&lt;/p&gt;

&lt;p&gt;The most quotable single result is the &lt;strong&gt;Frontend Code Arena&lt;/strong&gt;: in blind developer testing, K3 ranked &lt;strong&gt;first at 1,679 Elo&lt;/strong&gt;, ahead of Fable 5. That is a human-preference benchmark, not a synthetic one — real developers picked K3's UI code more often. For anyone shipping frontend work, that is the result to actually test against your own prompts.&lt;/p&gt;

&lt;p&gt;If you want context on how fast Chinese labs have been closing this gap, I traced the earlier jump in &lt;a href="https://umesh-malik.com/blog/deepseek-v4-release-challenge-us-ai-rivals" rel="noopener noreferrer"&gt;DeepSeek V4's challenge to US AI rivals&lt;/a&gt; — K3 is that trajectory reaching the frontier.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cost Story Is the Real Weapon
&lt;/h2&gt;

&lt;p&gt;Benchmarks get the headlines. &lt;strong&gt;Price gets the migration.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F101hlwptz53dp0tean9l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F101hlwptz53dp0tean9l.png" alt="Animated cost comparison bars showing Kimi K3 at 3 dollars input and 15 dollars output per million tokens versus Claude Fable 5 at 10 and 50 dollars, about 3.3 times cheaper" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Do the arithmetic on a real agentic workload. A long-horizon coding agent that consumes, say, 20M input and 4M output tokens across a run costs about &lt;strong&gt;$260 on Fable 5&lt;/strong&gt; and about &lt;strong&gt;$120 on K3&lt;/strong&gt; — and K3 actually &lt;em&gt;scores higher&lt;/em&gt; on SWE-Marathon. That is the entire pitch in one line: &lt;strong&gt;more agentic stamina, less than half the bill.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Interactive: The Same Prompt, Two Models
&lt;/h2&gt;

&lt;p&gt;Toggle between the two to see how the tradeoff plays out on a representative agentic-coding task.&lt;/p&gt;

&lt;p&gt;{#snippet oldContent()}&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frontier ceiling, premium bill.&lt;/strong&gt; Fable 5 is the stronger single-shot reasoner and edges ahead on frontier-difficulty subtasks. Adaptive reasoning effort lets it dial compute up on the hard steps. But at &lt;strong&gt;$10 / $50&lt;/strong&gt;, a 40-step run that burns millions of tokens gets expensive fast, and on the &lt;em&gt;long-horizon&lt;/em&gt; version of this task it actually scores lower than K3 (SWE-Marathon 35.0).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pick it when:&lt;/strong&gt; the task has genuinely frontier-hard steps, needs top-tier vision, or a single wrong answer is costly.&lt;/p&gt;

&lt;p&gt;{/snippet}&lt;br&gt;
  {#snippet newContent()}&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic stamina, one-third the cost.&lt;/strong&gt; K3 leads on SWE-Marathon (42.0), Terminal-Bench (88.3), and BrowseComp (91.2) — exactly the skills this loop stresses. At &lt;strong&gt;$3 / $15&lt;/strong&gt; the same 40-step run costs less than half as much, and open weights mean you can self-host to cut it further.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pick it when:&lt;/strong&gt; the workload is long, repetitive, high-volume, or agentic — and cost per run matters more than peak IQ.&lt;/p&gt;

&lt;p&gt;{/snippet}&lt;/p&gt;
&lt;h2&gt;
  
  
  Under the Hood: The Largest Open Model Ever
&lt;/h2&gt;

&lt;p&gt;K3's specs are a statement of intent as much as an engineering result.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsh9kiclug6cgblr2kyxw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsh9kiclug6cgblr2kyxw.png" alt="Kimi K3 architecture diagram showing 2.8 trillion parameters, a sparse mixture-of-experts with 16 of 896 experts active per token, Kimi Delta Attention, and a 1M context window" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The design choices tell you what Moonshot cares about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sparse mixture-of-experts&lt;/strong&gt; — 2.8T total parameters, but only &lt;strong&gt;16 of 896 experts&lt;/strong&gt; fire per token, so inference cost tracks a far smaller active model. This is how they hit frontier quality without frontier-class compute per request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kimi Delta Attention&lt;/strong&gt; — a linear-attention variant that keeps the &lt;strong&gt;1M-token context&lt;/strong&gt; affordable and flat-priced instead of ballooning cost at long context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native multimodal&lt;/strong&gt; — images and video in, not bolted on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open weights&lt;/strong&gt; — the whole thing is downloadable, which is why "largest open model ever" is not just a spec-sheet flex. It changes who can build on the frontier.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The honest caveat you must not skip&lt;/strong&gt;&lt;br&gt;
These benchmark numbers come from &lt;strong&gt;different software harnesses&lt;/strong&gt; — K3 ran through KimiCode, Fable 5 through Claude Code and Codex. Harness quality materially affects agentic scores, so treat every cross-model gap here as directional, not surgical. Before you migrate anything, re-run &lt;strong&gt;your&lt;/strong&gt; prompts through &lt;strong&gt;your&lt;/strong&gt; harness. Vendor benchmarks start the conversation; they don't end it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  How to Use Kimi K3
&lt;/h2&gt;

&lt;p&gt;You have four realistic paths, from zero-effort to full control.&lt;/p&gt;

&lt;p&gt;Because the API is OpenAI-compatible, switching an existing integration is often a base-URL and model-string change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// Point the standard SDK at Moonshot's OpenAI-compatible endpoint.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;MOONSHOT_API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://api.moonshot.ai/v1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;kimi-k3&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Refactor this module and run the tests until they pass.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Two launch-day quirks to plan for&lt;/strong&gt;&lt;br&gt;
At launch K3 fixes &lt;strong&gt;temperature at 1.0&lt;/strong&gt; and offers &lt;strong&gt;only max thinking effort&lt;/strong&gt; — you can't dial reasoning down for cheap, fast extraction yet. If your workload needs low-latency, low-effort calls, benchmark that specifically before assuming K3 is a universal Fable 5 replacement.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A Migration Plan That Won't Burn You
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Timeline
&lt;/h2&gt;

&lt;h2&gt;
  
  
  The Verdict
&lt;/h2&gt;

&lt;p&gt;The takeaway is not "Kimi K3 killed Claude Fable 5." It didn't. Fable 5 is still the sharper model on the hardest single tasks, and the intelligence index says so.&lt;/p&gt;

&lt;p&gt;The takeaway is that &lt;strong&gt;the frontier is no longer a single-vendor, closed-weights club, and the price of "good enough to ship" just fell by 70%.&lt;/strong&gt; For agent builders running long, repetitive, high-volume workloads, K3 is the most disruptive release of the summer — not because it's the smartest model, but because it makes the &lt;em&gt;second&lt;/em&gt;-smartest model cheap and open.&lt;/p&gt;

&lt;p&gt;Route your workloads. Send the agentic grinds to K3, keep the frontier-hard exams on Fable 5, and measure cost per successful task. That's how you turn this rivalry into a lower bill without giving up quality where it counts.&lt;/p&gt;

&lt;p&gt;For the other side of this matchup, read the &lt;a href="https://umesh-malik.com/blog/claude-fable-5-guide" rel="noopener noreferrer"&gt;Claude Fable 5 deep-dive&lt;/a&gt; and the &lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;GPT-5.6 Sol / Terra / Luna guide&lt;/a&gt; — the third model in the frontier race K3 just crashed.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/" rel="noopener noreferrer"&gt;Fortune: Moonshot's Kimi K3 pushes Chinese AI into Fable-level territory&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3" rel="noopener noreferrer"&gt;Tom's Hardware: Kimi K3 beats Claude Fable 5 in Frontend Code Arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/" rel="noopener noreferrer"&gt;MarkTechPost: Kimi K3 — a 2.8T open MoE with Kimi Delta Attention and 1M context&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openrouter.ai/moonshotai/kimi-k3" rel="noopener noreferrer"&gt;OpenRouter: Kimi K3 API pricing &amp;amp; benchmarks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/" rel="noopener noreferrer"&gt;Simon Willison: Kimi K3, and the pelican benchmark&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Explore more:&lt;/strong&gt; &lt;a href="https://umesh-malik.com/topics/llm-engineering" rel="noopener noreferrer"&gt;LLM Engineering — RAG, Fine-Tuning &amp;amp; Production LLMs&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/kimi-k3-vs-claude-fable-5" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/gpt-5-6-sol-vs-terra-vs-luna" rel="noopener noreferrer"&gt;GPT-5.6 Sol vs Terra vs Luna: Which One Should You Actually Use? (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;OpenAI GPT-5.6 Complete Guide: Sol, Terra, Luna Benchmarks, Pricing, and API (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-4-complete-guide" rel="noopener noreferrer"&gt;GPT-5.4 Guide: Benchmarks, Pricing, API &amp;amp; GPT-5.4 Pro&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>kimik3</category>
      <category>claude</category>
      <category>moonshotai</category>
    </item>
    <item>
      <title>Streaming HTML Out of Order Without JavaScript (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Tue, 14 Jul 2026 07:19:26 +0000</pubDate>
      <link>https://dev.to/umesh_malik/streaming-html-out-of-order-without-javascript-2026-4ppp</link>
      <guid>https://dev.to/umesh_malik/streaming-html-out-of-order-without-javascript-2026-4ppp</guid>
      <description>&lt;p&gt;For thirty years, HTML has streamed one way: top to bottom, in the exact order the server wrote it. If the slow database query lived in your page header, everything below it waited. &lt;strong&gt;That constraint is finally breaking — and the fix needs zero JavaScript.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Streaming HTML out of order&lt;/strong&gt; flips that: you send the fast parts of a page immediately, leave labeled holes where the slow parts go, and fill those holes later in the same response — in whatever order the data actually arrives. No client framework. No hydration. No &lt;code&gt;fetch()&lt;/code&gt; on the client. The browser's own HTML parser does the reordering.&lt;/p&gt;

&lt;p&gt;This started as a clever hack built on Declarative Shadow DOM. In 2026 it became a real standards-track feature — &lt;strong&gt;Declarative Partial Updates&lt;/strong&gt; — shipping behind a flag in Chrome 148. Here's how it works, where it came from, and whether you should care yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Streaming HTML out of order&lt;/strong&gt; means sending page sections in the order they become &lt;em&gt;ready&lt;/em&gt; on the server, not the order they appear in the document — so a slow widget never blocks the fast content around it.&lt;/li&gt;
&lt;li&gt;It's been possible &lt;strong&gt;without JavaScript since 2024&lt;/strong&gt; using &lt;strong&gt;Declarative Shadow DOM&lt;/strong&gt; (&lt;code&gt;&amp;lt;template shadowrootmode&amp;gt;&lt;/code&gt;) plus named &lt;code&gt;&amp;lt;slot&amp;gt;&lt;/code&gt; elements. You stream a shell with placeholder slots, then stream the real content later and the browser slots it into place.&lt;/li&gt;
&lt;li&gt;The 2026 evolution is &lt;strong&gt;Declarative Partial Updates (DPU)&lt;/strong&gt; — a &lt;a href="https://github.com/WICG/declarative-partial-updates" rel="noopener noreferrer"&gt;WICG proposal&lt;/a&gt; from Chrome's Barry Pollard and Noam Rosenthal — that removes the Shadow DOM requirement. You drop &lt;code&gt;&amp;lt;?marker&amp;gt;&lt;/code&gt; / &lt;code&gt;&amp;lt;?start&amp;gt;…&amp;lt;?end&amp;gt;&lt;/code&gt; processing instructions where content goes, then stream &lt;code&gt;&amp;lt;template for="…"&amp;gt;&lt;/code&gt; elements to fill them.&lt;/li&gt;
&lt;li&gt;The declarative half is tracked in a &lt;a href="https://github.com/whatwg/html/pull/11818" rel="noopener noreferrer"&gt;WHATWG HTML pull request&lt;/a&gt; (&lt;code&gt;&amp;lt;template for&amp;gt;&lt;/code&gt;) authored by Philip Jägenstedt. There's also a JS half — &lt;code&gt;setHTML&lt;/code&gt;, &lt;code&gt;streamHTML&lt;/code&gt;, &lt;code&gt;appendHTML&lt;/code&gt;, &lt;code&gt;streamAppendHTML&lt;/code&gt; — for updates &lt;em&gt;after&lt;/em&gt; initial load.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Status (mid-2026):&lt;/strong&gt; DPU is behind &lt;code&gt;chrome://flags/#enable-experimental-web-platform-features&lt;/code&gt; in &lt;strong&gt;Chrome 148&lt;/strong&gt;. Firefox and WebKit are interested but haven't shipped. It is &lt;strong&gt;not production-ready&lt;/strong&gt; — but Declarative Shadow DOM streaming &lt;em&gt;is&lt;/em&gt;, today, cross-browser.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What "streaming HTML out of order" actually means
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Streaming HTML out of order is the technique of sending an HTML document in chunks where later chunks patch content into positions defined earlier in the stream, so the browser can render each section as soon as its data is ready rather than waiting for document order.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Compare the two models directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In-order streaming (the last 30 years):&lt;/strong&gt; The server writes the response top to bottom. If your page is a header, a personalized recommendations widget backed by a 900ms query, and a footer, the bytes leave the server in that order. The user stares at a header while the query runs, because the footer literally cannot be sent before the widget above it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Out-of-order streaming:&lt;/strong&gt; The server sends the header, then a &lt;em&gt;placeholder&lt;/em&gt; for the widget ("Loading…"), then immediately the footer — the whole shell arrives in ~50ms. When the 900ms query finishes, the server sends the widget's real HTML as a trailing chunk, tagged to fill the placeholder. The browser swaps it in. The user saw a complete, laid-out page almost instantly, and the slow part filled in without blocking anything.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: Out-of-order streaming decouples &lt;em&gt;the order you author HTML&lt;/em&gt; from &lt;em&gt;the order you can afford to compute it&lt;/em&gt;. Slow data stops being a layout-blocking tax on everything below it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If that sounds like React Suspense or streaming SSR — it is the same &lt;em&gt;idea&lt;/em&gt;. The difference is the machinery. React ships a runtime that receives the out-of-order chunks and repositions DOM nodes with JavaScript. The techniques in this post push that job &lt;strong&gt;down into the browser itself&lt;/strong&gt;, so the reordering happens in the HTML parser with no script at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters right now
&lt;/h2&gt;

&lt;p&gt;Perceived performance is the whole game. Largest Contentful Paint, Time to First Byte, and how "alive" a page feels are dominated by how fast &lt;em&gt;something useful&lt;/em&gt; appears — not by when the slowest widget resolves. (If you're optimizing these numbers, my &lt;a href="https://umesh-malik.com/blog/core-web-vitals-optimization-guide" rel="noopener noreferrer"&gt;Core Web Vitals optimization guide&lt;/a&gt; covers the metrics this technique moves.)&lt;/p&gt;

&lt;p&gt;The old workaround was to make everything fast enough to send in order, or to punt the slow parts to client-side &lt;code&gt;fetch()&lt;/code&gt; after load — which trades a blocked render for a JavaScript bill, a loading spinner, and a layout shift. Out-of-order streaming gives you the best of both: &lt;strong&gt;one HTTP response, first paint at the speed of your fastest content, and no client runtime.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This lands at a moment when the whole industry is rediscovering server-driven HTML — HTMX, Astro, Qwik, React Server Components, and the &lt;a href="https://umesh-malik.com/blog/cloudflare-vinext-next-js-vite-revolution" rel="noopener noreferrer"&gt;Cloudflare viNext&lt;/a&gt; rebuild of Next.js all lean on streaming. A &lt;em&gt;native, framework-free&lt;/em&gt; primitive for out-of-order delivery is a big deal: it's the plumbing every one of those tools has been faking in userland.&lt;/p&gt;

&lt;h2&gt;
  
  
  The no-JavaScript technique that started it — Declarative Shadow DOM and slots
&lt;/h2&gt;

&lt;p&gt;Before any new proposal existed, developers found a way to stream out of order with &lt;strong&gt;zero JavaScript&lt;/strong&gt; using two features that already shipped: &lt;strong&gt;Declarative Shadow DOM (DSD)&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;&amp;lt;slot&amp;gt;&lt;/code&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here's the mechanism. A &lt;code&gt;&amp;lt;slot&amp;gt;&lt;/code&gt; inside a shadow root is a &lt;em&gt;placeholder&lt;/em&gt;. The browser fills it with any light-DOM child of the host element carrying a matching &lt;code&gt;slot="…"&lt;/code&gt; attribute — and crucially, it does this &lt;strong&gt;whenever that child appears in the stream&lt;/strong&gt;, even if that's much later than the slot itself.&lt;/p&gt;

&lt;p&gt;So you stream a shell first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;article&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;template&lt;/span&gt; &lt;span class="na"&gt;shadowrootmode=&lt;/span&gt;&lt;span class="s"&gt;"open"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;h1&amp;gt;&lt;/span&gt;My page&lt;span class="nt"&gt;&amp;lt;/h1&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;slot&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"reviews"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
      &lt;span class="nt"&gt;&amp;lt;p&amp;gt;&lt;/span&gt;Loading reviews…&lt;span class="nt"&gt;&amp;lt;/p&amp;gt;&lt;/span&gt;   &lt;span class="c"&gt;&amp;lt;!-- fallback shown immediately --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/slot&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;footer&amp;gt;&lt;/span&gt;Site footer&lt;span class="nt"&gt;&amp;lt;/footer&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/template&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/article&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The browser paints the heading, the "Loading reviews…" fallback, and the footer right away. Then — later in the &lt;em&gt;same response&lt;/em&gt;, after your slow query resolves — you stream the real content as a light-DOM child targeting that slot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;article&amp;gt;&lt;/span&gt;
  &lt;span class="c"&gt;&amp;lt;!-- ...the template above already streamed... --&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;div&lt;/span&gt; &lt;span class="na"&gt;slot=&lt;/span&gt;&lt;span class="s"&gt;"reviews"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;p&amp;gt;&lt;/span&gt;★★★★★ Actually great.&lt;span class="nt"&gt;&amp;lt;/p&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;p&amp;gt;&lt;/span&gt;★★★★☆ Pretty good.&lt;span class="nt"&gt;&amp;lt;/p&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/div&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/article&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The moment that &lt;code&gt;&amp;lt;div slot="reviews"&amp;gt;&lt;/code&gt; is parsed, the browser removes the fallback and slots the reviews into place. &lt;strong&gt;No script ran. No event fired. The HTML parser did it.&lt;/strong&gt; That's out-of-order streaming, and it's been &lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTML/Element/template#shadowrootmode" rel="noopener noreferrer"&gt;Baseline across Chrome, Firefox, and Safari since early 2024&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why DSD needed a rename to make this work&lt;/strong&gt;&lt;br&gt;
Declarative Shadow DOM first shipped in Chrome 90 (2021) using a &lt;code&gt;shadowroot&lt;/code&gt; attribute that attached the shadow root at the template's &lt;em&gt;closing&lt;/em&gt; tag — which defeats streaming, since you'd have to buffer the whole template. The fix, driven by Google's Mason Freed in a &lt;a href="https://groups.google.com/a/chromium.org/g/blink-dev/c/Ovz-6Dte-qA" rel="noopener noreferrer"&gt;Chromium "Intent to Prototype"&lt;/a&gt;, renamed it to &lt;code&gt;shadowrootmode&lt;/code&gt; and attached the root at the &lt;em&gt;opening&lt;/em&gt; tag. That one change is what unlocked streaming and got Firefox and WebKit on board.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The catch: &lt;strong&gt;this only works inside Shadow DOM.&lt;/strong&gt; Every out-of-order region needs a custom-element host and a shadow root, you inherit shadow-DOM style scoping whether you want it or not, and it's awkward for the common case of "just patch this one &lt;code&gt;&amp;lt;div&amp;gt;&lt;/code&gt; in my normal document." That limitation is exactly what the next proposal set out to remove.&lt;/p&gt;

&lt;h2&gt;
  
  
  Declarative Partial Updates — the native proposal
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Declarative Partial Updates (DPU)&lt;/strong&gt; is the web-platform proposal that turns out-of-order streaming into a first-class HTML feature with no Shadow DOM requirement. It was introduced in a &lt;a href="https://developer.chrome.com/blog/declarative-partial-updates" rel="noopener noreferrer"&gt;Chrome for Developers article&lt;/a&gt; on May 19, 2026 by &lt;strong&gt;Barry Pollard&lt;/strong&gt; and &lt;strong&gt;Noam Rosenthal&lt;/strong&gt;, with a full &lt;a href="https://github.com/WICG/declarative-partial-updates/blob/main/patching-explainer.md" rel="noopener noreferrer"&gt;WICG explainer&lt;/a&gt; and a &lt;a href="https://github.com/whatwg/html/pull/11818" rel="noopener noreferrer"&gt;WHATWG HTML pull request (#11818)&lt;/a&gt; for the &lt;code&gt;&amp;lt;template for&amp;gt;&lt;/code&gt; element authored by &lt;strong&gt;Philip Jägenstedt&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The explainer states the problem plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Streaming of HTML existed from the early days of the web… However, it always had the following major constraints: 1. HTML content is streamed in DOM order. 2. After the initial parsing of the document, streaming is no longer enabled."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;DPU attacks both constraints. Two mechanisms, two audiences.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mechanism 1 — declarative out-of-order streaming (no JavaScript)
&lt;/h3&gt;

&lt;p&gt;You mark a spot in the document with a &lt;strong&gt;processing instruction&lt;/strong&gt;, then fill it later with a &lt;code&gt;&amp;lt;template for&amp;gt;&lt;/code&gt; whose &lt;code&gt;for&lt;/code&gt; attribute matches the marker's name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Single-point insertion&lt;/strong&gt; — drop content at a marker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;div&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;&lt;/span&gt;&lt;span class="err"&gt;?&lt;/span&gt;&lt;span class="na"&gt;marker&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"placeholder"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/div&amp;gt;&lt;/span&gt;

&lt;span class="c"&gt;&amp;lt;!-- ...anything else streams here... --&amp;gt;&lt;/span&gt;

&lt;span class="nt"&gt;&amp;lt;template&lt;/span&gt; &lt;span class="na"&gt;for=&lt;/span&gt;&lt;span class="s"&gt;"placeholder"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  Here is some &lt;span class="nt"&gt;&amp;lt;em&amp;gt;&lt;/span&gt;HTML content&lt;span class="nt"&gt;&amp;lt;/em&amp;gt;&lt;/span&gt;!
&lt;span class="nt"&gt;&amp;lt;/template&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Temporary placeholder&lt;/strong&gt; — show fallback content until the real thing arrives, using a &lt;code&gt;&amp;lt;?start&amp;gt;&lt;/code&gt; / &lt;code&gt;&amp;lt;?end&amp;gt;&lt;/code&gt; pair to define the range to replace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;div&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;&lt;/span&gt;&lt;span class="err"&gt;?&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"reviews"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  Loading…
  &lt;span class="nt"&gt;&amp;lt;&lt;/span&gt;&lt;span class="err"&gt;?&lt;/span&gt;&lt;span class="na"&gt;end&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/div&amp;gt;&lt;/span&gt;

&lt;span class="c"&gt;&amp;lt;!-- ...footer, other sections, all stream immediately... --&amp;gt;&lt;/span&gt;

&lt;span class="nt"&gt;&amp;lt;template&lt;/span&gt; &lt;span class="na"&gt;for=&lt;/span&gt;&lt;span class="s"&gt;"reviews"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  Here is some &lt;span class="nt"&gt;&amp;lt;em&amp;gt;&lt;/span&gt;HTML content&lt;span class="nt"&gt;&amp;lt;/em&amp;gt;&lt;/span&gt;!
&lt;span class="nt"&gt;&amp;lt;/template&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the &lt;code&gt;&amp;lt;template for="reviews"&amp;gt;&lt;/code&gt; streams in, the browser replaces everything between &lt;code&gt;&amp;lt;?start name="reviews"&amp;gt;&lt;/code&gt; and &lt;code&gt;&amp;lt;?end&amp;gt;&lt;/code&gt; — the "Loading…" text — with the template's content. Notice there's &lt;strong&gt;no custom element and no shadow root&lt;/strong&gt;. This patches a plain &lt;code&gt;&amp;lt;div&amp;gt;&lt;/code&gt; in the normal document.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multiple / chunked updates&lt;/strong&gt; — a template can contain its &lt;em&gt;own&lt;/em&gt; nested marker, so you can stream a list one item at a time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;ul&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"results"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;&lt;/span&gt;&lt;span class="err"&gt;?&lt;/span&gt;&lt;span class="na"&gt;start&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"results"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  Loading…
  &lt;span class="nt"&gt;&amp;lt;&lt;/span&gt;&lt;span class="err"&gt;?&lt;/span&gt;&lt;span class="na"&gt;end&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/ul&amp;gt;&lt;/span&gt;

&lt;span class="nt"&gt;&amp;lt;template&lt;/span&gt; &lt;span class="na"&gt;for=&lt;/span&gt;&lt;span class="s"&gt;"results"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;li&amp;gt;&lt;/span&gt;Result One&lt;span class="nt"&gt;&amp;lt;/li&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;&lt;/span&gt;&lt;span class="err"&gt;?&lt;/span&gt;&lt;span class="na"&gt;marker&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"results"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/template&amp;gt;&lt;/span&gt;

&lt;span class="nt"&gt;&amp;lt;template&lt;/span&gt; &lt;span class="na"&gt;for=&lt;/span&gt;&lt;span class="s"&gt;"results"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;li&amp;gt;&lt;/span&gt;Result Two&lt;span class="nt"&gt;&amp;lt;/li&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;&lt;/span&gt;&lt;span class="err"&gt;?&lt;/span&gt;&lt;span class="na"&gt;marker&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"results"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/template&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each &lt;code&gt;&amp;lt;template for="results"&amp;gt;&lt;/code&gt; appends its &lt;code&gt;&amp;lt;li&amp;gt;&lt;/code&gt; and leaves a fresh &lt;code&gt;&amp;lt;?marker name="results"&amp;gt;&lt;/code&gt; for the next chunk. This is how you stream search results or a chat transcript row-by-row, declaratively.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why processing instructions instead of an attribute?&lt;/strong&gt;&lt;br&gt;
An earlier design used a &lt;code&gt;contentmethod&lt;/code&gt; attribute, but the WICG explainer notes it "doesn't support replacing arbitrary ranges of nodes, only an element or all of its children." Processing instructions let the server define an &lt;em&gt;arbitrary range&lt;/em&gt; — a start/end marker pair — to swap out, which matters for frameworks whose patch boundaries aren't known ahead of time.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Mechanism 2 — streaming updates after load (this half needs JavaScript)
&lt;/h3&gt;

&lt;p&gt;The second half of DPU modernizes DOM insertion for content that arrives &lt;em&gt;after&lt;/em&gt; the initial page — think live dashboards or infinite scroll. It replaces the messy &lt;code&gt;innerHTML&lt;/code&gt; / &lt;code&gt;insertAdjacentHTML&lt;/code&gt; landscape with a consistent, sanitizer-aware, Streams-integrated set of methods:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;el&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;querySelector&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;#content-to-update&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// Static: parse a string of HTML into the element&lt;/span&gt;
&lt;span class="nx"&gt;el&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setHTML&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;html&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;el&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;appendHTML&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;html&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// Streaming: pipe an HTTP response straight into the DOM&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/content.html&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pipeThrough&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextDecoderStream&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pipeTo&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;el&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;streamHTMLUnsafe&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are &lt;code&gt;Unsafe&lt;/code&gt; variants (&lt;code&gt;streamHTMLUnsafe&lt;/code&gt;, &lt;code&gt;streamAppendHTMLUnsafe&lt;/code&gt;) that skip sanitization, mirroring the &lt;code&gt;setHTMLUnsafe&lt;/code&gt; naming already standardized with the HTML Sanitizer API. This half obviously uses JavaScript — but the &lt;em&gt;first&lt;/em&gt; half, the out-of-order document streaming, does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Old slot technique vs Declarative Partial Updates
&lt;/h2&gt;

&lt;p&gt;Both give you no-JavaScript out-of-order streaming. Here's how to choose.&lt;/p&gt;

&lt;p&gt;…&amp;lt;?end&amp;gt;', tone: 'positive' }] },&lt;br&gt;
    { label: 'Post-load streaming updates', cells: [{ text: 'Not really its job', tone: 'neutral' }, { text: 'Built in — streamHTML / streamAppendHTML', tone: 'positive' }] },&lt;br&gt;
    { label: 'Production-ready in mid-2026', cells: [{ text: 'Yes', tone: 'positive' }, { text: 'No — explicitly for developer testing', tone: 'negative' }] }&lt;br&gt;
  ]}&lt;br&gt;
/&amp;gt;&lt;/p&gt;

&lt;p&gt;The honest recommendation: &lt;strong&gt;if you need out-of-order streaming in production today, use Declarative Shadow DOM with slots.&lt;/strong&gt; It's cross-browser, it's stable, and the Shadow DOM constraints are livable for coarse-grained regions. &lt;strong&gt;Reach for Declarative Partial Updates when you're prototyping the future&lt;/strong&gt; or building a framework's rendering layer that you can gate behind feature detection and a polyfill.&lt;/p&gt;
&lt;h2&gt;
  
  
  How to try Declarative Partial Updates today
&lt;/h2&gt;

&lt;p&gt;.", status: "done" },&lt;br&gt;
    { date: "Mid-2026", title: "Chrome 148 developer testing", description: "DPU available behind chrome://flags/#enable-experimental-web-platform-features. Firefox and WebKit interested; not yet shipped.", status: "active" },&lt;br&gt;
    { date: "Future", title: "Cross-browser standardization", description: "Pending WHATWG HTML + DOM spec merges and multi-implementer support before it's safe for production.", status: "upcoming" }&lt;br&gt;
  ]}&lt;br&gt;
/&amp;gt;&lt;/p&gt;

&lt;p&gt;To experiment right now:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Enable the flag.&lt;/strong&gt; Open &lt;code&gt;chrome://flags/#enable-experimental-web-platform-features&lt;/code&gt; in Chrome 148+ and turn it on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stream from a real server.&lt;/strong&gt; The effect only shows if bytes actually arrive over time — flush the shell, do your slow work, then flush the &lt;code&gt;&amp;lt;template for&amp;gt;&lt;/code&gt;. A buffered response that sends everything at once will look correct but won't demonstrate the streaming win.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use a polyfill for reach.&lt;/strong&gt; The proposal ships with polyfills so you can adopt the API shape today and let it fall through to native behavior as browsers catch up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feature-detect, don't assume.&lt;/strong&gt; Gate on the presence of the API (e.g. checking for &lt;code&gt;Element.prototype.streamHTML&lt;/code&gt;) before relying on it.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not ship this to users yet&lt;/strong&gt;&lt;br&gt;
Both the Chrome team and the WICG explainer are explicit: Declarative Partial Updates is for developer testing and feedback. The syntax — the marker and start/end processing instructions, plus the template-&lt;code&gt;for&lt;/code&gt; element — can still change before it's standardized, and only one browser exposes it. Prototype freely; don't put it on your production critical path.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  Common mistakes and pitfalls
&lt;/h2&gt;

&lt;p&gt;Most people who try out-of-order streaming trip on the same few things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Buffering the response and expecting magic.&lt;/strong&gt; If your framework or reverse proxy buffers the full HTML before sending, there is no "out of order" — everything arrives at once. You must &lt;em&gt;flush&lt;/em&gt; the shell early. Check your server and any nginx/CDN buffering in front of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assuming Suspense-style streaming is JavaScript-free.&lt;/strong&gt; React's out-of-order streaming is not the same class of thing — it ships a client runtime. If your goal is "works with JS disabled," you need DSD/slots or DPU, not a framework's streaming SSR.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fighting Shadow DOM styles.&lt;/strong&gt; With the slot technique, your global CSS doesn't pierce the shadow root. Plan for &lt;code&gt;::part()&lt;/code&gt;, inherited custom properties, or accept scoped styles — don't discover this after your design looks unstyled.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forgetting the fallback content.&lt;/strong&gt; The whole point of &lt;code&gt;&amp;lt;slot&amp;gt;&lt;/code&gt; children or the &lt;code&gt;&amp;lt;?start&amp;gt;…&amp;lt;?end&amp;gt;&lt;/code&gt; range is what the user sees &lt;em&gt;before&lt;/em&gt; the slow content arrives. Ship a real skeleton or "Loading…" state, not an empty hole that causes layout shift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reaching for DPU in production.&lt;/strong&gt; It's one browser, behind a flag, with a syntax that may change. Treat it as R&amp;amp;D.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Best practices
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Default to Declarative Shadow DOM + slots for production out-of-order streaming.&lt;/strong&gt; It's the only cross-browser, no-JavaScript option that's actually stable in 2026.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flush aggressively.&lt;/strong&gt; Send the shell, headers, and any above-the-fold static content in the first chunk. The perceived-performance win is entirely about getting &lt;em&gt;something&lt;/em&gt; painted fast.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Order your slow work by user value, not document order.&lt;/strong&gt; Stream the widget the user actually looks at first, even if it's lower on the page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep JavaScript optional, not required.&lt;/strong&gt; Use the declarative techniques for the initial render; layer the JS &lt;code&gt;streamHTML&lt;/code&gt; methods on top only for genuinely post-load updates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Follow the standard, not one browser.&lt;/strong&gt; Track the &lt;a href="https://github.com/whatwg/html/pull/11818" rel="noopener noreferrer"&gt;WHATWG PR&lt;/a&gt; and &lt;a href="https://github.com/WICG/declarative-partial-updates" rel="noopener noreferrer"&gt;WICG explainer&lt;/a&gt; for syntax changes, and gate any DPU usage behind feature detection + polyfill.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  A concrete example: a product page with a slow reviews panel
&lt;/h2&gt;

&lt;p&gt;Picture a product page. The product info comes from cache in 20ms. The reviews come from a service that takes 700ms. Personalized recommendations take 1.2s.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In-order:&lt;/strong&gt; first byte of &lt;em&gt;anything useful&lt;/em&gt; waits behind the slowest thing above it. Best case, you reorder your markup so slow things sit at the bottom — but then your layout is dictated by latency, not design.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Out of order:&lt;/strong&gt; you flush the full page shell — header, product info, empty reviews slot, empty recommendations slot, footer — in ~30ms. LCP fires on the product image almost immediately. At 700ms the reviews &lt;code&gt;&amp;lt;template for="reviews"&amp;gt;&lt;/code&gt; streams in and fills its slot. At 1.2s the recommendations do the same. One request, no client fetch, no spinner-driven layout shift, and the page is fully interactive-looking the entire time.&lt;/p&gt;

&lt;p&gt;That's the shape of nearly every modern page: fast core content plus a few slow, independent widgets. Out-of-order streaming is purpose-built for it — and it pairs naturally with server-first stacks like the ones I covered in the &lt;a href="https://umesh-malik.com/blog/cloudflare-vinext-next-js-vite-revolution" rel="noopener noreferrer"&gt;Cloudflare viNext teardown&lt;/a&gt; and &lt;a href="https://umesh-malik.com/blog/fastapi-spa-app-frontend-explained" rel="noopener noreferrer"&gt;FastAPI's native SPA support&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;/ &amp;lt;?start&amp;gt;/&amp;lt;?end&amp;gt; and .",&lt;br&gt;
      tag: 'Basics'&lt;br&gt;
    },&lt;br&gt;
    {&lt;br&gt;
      question: 'What is  in HTML?',&lt;br&gt;
      answer: "It's the element at the center of the Declarative Partial Updates proposal (WHATWG HTML PR #11818, by Philip Jägenstedt). Its for attribute matches a processing-instruction marker earlier in the stream; when the template is parsed, the browser moves its content into the marked location — enabling out-of-order streaming without JavaScript.",&lt;br&gt;
      tag: 'Syntax'&lt;br&gt;
    },&lt;br&gt;
    {&lt;br&gt;
      question: 'Which browsers support Declarative Partial Updates?',&lt;br&gt;
      answer: "As of mid-2026, only Chrome 148 exposes it, behind chrome://flags/#enable-experimental-web-platform-features. Firefox and WebKit have signaled interest but have not shipped. Declarative Shadow DOM streaming, by contrast, is Baseline across all three since early 2024.",&lt;br&gt;
      tag: 'Support'&lt;br&gt;
    },&lt;br&gt;
    {&lt;br&gt;
      question: 'Is this a replacement for React Suspense or streaming SSR?',&lt;br&gt;
      answer: "For the reordering itself, arguably yes — it does natively what frameworks do with a client runtime. But React Suspense also coordinates hydration and interactivity, which these HTML primitives don't. Expect frameworks to build on top of DPU rather than be replaced by it.",&lt;br&gt;
      tag: 'Comparison'&lt;br&gt;
    },&lt;br&gt;
    {&lt;br&gt;
      question: 'Why does my out-of-order stream render all at once instead of progressively?',&lt;br&gt;
      answer: "Almost always response buffering. If your app server, framework, or a reverse proxy/CDN buffers the full response before sending, the browser receives everything at once. You must flush the shell early and ensure nothing downstream re-buffers the stream.",&lt;br&gt;
      tag: 'Troubleshooting'&lt;br&gt;
    },&lt;br&gt;
    {&lt;br&gt;
      question: 'Should I use this in production?',&lt;br&gt;
      answer: "Use Declarative Shadow DOM + slots in production — it's stable and cross-browser. Keep Declarative Partial Updates to prototypes and framework R&amp;amp;D until it standardizes and ships in more than one browser.",&lt;br&gt;
      tag: 'Adoption'&lt;br&gt;
    }&lt;br&gt;
  ]}&lt;br&gt;
/&amp;gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;HTML's top-to-bottom, in-order delivery was a 30-year-old assumption, not a law of physics. Breaking it — sending pages in the order data is &lt;em&gt;ready&lt;/em&gt; rather than the order it's &lt;em&gt;written&lt;/em&gt; — is one of the most consequential quiet upgrades the web platform has shipped in a decade, and the best part is that its core doesn't cost a single kilobyte of JavaScript.&lt;/p&gt;

&lt;p&gt;The pragmatic move today: &lt;strong&gt;use Declarative Shadow DOM and slots for real out-of-order streaming now, and start prototyping with Declarative Partial Updates so you're fluent when it lands.&lt;/strong&gt; Flush your shells, kill your buffering, and stop letting your slowest widget dictate when the whole page appears.&lt;/p&gt;

&lt;p&gt;If this was useful, read the &lt;a href="https://umesh-malik.com/blog/core-web-vitals-optimization-guide" rel="noopener noreferrer"&gt;Core Web Vitals optimization guide&lt;/a&gt; next — it covers the exact metrics that out-of-order streaming is designed to move.&lt;/p&gt;
&lt;h2&gt;
  
  
  Sources and further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://developer.chrome.com/blog/declarative-partial-updates" rel="noopener noreferrer"&gt;Declarative partial updates&lt;/a&gt; — Chrome for Developers (Barry Pollard, Noam Rosenthal), May 19, 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/WICG/declarative-partial-updates/blob/main/patching-explainer.md" rel="noopener noreferrer"&gt;WICG Declarative Partial Updates explainer&lt;/a&gt; — problem statement, goals, and full syntax&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/whatwg/html/pull/11818" rel="noopener noreferrer"&gt;WHATWG HTML PR #11818 — &lt;code&gt;&amp;lt;template for&amp;gt;&lt;/code&gt; for out-of-order streaming&lt;/a&gt; — Philip Jägenstedt&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://groups.google.com/a/chromium.org/g/blink-dev/c/Ovz-6Dte-qA" rel="noopener noreferrer"&gt;Chromium Intent to Prototype: streaming declarative shadow DOM&lt;/a&gt; — the &lt;code&gt;shadowrootmode&lt;/code&gt; rename that unlocked streaming&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTML/Element/template#shadowrootmode" rel="noopener noreferrer"&gt;MDN — &lt;code&gt;&amp;lt;template&amp;gt;&lt;/code&gt; &lt;code&gt;shadowrootmode&lt;/code&gt;&lt;/a&gt; — Declarative Shadow DOM reference and Baseline status&lt;/li&gt;
&lt;/ul&gt;



&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering.&lt;/em&gt;&lt;/p&gt;



&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/streaming-html-out-of-order-without-javascript" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/react-server-components-guide" rel="noopener noreferrer"&gt;React Server Components in 2026: The Mental Model, the use client Boundary &amp;amp; When Not to Use Them&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/fastapi-spa-app-frontend-explained" rel="noopener noreferrer"&gt;FastAPI Finally Has Native SPA Support: app.frontend() Explained&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/cloudflare-vinext-next-js-vite-revolution" rel="noopener noreferrer"&gt;Cloudflare viNext: The $1,100 Next.js-on-Vite Rebuild&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;


</description>
      <category>html</category>
      <category>webplatform</category>
      <category>performance</category>
      <category>streaming</category>
    </item>
    <item>
      <title>OpenAI GPT-5.6 Complete Guide: Sol, Terra, Luna Benchmarks, Pricing, and API (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Sat, 11 Jul 2026 15:07:56 +0000</pubDate>
      <link>https://dev.to/umesh_malik/openai-gpt-56-complete-guide-sol-terra-luna-benchmarks-pricing-and-api-2026-5aj7</link>
      <guid>https://dev.to/umesh_malik/openai-gpt-56-complete-guide-sol-terra-luna-benchmarks-pricing-and-api-2026-5aj7</guid>
      <description>&lt;p&gt;&lt;strong&gt;GPT-5.6 is OpenAI's July 2026 model family — three tiers (Sol, Terra, and Luna) that share a 1.05M-token context window and range from $1 to $30 per million tokens.&lt;/strong&gt; Sol is the coding-and-reasoning flagship, Terra is the balanced everyday model, and Luna is the fastest and cheapest.&lt;/p&gt;

&lt;p&gt;OpenAI released &lt;strong&gt;GPT-5.6 on July 9, 2026&lt;/strong&gt;, and for the first time in the GPT-5 line the headline is not a single model. It is &lt;strong&gt;three&lt;/strong&gt;: &lt;strong&gt;Sol, Terra, and Luna&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That naming change is the whole story. Instead of shipping one frontier model and a &lt;code&gt;-pro&lt;/code&gt; step-up, OpenAI split the release into a &lt;strong&gt;capability-and-cost ladder&lt;/strong&gt; where each rung is tuned for a different budget and workload. Sol is the flagship. Terra is the everyday workhorse. Luna is the speed-and-price play. And crucially, they all share the same &lt;strong&gt;1.05M-token context window&lt;/strong&gt;, the same &lt;strong&gt;February 16, 2026 knowledge cutoff&lt;/strong&gt;, and the same platform features — so you can move up and down the ladder without rewriting your integration.&lt;/p&gt;

&lt;p&gt;The short answer: &lt;strong&gt;GPT-5.6 is OpenAI's most efficiency-focused release yet.&lt;/strong&gt; Sam Altman is on record calling the family "orders of magnitude more efficient and cost-effective than previous versions," and the coding numbers back the claim — Sol reportedly finishes agentic coding tasks with &lt;strong&gt;54% better token efficiency&lt;/strong&gt; than the previous generation. If you run models at scale, this release is less about a new capability ceiling and more about &lt;strong&gt;doing the same work for a fraction of the tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPT-5.6 launched July 9, 2026&lt;/strong&gt; as a three-model family: &lt;strong&gt;Sol&lt;/strong&gt; (flagship), &lt;strong&gt;Terra&lt;/strong&gt; (balanced), and &lt;strong&gt;Luna&lt;/strong&gt; (fast and cheap).&lt;/li&gt;
&lt;li&gt;All three share a &lt;strong&gt;1.05M-token context window&lt;/strong&gt;, &lt;strong&gt;128K max output&lt;/strong&gt;, and a &lt;strong&gt;February 16, 2026&lt;/strong&gt; knowledge cutoff.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sol is the best coding model in the family&lt;/strong&gt;, scoring &lt;strong&gt;80&lt;/strong&gt; on the Artificial Analysis Coding Agent Index — about &lt;strong&gt;2.8 points above&lt;/strong&gt; the previous frontier competitor — and &lt;strong&gt;91.9% on Terminal-Bench 2.1&lt;/strong&gt; with ultra thinking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sol is the first model to clear the halfway mark on Agent's Last Exam&lt;/strong&gt;, at roughly &lt;strong&gt;50.9%&lt;/strong&gt; in code mode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terra lands just above the previous frontier tier&lt;/strong&gt; on coding, and &lt;strong&gt;Luna outperforms the last generation's flagship&lt;/strong&gt; while being the cheapest option.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pricing per 1M tokens&lt;/strong&gt;: Sol &lt;code&gt;$5 / $30&lt;/code&gt;, Terra &lt;code&gt;$2.50 / $15&lt;/code&gt;, Luna &lt;code&gt;$1 / $6&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;New &lt;strong&gt;ultra thinking&lt;/strong&gt; and &lt;strong&gt;max&lt;/strong&gt; reasoning modes push the ceiling on hard, long-horizon tasks.&lt;/li&gt;
&lt;li&gt;OpenAI calls GPT-5.6 its &lt;strong&gt;strongest cybersecurity model yet&lt;/strong&gt;, tuned for defensive work like threat modeling, code review, and blue teaming.&lt;/li&gt;
&lt;li&gt;API names: &lt;strong&gt;&lt;code&gt;gpt-5.6-sol&lt;/code&gt;&lt;/strong&gt;, &lt;strong&gt;&lt;code&gt;gpt-5.6-terra&lt;/code&gt;&lt;/strong&gt;, &lt;strong&gt;&lt;code&gt;gpt-5.6-luna&lt;/code&gt;&lt;/strong&gt;; the alias &lt;strong&gt;&lt;code&gt;gpt-5.6&lt;/code&gt;&lt;/strong&gt; routes to Sol. Available in &lt;strong&gt;ChatGPT, Codex, the API, and GitHub Copilot&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6nr6ysyosd4d6gy9zfyr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6nr6ysyosd4d6gy9zfyr.png" alt="GPT-5.6 model family showing Sol, Terra, and Luna tiers with their coding scores, pricing, and shared 1.05M context window" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What GPT-5.6 Actually Is
&lt;/h2&gt;

&lt;p&gt;GPT-5.6 is not "GPT-5.5 but smarter." It is a &lt;strong&gt;repackaging of the frontier into three price points&lt;/strong&gt;, and that is a more interesting decision than another benchmark bump.&lt;/p&gt;

&lt;p&gt;Here is the mental model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sol&lt;/strong&gt; — the flagship. Built for complex work across &lt;strong&gt;coding, knowledge work, research, cybersecurity, science, computer use, and design&lt;/strong&gt;. This is the one you reach for when the task is genuinely hard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terra&lt;/strong&gt; — the middle tier. A deliberate &lt;strong&gt;balance of capability, speed, and cost&lt;/strong&gt; for everyday production work. It is the model most teams will actually run by default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Luna&lt;/strong&gt; — the floor. The &lt;strong&gt;fastest and lowest-cost&lt;/strong&gt; member of the family, aimed at high-volume, latency-sensitive, or cost-capped workloads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important part is what they have in common. Every tier gets the &lt;strong&gt;same 1.05M context window&lt;/strong&gt;, the &lt;strong&gt;same 128K max output&lt;/strong&gt;, and the &lt;strong&gt;same knowledge cutoff&lt;/strong&gt;. There is no "the cheap model also has a smaller brain for context" catch. You are trading raw reasoning depth for price and speed — not memory.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Naming note&lt;/strong&gt;&lt;br&gt;
The GPT-5.6 family uses celestial codenames — Sol (sun), Terra (earth), Luna (moon) — instead of the old &lt;code&gt;-mini&lt;/code&gt; / &lt;code&gt;-pro&lt;/code&gt; suffixes. It is a marketing rebrand, but it maps cleanly onto a real axis: Sol is the brightest and most expensive, Luna is the smallest and cheapest, Terra sits in between.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  1. The Real Headline Is Efficiency, Not a New Ceiling
&lt;/h2&gt;

&lt;p&gt;Most model launches lead with "we beat the benchmark." GPT-5.6 leads with &lt;strong&gt;"we beat it for less."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is a genuine shift. The standout claim is that &lt;strong&gt;Sol is 54% more token-efficient on AI coding tasks&lt;/strong&gt; than the previous generation. On head-to-head coding runs, OpenAI says Sol uses &lt;strong&gt;less than half the output tokens&lt;/strong&gt; of a comparable frontier model, finishes in &lt;strong&gt;less than half the time&lt;/strong&gt;, and costs about &lt;strong&gt;a third less&lt;/strong&gt; to complete the same task.&lt;/p&gt;

&lt;p&gt;For anyone paying a real API bill, that math matters more than a two-point benchmark win.&lt;/p&gt;

&lt;p&gt;Why does this land now? Because the bottleneck for most production LLM systems in 2026 is &lt;strong&gt;not&lt;/strong&gt; "the model cannot do it." It is "the model does it, but the token bill and latency make it uneconomical at scale." A model that produces the same answer with half the tokens changes which use cases are actually viable.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: GPT-5.6's most important number is not a benchmark score — it is the token count it takes to &lt;em&gt;reach&lt;/em&gt; that score. Efficiency is the feature.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Coding and Agents: Where Sol Actually Wins
&lt;/h2&gt;

&lt;p&gt;Sol is positioned as &lt;strong&gt;the best coding model in the family&lt;/strong&gt;, and the public numbers are strong.&lt;/p&gt;

&lt;p&gt;The Terminal-Bench and Agent's Last Exam numbers are the ones agent builders should care about. They measure &lt;strong&gt;long-horizon, multi-step task completion&lt;/strong&gt; — the model has to plan, run commands, read output, recover from errors, and keep going without a human babysitting each step. Clearing 50% on Agent's Last Exam is a milestone; most models still stall well before the halfway point.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsmcxcb83s5gptlhqbfd2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsmcxcb83s5gptlhqbfd2.png" alt="GPT-5.6 Sol coding and agent benchmark scores including Terminal-Bench 2.1 and Agent's Last Exam" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For a concrete sense of how Codex-class coding behaves inside a real agent loop, I walked through &lt;a href="https://umesh-malik.com/blog/figma-codex-react-2026" rel="noopener noreferrer"&gt;building frontend UIs with Codex and Figma&lt;/a&gt; — the same plan-run-inspect-iterate pattern these benchmarks are trying to measure.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where Terra and Luna land on coding&lt;/strong&gt;&lt;br&gt;
Terra performs &lt;strong&gt;just above&lt;/strong&gt; the previous frontier competitor on the coding index, and Luna &lt;strong&gt;outperforms the previous generation's flagship&lt;/strong&gt; while being the cheapest option in the lineup. In other words: even the budget tier of GPT-5.6 is roughly a frontier model from one generation ago.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Thinking Modes: Ultra and Max
&lt;/h2&gt;

&lt;p&gt;GPT-5.6 introduces higher-effort reasoning modes that let you dial compute up for the hardest work.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ultra thinking&lt;/strong&gt; — the top reasoning setting, used to post Sol's record &lt;strong&gt;91.9% on Terminal-Bench 2.1&lt;/strong&gt;. Reserve it for genuinely hard, long-horizon problems where extra deliberation pays for itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max mode&lt;/strong&gt; — a strong high-effort tier that still hit &lt;strong&gt;88.76%&lt;/strong&gt; on the same benchmark without going all the way to ultra.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical rule is the same as it has always been with reasoning models: &lt;strong&gt;effort is a cost dial, not a free upgrade.&lt;/strong&gt; Higher thinking modes consume more (billed) reasoning tokens and add latency. Use them where the task genuinely needs multi-step planning — agentic coding, deep research, complex analysis — and drop back to lower effort for extraction, formatting, and simple transforms.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Reasoning tokens still cost real money&lt;/strong&gt;&lt;br&gt;
As with every recent OpenAI reasoning model, the tokens spent "thinking" in ultra and max modes are billed as output and consume your context budget even though you never see them. Leave headroom, and measure incomplete responses before you cap &lt;code&gt;max_output_tokens&lt;/code&gt; aggressively.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Cybersecurity: OpenAI's "Strongest Yet"
&lt;/h2&gt;

&lt;p&gt;OpenAI describes GPT-5.6 as its &lt;strong&gt;strongest cybersecurity model to date&lt;/strong&gt;, explicitly tuned to help with &lt;strong&gt;defensive&lt;/strong&gt; security work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;threat modeling and architecture review&lt;/li&gt;
&lt;li&gt;security-focused code review and vulnerability triage&lt;/li&gt;
&lt;li&gt;patch generation and remediation guidance&lt;/li&gt;
&lt;li&gt;blue-team workflows and detection engineering&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a meaningful positioning choice. A model that is good at finding and fixing vulnerabilities is, by definition, also more capable in the offensive direction — so OpenAI pairs the capability with monitoring and access controls, and frames the sanctioned use cases around defense. If your security team has been waiting for a model strong enough to sit inside real review pipelines, Sol is the one to evaluate.&lt;/p&gt;

&lt;p&gt;If you are thinking about agents with this much capability touching production systems, the governance questions in &lt;a href="https://umesh-malik.com/blog/agentic-ai-enterprise-security-model" rel="noopener noreferrer"&gt;the agentic AI enterprise security model&lt;/a&gt; apply directly here.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The 1.05M Context Window Is Shared — but Read the Fine Print
&lt;/h2&gt;

&lt;p&gt;Every GPT-5.6 tier ships with a &lt;strong&gt;1,050,000-token context window&lt;/strong&gt; and &lt;strong&gt;128,000 max output tokens&lt;/strong&gt;. That is a real capability, and the fact that even Luna gets the full window is genuinely useful.&lt;/p&gt;

&lt;p&gt;But the same caveat from every large-context model still holds: &lt;strong&gt;a big window is not perfect recall.&lt;/strong&gt; Long-context retrieval quality degrades at the far edge of the window, and giant prompts carry hidden cost and latency. Treat 1M context as a tool for &lt;strong&gt;broad synthesis and large working memory&lt;/strong&gt;, not as a replacement for retrieval discipline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sol vs Terra vs Luna: The Decision
&lt;/h2&gt;

&lt;p&gt;If you only remember one section, make it this one.&lt;/p&gt;

&lt;p&gt;The simplest rule:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default to Terra.&lt;/strong&gt; It is the balanced everyday model and will be the right call for most production traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalate to Sol&lt;/strong&gt; when the task is genuinely hard — agentic coding, deep research, security review, or anything where a wrong answer is expensive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drop to Luna&lt;/strong&gt; for high-volume, latency-sensitive, or cost-capped work where "good and fast and cheap" beats "best."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I go much deeper on this — including the break-even math and a routing strategy — in the companion post: &lt;a href="https://umesh-malik.com/blog/gpt-5-6-sol-vs-terra-vs-luna" rel="noopener noreferrer"&gt;GPT-5.6 Sol vs Terra vs Luna: which one to actually use&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using GPT-5.6 in the API
&lt;/h2&gt;

&lt;p&gt;The models are available as &lt;code&gt;gpt-5.6-sol&lt;/code&gt;, &lt;code&gt;gpt-5.6-terra&lt;/code&gt;, and &lt;code&gt;gpt-5.6-luna&lt;/code&gt;, with the alias &lt;code&gt;gpt-5.6&lt;/code&gt; routing to Sol. Here is a minimal call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-terra&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// default workhorse; swap to sol/luna as needed&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Summarize this incident timeline and propose three remediation steps.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;reasoning&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;effort&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;medium&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output_text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Escalating a single hard request to Sol with a higher thinking mode is a one-line change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hard&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-sol&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Audit this auth module for vulnerabilities and produce a patch.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;reasoning&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;effort&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;high&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="c1"&gt;// dial up for long-horizon, high-stakes work&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Migration and Rollout
&lt;/h2&gt;

&lt;h2&gt;
  
  
  What GPT-5.6 Still Does Not Solve
&lt;/h2&gt;

&lt;p&gt;The release is strong, but read the tradeoffs before you over-index on the launch numbers.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The knowledge cutoff is February 16, 2026
&lt;/h3&gt;

&lt;p&gt;Impressive, but still a cutoff. For genuinely current facts you still need web search or your own retrieval layer. Do not assume the model "knows" anything after mid-February 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. A 1M window is not 1M of perfect recall
&lt;/h3&gt;

&lt;p&gt;Shared context across all three tiers is great, but far-edge retrieval still degrades. Keep your retrieval discipline.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Higher thinking modes cost real time and money
&lt;/h3&gt;

&lt;p&gt;Ultra and max deliver the record scores, but they are slower and more expensive. They are not a default; they are an escalation.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Cheaper does not mean free to misroute
&lt;/h3&gt;

&lt;p&gt;The whole point of three tiers is routing. If you send everything to Sol out of caution, you lose the entire cost advantage of the family. If you send everything to Luna to save money, you will pay for it in quality failures. The value is in matching the tier to the task.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Powerful cyber capability is a double-edged surface
&lt;/h3&gt;

&lt;p&gt;A model strong enough to be OpenAI's best defensive security tool is also more capable in the wrong hands. Treat security-related agent workflows as privileged, with logging, scope limits, and human confirmation.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Final Take
&lt;/h2&gt;

&lt;p&gt;The most important thing to understand about GPT-5.6 is that OpenAI stopped competing purely on the capability ceiling and started competing on &lt;strong&gt;capability per dollar&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Splitting the release into Sol, Terra, and Luna — all sharing the same 1.05M context and platform features — turns model selection into a routing problem instead of an all-or-nothing bet. The flagship is genuinely strong on coding and agents. But the release's real weapon is that even the cheap tier is roughly a last-generation frontier model, and the flagship finishes the same work with half the tokens.&lt;/p&gt;

&lt;p&gt;If you run LLMs at any real scale, evaluate GPT-5.6 with a cost-per-successful-task lens, not just a leaderboard lens. Then set Terra as your default, escalate to Sol where it earns its price, and push volume to Luna. That is the whole game this release is built around.&lt;/p&gt;

&lt;p&gt;Next, read the companion deep-dive on &lt;a href="https://umesh-malik.com/blog/gpt-5-6-sol-vs-terra-vs-luna" rel="noopener noreferrer"&gt;choosing between Sol, Terra, and Luna&lt;/a&gt;, and the breakdown of &lt;a href="https://umesh-malik.com/blog/chatgpt-apps-sdk-super-app-guide" rel="noopener noreferrer"&gt;the ChatGPT super-app reform and Apps SDK&lt;/a&gt; that shipped alongside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/" rel="noopener noreferrer"&gt;TechCrunch: OpenAI launches its new family of models with GPT-5.6&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/gpt-5-6/" rel="noopener noreferrer"&gt;OpenAI: GPT-5.6&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna" rel="noopener noreferrer"&gt;OpenAI Help Center: A preview of GPT-5.6 Sol, Terra, and Luna&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.blog/changelog/2026-07-09-openais-gpt-5-6-sol-terra-and-luna-are-now-available-in-github-copilot/" rel="noopener noreferrer"&gt;GitHub Changelog: GPT-5.6 Sol, Terra, and Luna in GitHub Copilot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jul/9/gpt-5-6/" rel="noopener noreferrer"&gt;Simon Willison: The new GPT-5.6 family — Luna, Terra, Sol&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Explore more:&lt;/strong&gt; &lt;a href="https://umesh-malik.com/topics/llm-engineering" rel="noopener noreferrer"&gt;LLM Engineering — RAG, Fine-Tuning &amp;amp; Production LLMs&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-4-complete-guide" rel="noopener noreferrer"&gt;GPT-5.4 Guide: Benchmarks, Pricing, API &amp;amp; GPT-5.4 Pro&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/gpt-5-6-sol-vs-terra-vs-luna" rel="noopener noreferrer"&gt;GPT-5.6 Sol vs Terra vs Luna: Which One Should You Actually Use? (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-3-instant-fewer-refusals-better-answers" rel="noopener noreferrer"&gt;OpenAI GPT-5.3 Instant: 26.8% Fewer Hallucinations, Reduced Refusals, and Better Web Answers&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>chatgpt</category>
      <category>gpt56</category>
    </item>
    <item>
      <title>GPT-5.6 Sol vs Terra vs Luna: Which One Should You Actually Use? (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Sat, 11 Jul 2026 15:07:56 +0000</pubDate>
      <link>https://dev.to/umesh_malik/gpt-56-sol-vs-terra-vs-luna-which-one-should-you-actually-use-2026-5fae</link>
      <guid>https://dev.to/umesh_malik/gpt-56-sol-vs-terra-vs-luna-which-one-should-you-actually-use-2026-5fae</guid>
      <description>&lt;p&gt;&lt;strong&gt;GPT-5.6 Sol vs Terra vs Luna&lt;/strong&gt; is the wrong way to frame the choice — the real question is not which model is best, but which one to use &lt;em&gt;for this request&lt;/em&gt;. OpenAI shipping &lt;strong&gt;GPT-5.6 as three models is not a marketing gimmick. It is a routing problem handed to you.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The old question was "should I pay for the pro model?" The new question is sharper: &lt;strong&gt;for this specific request, what is the cheapest tier that still gets it right?&lt;/strong&gt; Get that wrong in the expensive direction and you burn money sending trivial requests to the flagship. Get it wrong in the cheap direction and you ship quality failures to users. This post is about getting it right.&lt;/p&gt;

&lt;p&gt;The short answer: &lt;strong&gt;make Terra your default, escalate to Sol only when the task is genuinely hard, and push high-volume work to Luna.&lt;/strong&gt; That single rule will beat both "always use the best model" and "always use the cheapest" on cost-per-successful-task — which is the only metric that actually matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default to Terra.&lt;/strong&gt; It is half the price of Sol, keeps the full &lt;strong&gt;1.05M context&lt;/strong&gt;, and is strong enough for most production work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalate to Sol&lt;/strong&gt; for hard, agentic, or high-stakes tasks — coding agents, deep research, security review. Sol scores &lt;strong&gt;80&lt;/strong&gt; on the Artificial Analysis Coding Agent Index and uses &lt;strong&gt;far fewer tokens&lt;/strong&gt; to finish, which partly offsets its higher price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drop to Luna&lt;/strong&gt; for high-volume, latency-sensitive, or cost-capped paths. At &lt;strong&gt;$1 / $6&lt;/strong&gt; it is &lt;strong&gt;5× cheaper&lt;/strong&gt; than Sol and still beats the last generation's flagship.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sol's per-token price is 2× Terra and 5× Luna&lt;/strong&gt;, but Sol's &lt;strong&gt;token efficiency&lt;/strong&gt; means the effective gap on a completed task is smaller than the sticker price suggests.&lt;/li&gt;
&lt;li&gt;The winning pattern is &lt;strong&gt;difficulty-based routing&lt;/strong&gt;: classify the request, send it to the cheapest capable tier, and escalate on failure or low confidence.&lt;/li&gt;
&lt;li&gt;Because all three share one API surface and context window, &lt;strong&gt;mixing tiers in a single app is trivial&lt;/strong&gt; — it is a model-string change, not a rewrite.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fims7w0rl7xslykd5d6w8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fims7w0rl7xslykd5d6w8.png" alt="GPT-5.6 routing map showing how to send each request to Luna, Terra, or Sol based on difficulty and stakes" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pricing Math (Do This First)
&lt;/h2&gt;

&lt;p&gt;Everything downstream depends on the raw numbers, so start here.&lt;/p&gt;

&lt;p&gt;Two things jump out.&lt;/p&gt;

&lt;p&gt;First, the &lt;strong&gt;gaps are clean multiples&lt;/strong&gt;: Terra is exactly 2× cheaper than Sol, and Luna is 5× cheaper. That makes routing decisions easy to reason about — moving a request from Sol to Terra literally halves its token cost.&lt;/p&gt;

&lt;p&gt;Second, and this is the part people miss: &lt;strong&gt;Sol uses fewer tokens to finish the same task.&lt;/strong&gt; OpenAI cites roughly &lt;strong&gt;54% better token efficiency&lt;/strong&gt; on coding work, with Sol using &lt;strong&gt;less than half the output tokens&lt;/strong&gt; of a comparable frontier model. So the effective cost gap between Sol and Terra on a &lt;em&gt;completed&lt;/em&gt; task is smaller than the 2× per-token headline — sometimes much smaller on long agentic runs where Sol simply finishes in fewer steps.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 &lt;strong&gt;Key insight&lt;/strong&gt;: Compare models on &lt;strong&gt;cost per successful task&lt;/strong&gt;, not cost per token. A pricier model that finishes in half the tokens and half the retries can be cheaper in practice.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Cost Per Task, Not Cost Per Token
&lt;/h2&gt;

&lt;p&gt;Here is the trap. A naive cost model says "Sol is 2× Terra, so always prefer Terra." But real workloads have &lt;strong&gt;retries, failures, and multi-step loops&lt;/strong&gt;, and those wreck the naive math.&lt;/p&gt;

&lt;p&gt;Consider an agentic coding task:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On &lt;strong&gt;Luna&lt;/strong&gt;, it might take 4 attempts and still fail once, burning tokens on dead ends.&lt;/li&gt;
&lt;li&gt;On &lt;strong&gt;Terra&lt;/strong&gt;, it succeeds in 2 attempts.&lt;/li&gt;
&lt;li&gt;On &lt;strong&gt;Sol&lt;/strong&gt;, it succeeds first try, in half the tokens, with ultra thinking.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Suddenly "the expensive model" can be the cheapest path to a &lt;em&gt;shipped&lt;/em&gt; result, because you paid once instead of three times. This is why routing beats a fixed choice: &lt;strong&gt;the right tier depends on how hard the request is, and you often do not know until you try.&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The cheap-model tax&lt;/strong&gt;&lt;br&gt;
Sending a hard task to Luna to save money frequently costs more, because you pay for repeated failed attempts, longer loops, and eventually a human cleanup. Cheap-per-token is not cheap-per-outcome. Measure the outcome.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Coding: When Sol Actually Earns Its Price
&lt;/h2&gt;

&lt;p&gt;Coding is where the tiers separate most clearly.&lt;/p&gt;

&lt;p&gt;The practical read:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use Sol&lt;/strong&gt; when the coding task is open-ended and agentic — "understand this repo, plan the change, edit across files, run the tests, fix what breaks." That is where its long-horizon strength and token efficiency compound.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use Terra&lt;/strong&gt; for scoped, everyday engineering — a well-defined feature, a code review, a bug with a clear repro. It is the right default for most day-to-day work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use Luna&lt;/strong&gt; for short, mechanical edits and high-volume operations — bulk refactors with a clear pattern, autocomplete-style suggestions, or anything where speed and price dominate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want to see the agentic-coding loop that Sol is optimized for, my walkthrough of &lt;a href="https://umesh-malik.com/blog/figma-codex-react-2026" rel="noopener noreferrer"&gt;building UIs with Codex and Figma&lt;/a&gt; and the field notes on &lt;a href="https://umesh-malik.com/blog/claude-code-auto-mode-production-field-report" rel="noopener noreferrer"&gt;Claude Code auto mode in production&lt;/a&gt; both show what "the model keeps going without a babysitter" actually looks like in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Routing Strategy
&lt;/h2&gt;

&lt;p&gt;This is the part worth stealing. Do not pick one model — pick a &lt;strong&gt;router&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Because all three models share the same API surface, implementing this is genuinely a model-string swap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// A crude but effective starting router.&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;pickModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;difficulty&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;highStakes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;highStakes&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;difficulty&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-sol&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;difficulty&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-terra&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-luna&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;difficulty&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;highStakes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;pickModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="c1"&gt;// Escalate one tier if the cheap answer fails your validator.&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;isValid&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-sol&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-luna&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-terra&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-5.6-sol&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output_text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Start simple, then measure&lt;/strong&gt;&lt;br&gt;
You do not need a fancy ML router on day one. Length thresholds, task-type rules, and "escalate on validation failure" will already beat a fixed single-model choice. Add sophistication only where the logs show it pays.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  GPT-5.6 Sol vs Terra vs Luna: The Decisions, Made Explicit
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Migration Checklist
&lt;/h2&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Final Take
&lt;/h2&gt;

&lt;p&gt;GPT-5.6 is the first OpenAI release where &lt;strong&gt;the smartest move is not choosing a model — it is choosing a router.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Make Terra your default, escalate to Sol when the task is genuinely hard, and push volume to Luna. Measure cost per successful task, not cost per token, and let the data move your thresholds. Do that, and you get most of the flagship's quality on the requests that need it while paying Luna-and-Terra prices for everything else.&lt;/p&gt;

&lt;p&gt;That is the entire point of splitting the frontier into three tiers. Use it.&lt;/p&gt;

&lt;p&gt;For the full release picture — benchmarks, cybersecurity positioning, and the API — start with the &lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;GPT-5.6 complete guide&lt;/a&gt;. And if you are building on the platform, the &lt;a href="https://umesh-malik.com/blog/chatgpt-apps-sdk-super-app-guide" rel="noopener noreferrer"&gt;ChatGPT Apps SDK and super-app reform&lt;/a&gt; shipped in the same window.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/" rel="noopener noreferrer"&gt;TechCrunch: OpenAI launches its new family of models with GPT-5.6&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/gpt-5-6/" rel="noopener noreferrer"&gt;OpenAI: GPT-5.6&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna" rel="noopener noreferrer"&gt;OpenAI Help Center: A preview of GPT-5.6 Sol, Terra, and Luna&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Explore more:&lt;/strong&gt; &lt;a href="https://umesh-malik.com/topics/llm-engineering" rel="noopener noreferrer"&gt;LLM Engineering — RAG, Fine-Tuning &amp;amp; Production LLMs&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/gpt-5-6-sol-vs-terra-vs-luna" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;OpenAI GPT-5.6 Complete Guide: Sol, Terra, Luna Benchmarks, Pricing, and API (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-4-complete-guide" rel="noopener noreferrer"&gt;GPT-5.4 Guide: Benchmarks, Pricing, API &amp;amp; GPT-5.4 Pro&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-3-instant-fewer-refusals-better-answers" rel="noopener noreferrer"&gt;OpenAI GPT-5.3 Instant: 26.8% Fewer Hallucinations, Reduced Refusals, and Better Web Answers&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>gpt56</category>
      <category>api</category>
    </item>
    <item>
      <title>ChatGPT Apps SDK and the Super App Reform: How Apps in ChatGPT Work (2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Sat, 11 Jul 2026 15:07:24 +0000</pubDate>
      <link>https://dev.to/umesh_malik/chatgpt-apps-sdk-and-the-super-app-reform-how-apps-in-chatgpt-work-2026-4eod</link>
      <guid>https://dev.to/umesh_malik/chatgpt-apps-sdk-and-the-super-app-reform-how-apps-in-chatgpt-work-2026-4eod</guid>
      <description>&lt;p&gt;OpenAI just turned ChatGPT into a platform — and the &lt;strong&gt;ChatGPT Apps SDK&lt;/strong&gt; is how developers ship to it. The company stopped treating ChatGPT as a chatbot and started treating it as an &lt;strong&gt;operating system&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The 2026 reform — a "super app" redesign reportedly codenamed &lt;strong&gt;Aria&lt;/strong&gt; — folds AI agents, the Codex coding tool, image generation, and a growing catalog of &lt;strong&gt;third-party apps&lt;/strong&gt; directly into the main ChatGPT interface. For its ~900 million weekly users, the chat box is becoming a launcher: you ask for a playlist and Spotify renders inside the conversation, you ask for a place to stay and Booking.com shows up as an interactive card, you sketch an idea and Canva opens inline.&lt;/p&gt;

&lt;p&gt;The thing that makes this real for developers is the &lt;strong&gt;Apps SDK&lt;/strong&gt; — and the most important detail is that it is &lt;strong&gt;built on the Model Context Protocol (MCP)&lt;/strong&gt;. OpenAI did not invent a proprietary plugin format this time. It standardized on the same open protocol the rest of the industry is converging on, and added a UI layer on top.&lt;/p&gt;

&lt;p&gt;The short answer: &lt;strong&gt;ChatGPT is now a distribution platform, and the Apps SDK is how you ship to it.&lt;/strong&gt; If you build software, "being inside ChatGPT" is about to be a channel you have to think about — the way "being in the App Store" became one in 2008.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI reworked ChatGPT into a &lt;strong&gt;super app&lt;/strong&gt;: agents, Codex, image generation, and &lt;strong&gt;third-party apps&lt;/strong&gt; live inside the main interface, reaching ~900M weekly users.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;Apps SDK&lt;/strong&gt; is how developers build those apps, and it is &lt;strong&gt;built on MCP (Model Context Protocol)&lt;/strong&gt; — an open standard, not a closed plugin format.&lt;/li&gt;
&lt;li&gt;An app is an &lt;strong&gt;MCP server&lt;/strong&gt; that lists tools, executes tool calls, and returns &lt;strong&gt;widgets&lt;/strong&gt; — web UI components ChatGPT renders &lt;strong&gt;inline&lt;/strong&gt; in the conversation via a sandboxed iframe.&lt;/li&gt;
&lt;li&gt;The UI talks to ChatGPT over a &lt;strong&gt;JSON-RPC bridge (postMessage)&lt;/strong&gt; and can request &lt;strong&gt;inline, picture-in-picture, or fullscreen&lt;/strong&gt; display via &lt;code&gt;window.openai.requestDisplayMode&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Launch partners&lt;/strong&gt; included Booking.com, Canva, Coursera, Figma, Expedia, Spotify, and Zillow.&lt;/li&gt;
&lt;li&gt;Apps are available to logged-in users &lt;strong&gt;outside the EEA, Switzerland, and the UK&lt;/strong&gt; on &lt;strong&gt;Free, Go, Plus, and Pro&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Developers submit apps through a &lt;strong&gt;submission portal&lt;/strong&gt;; approved ones land in the &lt;strong&gt;App Directory&lt;/strong&gt; to be discovered and shared.&lt;/li&gt;
&lt;li&gt;Because it is MCP-based, the &lt;strong&gt;same app can run in other MCP Apps–compatible hosts&lt;/strong&gt; — you build the UI once.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8brwvyu75f91vjzkr1m4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8brwvyu75f91vjzkr1m4.png" alt="ChatGPT super app anatomy showing agents, Codex, image generation, and third-party apps embedded in one interface" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the "Super App" Reform Actually Is
&lt;/h2&gt;

&lt;p&gt;For years, ChatGPT was a text box with a model behind it. The reform changes the &lt;strong&gt;shape of the product&lt;/strong&gt;: the conversation becomes a surface where different capabilities and apps render directly.&lt;/p&gt;

&lt;p&gt;Concretely, the redesigned ChatGPT pulls several things into one place:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agents&lt;/strong&gt; that can take multi-step actions on your behalf&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Codex&lt;/strong&gt;, the coding tool, available in-line for engineering work&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image generation&lt;/strong&gt; as a first-class capability, not a mode you switch to&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Third-party apps&lt;/strong&gt; — Spotify, Canva, Figma, Booking.com and more — that render interactive UI inside the chat&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Codex and Figma sitting side by side in that same launch lineup is not a coincidence — see &lt;a href="https://umesh-malik.com/blog/figma-codex-react-2026" rel="noopener noreferrer"&gt;how to turn Figma designs into production React with Codex&lt;/a&gt; for that workflow in practice.&lt;/p&gt;

&lt;p&gt;The strategic move is obvious once you name it: OpenAI wants ChatGPT to be the place you &lt;strong&gt;start&lt;/strong&gt; a task, not the place you go to get text you then paste elsewhere. That is the definition of a super app — one interface that handles many jobs — and it is the same playbook WeChat ran in China a decade ago.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why this is not 'plugins, again'&lt;/strong&gt;&lt;br&gt;
The 2023 plugin era failed partly because it was a closed, bespoke format with weak UI and weak discovery. The 2026 reform fixes all three: it is built on the &lt;strong&gt;open MCP standard&lt;/strong&gt;, apps render &lt;strong&gt;rich inline UI&lt;/strong&gt;, and there is a real &lt;strong&gt;App Directory&lt;/strong&gt; for discovery. Same ambition, much better foundation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why Building on MCP Is the Real Story
&lt;/h2&gt;

&lt;p&gt;The single most important technical decision in this release is that the &lt;strong&gt;Apps SDK builds on the Model Context Protocol.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;MCP is the open standard for connecting models to external tools and data. If you have already built an MCP server — for Claude, for an internal agent, for anything — you are most of the way to a ChatGPT app. And because ChatGPT implements the same &lt;strong&gt;iframe-and-bridge model&lt;/strong&gt; that the broader &lt;strong&gt;MCP Apps&lt;/strong&gt; effort defines, a UI you build for ChatGPT can run in other MCP Apps–compatible hosts too.&lt;/p&gt;

&lt;p&gt;That is a genuinely different bet than a proprietary plugin store. It means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You build once.&lt;/strong&gt; The tool contract and UI bridge are standardized.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You are not fully locked in.&lt;/strong&gt; MCP is an open protocol with multiple hosts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The ecosystem compounds.&lt;/strong&gt; Every MCP investment across the industry now points at ChatGPT distribution too.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are new to MCP, start with my &lt;a href="https://umesh-malik.com/blog/how-to-build-mcp-server" rel="noopener noreferrer"&gt;guide to building an MCP server&lt;/a&gt; and the walkthrough on &lt;a href="https://umesh-malik.com/blog/deploy-mcp-server-cloudflare-workers" rel="noopener noreferrer"&gt;deploying an MCP server on Cloudflare Workers&lt;/a&gt; — the same server you build there is the foundation for a ChatGPT app.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the ChatGPT Apps SDK Works
&lt;/h2&gt;

&lt;p&gt;An app in ChatGPT has two halves: an &lt;strong&gt;MCP server&lt;/strong&gt; (the logic) and an optional &lt;strong&gt;UI component&lt;/strong&gt; (the interface). This is the core of the ChatGPT Apps SDK — here is the anatomy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fijit1x8ou3x5ynexv8vv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fijit1x8ou3x5ynexv8vv.png" alt="ChatGPT Apps SDK architecture: an MCP server exposes tools and widgets, and ChatGPT renders the UI in a sandboxed iframe connected by a JSON-RPC bridge" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At minimum, your MCP server implements three capabilities:&lt;/p&gt;

&lt;p&gt;The UI half runs inside a &lt;strong&gt;sandboxed iframe&lt;/strong&gt; and communicates with ChatGPT over a standard bridge — &lt;strong&gt;&lt;code&gt;ui/*&lt;/code&gt; JSON-RPC messages over &lt;code&gt;postMessage&lt;/code&gt;&lt;/strong&gt;. Your component can read state, call your tools, and negotiate how it is displayed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Inside your app's UI bundle, running in the ChatGPT iframe.&lt;/span&gt;

&lt;span class="c1"&gt;// Ask ChatGPT to present the widget larger for an interactive task.&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;requestDisplayMode&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;fullscreen&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt; &lt;span class="c1"&gt;// or 'inline' | 'pip'&lt;/span&gt;

&lt;span class="c1"&gt;// Call one of your MCP server's tools from the UI and render the result.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;callTool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;search_listings&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Lisbon&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;checkIn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2026-08-01&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="nf"&gt;renderListings&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;structuredContent&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The bridge is the whole trick&lt;/strong&gt;&lt;br&gt;
Your UI never talks to the network directly for app logic — it talks to ChatGPT over the bridge, and ChatGPT relays to your MCP server. That is what keeps apps sandboxed, consistent, and portable across MCP Apps hosts. Build the UI once; the host handles presentation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The display modes
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;window.openai.requestDisplayMode&lt;/code&gt; lets a widget negotiate how it appears:&lt;/p&gt;

&lt;h2&gt;
  
  
  Building an App: The Shape of the Work
&lt;/h2&gt;

&lt;p&gt;You do not need to learn a new framework. If you can write an MCP server and a small web UI, you can build a ChatGPT app.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Availability caveat&lt;/strong&gt;&lt;br&gt;
At launch, apps in ChatGPT were available to logged-in users &lt;strong&gt;outside the EEA, Switzerland, and the UK&lt;/strong&gt;, on Free, Go, Plus, and Pro. If your audience is primarily European, factor that rollout gap into any launch plan — the distribution surface is not yet global.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Opportunity — and the Risk
&lt;/h2&gt;

&lt;p&gt;This is a real distribution channel, but building on someone else's platform always cuts both ways.&lt;/p&gt;

&lt;p&gt;For brands, the mental shift is the same one search and app stores forced: &lt;strong&gt;users increasingly start inside an aggregator.&lt;/strong&gt; If people begin their travel, design, or media tasks in ChatGPT, being absent from the App Directory is a discovery problem — the same way being absent from Google once was.&lt;/p&gt;

&lt;h2&gt;
  
  
  Readiness Checklist
&lt;/h2&gt;

&lt;h2&gt;
  
  
  The Bigger Picture: This Pairs With GPT-5.6
&lt;/h2&gt;

&lt;p&gt;The super app reform did not ship in a vacuum. It arrived in the same window as the &lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;GPT-5.6 model family&lt;/a&gt;, and the two reinforce each other. A more capable, more efficient model makes agentic, multi-app workflows economically viable; a super app surface gives that model somewhere to &lt;em&gt;do&lt;/em&gt; things. Cheaper tokens plus a distribution platform is how "AI agent that gets real work done" stops being a demo.&lt;/p&gt;

&lt;p&gt;If you are building agents that call many tools, the same discipline I wrote about in &lt;a href="https://umesh-malik.com/blog/agentic-ai-enterprise-security-model" rel="noopener noreferrer"&gt;the agentic AI enterprise security model&lt;/a&gt; applies: capability without governance is a liability, and apps inside ChatGPT are no exception.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Final Take
&lt;/h2&gt;

&lt;p&gt;The ChatGPT super app reform is OpenAI's clearest bet yet that the &lt;strong&gt;interface&lt;/strong&gt;, not just the model, is the product.&lt;/p&gt;

&lt;p&gt;By building the Apps SDK on MCP, OpenAI turned "being inside ChatGPT" into a channel you can reach with mostly standard, reusable work — and turned ChatGPT into a distribution platform sitting in front of ~900 million weekly users. That is a genuine opportunity and a genuine dependency at the same time.&lt;/p&gt;

&lt;p&gt;The right posture is pragmatic: &lt;strong&gt;reuse your MCP investment, ship an app as a channel, design for portability, and keep your product's center of gravity your own.&lt;/strong&gt; Do that, and the super app is upside without the classic platform trap.&lt;/p&gt;

&lt;p&gt;Next, read the &lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;GPT-5.6 complete guide&lt;/a&gt; that shipped alongside it, and — if you are building the backend — my &lt;a href="https://umesh-malik.com/blog/how-to-build-mcp-server" rel="noopener noreferrer"&gt;MCP server guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/introducing-apps-in-chatgpt/" rel="noopener noreferrer"&gt;OpenAI: Introducing apps in ChatGPT and the new Apps SDK&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.openai.com/apps-sdk" rel="noopener noreferrer"&gt;OpenAI Developers: Apps SDK&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.openai.com/apps-sdk/build/chatgpt-ui" rel="noopener noreferrer"&gt;OpenAI Developers: Build your ChatGPT UI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/developers-can-now-submit-apps-to-chatgpt/" rel="noopener noreferrer"&gt;OpenAI: Developers can now submit apps to ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.modelcontextprotocol.io/posts/2026-01-26-mcp-apps/" rel="noopener noreferrer"&gt;Model Context Protocol Blog: MCP Apps — bringing UI capabilities to MCP clients&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Explore more:&lt;/strong&gt; &lt;a href="https://umesh-malik.com/topics/ai-coding-agents" rel="noopener noreferrer"&gt;AI Coding Agents &amp;amp; Developer Experience&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/chatgpt-apps-sdk-super-app-guide" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-6-sol-terra-luna-guide" rel="noopener noreferrer"&gt;OpenAI GPT-5.6 Complete Guide: Sol, Terra, Luna Benchmarks, Pricing, and API (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/openai-gpt-5-4-complete-guide" rel="noopener noreferrer"&gt;GPT-5.4 Guide: Benchmarks, Pricing, API &amp;amp; GPT-5.4 Pro&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/gpt-5-6-sol-vs-terra-vs-luna" rel="noopener noreferrer"&gt;GPT-5.6 Sol vs Terra vs Luna: Which One Should You Actually Use? (2026)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>chatgpt</category>
      <category>appssdk</category>
    </item>
    <item>
      <title>How to Build Enterprise-Grade AI Agents for Free (MaxKB, 2026)</title>
      <dc:creator>Umesh Malik</dc:creator>
      <pubDate>Tue, 07 Jul 2026 19:57:26 +0000</pubDate>
      <link>https://dev.to/umesh_malik/how-to-build-enterprise-grade-ai-agents-for-free-maxkb-2026-1nhc</link>
      <guid>https://dev.to/umesh_malik/how-to-build-enterprise-grade-ai-agents-for-free-maxkb-2026-1nhc</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How to build enterprise-grade AI agents for free:&lt;/strong&gt; self-host &lt;a href="https://github.com/1Panel-dev/MaxKB" rel="noopener noreferrer"&gt;MaxKB&lt;/a&gt; — an open-source (GPLv3, ~22k GitHub stars) agent platform — and point it at a local model like DeepSeek or Llama via Ollama, so you pay zero API tokens and no data ever leaves your server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Enterprise-grade" isn't a checkbox — it's five things:&lt;/strong&gt; answer precision, cost control, data-sovereign security, access control, and observability. Free tools can nail four of them out of the box.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The one honest catch:&lt;/strong&gt; MaxKB's Community edition is free forever but capped (2 users, 5 apps, 50 knowledge bases). SSO, LDAP, and RBAC live in the paid Pro tier ($1,920/yr) — you can replace them yourself for $0 with more effort.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Precision comes from your RAG pipeline, not the model.&lt;/strong&gt; Chunking, hybrid search, a reranker, and a "cite or refuse" prompt matter more than which LLM you pick.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MaxKB vs Dify vs n8n:&lt;/strong&gt; pick MaxKB for a knowledge-grounded Q&amp;amp;A agent, Dify for a broad LLM app builder, n8n when the agent is one step in a bigger automation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Most "Free AI Agent" Guides Are Lying to You
&lt;/h2&gt;

&lt;p&gt;Here's how to build enterprise-grade AI agents for free in 2026: self-host the open-source platform MaxKB, point it at a local model, and you get a document-grounded, tool-using agent with $0 API cost and zero data leaving your network.&lt;/p&gt;

&lt;p&gt;Most other "free AI agent" tutorials are either toys or bait. The toy version wires ChatGPT to a prompt and calls it an "agent." The bait version is free until step 7, when you hit a paywall, a per-token API meter, or a "contact sales" wall right as it gets useful. Neither gives you something you'd actually put in front of customers or run your internal knowledge base on.&lt;/p&gt;

&lt;p&gt;This guide is the version I wish existed. We're going to stand up a real, &lt;strong&gt;enterprise-grade AI agent&lt;/strong&gt; — one that answers from &lt;em&gt;your&lt;/em&gt; documents with citations, takes actions through tools, runs entirely on infrastructure you control, and costs &lt;strong&gt;$0 in API fees&lt;/strong&gt; — and I'll be honest about exactly where "free" stops and money starts. No hand-waving.&lt;/p&gt;

&lt;p&gt;The vehicle is &lt;strong&gt;MaxKB&lt;/strong&gt;, and by the end you'll have a working agent plus a clear-eyed view of the five things that separate a demo from something you can trust in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is an Enterprise-Grade AI Agent?
&lt;/h2&gt;

&lt;p&gt;An &lt;strong&gt;enterprise-grade AI agent&lt;/strong&gt; is an AI system that answers or acts on your organization's own data with measurable accuracy, keeps that data under your control, enforces who can do what, stays observable, and does all of it at a cost you can predict. "Enterprise-grade" is about &lt;em&gt;trust and control&lt;/em&gt; — not about how big the model is.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The five-question test&lt;/strong&gt;&lt;br&gt;
Point any agent at five questions: is it accurate on &lt;em&gt;our&lt;/em&gt; data, is the cost predictable, does our data stay ours, can we control who does what, and can we see &lt;em&gt;why&lt;/em&gt; it answered the way it did? Five yeses is enterprise-grade. Anything less is a demo.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;People throw the phrase around like it means "expensive." It doesn't. A $30/month SaaS chatbot can be less enterprise-grade than a well-configured open-source stack running on a $40 VPS. What makes an agent enterprise-grade is whether it holds up on &lt;strong&gt;five pillars&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;Keep these five in mind. Everything below maps back to them. The good news: a free, self-hosted stack wins pillars 1, 2, 3, and 5 outright. Pillar 4 is the one place "free" gets an asterisk, and we'll deal with it head-on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The $0 Stack: MaxKB + a Local Model
&lt;/h2&gt;

&lt;p&gt;The cheapest enterprise-grade agent in 2026 is &lt;strong&gt;MaxKB running against a local LLM, both self-hosted on one machine&lt;/strong&gt;. That's the whole trick. MaxKB gives you the agent platform; a local model kills the API bill; self-hosting solves data sovereignty.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MaxKB&lt;/strong&gt; (short for &lt;em&gt;Max Knowledge Brain&lt;/em&gt;, built by the 1Panel team) is an open-source platform for building enterprise-grade agents. Its stack is boringly solid — Vue frontend, Django backend, LangChain under the hood, and &lt;strong&gt;PostgreSQL + pgvector&lt;/strong&gt; as the vector store — which means no exotic dependencies and one-command deployment.&lt;/p&gt;

&lt;p&gt;Out of the box it gives you a full &lt;strong&gt;RAG pipeline&lt;/strong&gt; (upload docs or crawl a site → automatic chunking → vectorization), a &lt;strong&gt;visual workflow engine&lt;/strong&gt;, and &lt;strong&gt;MCP tool-use&lt;/strong&gt; so the agent can call external tools. Crucially, it's &lt;strong&gt;model-agnostic&lt;/strong&gt; — it'll talk to OpenAI, Claude, and Gemini &lt;em&gt;or&lt;/em&gt; to local models like DeepSeek, Qwen, and Llama.&lt;/p&gt;

&lt;p&gt;That last part is the money-saver. Point MaxKB at a local model served by &lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; and your per-token cost drops to exactly zero.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp2abm9yqgqxn1ase2nck.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp2abm9yqgqxn1ase2nck.png" alt="The $0 enterprise AI agent stack: users hit a self-hosted MaxKB instance running the RAG pipeline, workflow engine and MCP tools, which calls a local LLM via Ollama and a pgvector database — all on one server you control, so no data leaves and there are no API token costs" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Build Enterprise-Grade AI Agents for Free, Step by Step
&lt;/h2&gt;

&lt;p&gt;Here's the honest, end-to-end walkthrough. You need a machine with Docker installed — a laptop works for testing; for a local model you'll want at least 16GB of RAM (more if you run larger models). Every step below is free.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1 — Run MaxKB
&lt;/h3&gt;

&lt;p&gt;One command. This maps a data volume so your knowledge bases survive restarts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;maxkb &lt;span class="nt"&gt;--restart&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;always &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; ~/.maxkb:/var/lib/postgresql/data &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; ~/.python-packages:/opt/maxkb/app/sandbox/python-packages &lt;span class="se"&gt;\&lt;/span&gt;
  1panel/maxkb
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open &lt;code&gt;http://localhost:8080&lt;/code&gt; and log in with the default credentials shown in the &lt;a href="https://docs.maxkb.pro/" rel="noopener noreferrer"&gt;MaxKB docs&lt;/a&gt;. Change the password immediately — that's the first line of your security checklist.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2 — Serve a local model for $0
&lt;/h3&gt;

&lt;p&gt;Install Ollama, then pull a model. DeepSeek and &lt;a href="https://umesh-malik.com/blog/local-llm-coding-revolution-qwen3-coder-desktop" rel="noopener noreferrer"&gt;Qwen&lt;/a&gt; punch far above their weight for RAG in 2026:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# install ollama, then:&lt;/span&gt;
ollama pull deepseek-r1:7b        &lt;span class="c"&gt;# reasoning model, runs on modest hardware&lt;/span&gt;
ollama pull nomic-embed-text      &lt;span class="c"&gt;# embeddings for the RAG pipeline&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ollama now serves an OpenAI-compatible endpoint at &lt;code&gt;http://localhost:11434&lt;/code&gt;. This is the whole reason your token bill is zero — inference happens on your hardware, not someone's metered API.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3 — Connect the model in MaxKB
&lt;/h3&gt;

&lt;p&gt;In the MaxKB UI go to &lt;strong&gt;Model Settings → Add Model&lt;/strong&gt;, choose the Ollama provider, and point it at your Ollama host. If MaxKB runs in Docker and Ollama runs on the host machine, use &lt;code&gt;http://host.docker.internal:11434&lt;/code&gt; as the base URL. Add both the chat model (&lt;code&gt;deepseek-r1:7b&lt;/code&gt;) and the embedding model (&lt;code&gt;nomic-embed-text&lt;/code&gt;). No API key required.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Precision tip&lt;/strong&gt;&lt;br&gt;
Use a dedicated &lt;strong&gt;embedding&lt;/strong&gt; model (like &lt;code&gt;nomic-embed-text&lt;/code&gt;) for the knowledge base, not your chat model. Retrieval quality — and therefore answer precision — depends far more on the embedding model than on the chat model. This is the single most common mistake people make and never diagnose.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Step 4 — Build the knowledge base
&lt;/h3&gt;

&lt;p&gt;Create a knowledge base, then either upload documents or paste a URL to crawl. MaxKB handles splitting, vectorizing, and indexing into pgvector automatically. Two settings decide your precision:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chunk size&lt;/strong&gt; — too big and retrieval pulls in noise; too small and it loses context. Start around 500–800 tokens with overlap, then tune against real questions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Segment cleanup&lt;/strong&gt; — strip navigation boilerplate and repeated headers before indexing. Garbage in the index is the number-one cause of confidently wrong answers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 5 — Create the agent and give it a tool
&lt;/h3&gt;

&lt;p&gt;Create an application, attach your knowledge base, and set a system prompt that enforces grounding — the "cite or refuse" rule:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Answer only from the provided knowledge base.
Cite the source document for every claim.
If the answer isn't in the knowledge base, say
"I don't have that information" — never guess.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then wire an &lt;strong&gt;MCP tool&lt;/strong&gt; so the agent can &lt;em&gt;act&lt;/em&gt;, not just answer — look up an order, create a ticket, query a database. MaxKB's function library and MCP support let you register tools the workflow can call. If you're new to MCP, start with &lt;a href="https://umesh-malik.com/blog/how-to-build-mcp-server" rel="noopener noreferrer"&gt;how to build an MCP server&lt;/a&gt; and &lt;a href="https://umesh-malik.com/blog/deploy-mcp-server-cloudflare-workers" rel="noopener noreferrer"&gt;deploying one on Cloudflare Workers&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6 — Embed it
&lt;/h3&gt;

&lt;p&gt;MaxKB generates an embeddable chat widget and a REST API. Paste the widget script into any page, or call the API from your backend. Zero front-end code required — this is the "zero-coding integration" MaxKB is built around.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five Pillars, Judged Honestly
&lt;/h2&gt;

&lt;p&gt;A running agent isn't the same as an enterprise-grade one. Let's grade the free stack against the five pillars — including where it falls short.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pillar 1 — Effectiveness &amp;amp; precision
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Precision comes from the retrieval layer, not the model.&lt;/strong&gt; This is the most important sentence in this article. Teams burn weeks swapping models when their real problem is a bad chunking strategy or no reranking. RAG works by retrieving relevant chunks and forcing the model to answer &lt;em&gt;from those chunks&lt;/em&gt;, which is what crushes hallucinations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg0p53dj4ft2iwszy3j2y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg0p53dj4ft2iwszy3j2y.png" alt="The RAG precision pipeline that keeps an agent accurate: documents are chunked and embedded, stored in a pgvector index, retrieved as top-k matches, reranked and filtered, then passed to the LLM for a grounded, cited answer — precision comes from the retrieval layer, not the model" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The precision levers, in order of impact:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Chunking&lt;/strong&gt; — right size + overlap, boilerplate stripped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hybrid search&lt;/strong&gt; — combine keyword and vector search so exact terms (part numbers, names) aren't lost to fuzzy semantics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reranking&lt;/strong&gt; — reorder the top-k so the &lt;em&gt;best&lt;/em&gt; chunk lands in the model's context, not just a &lt;em&gt;relevant&lt;/em&gt; one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A grounding prompt&lt;/strong&gt; — "cite or refuse," as above.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nail those four and a 7B local model will out-answer a frontier model with a sloppy pipeline. If you want the full theory, read &lt;a href="https://umesh-malik.com/blog/build-rag-pipeline-from-scratch" rel="noopener noreferrer"&gt;building a RAG pipeline from scratch&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pillar 2 — Cost control
&lt;/h3&gt;

&lt;p&gt;This is where self-hosting quietly wins. A metered API bills you more as you succeed; a per-seat SaaS bills you more as your team grows. A self-hosted local model turns both into &lt;strong&gt;one flat server bill&lt;/strong&gt; that barely moves.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgq19hzvu8saoqm2x7ghx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgq19hzvu8saoqm2x7ghx.png" alt="Cost as you scale, three ways: a self-hosted MaxKB plus local model stays roughly flat at a fixed server cost, a metered LLM API rises with token usage, and a per-seat SaaS agent platform rises steeply with team size" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The tradeoff is real and worth stating: self-hosting trades a variable &lt;em&gt;money&lt;/em&gt; cost for a fixed &lt;em&gt;operational&lt;/em&gt; cost — you run the server, you patch it, you own uptime. For a small internal agent that's a rounding error. At scale it's a massive saving.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pillar 3 — Security &amp;amp; data sovereignty
&lt;/h3&gt;

&lt;p&gt;The strongest argument for this stack. When the model runs locally and MaxKB runs on your server, &lt;strong&gt;no document ever touches a third-party cloud&lt;/strong&gt;. For anyone handling PII, health, financial, or regulated data, that alone can be the difference between "allowed" and "not allowed." You're not sending your knowledge base to an API you don't control.&lt;/p&gt;

&lt;p&gt;That's the &lt;em&gt;architecture&lt;/em&gt; being secure. You still have to &lt;em&gt;harden the deployment&lt;/em&gt;:&lt;/p&gt;

&lt;p&gt;For the broader threat model of agents that take actions, see &lt;a href="https://umesh-malik.com/blog/agentic-ai-enterprise-security-model" rel="noopener noreferrer"&gt;the agentic AI enterprise security model&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pillar 4 — Access control (the honest asterisk)
&lt;/h3&gt;

&lt;p&gt;Here's where free ends. MaxKB's &lt;strong&gt;Community edition caps you at 2 users&lt;/strong&gt;, and &lt;strong&gt;SSO, LDAP, and RBAC are Pro-tier features&lt;/strong&gt;. If you need "marketing can only see the marketing knowledge base, support leads can edit agents, everyone logs in with Okta" — that's the paid tier, or DIY work you take on yourself (an auth proxy in front, separate instances per team). I'd rather tell you that now than let you discover it at rollout.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pillar 5 — Observability
&lt;/h3&gt;

&lt;p&gt;MaxKB logs conversations and lets you inspect what was retrieved for a given answer, which is enough to debug precision and iterate on chunking. It's not a full LLM-observability suite — if you need deep tracing and eval dashboards you'll add tooling — but for "why did the agent say that?" you have what you need for free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Does "Free" Actually End?
&lt;/h2&gt;

&lt;p&gt;No dodging it. Here's the exact line between $0 and paid, so you can plan.&lt;/p&gt;

&lt;p&gt;My take: &lt;strong&gt;start on Community.&lt;/strong&gt; It is genuinely free and genuinely capable. Only pay when access control across multiple teams becomes a real, present need — not a hypothetical one. Most people building their first agent are nowhere near the 2-user wall.&lt;/p&gt;

&lt;h2&gt;
  
  
  MaxKB vs Dify vs n8n: Which Free AI Agent Platform Should You Use?
&lt;/h2&gt;

&lt;p&gt;"Free AI agent platform" returns a dozen tools that do overlapping-but-different things. Here's how the honest contenders compare — and none of these are the same product.&lt;/p&gt;

&lt;p&gt;The key insight: &lt;strong&gt;Ollama isn't a competitor to MaxKB — it's a component of the stack.&lt;/strong&gt; MaxKB and Dify compete; n8n plays a different game (orchestration). Here's how I'd choose:&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Mistakes That Kill Free Agents
&lt;/h2&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Bottom Line
&lt;/h2&gt;

&lt;p&gt;Free and enterprise-grade are not opposites in 2026 — that's the myth this guide exists to kill. Self-host MaxKB, point it at a local model, invest in your retrieval pipeline, harden the box, and you have an agent that answers from your data with citations, keeps that data on your own hardware, and costs nothing per token. The only honest asterisk is multi-team access control, and you now know exactly when that bill arrives.&lt;/p&gt;

&lt;p&gt;Start on the Community edition today. Build the narrow, useful agent — the one that answers your support questions or your internal docs — get the retrieval right, and let the results decide whether you ever need to pay.&lt;/p&gt;

&lt;p&gt;If this was useful, read &lt;a href="https://umesh-malik.com/blog/build-rag-pipeline-from-scratch" rel="noopener noreferrer"&gt;building a RAG pipeline from scratch&lt;/a&gt; next to push your agent's precision further, or &lt;a href="https://umesh-malik.com/blog/autonomous-ai-agents-production-gap-2026" rel="noopener noreferrer"&gt;why 77% of AI agents never reach production&lt;/a&gt; to make sure yours is in the 23% that do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/1Panel-dev/MaxKB" rel="noopener noreferrer"&gt;MaxKB — 1Panel-dev/MaxKB on GitHub&lt;/a&gt; (open-source, GPLv3)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.maxkb.pro/" rel="noopener noreferrer"&gt;MaxKB official documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://maxkb.pro/" rel="noopener noreferrer"&gt;MaxKB project site&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama — run local models&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Written for &lt;a href="https://umesh-malik.com" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt; — no-fluff technical writing on AI, Web Dev, and Engineering.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://umesh-malik.com/blog/build-enterprise-ai-agents-free" rel="noopener noreferrer"&gt;umesh-malik.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep reading on umesh-malik.com:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/autonomous-ai-agents-production-gap-2026" rel="noopener noreferrer"&gt;Why 77% of Autonomous AI Agents Never Reach Production (2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/rag-vs-fine-tuning-llms-2026" rel="noopener noreferrer"&gt;RAG vs Fine-Tuning for LLMs in 2026: A Production Decision Framework With Real Tradeoffs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://umesh-malik.com/blog/rag-chatbot-nextjs-guide" rel="noopener noreferrer"&gt;Build a RAG Chatbot in Next.js: Retrieval, Streaming &amp;amp; Citations (2026)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aiagents</category>
      <category>maxkb</category>
      <category>opensource</category>
      <category>rag</category>
    </item>
  </channel>
</rss>
