<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nomad</title>
    <description>The latest articles on DEV Community by Nomad (@technomadcode).</description>
    <link>https://dev.to/technomadcode</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166490%2F3ef39e8a-a67d-4016-aa6b-79f3339ec6f8.png</url>
      <title>DEV Community: Nomad</title>
      <link>https://dev.to/technomadcode</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/technomadcode"/>
    <language>en</language>
    <item>
      <title>I made a free, open-source toolkit for building apps with AI: prompts that interview you before they write your plan, and a step-by-step app starter for Claude Code. 1,000+ stars so far. https://github.com/TechNomadCode/AI-Product-Development-Toolkit</title>
      <dc:creator>Nomad</dc:creator>
      <pubDate>Wed, 07 Oct 2026 08:20:57 +0000</pubDate>
      <link>https://dev.to/technomadcode/i-made-a-free-open-source-toolkit-for-building-apps-with-ai-prompts-that-interview-you-before-2hf</link>
      <guid>https://dev.to/technomadcode/i-made-a-free-open-source-toolkit-for-building-apps-with-ai-prompts-that-interview-you-before-2hf</guid>
      <description>&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
        &lt;div class="c-embed__cover"&gt;
          &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" class="c-link align-middle" rel="noopener noreferrer"&gt;
            &lt;img alt="" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Frepository-images.githubusercontent.com%2F970875123%2F8f86e2a7-2456-45c7-bec6-54e5cde31f77" height="640" class="m-0" width="1280"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="c-embed__body"&gt;
        &lt;h2 class="fs-xl lh-tight"&gt;
          &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer" class="c-link"&gt;
            GitHub - TechNomadCode/AI-Product-Development-Toolkit: Plan, build and launch products with AI: guided planning prompts, an AI App Starter for Next.js, Supabase and Vercel, and agent configurations. · GitHub
          &lt;/a&gt;
        &lt;/h2&gt;
          &lt;p class="truncate-at-3"&gt;
            Plan, build and launch products with AI: guided planning prompts, an AI App Starter for Next.js, Supabase and Vercel, and agent configurations. - TechNomadCode/AI-Product-Development-Toolkit
          &lt;/p&gt;
        &lt;div class="color-secondary fs-s flex items-center"&gt;
            &lt;img alt="favicon" class="c-embed__favicon m-0 mr-2 radius-0" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.githubassets.com%2Ffavicons%2Ffavicon.svg" width="32" height="32"&gt;
          github.com
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;


</description>
      <category>ai</category>
      <category>github</category>
      <category>opensource</category>
      <category>tools</category>
    </item>
    <item>
      <title>What Independent Benchmarks Say About Opus 5.5</title>
      <dc:creator>Nomad</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:42:16 +0000</pubDate>
      <link>https://dev.to/technomadcode/what-independent-benchmarks-say-about-opus-55-21jh</link>
      <guid>https://dev.to/technomadcode/what-independent-benchmarks-say-about-opus-55-21jh</guid>
      <description>&lt;p&gt;Two independent benchmark results for Opus 5.5 came out this week, and they don't really agree. One has it in first place. The other has it fast and cheap, but third on secure code once you take out the answers it memorized.&lt;/p&gt;

&lt;p&gt;A few days ago I &lt;a href="https://promptquick.ai/blog/opus-5-5-improvements-and-prompting" rel="noopener noreferrer"&gt;read through Anthropic's Opus 5.5 docs&lt;/a&gt; to see what they say got better. That's Anthropic talking about their own model, though. I wanted to see what people who didn't build it found, so I read the first two outside evaluations I could find.&lt;/p&gt;

&lt;p&gt;Quick disclaimer: I didn't run any of these tests. All the numbers here come from &lt;a href="https://artificialanalysis.ai/" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; (via &lt;a href="https://www.heise.de/en/news/External-benchmarks-Claude-Opus-5-5-overtakes-OpenAI-s-Astra-and-Fable-5-1-11466259.html" rel="noopener noreferrer"&gt;heise online&lt;/a&gt;) and &lt;a href="https://www.endorlabs.com/learn/opus-5-5-6x-cheaper-and-2x-faster-than-fable-5-1-but-memorization-keeps-it-off-the-top-spot" rel="noopener noreferrer"&gt;Endor Labs&lt;/a&gt;. The charts are theirs too, credited under each one. The opinions are mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Artificial Analysis: first place
&lt;/h2&gt;

&lt;p&gt;Artificial Analysis has an Intelligence Index that combines ten different tests, things like programming and knowledge questions. According to &lt;a href="https://www.heise.de/en/news/External-benchmarks-Claude-Opus-5-5-overtakes-OpenAI-s-Astra-and-Fable-5-1-11466259.html" rel="noopener noreferrer"&gt;heise's report&lt;/a&gt;, Opus 5.5 is now on top with 58 points. OpenAI's GPT-6 Astra and Anthropic's Fable 5.1 got 53 each. All three were run at their highest setting.&lt;/p&gt;

&lt;p&gt;Two of those tests are about code. On Terminal-Bench 4.0 (coding agents working in a terminal), Opus 5.5 and Astra are pretty much tied. On SciCode (scientific programming), Opus is 11 points ahead of Astra and 4 ahead of Fable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F219tlvghwunplpsyvykf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F219tlvghwunplpsyvykf.png" alt="Two bar charts from Artificial Analysis. Terminal-Bench 4.0, agentic coding and terminal use: Claude Opus 5.5 60%, GPT-6 Astra 59%, Claude Fable 5.1 52%. SciCode, coding: Claude Opus 5.5 67%, Claude Fable 5.1 63%, GPT-6 Astra 56%. Higher is better." width="800" height="264"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Terminal-Bench 4.0 and SciCode, all three models at their highest setting. Chart: &lt;a href="https://artificialanalysis.ai/" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;, via &lt;a href="https://www.heise.de/en/news/External-benchmarks-Claude-Opus-5-5-overtakes-OpenAI-s-Astra-and-Fable-5-1-11466259.html" rel="noopener noreferrer"&gt;heise online&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One point on Terminal-Bench doesn't mean much. Eleven points on SciCode does, and five points on the whole index is a real lead. Just keep in mind these are all max-effort runs, which matters for the next part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheaper tokens, pricier tasks
&lt;/h2&gt;

&lt;p&gt;Artificial Analysis also worked out what one task from the index costs on average. Opus tokens are cheaper than Astra's, but Astra uses a lot fewer of them, so per task Astra comes out cheaper. Fable 5.1 costs more per token and more per task.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft5hgpz3uz9jzz9ls21ns.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft5hgpz3uz9jzz9ls21ns.png" alt="Stacked bar chart from Artificial Analysis: weighted average cost in US dollars per Intelligence Index task, split by token type. GPT-6 Astra at max: $3.26. Claude Opus 5.5 at max with fallback: $5.98. Claude Fable 5.1 at max with fallback: $7.63. Lower is better." width="800" height="304"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Average cost per Intelligence Index task: Astra $3.26, Opus 5.5 $5.98, Fable 5.1 $7.63. Chart: &lt;a href="https://artificialanalysis.ai/" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;, via &lt;a href="https://www.heise.de/en/news/External-benchmarks-Claude-Opus-5-5-overtakes-OpenAI-s-Astra-and-Fable-5-1-11466259.html" rel="noopener noreferrer"&gt;heise online&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Heise's summary seems fair to me: if you mostly care about performance, try Opus 5.5. If you're watching costs and still want similar results, look at Astra.&lt;/p&gt;

&lt;p&gt;One thing I'd add. Opus 5.5 defaults to medium effort, not max, and Anthropic says medium matched or beat Opus 5 on high in their own tests. I haven't seen anyone publish index numbers for Opus 5.5 on medium, so I can't say how that plays out. But the max-effort cost probably isn't what most people will actually pay.&lt;/p&gt;

&lt;h2&gt;
  
  
  Endor Labs: secure code in real projects
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.endorlabs.com/learn/opus-5-5-6x-cheaper-and-2x-faster-than-fable-5-1-but-memorization-keeps-it-off-the-top-spot" rel="noopener noreferrer"&gt;Endor Labs&lt;/a&gt; tests something much narrower. Their Agent Security League has a coding agent work on real open-source projects, in code that was once part of a security fix. The agent isn't told that. It's only asked to follow security best practices.&lt;/p&gt;

&lt;p&gt;Then every patch gets run against two sets of tests. &lt;strong&gt;FuncPass&lt;/strong&gt; means the functional tests pass, so the code works. &lt;strong&gt;SecPass&lt;/strong&gt; means the hidden security tests from the original fix pass too, so the code works and is safe. They ran Opus 5.5 in Claude Code, same setup as for Opus 5 and Fable 5.1, so the model is the main thing that changes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;FuncPass&lt;/th&gt;
&lt;th&gt;SecPass&lt;/th&gt;
&lt;th&gt;Memorized&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fable 5.1&lt;/td&gt;
&lt;td&gt;87.2%&lt;/td&gt;
&lt;td&gt;37.4%&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5.5&lt;/td&gt;
&lt;td&gt;68.7%&lt;/td&gt;
&lt;td&gt;33.5%&lt;/td&gt;
&lt;td&gt;51&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5&lt;/td&gt;
&lt;td&gt;73.7%&lt;/td&gt;
&lt;td&gt;32.4%&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Claude Code on the Agent Security League, with memorized solutions removed. “Memorized” counts the confirmed cases. Source: &lt;a href="https://www.endorlabs.com/learn/opus-5-5-6x-cheaper-and-2x-faster-than-fable-5-1-but-memorization-keeps-it-off-the-top-spot" rel="noopener noreferrer"&gt;Endor Labs&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On their full leaderboard, that puts Opus 5.5 third for secure code, behind Claude Code with Fable 5.1 and OpenAI's Codex with GPT-6 Astra (34.6%). For code that just has to work, it's nineteenth. Endor points out that this doesn't line up with Anthropic calling it a new state of the art for coding, at least not on their benchmark.&lt;/p&gt;

&lt;p&gt;The number that stuck with me isn't even about Opus 5.5. Fable 5.1, at the top of their board, writes working and secure code on about 37% of these tasks. So whatever model you use, you still need to check the security side yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The memorization thing
&lt;/h2&gt;

&lt;p&gt;This is the part that makes Endor's results worth reading closely. Their tasks come from real, public security fixes, so a model might have seen the answer during training. Endor counts recalling a known fix as cheating, whether it comes from git history, the web or the model's own memory. A pipeline flags suspicious runs, an AI judge reviews each one, and confirmed cases don't count.&lt;/p&gt;

&lt;p&gt;For Opus 5.5 they confirmed 51 of those, more than for any other model they've tested. If those counted, it would score 94.4% FuncPass and 52.5% SecPass and be first on security by more than 11 points. Without them it's 68.7% and 33.5%. For comparison, taking out memorized answers cost Fable 5.1 6.7 and 3.9 points.&lt;/p&gt;

&lt;p&gt;In one case the model gave itself away by naming the vulnerability's CVE number (its public ID) before it made its first edit. Endor says they added a check for exactly that afterwards.&lt;/p&gt;

&lt;p&gt;They also found the memorized tasks were the quick ones: 3.0 minutes on average against 4.4 for the rest. Makes sense. If the model recognizes the task, it doesn't need to read much before writing the fix.&lt;/p&gt;

&lt;p&gt;My take: a benchmark built from public data can end up measuring memory as much as skill, and it takes real work to separate the two. Most benchmarks don't publish this kind of breakdown. It's one more reason to try a model on your own code, which it can't have seen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it does well: speed and cost
&lt;/h2&gt;

&lt;p&gt;On speed and cost, Endor found Anthropic's launch claims hold up. Opus 5.5 finished a task in 2.2 minutes at the median and 4.0 on average. Fable 5.1 averaged 9.5 minutes, and Opus 5 averaged 14.1 with 15 timeouts. Opus 5.5 had none.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqm6tb58xt4w10qzn99pn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqm6tb58xt4w10qzn99pn.png" alt="Bar chart from Endor Labs, prediction duration distribution. Opus 5.5 finishes 157 coding tasks in under five minutes, against 74 for Fable 5.1 and 53 for Opus 5. Opus 5.5 has no task above 40 minutes and a duration record for every run; Opus 5 has runs up to 60 minutes and 5 without a record." width="799" height="453"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;How long each coding task took, in Claude Code. Chart: &lt;a href="https://www.endorlabs.com/learn/opus-5-5-6x-cheaper-and-2x-faster-than-fable-5-1-but-memorization-keeps-it-off-the-top-spot" rel="noopener noreferrer"&gt;Endor Labs&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The cost gap is bigger still. The full Opus 5.5 run cost $116, or $0.33 per task at the median. The same run cost $672 with Fable 5.1 and about $1,116 with Opus 5. According to Endor, about half the gap to Fable comes from the lower price per token and the rest from using fewer of them: about a third of Fable's output tokens, and fewer than half its tool calls (16.6 per task against 42.4).&lt;/p&gt;

&lt;p&gt;That lines up with what Anthropic's docs promised, and honestly it's the part I care about most day to day. Fewer steps means less to read, less to check, and less room for the task to drift.&lt;/p&gt;

&lt;h2&gt;
  
  
  So do they disagree?
&lt;/h2&gt;

&lt;p&gt;Not as much as it looks. They measure different things. Artificial Analysis looks at general ability across ten tests, at max effort, with a model from OpenAI in the mix. Endor looks at one hard, narrow job, secure fixes in real projects, and makes a point of throwing out answers the model memorized.&lt;/p&gt;

&lt;p&gt;Where they overlap, they agree: Opus 5.5 is cheaper to run than Fable 5.1 and good at coding. Whether it's the best depends on the task, the effort setting, and whether you count remembered answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm taking from it
&lt;/h2&gt;

&lt;p&gt;If speed and cost matter to you, and for most everyday coding they do, Opus 5.5 seems like a reasonable default. Both evaluations back that up.&lt;/p&gt;

&lt;p&gt;If you work on security-sensitive code, nothing on Endor's board is close to safe to use without review. Keep your tests, and have someone look at anything that touches logins, user input or permissions.&lt;/p&gt;

&lt;p&gt;And if you're choosing between models, a few of your own tasks where you know what good looks like will tell you more than a leaderboard. Leaderboards are a fine place to start, just not to stop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heise.de/en/news/External-benchmarks-Claude-Opus-5-5-overtakes-OpenAI-s-Astra-and-Fable-5-1-11466259.html" rel="noopener noreferrer"&gt;External benchmarks: Claude Opus 5.5 overtakes OpenAI's Astra and Fable 5.1&lt;/a&gt;, heise online, September 25, 2026. Reports the Artificial Analysis results; both Artificial Analysis charts above are from this article.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://artificialanalysis.ai/" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;, publisher of the Intelligence Index.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.endorlabs.com/learn/opus-5-5-6x-cheaper-and-2x-faster-than-fable-5-1-but-memorization-keeps-it-off-the-top-spot" rel="noopener noreferrer"&gt;Opus 5.5: 6x cheaper and 2x faster than Fable 5.1, but only 33.5% of code is secure&lt;/a&gt;, Endor Labs, September 24, 2026. Source of the Agent Security League results and the task-duration chart.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://promptquick.ai/blog/opus-5-5-independent-benchmarks" rel="noopener noreferrer"&gt;promptquick.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>benchmark</category>
    </item>
    <item>
      <title>What Anthropic Says Has Improved in Opus 5.5</title>
      <dc:creator>Nomad</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:42:11 +0000</pubDate>
      <link>https://dev.to/technomadcode/what-anthropic-says-has-improved-in-opus-55-1mod</link>
      <guid>https://dev.to/technomadcode/what-anthropic-says-has-improved-in-opus-55-1mod</guid>
      <description>&lt;p&gt;I went back to Anthropic's docs to see what 5.5 changes about the verbosity, repeated checks, and unnecessary delegation I wrote about with Opus 5.&lt;/p&gt;

&lt;p&gt;I previously &lt;a href="https://www.reddit.com/r/ClaudeAI/comments/1vd57c0/claudemd_for_opus_5_based_on_anthropics_official/" rel="noopener noreferrer"&gt;shared a CLAUDE.md file on Reddit&lt;/a&gt; after reading Anthropic's documentation and noticing how many of the complaints about Opus 5 had explanations sitting right there. The verbosity. The extra verification. The horrible army of subagents. People were describing frustrating behavior, and Anthropic was describing ways to tune it.&lt;/p&gt;

&lt;p&gt;That was what made the docs interesting to me. A model can be good at difficult coding tasks and still make the person working with it want to close the laptop. Having to fight its habits becomes part of the job.&lt;/p&gt;

&lt;p&gt;So I went through the &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5" rel="noopener noreferrer"&gt;Opus 5.5 prompting guide&lt;/a&gt; to see what changed. This time, some of the improvements Anthropic describes are in the same areas that made 5 annoying. There are also a few passages that make me reluctant to declare the whole problem solved.&lt;/p&gt;

&lt;p&gt;This is a reading of their documentation, with my interpretation of what matters in practice. It is not a benchmark of my configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Opus 5 complaints were not imaginary
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5" rel="noopener noreferrer"&gt;Opus 5 guide&lt;/a&gt; describes longer replies, a greater tendency to delegate, and task expansion. On checking its own work, it says the model:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;verifies its own work without being told to&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The same section warns that adding generic final checks or automatic reviewer agents can cause excessive verification. That is quite a thing to discover after trying to make an agent more reliable by telling it to double-check everything.&lt;/p&gt;

&lt;p&gt;It also explains why the advice could feel backwards. You add an instruction because you want careful work. Then the model changes, and the instruction contributes to the behavior you are trying to stop. Meanwhile, you are paying for the extra work and reading the extra explanation.&lt;/p&gt;

&lt;p&gt;Some of those habits sucked. Calling them model defaults does not make them pleasant.&lt;/p&gt;

&lt;p&gt;But removing an automatic extra review is different from removing your project's actual tests. If a payment change needs an integration test, it still needs that test. The useful distinction is whether a check can establish something important, or whether another pass has become a ritual.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fewer tokens would be a much better improvement than more enthusiasm
&lt;/h2&gt;

&lt;p&gt;Anthropic reports over 30% faster output-token generation, usually fewer tokens per task, and results at medium effort that matched or beat Opus 5 at high in its coding and knowledge-work tests. The phrase that caught my attention was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;in fewer steps and with fewer tokens&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is from the &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#capabilities-relevant-to-prompting" rel="noopener noreferrer"&gt;5.5 guide's capability discussion&lt;/a&gt;. It is much closer to what I want from an upgrade than another promise that the model will work harder.&lt;/p&gt;

&lt;p&gt;I want the right work done. If an agent reaches a sound result with less wandering, there is less output to inspect and fewer opportunities for the task to drift. Spending longer on something is easy to mistake for doing it better, especially when every extra step comes with a confident explanation.&lt;/p&gt;

&lt;p&gt;There are three different things to watch here: the speed at which tokens appear, the amount of work the agent performs, and the length of the answer you read. A faster stream of tokens does not automatically make a task finish 30% sooner. Tool calls and waiting still exist. Fewer tokens across a task also does not guarantee a shorter final reply.&lt;/p&gt;

&lt;p&gt;I would judge this improvement by completed work, the evidence behind it, and how much intervention it needed. A fast answer that leaves the difficult part for me is not a win.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your old effort setting may be part of the problem
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/build-with-claude/effort#recommended-effort-levels-for-claude-opus-5-5" rel="noopener noreferrer"&gt;effort documentation&lt;/a&gt; identifies a concrete change: Opus 5.5 defaults to medium, while 5 defaults to high. Thinking is always enabled on 5.5, so effort is the main control to reconsider.&lt;/p&gt;

&lt;p&gt;There is a catch. The &lt;a href="https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5#behavior-differences" rel="noopener noreferrer"&gt;5.5 behavior notes&lt;/a&gt; say the same effort label can produce more thinking per turn than it did on 5, especially at xhigh and max.&lt;/p&gt;

&lt;p&gt;That means leaving the old setting in place is not necessarily keeping the old behavior. If you turn everything up, see a long run, and conclude the new model is just as wasteful, you might be comparing two different amounts of work under the same label.&lt;/p&gt;

&lt;p&gt;I would start with medium on a few tasks I understand well. Then I would raise it where the results justify doing so. A tricky migration and a small copy change do not need to earn their credibility by consuming the same amount of time.&lt;/p&gt;

&lt;p&gt;In Claude Code, this is an actual setting you inspect and change through the &lt;a href="https://code.claude.com/docs/en/model-config#adjust-effort-level" rel="noopener noreferrer"&gt;effort controls&lt;/a&gt;. Writing a wish for less thinking into CLAUDE.md is not the same operation. I still want a concise communication preference in the file, because the reasoning budget and the answer I have to read are different concerns.&lt;/p&gt;

&lt;h2&gt;
  
  
  A better reviewer should create less work for the human
&lt;/h2&gt;

&lt;p&gt;Early testers reported more real review findings and fewer false alarms, according to the &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#capabilities-relevant-to-prompting" rel="noopener noreferrer"&gt;capability discussion&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I am particularly interested in the false alarms. A review that produces twenty findings is not automatically better than one that produces five. Every finding asks someone to investigate it, understand it, and decide whether changing the code would help.&lt;/p&gt;

&lt;p&gt;For me, a useful finding has a failure scenario I can follow. What triggers it? What goes wrong? Where does the code allow that to happen? If I cannot answer those questions after reading the finding, more severity labels and longer explanations do not help much.&lt;/p&gt;

&lt;p&gt;Early tester reports are encouraging, but they are not a controlled result from my own repositories. I would want to see whether 5.5 catches a real problem with less noise around it. That would address a much more expensive kind of verbosity than a long final paragraph.&lt;/p&gt;

&lt;h2&gt;
  
  
  Better communication can still look like silence
&lt;/h2&gt;

&lt;p&gt;The guide describes reports that:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;say plainly what it did&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is its &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#capabilities-relevant-to-prompting" rel="noopener noreferrer"&gt;communication claim&lt;/a&gt;. There is also an application detail that is easy to miss if you only skim for model improvements.&lt;/p&gt;

&lt;p&gt;In an API integration, 5.5's progress notes can arrive in thinking blocks. With the default display setting, their text is hidden. A client that only displays ordinary text can therefore appear silent between tool calls. Anthropic documents an updates display mode, currently in beta, for exposing those notes. See &lt;a href="https://platform.claude.com/docs/en/build-with-claude/thinking#progress-updates-between-tool-calls" rel="noopener noreferrer"&gt;progress updates between tool calls&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That is something the application has to handle. It is not evidence, by itself, that the model has become less communicative, and a paragraph in CLAUDE.md cannot make an application render content it is ignoring.&lt;/p&gt;

&lt;p&gt;This is why I am cautious about taking a symptom and immediately adding another rule. If a model is doing the work but the interface is not showing the updates, telling it to narrate more may be aimed at the wrong problem.&lt;/p&gt;

&lt;p&gt;As a user, I want enough information to understand what is happening. A short explanation when a finding changes the approach is useful. A running commentary on every ordinary file read usually is not. Those are the preferences I want to express, without turning progress reporting into a second task.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part where it still stops before the work is done
&lt;/h2&gt;

&lt;p&gt;This passage is less exciting than the performance claims:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;some of those updates end the turn&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is from &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#unattended-agentic-runs" rel="noopener noreferrer"&gt;unattended agentic runs&lt;/a&gt;. An unattended loop can stop prematurely. The continuation advice limits automatic follow-ups to two or three, keeps risky actions approval-gated, and excludes interactive use from its long system-prompt example.&lt;/p&gt;

&lt;p&gt;The distinction matters because &lt;a href="https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons#end-turn" rel="noopener noreferrer"&gt;an ended response&lt;/a&gt; tells you that generation stopped. It is not a certificate that every part of your project is complete.&lt;/p&gt;

&lt;p&gt;Imagine asking for a migration of four endpoints. The agent finishes two and writes a tidy update announcing the other two. If the surrounding program treats that update as completion, the remaining work sits there. A very readable status report has still left you with half a migration.&lt;/p&gt;

&lt;p&gt;My preference is straightforward: finish what is authorized, make unfinished work visible, and identify real blockers. I do not want to type continue because the agent reached a convenient paragraph break. I also do not want an automatic loop repeatedly attempting an action that needs my approval.&lt;/p&gt;

&lt;p&gt;This is one of the clearest reasons to read the limitations alongside the improvement claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  About that army of subagents
&lt;/h2&gt;

&lt;p&gt;For agent teams, the guide reports faster completion with elapsed-time or deadline signals, while warning that verification may suffer under time pressure. See &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#time-signals-for-multi-agent-harnesses" rel="noopener noreferrer"&gt;time signals for multi-agent harnesses&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That addresses how a team works once delegation makes sense. I would still decide whether the task needs a team in the first place.&lt;/p&gt;

&lt;p&gt;Independent investigations across a large codebase are a reasonable candidate. A small change with two obvious callers probably does not need several agents reporting their interpretations back to another agent. Coordination is work too, and faster completion does not by itself establish lower total cost.&lt;/p&gt;

&lt;p&gt;There is also a difference between giving an agent a target and enforcing a limit. Anthropic's separate &lt;a href="https://platform.claude.com/docs/en/build-with-claude/task-budgets#task-budgets-are-advisory-not-enforced" rel="noopener noreferrer"&gt;task-budget documentation&lt;/a&gt; makes that distinction explicit: advisory budgets can be exceeded, and a per-request token limit is not a limit on an entire multi-request job.&lt;/p&gt;

&lt;p&gt;I am keeping the delegation boundaries in my configuration. Better teamwork would be welcome, but it does not answer the original complaint about involving a team in work that did not need one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Better eyesight does not automatically give it better taste
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5#behavior-differences" rel="noopener noreferrer"&gt;5.5 model notes&lt;/a&gt; describe more accurate interpretation of charts, diagrams, and screenshots, and suggest reassessing older visual workarounds. Image tools can still help with demanding inputs.&lt;/p&gt;

&lt;p&gt;That matters for development. Misreading a layout is a different failure from writing the wrong CSS to reproduce it. If the first interpretation is wrong, competent implementation can take you further in the wrong direction. I would check the model's reading of a reference before judging the page it builds from it.&lt;/p&gt;

&lt;p&gt;Taste is another issue. On generic frontend styling, Anthropic says a vague request to avoid the usual AI look:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;mostly swaps one default for another&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is from &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#frontend-design-defaults" rel="noopener noreferrer"&gt;frontend design defaults&lt;/a&gt;. It explains why I prefer a real reference and concrete feedback to an instruction that amounts to make it less annoying.&lt;/p&gt;

&lt;p&gt;Name the element that is wrong. Point to the spacing, the button shape, the headline treatment, or the hierarchy. Otherwise the model is still guessing what your taste means, even if it can now inspect the result more accurately.&lt;/p&gt;

&lt;h2&gt;
  
  
  A pasted document still should not become the boss
&lt;/h2&gt;

&lt;p&gt;The guide recommends application-generated boundaries around pasted source content and treats them as one defense against embedded instructions. See &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#mark-pasted-text-in-user-messages" rel="noopener noreferrer"&gt;pasted text in user messages&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For my workflow, the principle is simple. If I give an agent an article to analyze, commands inside that article are part of the material. They are not automatically my instructions. This matters even with ordinary, non-malicious material: a setup guide can contain destructive commands that I want explained rather than executed.&lt;/p&gt;

&lt;p&gt;A reminder in my rules file makes that expectation explicit. It does not implement a security boundary in the application. I want the model's judgment and the surrounding permission controls to support the same distinction.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed in my CLAUDE.md
&lt;/h2&gt;

&lt;p&gt;The advice to simplify context did not suddenly arrive with 5.5. Anthropic's earlier &lt;a href="https://claude.dev/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models/" rel="noopener noreferrer"&gt;context engineering article&lt;/a&gt; already described removing over 80% of Claude Code's system prompt for models including Opus 5 without measurable loss on its coding evaluations. It recommended concise project context and loading specialized guidance when needed.&lt;/p&gt;

&lt;p&gt;I have made a &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit/tree/main/agent-configs/claude-code-desktop/claude-opus-5-5" rel="noopener noreferrer"&gt;separate Opus 5.5 configuration&lt;/a&gt; in my free AI Product Development Toolkit. It keeps the preferences that matter to me: stay within scope, communicate plainly, preserve useful checks, and ask before destructive or shared actions. The hard sentence limit and some repeated instructions are gone.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/TechNomadCode" rel="noopener noreferrer"&gt;
        TechNomadCode
      &lt;/a&gt; / &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;
        AI-Product-Development-Toolkit
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Plan, build and launch products with AI: guided planning prompts, an AI App Starter for Next.js, Supabase and Vercel, and agent configurations.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;AI Product Development Toolkit&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Plan, build and launch products with AI, without the mess.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit/./LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/b8cadaa967891081f8f165695470689986c028821dd8a040132f6e661795dc0d/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d626c7565" alt="MIT license"&gt;&lt;/a&gt;
&lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/bb13ea0cda91b4b727d9cf100cbd4c1a206302baa3355b13f087a1819d29cce8/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436c61756465253230436f64652d443937373537"&gt;&lt;img src="https://camo.githubusercontent.com/bb13ea0cda91b4b727d9cf100cbd4c1a206302baa3355b13f087a1819d29cce8/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436c61756465253230436f64652d443937373537" alt="Works with Claude Code"&gt;&lt;/a&gt;
&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#tested-so-far" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/8d09579d9b5797997037f0d7750b9116a5e11e3a35ea9513ffc18a44e6e08dc5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436f6465782532302862657461292d343132393931" alt="Works with Codex (beta)"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Use it when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#planning-prompts" rel="noopener noreferrer"&gt;Planning prompts&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Guided prompts that interview you, then write your plan. Any AI chat.&lt;/td&gt;
&lt;td&gt;You have an idea and want a solid plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#ai-app-starter" rel="noopener noreferrer"&gt;AI App Starter&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A project your coding agent builds with you, step by step, up to launch.&lt;/td&gt;
&lt;td&gt;You want to build and launch a web app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#agent-configurations" rel="noopener noreferrer"&gt;Agent configurations&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Working instructions for Claude Code.&lt;/td&gt;
&lt;td&gt;You code with Claude Code every day&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The planning prompts and the AI App Starter also install as &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#skills" rel="noopener noreferrer"&gt;skills&lt;/a&gt; for Claude Code and Codex, with one line, together with a skill that makes short product videos in code.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Planning prompts&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Most prompts tell the AI to write a document straight away, so it fills every gap with guesses. These prompts make it &lt;strong&gt;interview you first&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It asks before it writes.&lt;/strong&gt; A few targeted…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;My Context7 requirement stays. For dependency-related work, I want current documentation checked before the first edit, including implementation patterns and configuration. Remembering an API is a reason to verify it. Only when Context7 has no relevant coverage should the agent use curl to retrieve official documentation and read it directly; web search and web fetch are excluded for those lookups.&lt;/p&gt;

&lt;p&gt;That is my preference, not an Anthropic requirement. Read the configuration and its README before adopting it, and retain the instructions that your own project actually needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What would convince me that 5.5 is better to work with
&lt;/h2&gt;

&lt;p&gt;I would compare a few tasks where I can judge the outcome: a small bug fix, a review with a reproducible failure, a change that crosses several files, and a frontend task with a clear reference.&lt;/p&gt;

&lt;p&gt;Then I would look at the corrections I had to make. Did the agent stay within the requested scope? Did it finish the awkward parts? Were its findings real? Could I understand the result without digging through a long report? Did the time and token use make sense for the work?&lt;/p&gt;

&lt;p&gt;Those are the improvements that would change my working day. More intelligence is welcome, but I do not want to spend the benefit supervising a more elaborate process.&lt;/p&gt;

&lt;p&gt;It was funny reading the Opus 5 docs because they explained so many of the things people disliked. Reading the 5.5 guide is more encouraging. It gives me specific claims to test in those same areas, alongside some very useful reasons to remain skeptical.&lt;/p&gt;

&lt;p&gt;If you went back to 4.8 or Fable because Opus 5 was exhausting, that seems like a reasonable basis for giving 5.5 another look. Whether it earns a place in your workflow should depend on what happens when you use it.&lt;/p&gt;

&lt;p&gt;The configuration is free in the toolkit. For the everyday prompts that do not need a full coding setup, the Prompt Rulebook sample is another place to start.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://promptquick.ai/signup" class="crayons-btn crayons-btn--primary" rel="noopener noreferrer"&gt;Get the Free Sample&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" class="crayons-btn crayons-btn--primary" rel="noopener noreferrer"&gt;Open the Toolkit on GitHub&lt;/a&gt;
&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://promptquick.ai/blog/opus-5-5-improvements-and-prompting" rel="noopener noreferrer"&gt;promptquick.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>promptengineering</category>
      <category>llm</category>
    </item>
    <item>
      <title>How I Turn Ideas Into Products With AI Prompts</title>
      <dc:creator>Nomad</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:41:51 +0000</pubDate>
      <link>https://dev.to/technomadcode/how-i-turn-ideas-into-products-with-ai-prompts-327h</link>
      <guid>https://dev.to/technomadcode/how-i-turn-ideas-into-products-with-ai-prompts-327h</guid>
      <description>&lt;p&gt;I'm an engineer. Before I build anything, I want to know what I'm building, who it's for and why.&lt;/p&gt;

&lt;p&gt;So when AI tools showed up, I never just typed one line at them. I gave them what I'd give a new teammate: the full picture.&lt;/p&gt;

&lt;p&gt;What surprised me was how rarely other people did. I kept seeing the same thing: someone types one line into ChatGPT, gets something generic back, and decides AI just isn't that good.&lt;/p&gt;

&lt;p&gt;It is. It was just guessing.&lt;/p&gt;

&lt;p&gt;The people getting great results from AI aren't using secret tools. They're just better at explaining what they want. That skill has a fancy name now, &lt;em&gt;context engineering&lt;/em&gt;, but the idea behind it is simple, and anyone can learn it. I want to help more people get good at it, so this post is about the prompt I use to turn a rough idea into a clear plan for a real product, and why it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI is guessing, and it's a very confident guesser
&lt;/h2&gt;

&lt;p&gt;Take a request like “Build me a habit tracker with reminders.” The AI will happily come back with a full plan, a feature list and a pile of code. To get there, it fills every gap with the most likely answer. Who's the app for? It guesses. What's the most important feature? It guesses. Phone or website? Free or paid? Guess, guess, guess.&lt;/p&gt;

&lt;p&gt;The trouble is that it never tells you it's guessing. It answers in the same calm, confident voice it uses for everything. So you get something that looks finished but sits on a pile of assumptions you never agreed to.&lt;/p&gt;

&lt;p&gt;That's what “context” means here: everything the AI needs to know to stop guessing. Your goal, your users, your limits, what you've already decided. Give it that, and the answers change completely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vibe prompting is fine, until it isn't
&lt;/h2&gt;

&lt;p&gt;There's a popular way of working with AI that people call &lt;em&gt;vibe prompting&lt;/em&gt;. You type whatever comes to mind, see what comes back, and nudge it until it looks right. No plan, just vibes.&lt;/p&gt;

&lt;p&gt;For small things, it works fine. Rewriting an email, brainstorming names, explaining a confusing paragraph. (That's also what the &lt;a href="https://promptquick.ai/" rel="noopener noreferrer"&gt;Prompt Rulebook&lt;/a&gt; is for: quick copy-paste rules that make those everyday prompts better.)&lt;/p&gt;

&lt;p&gt;But try to vibe your way from an idea to an actual product, and it falls apart. The AI forgets things you told it an hour earlier. It adds features you never asked for. It contradicts itself between messages. And because nothing is written down, you can't even tell whether the problem is the AI or you changing your mind.&lt;/p&gt;

&lt;p&gt;The fix is to flip it around. Instead of telling the AI what to build, you let the AI ask you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The prompt I start every project with
&lt;/h2&gt;

&lt;p&gt;It's a free prompt from my &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;AI Product Development Toolkit&lt;/a&gt; on GitHub. Its official name is the &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit/blob/main/prompt-templates/prd-generation/prd-generation.md" rel="noopener noreferrer"&gt;PRD prompt&lt;/a&gt;. PRD stands for Product Requirements Document, which is just a written plan: what you're building, who it's for and what it needs to do. I think of it as the brain-dump interview, because that's what it feels like.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/TechNomadCode" rel="noopener noreferrer"&gt;
        TechNomadCode
      &lt;/a&gt; / &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;
        AI-Product-Development-Toolkit
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Plan, build and launch products with AI: guided planning prompts, an AI App Starter for Next.js, Supabase and Vercel, and agent configurations.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;AI Product Development Toolkit&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Plan, build and launch products with AI, without the mess.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit/./LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/b8cadaa967891081f8f165695470689986c028821dd8a040132f6e661795dc0d/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d626c7565" alt="MIT license"&gt;&lt;/a&gt;
&lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/bb13ea0cda91b4b727d9cf100cbd4c1a206302baa3355b13f087a1819d29cce8/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436c61756465253230436f64652d443937373537"&gt;&lt;img src="https://camo.githubusercontent.com/bb13ea0cda91b4b727d9cf100cbd4c1a206302baa3355b13f087a1819d29cce8/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436c61756465253230436f64652d443937373537" alt="Works with Claude Code"&gt;&lt;/a&gt;
&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#tested-so-far" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/8d09579d9b5797997037f0d7750b9116a5e11e3a35ea9513ffc18a44e6e08dc5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436f6465782532302862657461292d343132393931" alt="Works with Codex (beta)"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Use it when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#planning-prompts" rel="noopener noreferrer"&gt;Planning prompts&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Guided prompts that interview you, then write your plan. Any AI chat.&lt;/td&gt;
&lt;td&gt;You have an idea and want a solid plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#ai-app-starter" rel="noopener noreferrer"&gt;AI App Starter&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A project your coding agent builds with you, step by step, up to launch.&lt;/td&gt;
&lt;td&gt;You want to build and launch a web app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#agent-configurations" rel="noopener noreferrer"&gt;Agent configurations&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Working instructions for Claude Code.&lt;/td&gt;
&lt;td&gt;You code with Claude Code every day&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The planning prompts and the AI App Starter also install as &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#skills" rel="noopener noreferrer"&gt;skills&lt;/a&gt; for Claude Code and Codex, with one line, together with a skill that makes short product videos in code.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Planning prompts&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Most prompts tell the AI to write a document straight away, so it fills every gap with guesses. These prompts make it &lt;strong&gt;interview you first&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It asks before it writes.&lt;/strong&gt; A few targeted…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;You copy it, paste it into a good AI chat tool like ChatGPT, Claude or Gemini, and here's what happens.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. You dump everything you've got
&lt;/h3&gt;

&lt;p&gt;The prompt has a space where you paste your notes. All of them. Half-sentences, features you're unsure about, problems you've noticed, the thing your friend said over coffee. It doesn't need to be tidy. That's the point.&lt;/p&gt;

&lt;p&gt;If you're staring at the empty space, run through these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who is it for?&lt;/li&gt;
&lt;li&gt;What problem does it solve for them?&lt;/li&gt;
&lt;li&gt;Which features are you already picturing?&lt;/li&gt;
&lt;li&gt;What are your limits? Time, money, skills.&lt;/li&gt;
&lt;li&gt;What are you unsure about? Write that down too. The AI can help you decide.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Say you want to make a simple app for your book club. Your brain dump might be four lines: people forget which book we're reading, we never agree on a date, it would be nice to vote on the next book, and maybe share notes.&lt;/p&gt;

&lt;p&gt;That's plenty to start.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The AI interviews you
&lt;/h3&gt;

&lt;p&gt;This is where it gets different. The prompt tells the AI not to write anything yet. First it has to ask you the one to three most important questions.&lt;/p&gt;

&lt;p&gt;For the book club, you'd get questions like these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How many people are in the club?&lt;/li&gt;
&lt;li&gt;Does everyone use a smartphone, or would some people prefer email?&lt;/li&gt;
&lt;li&gt;When you say “vote”, is it one vote each, or ranking a few books?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Questions you probably hadn't thought about yet. And every answer you give makes the next question sharper.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. It checks in before it moves on
&lt;/h3&gt;

&lt;p&gt;The prompt also tells the AI to stop and check with you whenever it's about to change topic or interpret something you said. You'll see things like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;My understanding is that members vote once a month and the book with the most votes wins. Is that right?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the part that saves you. A wrong assumption gets caught while it's still one sentence long, not after it has shaped the whole plan.&lt;/p&gt;

&lt;p&gt;On top of that, the AI has to say its assumptions out loud, ask for numbers where numbers matter (“How many members do you expect in the first year?”), and point out when two things you said don't match.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. You get a plan you actually agree with
&lt;/h3&gt;

&lt;p&gt;Only when you agree that enough has been covered does the AI offer to write the document. It comes out in clear sections: the goals, who it's for, what people need to be able to do, what success looks like, and the questions you still need to answer.&lt;/p&gt;

&lt;p&gt;For the book club, the goals part might read something like this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every member knows the current book and the next meeting date without asking in the group chat. The next book is picked by vote within one week, and at least 8 of 10 members vote.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Notice the numbers. They came from the interview, when the AI asked what success would look like. A one-line prompt would never have produced them, because the AI had no way to know.&lt;/p&gt;

&lt;p&gt;It's a draft, not gospel. But it's &lt;em&gt;your&lt;/em&gt; draft, built from your answers. And from then on, every AI conversation about the project can start from that document instead of one vague sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can steal from this prompt, even if you never build an app
&lt;/h2&gt;

&lt;p&gt;You don't need to be building a product to use the ideas behind this prompt. Here's what makes it work, and all of it carries over to everyday prompting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Give the AI your mess.&lt;/strong&gt; Messy notes beat a neat one-liner. More context beats prettier context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make it ask before it answers.&lt;/strong&gt; Add “Before you start, ask me the questions you need answered to do this well” to almost any prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make it say its assumptions out loud.&lt;/strong&gt; “List any assumptions you're making” turns hidden guesses into things you can correct.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One topic at a time.&lt;/strong&gt; Big, sprawling conversations drift. Small steps that you confirm don't.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's context engineering in plain English: give the AI what it needs to know, in an order it can handle, and check that it understood you.&lt;/p&gt;

&lt;h2&gt;
  
  
  A fair warning: this is not a shortcut
&lt;/h2&gt;

&lt;p&gt;I want to be straight with you here, because the internet is full of “build an app in 10 minutes with AI” promises.&lt;/p&gt;

&lt;p&gt;These templates are the opposite of that. They're systematic, and that takes time. The interview can go on for a while, and it will ask questions that make you stop and think. Sometimes you'll realize you don't know the answer yet. That's useful too.&lt;/p&gt;

&lt;p&gt;If you use the toolkit, be ready for a deep dive. Treat it as a way to think your idea through properly, not as a magic button. The effort is still yours.&lt;/p&gt;

&lt;p&gt;But the results can be wonderful. You come out with a clear plan, fewer surprises, and an AI that actually understands what you're building. For me, that trade is worth it.&lt;/p&gt;

&lt;p&gt;One more thing: whatever the AI writes is a draft. Read it, fix it, and leave out private information you wouldn't want to share with an AI tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the toolkit, in one line each
&lt;/h2&gt;

&lt;p&gt;The brain-dump interview is where I start, but the toolkit has more prompts that pick up where it leaves off. They all work the same way: the AI asks, you answer, it checks in.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ask real people first:&lt;/strong&gt; builds a fair survey, so you can check that people want your idea before you build it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Map the screens:&lt;/strong&gt; turns your plan into the steps and screens someone goes through when they use your product.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Find the smallest useful version:&lt;/strong&gt; cuts your idea down to the smallest first version worth building.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan your testing:&lt;/strong&gt; makes a checklist to find out whether what you built actually works.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They're all free on &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;And if you build with Claude Code, the &lt;a href="https://promptquick.ai/blog/ai-app-starter-nextjs-supabase-vercel" rel="noopener noreferrer"&gt;AI App Starter&lt;/a&gt; gives it a step-by-step process from idea to launch. Its first step is this same interview.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to start
&lt;/h2&gt;

&lt;p&gt;If there's an idea you keep coming back to, open the PRD prompt, paste in your notes, and let it interview you. Give it an evening. See what comes out.&lt;/p&gt;

&lt;p&gt;And if you want to get better at the everyday prompts, the quick ones you type ten times a day, that's why I wrote the Prompt Rulebook. The free sample is a good place to start.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://promptquick.ai/signup" class="crayons-btn crayons-btn--primary" rel="noopener noreferrer"&gt;Get the Free Sample&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" class="crayons-btn crayons-btn--primary" rel="noopener noreferrer"&gt;Open the Toolkit on GitHub&lt;/a&gt;
&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://promptquick.ai/blog/turn-ideas-into-products-with-ai" rel="noopener noreferrer"&gt;promptquick.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>promptengineering</category>
      <category>opensource</category>
    </item>
    <item>
      <title>What Anthropic and Artificial Analysis Say About Sonnet 5.5</title>
      <dc:creator>Nomad</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:37:23 +0000</pubDate>
      <link>https://dev.to/technomadcode/what-anthropic-and-artificial-analysis-say-about-sonnet-55-254p</link>
      <guid>https://dev.to/technomadcode/what-anthropic-and-artificial-analysis-say-about-sonnet-55-254p</guid>
      <description>&lt;p&gt;Sonnet 5.5 came out on September 28. In Anthropic's own table it sits a few points behind Opus 5.5 on almost every test. The more interesting parts are a footnote under that table and what Artificial Analysis found when it ran the model at max effort.&lt;/p&gt;

&lt;p&gt;Over the last week I wrote about &lt;a href="https://promptquick.ai/blog/opus-5-5-improvements-and-prompting" rel="noopener noreferrer"&gt;what Anthropic's docs say improved in Opus 5.5&lt;/a&gt; and then about &lt;a href="https://promptquick.ai/blog/opus-5-5-independent-benchmarks" rel="noopener noreferrer"&gt;the first independent benchmarks for it&lt;/a&gt;. Same routine here: what Anthropic says, then what an outside tester found, then what the prompting guide says to do about it.&lt;/p&gt;

&lt;p&gt;Quick disclaimer: I didn't run any of these tests. The numbers come from Anthropic and &lt;a href="https://artificialanalysis.ai/models/claude-sonnet-5-5" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;, and I link each one to where I found it. I drew the charts myself from their numbers. The opinions are mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Anthropic says
&lt;/h2&gt;

&lt;p&gt;Anthropic's &lt;a href="https://x.com/claudeai/status/2104633119412523294" rel="noopener noreferrer"&gt;thread on X&lt;/a&gt; and &lt;a href="https://www.anthropic.com/claude-sonnet-5-5" rel="noopener noreferrer"&gt;launch post&lt;/a&gt; say Sonnet 5.5 is 30%+ faster than Sonnet 5 and costs up to 30% less per task. They call it a faster, lower-cost complement to Opus 5.5, strongest at well-scoped everyday work, bug fixes, and documents, slides and spreadsheets.&lt;/p&gt;

&lt;p&gt;It also says outright that benchmark scores only show one side of a model, and that in its own testing and its outside testers' testing, Opus 5.5 is still clearly stronger at complex, open-ended work that needs sustained judgment. So the table below isn't “Sonnet equals Opus.”&lt;/p&gt;

&lt;p&gt;The price per token didn't change from Sonnet 5: $2 per million input tokens and $10 per million output tokens. Opus 5.5 is $4 and $20. So the “up to 30% less” isn't a price cut. Anthropic says it comes from Sonnet 5.5 needing far fewer tokens to do the same work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anthropic's table
&lt;/h2&gt;

&lt;p&gt;Here is the benchmark table from the launch post, retyped. It's all Anthropic's reporting, even though two rows are Artificial Analysis's results and one column is a competitor.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Sonnet 5.5&lt;/th&gt;
&lt;th&gt;Sonnet 5&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;th&gt;GPT-6 Sol&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 4.0 (agentic coding)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;70.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10.3%&lt;/td&gt;
&lt;td&gt;66.4%&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierCode 1.1 (agentic coding, main set)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;52.1% (xhigh), 46.2% (max)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;42.4%&lt;/td&gt;
&lt;td&gt;54.4%&lt;/td&gt;
&lt;td&gt;49.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CursorBench 4.0 (agentic coding)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;55.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;34.1%&lt;/td&gt;
&lt;td&gt;57.8%&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDPval-AA v2.1 (knowledge work, Elo)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1844&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1449&lt;/td&gt;
&lt;td&gt;1846&lt;/td&gt;
&lt;td&gt;1487&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AA-Briefcase v1.1 (knowledge work, Elo)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1811&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1359&lt;/td&gt;
&lt;td&gt;1822&lt;/td&gt;
&lt;td&gt;1483&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Humanity's Last Exam (reasoning, with tools)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;64.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;54.9%&lt;/td&gt;
&lt;td&gt;67.7%&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OSWorld 2.1 (computer use, partial)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;80.1%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;57.0%&lt;/td&gt;
&lt;td&gt;81.8%&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chartography (chart reading, no tools)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;61.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;15.6%&lt;/td&gt;
&lt;td&gt;64.4%&lt;/td&gt;
&lt;td&gt;53.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Anthropic's numbers, retyped from its launch post. Opus 5.5's Terminal-Bench figure is at xhigh effort, its best. Sonnet 5.5's FrontierCode is shown at xhigh and max. Source: &lt;a href="https://www.anthropic.com/claude-sonnet-5-5" rel="noopener noreferrer"&gt;Anthropic&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Against Sonnet 5, every row is up, and some are up a lot. Terminal-Bench 4.0 goes from 10.3% to 70.6%. Chartography, a chart-reading test, goes from 15.6% to 61.6%. Anthropic also says it's the first Sonnet to beat Pokémon Red from screenshots alone.&lt;/p&gt;

&lt;p&gt;Against Opus 5.5, Sonnet 5.5 is ahead on Terminal-Bench 4.0 and behind on the other seven, mostly by two or three points. Against GPT-6 Sol it's ahead on the three rows where the comparison is simple. FrontierCode is the odd one, and it's the row I keep coming back to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Higher effort scored lower
&lt;/h2&gt;

&lt;p&gt;Anthropic lists Sonnet 5.5 twice on FrontierCode: 52.1% at xhigh effort and 46.2% at max. More effort, lower score. Footnote 2 says why.&lt;/p&gt;

&lt;p&gt;FrontierCode checks whether a code change could be merged without human edits. It penalizes changes outside the task's scope, even good ones. According to Anthropic, at max effort Sonnet 5.5 more often ran Claude Code's code-review skill, which splits the review across many subagents. In two cases Cognition looked at, that led to a timeout or to extra edits beyond the task.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2ivhkujtxzkkbervnipx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2ivhkujtxzkkbervnipx.png" alt=" " width="800" height="244"&gt;&lt;/a&gt;&lt;em&gt;FrontierCode 1.1 (main set). Higher is better. Anthropic's numbers, from its &lt;a href="https://www.anthropic.com/claude-sonnet-5-5" rel="noopener noreferrer"&gt;launch post&lt;/a&gt;; chart by me.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If you read my &lt;a href="https://promptquick.ai/blog/opus-5-5-improvements-and-prompting" rel="noopener noreferrer"&gt;post on the Opus 5.5 docs&lt;/a&gt;, you can see why this jumped out. Extra review passes and armies of subagents were two of the big complaints about Opus 5, and here they are again, this time costing points on a benchmark.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5" rel="noopener noreferrer"&gt;prompting guide&lt;/a&gt; describes the same behavior. At xhigh and max, it says, the model can start its own rounds of review and verification once a task is done, sometimes with subagents, and make related fixes it noticed along the way. The fix Anthropic suggests is a short instruction in the system prompt: when the work and its checks are done, stop and report, don't start extra review rounds, and don't launch reviewer subagents unless the user asked for a review. In Anthropic's coding tests at max effort, that cut session cost by about a third with no change in quality. That's their test, not mine, and they say it makes the behavior less frequent, not gone.&lt;/p&gt;

&lt;p&gt;My takeaway: on this model, max isn't a free upgrade. If a task has a clear scope, more effort can mean more work you didn't ask for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Artificial Analysis found
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://artificialanalysis.ai/models/claude-sonnet-5-5" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; ran Sonnet 5.5 at max effort on its Intelligence Index, which combines ten tests. It scored 56. Opus 5.5 got 58, and GPT-6 Astra and Fable 5.1 got 53 each, the same numbers as in &lt;a href="https://promptquick.ai/blog/opus-5-5-independent-benchmarks" rel="noopener noreferrer"&gt;my last post&lt;/a&gt;. Sonnet 5 scored 38, according to &lt;a href="https://officechai.com/ai/claude-sonnet-5-5-scores-56-on-artificial-analysis-intelligence-index-3-points-ahead-of-gpt-6-astra/" rel="noopener noreferrer"&gt;OfficeChai's write-up&lt;/a&gt; of the results.&lt;/p&gt;

&lt;p&gt;One of the ten tests is Terminal-Bench 4.0, so that's where I can put Anthropic's number next to an outside one. Per OfficeChai, Artificial Analysis has Sonnet 5.5 at 63.6% against 59.6% for Opus 5.5 and 59.1% for Astra. The numbers are lower than Anthropic's, but they point the same way.&lt;/p&gt;

&lt;p&gt;The catch is tokens. Per OfficeChai, Artificial Analysis says Sonnet 5.5 used more output tokens per task than any model it has measured. The model's page has the total: 410 million output tokens to run the whole index, against a median of 88 million. OfficeChai puts it at about 193,000 tokens per task, roughly 60% more than Opus 5.5 or Sonnet 5 at their own max settings.&lt;/p&gt;

&lt;p&gt;Tokens are what you pay for, so the cost per task looks like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4epdceifcunkp5j48lbo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4epdceifcunkp5j48lbo.png" alt=" " width="800" height="206"&gt;&lt;/a&gt;&lt;em&gt;Average cost per Intelligence Index task, all at max effort. Lower is better. Data: &lt;a href="https://artificialanalysis.ai/models/claude-sonnet-5-5" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;; chart by me.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Sonnet 5.5 costs half as much per token as Opus 5.5, but at max effort a task came out at $7.60 against $5.98 for Opus 5.5. OfficeChai adds that it's about 50% more than Sonnet 5 costs at max.&lt;/p&gt;

&lt;p&gt;That doesn't have to contradict Anthropic's claim of up to 30% less per task, which is measured against Sonnet 5 in Anthropic's own testing. The guide says effort levels were recalibrated on 5.5, so a level doesn't give the same amount of thinking it gave on Sonnet 5. Max on 5.5 may simply be more thinking than max on 5. Anthropic's post also says Sonnet 5.5 works best next to Opus 5.5 at lower effort, where it costs less per task, and that at higher settings it can perform comparably at a similar cost.&lt;/p&gt;

&lt;p&gt;Two more things about these numbers. First, Artificial Analysis ran a pre-release build with a bug that could hurt responses that use structured outputs. Anthropic says it's fixed and expects the effect on the scores, if any, to be small and to understate the model. Per OfficeChai, Artificial Analysis plans to re-run the affected tests, so treat these as first results. Second, its page showed no speed number for Sonnet 5.5 when I looked, so nothing outside Anthropic has checked the 30%+ faster claim yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What customers say about tokens
&lt;/h2&gt;

&lt;p&gt;Anthropic's post also quotes a lot of customers, and several talk about tokens. Slack says about 14% fewer output tokens on its Slackbot tests, with no prompt changes. Balyasny Asset Management reports about 121,000 tokens per answer on its finance tasks, against 497,000 for Sonnet 5. Lovable says a third fewer tool calls. Anthropic picked those quotes for its own launch post, so I read them as things that can happen, not things that will happen to you.&lt;/p&gt;

&lt;p&gt;This is where the effort setting matters. Anthropic's charts say that at low or medium effort Sonnet 5.5 beats Sonnet 5's best score on several benchmarks for about a tenth of the cost per task. The default in Claude Code and the Claude apps is medium, and the API defaults to high. Per OfficeChai, Artificial Analysis found high effort was Sonnet 5.5's most cost-competitive setting. So the max-effort bill above probably isn't the one you'll get.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the prompting guide says to do
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5" rel="noopener noreferrer"&gt;Sonnet 5.5 prompting guide&lt;/a&gt; opens by saying existing Sonnet 5 prompts should keep working. Then it lists what behaves differently. These are the parts I'd pay attention to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Redo your effort test.&lt;/strong&gt; A level doesn't give the same amount of thinking it gave on Sonnet 5. Anthropic says to start at high (the API default) unless the work is agentic or latency-sensitive. For agentic coding, medium for well-specified tasks and high for harder ones. For chat, medium or low. Keep xhigh and max for work where you've measured a better result. In Claude Code, &lt;code&gt;/effort&lt;/code&gt; changes it (&lt;a href="https://code.claude.com/docs/en/model-config#adjust-effort-level" rel="noopener noreferrer"&gt;docs&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't ask it to think less.&lt;/strong&gt; The guide says telling it to think less in the system prompt doesn't reliably work. Lower the effort instead. From medium up it thinks briefly before nearly every reply, even a greeting, and at low it skips thinking on most simple requests. Same lesson as with Opus 5.5: a line in your rules file isn't a setting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low effort has habits.&lt;/strong&gt; At low it can report a change as done without running a real check. At low and medium, on long tasks, it's more likely to stop and check in before finishing. Anthropic gives a short prompt for each, and says the keep-going one makes sessions longer and pricier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It adds things you didn't ask for.&lt;/strong&gt; Tests, docs and small supporting files that match your repo, at every effort level and more at higher ones. Anthropic thinks most teams will like that. If you don't, one short instruction tells it to mention them at the end instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open-ended asks get built.&lt;/strong&gt; Something like “show me what you can do with this” can turn into a slide deck. If you want ideas first, say so.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It may not search when it should.&lt;/strong&gt; When it has a search tool, the guide says it sometimes answers from what it learned in training where a search would catch something that changed, like what's allowed, required or charged. The fix is to tell it to search for those, even when it feels sure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Switching models drops its thinking.&lt;/strong&gt; Sonnet 5.5 can't read thinking from Opus 5, Opus 5.5 or the Fable and Mythos models, and no other model can read Sonnet 5.5's. Move a conversation between them and the turns after the switch run without that earlier reasoning. The request still works, and the dropped blocks aren't billed. It's the same for accounts: Sonnet 5.5's thinking only works in the account that produced it, and Anthropic specifically mentions switching accounts in the middle of a Claude Code session.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you use the API, three things to check. Sending &lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; now returns a 400 error, and you use &lt;code&gt;between_tools&lt;/code&gt; instead, at high effort or below. Forced tool use, meaning &lt;code&gt;tool_choice&lt;/code&gt; set to &lt;code&gt;any&lt;/code&gt; or to a named tool, also returns a 400. And prompts that ask the model to put its reasoning in the reply can trigger a &lt;code&gt;reasoning_extraction&lt;/code&gt; refusal. The &lt;a href="https://platform.claude.com/docs/en/models/sonnet-5-5/whats-new-sonnet-5-5" rel="noopener noreferrer"&gt;what's new page&lt;/a&gt; has the full list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does my CLAUDE.md need to change?
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://promptquick.ai/blog/opus-5-5-improvements-and-prompting" rel="noopener noreferrer"&gt;my Opus 5.5 post&lt;/a&gt; I explained the CLAUDE.md I use. I went through the Sonnet 5.5 guide with that file open, to see if Sonnet needs its own version.&lt;/p&gt;

&lt;p&gt;Mostly it doesn't. The main problems the guide describes are already covered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;My scope rule (“make routine choices yourself; ask when missing information materially changes the result or blocks safe progress”) covers the check-in habit.&lt;/li&gt;
&lt;li&gt;“Create review subagents only when the user requests them” covers the xhigh and max behavior from the footnote.&lt;/li&gt;
&lt;li&gt;My rule to look up current docs before editing, even when the model thinks it knows the API, is close to the guide's advice about searching for things that may have changed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three things I'd add if Sonnet 5.5 were my main model:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Don't add tests, docs or files that weren't asked for. Mention them at the end.&lt;/li&gt;
&lt;li&gt;Before saying the work is done, run a real check on the change. If none can run, say which one didn't.&lt;/li&gt;
&lt;li&gt;When I ask for ideas or a plan, give that and stop.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One thing I wouldn't add is a line telling it to think less, since the guide says that doesn't work. Effort is the setting for that. This is my reading of the docs, not something I've tested.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd wait for
&lt;/h2&gt;

&lt;p&gt;A few things would make me more confident about the numbers. Artificial Analysis re-running its tests after the bug fix. Someone measuring speed, since the 30%+ claim has no outside check yet. And a secure-code test like the Endor Labs one from &lt;a href="https://promptquick.ai/blog/opus-5-5-independent-benchmarks" rel="noopener noreferrer"&gt;my last post&lt;/a&gt;, where Opus 5.5 was quick and cheap but third on secure code. I couldn't find a Sonnet 5.5 result from them yet.&lt;/p&gt;

&lt;p&gt;Until then, the most useful test is a few of your own tasks at medium effort, where you know what a good result looks like. If it finishes them with less fuss than what you use now, that tells you more than any table here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.anthropic.com/claude-sonnet-5-5" rel="noopener noreferrer"&gt;Introducing Claude Sonnet 5.5&lt;/a&gt;, Anthropic, September 28, 2026. Source of the benchmark table and its footnotes, the pricing, the effort defaults and the customer quotes.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://x.com/claudeai/status/2104633119412523294" rel="noopener noreferrer"&gt;Anthropic's launch thread on X&lt;/a&gt;, September 28, 2026. I could only read its first two posts without an account, so the launch post above is where I got the rest.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5" rel="noopener noreferrer"&gt;Prompting Claude Sonnet 5.5&lt;/a&gt; and &lt;a href="https://platform.claude.com/docs/en/models/sonnet-5-5/whats-new-sonnet-5-5" rel="noopener noreferrer"&gt;What's new in Claude Sonnet 5.5&lt;/a&gt;, Anthropic's docs. Source of everything I attribute to “the guide.”&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://artificialanalysis.ai/models/claude-sonnet-5-5" rel="noopener noreferrer"&gt;Claude Sonnet 5.5&lt;/a&gt;, Artificial Analysis. Source of the Intelligence Index scores, the cost per task and the 410 million output tokens, as they stood on the day I wrote this.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://officechai.com/ai/claude-sonnet-5-5-scores-56-on-artificial-analysis-intelligence-index-3-points-ahead-of-gpt-6-astra/" rel="noopener noreferrer"&gt;Claude Sonnet 5.5 Scores 56 On Artificial Analysis Intelligence Index, 3 Points Ahead Of GPT-6 Astra&lt;/a&gt;, OfficeChai. Source of Sonnet 5's score of 38, the per-task token count, the Terminal-Bench numbers from Artificial Analysis and its plan to re-run the tests.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://code.claude.com/docs/en/model-config#adjust-effort-level" rel="noopener noreferrer"&gt;Model configuration&lt;/a&gt;, Claude Code docs. Source of the &lt;code&gt;/effort&lt;/code&gt; command.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://promptquick.ai/blog/sonnet-5-5-benchmarks-and-prompting-guide" rel="noopener noreferrer"&gt;promptquick.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>promptengineering</category>
    </item>
    <item>
      <title>AI App Starter: A Custom Step-by-Step Process for Apps With Claude Code</title>
      <dc:creator>Nomad</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:34:01 +0000</pubDate>
      <link>https://dev.to/technomadcode/ai-app-starter-a-custom-step-by-step-process-for-apps-with-claude-code-ph4</link>
      <guid>https://dev.to/technomadcode/ai-app-starter-a-custom-step-by-step-process-for-apps-with-claude-code-ph4</guid>
      <description>&lt;p&gt;I've built apps with Next.js, Supabase and Vercel together with AI several times, and it usually got messy. The biggest problem was structure. Every new chat session felt different. So did every planning phase, and every new phase of development. It wasted time, the workflow was unpredictable, and the project's state and documentation ended up a mess.&lt;/p&gt;

&lt;p&gt;A good start is half the work. So I made the AI App Starter: the same process every time, for every session and every phase, from the first idea to launch. You start it with one line. Open Claude Code in an empty folder and paste this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Setup line:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Set up this folder with the AI App Starter: clone https://github.com/TechNomadCode/AI-Product-Development-Toolkit with git clone --depth 1 into a temporary folder, then follow starters/nextjs-supabase-vercel/SETUP.md from it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;You can also add it as a skill: install the toolkit's &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#skills" rel="noopener noreferrer"&gt;skills&lt;/a&gt; once, then run &lt;code&gt;/nextjs-supabase-vercel&lt;/code&gt; in Claude Code or &lt;code&gt;$nextjs-supabase-vercel&lt;/code&gt; in Codex.&lt;/p&gt;

&lt;p&gt;The agent copies the starter into your folder. It holds no app code, only instructions and rules: what to do at each stage of building a Next.js, Supabase and Vercel app, what to ask you and when, and what it may never do. From then on, the agent leads and you answer.&lt;/p&gt;

&lt;p&gt;You need Node.js, Git and Docker. The starter is free and made for Claude Code; Codex support is in beta and not tested yet.&lt;/p&gt;
&lt;h2&gt;
  
  
  It starts with a conversation
&lt;/h2&gt;

&lt;p&gt;First, the agent tells you where you stand. It works only on your computer, with a local database and stand-ins for payment and email, and it never uses the keys to your hosting or payment accounts, not even test keys. Those steps are yours, and it prepares them for you.&lt;/p&gt;

&lt;p&gt;Then it asks whether you already have a plan: a PRD, a scope, a few notes. If you do, it asks only about the gaps. If you don't, it interviews you with the planning prompts from the same toolkit (I wrote about the first one in &lt;a href="https://promptquick.ai/blog/turn-ideas-into-products-with-ai" rel="noopener noreferrer"&gt;How I Turn Ideas Into Products With AI Prompts&lt;/a&gt;), one to three questions at a time, and writes nothing until you agree.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj5o9buqt6094igk8r52v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj5o9buqt6094igk8r52v.png" alt="The agent's first reply after setup in the Claude Code desktop app: setup is done, it works only on this computer and never uses real account logins or API keys, the four steps of working together, and a first question asking whether you already have a PRD or notes." width="547" height="883"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The first reply after setup, in the Claude Code desktop app. This run used the &lt;code&gt;/nextjs-supabase-vercel&lt;/code&gt; skill.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Next come the practical choices: who uses the app, what it stores, which languages, and which services handle payments or email. Each question comes with a suggested answer, so a plain “yes” is often enough. Your domain, tax or legal pages wait until they're due.&lt;/p&gt;
&lt;h2&gt;
  
  
  The first real feature
&lt;/h2&gt;

&lt;p&gt;Then the building starts, with the smallest real feature that stores something and shows it. In the test run, a small web shop, that was browsing the catalogue. The agent proposes it and waits for your OK.&lt;/p&gt;

&lt;p&gt;It builds that feature from the database to the screen, with tests in four layers and automatic checks on GitHub. A guard refuses to run anything against a database that isn't local or a service that isn't a stand-in. And the database starts locked, with row-level security on every table.&lt;/p&gt;

&lt;p&gt;It also prepares the way online. Staging, a test copy of your app, updates by itself when you merge on GitHub. Production, your live app, changes only when you start it, and in both the database is updated before the new code runs. The agent never runs these steps itself: pushing and deploying are blocked, and there are no keys on your computer.&lt;/p&gt;

&lt;p&gt;So this stage ends with a short list for you: create your GitHub, Vercel and Supabase accounts, follow the setup steps in your app's new README, push and merge. Staging then gets the database changes first, and the code after.&lt;/p&gt;
&lt;h2&gt;
  
  
  One piece at a time
&lt;/h2&gt;

&lt;p&gt;From here, every session works the same way. You ask “what's next?”, and the agent tells you where things stand without changing anything. You say “continue”, or “do X instead”.&lt;/p&gt;

&lt;p&gt;It builds one piece of the app per session. Its tests follow the requirements and are never weakened to get a pass, and it checks the current official docs before it uses a library or service. When the piece is done, it saves, tells you what to review and push, and suggests a fresh session. The plan always says where you are and what's next, and every session starts by reading it. Decisions and known problems each have one file too.&lt;/p&gt;

&lt;p&gt;New ideas fit in along the way. Ask for a feature halfway through, and it goes into the plan with a question: now, or later?&lt;/p&gt;

&lt;p&gt;Work on sign-in, permissions, payments or existing data gets a second look. The agent gives you one line to paste into a second agent, such as Codex, or into a fresh session, which reviews the work without changing anything. After one round of fixes and one more check, anything left is your call.&lt;/p&gt;

&lt;p&gt;And when a piece depends on a real service like payments, the agent says its tests used a stand-in, and tells you how to try the real thing on staging with the provider's test keys.&lt;/p&gt;
&lt;h2&gt;
  
  
  Launch
&lt;/h2&gt;

&lt;p&gt;The last stage takes one part per session: production, monitoring, security, search engines, privacy and legal pages, a final readiness check, and the switch to your domain. The questions it saved come back now. It won't decide law or tax for you; it shows you the options with official sources. You do the steps on the hosting sites, and it checks each one.&lt;/p&gt;
&lt;h2&gt;
  
  
  What's tested, and what isn't
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tested:&lt;/strong&gt; Claude Code desktop, from setup through planning, decisions, the first feature, and the first pieces of a web shop: a catalogue, product photos and a cart. The automatic checks passed on GitHub on the first push.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tested yet:&lt;/strong&gt; Codex, the hosted staging and production pipeline, and the launch stage.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'll update this post as those get tested.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where to get it
&lt;/h2&gt;

&lt;p&gt;The AI App Starter is part of the &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;AI Product Development Toolkit&lt;/a&gt; on GitHub, free under the MIT license. Found a problem? &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit/issues" rel="noopener noreferrer"&gt;Open an issue&lt;/a&gt;.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/TechNomadCode" rel="noopener noreferrer"&gt;
        TechNomadCode
      &lt;/a&gt; / &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;
        AI-Product-Development-Toolkit
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Plan, build and launch products with AI: guided planning prompts, an AI App Starter for Next.js, Supabase and Vercel, and agent configurations.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;AI Product Development Toolkit&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Plan, build and launch products with AI, without the mess.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit/./LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/b8cadaa967891081f8f165695470689986c028821dd8a040132f6e661795dc0d/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d626c7565" alt="MIT license"&gt;&lt;/a&gt;
&lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/bb13ea0cda91b4b727d9cf100cbd4c1a206302baa3355b13f087a1819d29cce8/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436c61756465253230436f64652d443937373537"&gt;&lt;img src="https://camo.githubusercontent.com/bb13ea0cda91b4b727d9cf100cbd4c1a206302baa3355b13f087a1819d29cce8/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436c61756465253230436f64652d443937373537" alt="Works with Claude Code"&gt;&lt;/a&gt;
&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#tested-so-far" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/8d09579d9b5797997037f0d7750b9116a5e11e3a35ea9513ffc18a44e6e08dc5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d436f6465782532302862657461292d343132393931" alt="Works with Codex (beta)"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Use it when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#planning-prompts" rel="noopener noreferrer"&gt;Planning prompts&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Guided prompts that interview you, then write your plan. Any AI chat.&lt;/td&gt;
&lt;td&gt;You have an idea and want a solid plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#ai-app-starter" rel="noopener noreferrer"&gt;AI App Starter&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A project your coding agent builds with you, step by step, up to launch.&lt;/td&gt;
&lt;td&gt;You want to build and launch a web app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#agent-configurations" rel="noopener noreferrer"&gt;Agent configurations&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Working instructions for Claude Code.&lt;/td&gt;
&lt;td&gt;You code with Claude Code every day&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The planning prompts and the AI App Starter also install as &lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit#skills" rel="noopener noreferrer"&gt;skills&lt;/a&gt; for Claude Code and Codex, with one line, together with a skill that makes short product videos in code.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Planning prompts&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Most prompts tell the AI to write a document straight away, so it fills every gap with guesses. These prompts make it &lt;strong&gt;interview you first&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It asks before it writes.&lt;/strong&gt; A few targeted…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;a href="https://github.com/TechNomadCode/AI-Product-Development-Toolkit/tree/main/starters/nextjs-supabase-vercel" class="crayons-btn crayons-btn--primary" rel="noopener noreferrer"&gt;Open the AI App Starter&lt;/a&gt;
&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://promptquick.ai/blog/ai-app-starter-nextjs-supabase-vercel" rel="noopener noreferrer"&gt;promptquick.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>nextjs</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How PromptQuick.ai Is Built With an AI Coding Agent</title>
      <dc:creator>Nomad</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:30:18 +0000</pubDate>
      <link>https://dev.to/technomadcode/how-promptquickai-is-built-with-an-ai-coding-agent-25m3</link>
      <guid>https://dev.to/technomadcode/how-promptquickai-is-built-with-an-ai-coding-agent-25m3</guid>
      <description>&lt;p&gt;PromptQuick.ai looks like a simple site: a few pages, a sign-up form and a buy button. The code behind it is private, so you can't see how it works. This post shows what happens when you sign up for the free sample, when you buy the Rulebook, and when a change goes live. And who does what, since an AI coding agent now writes the code.&lt;/p&gt;

&lt;p&gt;I built the first version in April 2025 with help from AI chat tools, and the full Rulebook went on sale that May. In September 2026 I rebuilt most of it, this time with Claude Code, Anthropic's coding agent, working in the project folder on my computer. Since 25 September, every change to the code names it as co-author.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it runs on
&lt;/h2&gt;

&lt;p&gt;The site is Next.js 16 with React 19 and TypeScript, styled with Tailwind CSS 4 and hosted on Vercel. The database is Postgres 17 on Supabase. A few services do the specialised work: Lemon Squeezy runs the checkout, Resend sends the email, Google reCAPTCHA checks the sign-up form for bots, Upstash limits how often one visitor can call the form and the counters, and Cloudflare hosts the PDFs.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you sign up for the free sample
&lt;/h2&gt;

&lt;p&gt;When you send the form, the server first checks the reCAPTCHA and your address. Then one database function does two things: it stores your address and adds a “send the welcome email” job to a queue. Both happen in one transaction, so a sign-up can't exist without its email.&lt;/p&gt;

&lt;p&gt;You get the same answer whether your address is new or signed up before, so the form never tells anyone who's on the list. A repeat sign-up gets no second email.&lt;/p&gt;

&lt;p&gt;Right after answering you, the site tries to send the email, so it usually arrives within seconds. That's only a shortcut. Every minute, the database's own scheduler calls a worker that sends whatever is still waiting. When a send fails, the worker tries again after 1 minute, then 2, then 4, up to 2 hours apart, 8 attempts in all.&lt;/p&gt;

&lt;p&gt;It's the database's scheduler, not the host's, because it runs every minute on any Supabase or Vercel plan. Vercel's own scheduled jobs run once a day on its free plan.&lt;/p&gt;

&lt;p&gt;Retrying email has a catch: you don't want two copies. So the first attempt saves the exact message, and every attempt sends that same message with the same idempotency key, a label that tells Resend “this is the email you may have seen already”. Resend sends it once. Resend remembers that label for a day, so if the worker still can't tell by then whether an email went out, it stops and leaves the decision to me instead of guessing. When an email fails for good, I get an alert.&lt;/p&gt;

&lt;p&gt;This fixed a real gap. In the 2025 version the site made one attempt, after answering the form. If that failed, nothing tried again, and because a repeat sign-up gets no second email, that person never got the sample.&lt;/p&gt;

&lt;p&gt;The saved message contains your address, so a daily job clears it 30 days later.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you buy the Rulebook
&lt;/h2&gt;

&lt;p&gt;Lemon Squeezy does the selling: the checkout, the payment, and the receipt email with the download link. The site never sees your card and never calculates a price.&lt;/p&gt;

&lt;p&gt;After a purchase, Lemon Squeezy sends the site a notice, called a webhook. The site checks its signature first, a code only Lemon Squeezy and the site can compute from that exact message, and refuses anything that doesn't match. But a valid signature only proves who sent it. So the site also checks that the order is paid, not refunded, and for the Rulebook in this store. Only then does it record the buyer. If the same notice arrives twice, the second one changes nothing.&lt;/p&gt;

&lt;p&gt;There's also a test copy of the site, which I'll get to below. Purchases there use Lemon Squeezy's test mode, and the site accepts a test order only when it's connected to a separate test database. A test purchase can't end up among the real buyers.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a change goes live
&lt;/h2&gt;

&lt;p&gt;Everything starts on my computer. The agent works against a local copy of the database and fake stand-ins for email, payments and the bot check. A guard refuses to run anything against a hosted database or a real provider, and the computer holds no keys to the live systems at all.&lt;/p&gt;

&lt;p&gt;Changes reach the main branch through pull requests. Each one goes through five automatic checks: formatting, code style, types, unit tests and a full build; the database's own tests; browser tests that click through the site and run an accessibility check on every page; an audit of the packages the site depends on; and a scan of the whole history for leaked secrets. I merge once they pass.&lt;/p&gt;

&lt;p&gt;When a change passes and lands on the main branch, staging updates by itself. That's the test copy: the same site with its own database. The live site at promptquick.ai only changes when I start a release by hand, and only from a change that passed the checks. In both, the database is updated first and the new code goes out after. If the database step fails, the code stays where it was.&lt;/p&gt;

&lt;p&gt;Only promptquick.ai itself can show up in search engines. Staging and every other copy tell them to stay away.&lt;/p&gt;

&lt;p&gt;One switch puts the site in maintenance mode. It pauses everything that needs the database, like the sign-up form and the counters, while the pages and the checkout stay up. I used it on 27 September to upgrade the live database from Postgres 15 to 17, after testing the switch on staging first.&lt;/p&gt;

&lt;p&gt;I chose a partial switch over one maintenance page for the whole site. The checkout runs on Lemon Squeezy, so people can still buy. Only recording the buyer waits: I resend the purchase notices from that window from Lemon Squeezy's log afterwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who does what
&lt;/h2&gt;

&lt;p&gt;Claude Code writes the code, but it works inside rules I set, written in files it reads at the start of every session:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A plan says where the project stands and what's next. Every session reads it first, takes one piece of work, and updates it before stopping.&lt;/li&gt;
&lt;li&gt;Tests follow the requirements, and it may never weaken a test to get a pass.&lt;/li&gt;
&lt;li&gt;Before it uses a library or a service, it looks up the current documentation instead of trusting what it remembers.&lt;/li&gt;
&lt;li&gt;It can't push code, deploy, or touch anything live. Those commands are blocked in its settings, and there are no keys on the machine to use anyway.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That way of working comes from the AI App Starter, which I wrote about in &lt;a href="https://promptquick.ai/blog/ai-app-starter-nextjs-supabase-vercel" rel="noopener noreferrer"&gt;its own post&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;My part is the rest. I decide what gets built and how it should look and behave, and those decisions go into a file the agent follows. I review the work, push it, and start every release. Everything on the live systems is mine too: setting up hosting and the scheduler, upgrading the live database, uploading the PDFs and sending the update emails.&lt;/p&gt;

&lt;p&gt;Work on payments, on deleting data or on the live database gets a second look. A separate agent session reviews it without changing anything, and I decide what to do with what it finds.&lt;/p&gt;

&lt;p&gt;Most of this never shows. You fill in a form and get an email, or you pay and get a PDF. What took the time is making sure that still happens when a step along the way fails.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://promptquick.ai/blog/how-promptquick-ai-is-built" rel="noopener noreferrer"&gt;promptquick.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>nextjs</category>
      <category>supabase</category>
    </item>
  </channel>
</rss>
