<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: claude</title>
    <description>The latest articles tagged 'claude' on DEV Community.</description>
    <link>https://dev.to/t/claude</link>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tag/claude"/>
    <language>en</language>
    <item>
      <title>How I Use SDD (Spec-Driven Development)</title>
      <dc:creator>Davi Orlandi</dc:creator>
      <pubDate>Mon, 05 Oct 2026 13:17:04 +0000</pubDate>
      <link>https://dev.to/dvorlandi/how-i-use-sdd-spec-driven-development-2g9g</link>
      <guid>https://dev.to/dvorlandi/how-i-use-sdd-spec-driven-development-2g9g</guid>
      <description>&lt;p&gt;If you follow tech the way I do, you've probably already felt lost with the flood of innovations launching every week since AI took off (or at least opened Twitter/LinkedIn and thought "okay, I became obsolete in 3 days"). Today I'll share a bit of my experience with one of them: Spec-Driven Development.&lt;/p&gt;

&lt;h1&gt;
  
  
  What is SDD?
&lt;/h1&gt;

&lt;p&gt;SDD, or Spec-Driven Development, is a software development framework where the &lt;strong&gt;specification comes before the code&lt;/strong&gt;. Instead of generating code unchecked, we define functional and technical specifications that guide development, serve as documentation and history, and help enrich the context LLMs use throughout the process.&lt;/p&gt;

&lt;p&gt;There are many ways to use SDD. You can rely on ready-made setups like &lt;a href="https://github.com/gotalab/cc-sdd" rel="noopener noreferrer"&gt;cc-sdd&lt;/a&gt; for Claude Code, or &lt;a href="https://github.com/madebyaris/spec-kit-command-cursor" rel="noopener noreferrer"&gt;Spec Kit Command&lt;/a&gt; for Cursor, which already ship with ready-to-use commands. You can also build your own commands on top of these setups to fit your day-to-day workflow. It's common to have agents, skills, or commands specialized for each company's workflows — "How to deploy the backend", "Security acceptance criteria", and so on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Workflow
&lt;/h2&gt;

&lt;p&gt;SDD implementations vary by use case, but most go through these phases:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Feeuyo5hgq0b2yomvcgyk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Feeuyo5hgq0b2yomvcgyk.png" alt=" " width="800" height="205"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Functional Spec
&lt;/h3&gt;

&lt;p&gt;This is where we describe requirements functionally. Use as little technical language as possible (yes, that's harder than it sounds).&lt;/p&gt;

&lt;p&gt;You can start from an idea, a task from tools like Jira via MCP, or even a bug. Example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Users complete onboarding but receive no follow-up. That reduces engagement and doesn't guide them through the next steps inside the platform.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Since this is a discovery moment, the AI starts asking questions:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI: Which users should we implement this for?&lt;/p&gt;

&lt;p&gt;You: All newly registered users who finished onboarding.&lt;/p&gt;

&lt;p&gt;AI: Exactly when should the email be sent?&lt;/p&gt;

&lt;p&gt;You: Immediately after confirming that onboarding completed successfully.&lt;/p&gt;

&lt;p&gt;AI: What are the functional requirements?&lt;/p&gt;

&lt;p&gt;You:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Detect the moment onboarding completes&lt;/li&gt;
&lt;li&gt;Trigger the email only once per user&lt;/li&gt;
&lt;li&gt;Personalize the email with basic data (name, etc.)&lt;/li&gt;
&lt;li&gt;Implement retry if the email service fails&lt;/li&gt;
&lt;li&gt;Log successes and errors&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;

&lt;p&gt;These refinements produce a Functional Spec — everything needed to describe the technical details of the implementation.&lt;/p&gt;

&lt;p&gt;That document should also include clear acceptance criteria, which serve as the validation baseline throughout the flow.&lt;/p&gt;

&lt;p&gt;The spec must be editable and have an approval mechanism. In almost every case, you'll iterate with the LLM until you land on something satisfactory.&lt;/p&gt;

&lt;p&gt;This phase can be started with commands like:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/sdd.spec functional Envio de email de onboarding&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Some setups use flags such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;--clarify&lt;/code&gt; when something isn't clear&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--approve&lt;/code&gt; when you consider the spec ready&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Technical Spec
&lt;/h3&gt;

&lt;p&gt;With the functional specification defined, translate the business requirements into technical details. You can start with:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/sdd.spec technical&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI: What technology stack will you use?&lt;/p&gt;

&lt;p&gt;You: Node.js with Express for the service, integration with an email provider (e.g. SES or similar), and a queue (e.g. SQS or PubSub) for asynchronous processing.&lt;/p&gt;

&lt;p&gt;AI: What API design approach?&lt;/p&gt;

&lt;p&gt;You: Event-driven architecture. The trigger will be based on the &lt;code&gt;user.onboarding.completed&lt;/code&gt; event, with asynchronous communication.&lt;/p&gt;

&lt;p&gt;AI: Which services or platforms will be integrated?&lt;/p&gt;

&lt;p&gt;You: An email service for sending, a messaging system for the queue, and possibly a template service. Also integration with the auth system to keep data consistent.&lt;/p&gt;

&lt;p&gt;AI: What are the performance requirements?&lt;/p&gt;

&lt;p&gt;You:&lt;br&gt;
Event processing latency under 5 seconds&lt;br&gt;
Ability to handle spikes without losing events&lt;br&gt;
Failure rate under 1 percent with automatic retry&lt;/p&gt;

&lt;p&gt;AI: How will reliability be ensured?&lt;/p&gt;

&lt;p&gt;You:&lt;br&gt;
Use a queue to decouple processing&lt;br&gt;
Retry with exponential backoff&lt;br&gt;
Dead-letter queue for unprocessed events&lt;br&gt;
Idempotency to avoid duplicates&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Just like in the functional phase, this step produces a technical document that should be reviewed and approved.&lt;/p&gt;

&lt;p&gt;Don't let technical details rewrite the functional requirements — though in the real world, technology often ends up shaping business specs anyway.&lt;/p&gt;




&lt;h3&gt;
  
  
  Tasks
&lt;/h3&gt;

&lt;p&gt;By now we've turned an idea into two clear things: what to do and how to do it. Next comes execution.&lt;/p&gt;

&lt;p&gt;Translate the spec into clear, actionable tasks — exactly what the LLM or the developer should do, with no ambiguity. Start with:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/sdd.plan&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The AI then generates something like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI: Here are the suggested tasks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;TASK-001: Implement consumer for the &lt;code&gt;user.onboarding.completed&lt;/code&gt; event&lt;/li&gt;
&lt;li&gt;TASK-002: Create email-sending service with provider integration&lt;/li&gt;
&lt;li&gt;TASK-003: Implement idempotency logic to avoid duplicates&lt;/li&gt;
&lt;li&gt;TASK-004: Create email template with personalization&lt;/li&gt;
&lt;li&gt;TASK-005: Wire the full flow from event to send&lt;/li&gt;
&lt;li&gt;TASK-006: Add logs and metrics&lt;/li&gt;
&lt;li&gt;AUTO-TASK-001: Create unit tests&lt;/li&gt;
&lt;li&gt;AUTO-TASK-002: Create integration tests&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;These tasks can be saved in files such as &lt;code&gt;tasks.json&lt;/code&gt; or individually as &lt;code&gt;TASK-001.md&lt;/code&gt;, containing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Status such as wip, done, or canceled&lt;/li&gt;
&lt;li&gt;Assignee&lt;/li&gt;
&lt;li&gt;Dependencies&lt;/li&gt;
&lt;li&gt;Possibility of parallelization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can refine with:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/sdd.plan --refine&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;And then approve with:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/sdd.plan --approve&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;This phase matters because it turns the spec into something executable. There should be no ambiguity here — only clear tasks.&lt;/p&gt;




&lt;h3&gt;
  
  
  Implementation
&lt;/h3&gt;

&lt;p&gt;With everything approved, enter the implementation phase:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/sdd.build&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;During this phase, the AI executes the tasks and continuously validates the acceptance criteria from the functional spec, keeping the implementation aligned with the plan.&lt;/p&gt;

&lt;p&gt;Because the tasks are already well defined and decoupled, some tools let you parallelize executions that have no dependencies.&lt;/p&gt;

&lt;p&gt;One example is &lt;a href="https://code.claude.com/docs/en/agent-teams" rel="noopener noreferrer"&gt;Claude Agent Teams&lt;/a&gt;, which let you run multiple agents in parallel, each responsible for part of the system. Instead of a single AI doing everything sequentially, you can have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One agent implementing services&lt;/li&gt;
&lt;li&gt;Another writing tests&lt;/li&gt;
&lt;li&gt;Another handling integration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All at the same time, respecting the dependencies defined in the tasks.&lt;/p&gt;

&lt;p&gt;This works especially well with SDD because the tasks are already well scoped — you don't need to manually coordinate who does what; the structure of the spec enables that split.&lt;/p&gt;

&lt;p&gt;Following the flow:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI: Creating branch &lt;code&gt;feature/email-onboarding&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;AI: Starting task execution.&lt;/p&gt;

&lt;p&gt;AI: Running TASK-001: Event consumer&lt;/p&gt;

&lt;p&gt;AI: Consumer created successfully. Want to validate with tests?&lt;/p&gt;

&lt;p&gt;You: Yes.&lt;/p&gt;

&lt;p&gt;AI: Tests ran successfully.&lt;/p&gt;

&lt;p&gt;AI: Running TASK-002: Email-sending service&lt;/p&gt;

&lt;p&gt;AI: Provider integration complete. Retry implemented.&lt;/p&gt;

&lt;p&gt;AI: Tests passing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This process continues until all tasks are done.&lt;/p&gt;




&lt;h3&gt;
  
  
  Finish
&lt;/h3&gt;

&lt;p&gt;After implementation, wrap up the feature: make sure everything defined in the specs was delivered, validated, and is ready for production.&lt;/p&gt;

&lt;p&gt;You can start with:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/sdd.finish&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The AI then runs final validations:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI: Checking task status...&lt;/p&gt;

&lt;p&gt;AI: All tasks are complete.&lt;/p&gt;

&lt;p&gt;AI: Validating tests...&lt;/p&gt;

&lt;p&gt;AI: All tests are passing with adequate coverage.&lt;/p&gt;

&lt;p&gt;AI: Validating consistency between specs and implementation...&lt;/p&gt;

&lt;p&gt;AI: No divergences found.&lt;/p&gt;

&lt;p&gt;AI: Do you want to finish the feature?&lt;/p&gt;

&lt;p&gt;You: Yes.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;At that point, a few things usually happen:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Mark the feature as complete&lt;/li&gt;
&lt;li&gt;Ensure all specs are up to date&lt;/li&gt;
&lt;li&gt;Consolidate documentation&lt;/li&gt;
&lt;li&gt;Prepare for deploy or merge&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Tips &amp;amp; tricks
&lt;/h2&gt;

&lt;p&gt;Don't limit yourself to ready-made setups — you can integrate this flow with day-to-day tools using MCPs like Jira or Linear to generate specs from tasks, GitHub to manage issues, branches, and comments, or Playwright to test a generated frontend.&lt;/p&gt;

&lt;p&gt;It's also common to have specialized agents for specific stages — code review, security validation, performance analysis. That makes the flow safer without adding manual complexity for the developer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;You're no longer deciding what to do while writing code. That was already settled in the earlier phases. In a world where LLMs take an active part in development, this shrinks scope, improves delivery quality, and makes the process far more reproducible.&lt;/p&gt;

&lt;p&gt;In the end, it isn't about writing more code. It's about thinking better before you code (and letting the AI suffer a bit more in your place).&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>A Beginner Vibe-Codes a Game — Zero Coding, 3 Days, $45</title>
      <dc:creator>Johmisking</dc:creator>
      <pubDate>Mon, 05 Oct 2026 12:37:15 +0000</pubDate>
      <link>https://dev.to/jaehyun_cho_0dff271e0d2e5/a-beginner-vibe-codes-a-game-zero-coding-3-days-45-371f</link>
      <guid>https://dev.to/jaehyun_cho_0dff271e0d2e5/a-beginner-vibe-codes-a-game-zero-coding-3-days-45-371f</guid>
      <description>&lt;p&gt;I had zero coding experience. I didn't even know what a "game engine" was. Still, by describing what I wanted to an AI in plain words, I made a game in 3 days, and getting it to launch cost &lt;strong&gt;$45&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F87typrs1nemgpmsa1isi.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F87typrs1nemgpmsa1isi.jpg" alt="Sperm Race Google Play feature graphic" width="800" height="391"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is part 1 of the dev log for Sperm Race, a silly sperm-racing game.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Google Play&lt;/strong&gt;: &lt;a href="https://play.google.com/store/apps/details?id=com.spermrace.game&amp;amp;hl=en" rel="noopener noreferrer"&gt;Sperm Race&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Web version (itch.io)&lt;/strong&gt;: &lt;a href="https://johnisking.itch.io/sperm-race" rel="noopener noreferrer"&gt;Sperm Race&lt;/a&gt; — it came out on itch.io first, then on Google Play.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  "Let's make an easy game" took me all the way back to sperm
&lt;/h2&gt;

&lt;p&gt;My goal at the start was simple: &lt;strong&gt;"Just make an easy game."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But the more I thought about what "easy" means, the more primal my ideas got. A game with one rule. A game that needs no explanation. A game everyone already knows. Following that line all the way back, I ended up at &lt;strong&gt;sperm&lt;/strong&gt;. Swim forward, reach the egg, done. When you think about it, it's the one race every one of us has already won once.&lt;/p&gt;

&lt;p&gt;The moment that thought hit me, I laughed out loud. &lt;strong&gt;"Wait, this is funny."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That one reaction set the direction. Not a serious game, but a goofy one that makes you smirk the second you see it.&lt;/p&gt;

&lt;h2&gt;
  
  
  I didn't know what a game engine was
&lt;/h2&gt;

&lt;p&gt;I had never written code. I didn't know engines like Unity or Godot existed, let alone what they do.&lt;/p&gt;

&lt;p&gt;So I started without one. I told Claude (Sonnet) in Korean, "I want to make a game like this," and what came back was a single HTML file that ran right in the browser. Later I wrapped that same file into an Android app and put it on Google Play. Building by describing what you want to an AI instead of writing code yourself has a name now: &lt;strong&gt;vibe coding&lt;/strong&gt;. Looking back, starting without knowing anything was actually faster. No engine to install, nothing to study, and I could play it on day one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first prompt: "I want to make a game where a sperm swims to meet the egg"
&lt;/h2&gt;

&lt;p&gt;That one sentence produced the first version. A sperm appeared on screen, and I could move it with my finger.&lt;/p&gt;

&lt;p&gt;The problem was the tail. It was stiff as a stick, more matchstick than sperm. So the second thing I said was:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Add a wave to the tail."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fru22r62umdiq4fj8q45d.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fru22r62umdiq4fj8q45d.gif" alt="Left: the first version with a stiff tail. Right: the next version with a wave" width="800" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the left is the stiff tail from the first version; on the right is the next version with the wave. Just making the tail wiggle finally made it look like a living sperm. That's how the main character was born.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three days of adding, cutting and adding again
&lt;/h2&gt;

&lt;p&gt;After that I couldn't stop. Every time I played a build, the next idea showed up. If it came to mind, I added it. If it didn't work, I cut it. Then I added again.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;"Make a map."&lt;/strong&gt; Now there was somewhere to swim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Add obstacles too."&lt;/strong&gt; Once there was something to dodge, it became a game. This is today's Adventure mode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Add sperm racing, like real car racing."&lt;/strong&gt; A Racing mode with 3 laps around a track.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Make a mode where they run on their own like horse racing, and I place obstacles."&lt;/strong&gt; Pick a sperm to root for and help it win with obstacles. This is Derby mode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Make a survival mode where you run away, with cancer as the motif."&lt;/strong&gt; Dodge cancer cells in a maze and survive for 1 minute 30 seconds.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The file dates left in my Downloads folder show how fast it went. I got the first HTML file in the early hours of August 19, and by the next afternoon I had downloaded a new version &lt;strong&gt;16 times&lt;/strong&gt;. That's "fix this → download the new file → play" four or five times every half day. Meanwhile the code grew from 47 KB to 149 KB. By August 21 I was already taking screenshots for the store.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rx4u63lk7ucnxas9ewd.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rx4u63lk7ucnxas9ewd.jpg" alt="Sperm Race main menu (current version — the Play with friends button was added after launch)" width="525" height="1058"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Graphics and sound effects, all in code
&lt;/h2&gt;

&lt;p&gt;There isn't a single image file in this game. The character, backgrounds, tracks and mazes were all drawn in code by Sonnet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnun9z96iunr0yz1c2rdx.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnun9z96iunr0yz1c2rdx.jpg" alt="Four modes drawn entirely in code: Adventure, Racing, Derby, Survival" width="799" height="241"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;From left: Adventure, Racing, Derby, Survival. The sperm is one ellipse and a wiggling line, the track is a thick curve, and the maze is glowing straight lines. For a silly game, that simplicity actually worked.&lt;/p&gt;

&lt;p&gt;Same with sound effects. There are no sound files; Sonnet synthesized the sounds in code. The only audio file in the game folder is free background music from Pixabay.&lt;/p&gt;

&lt;h2&gt;
  
  
  The testers' first reaction: "That's so original"
&lt;/h2&gt;

&lt;p&gt;Before launch I showed it to testers. What came back was: &lt;strong&gt;"That's so original."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The idea that started with "wait, this is funny" wasn't funny only to me.&lt;/p&gt;

&lt;p&gt;Here's what was in the game at launch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;4 modes: Adventure (120 stages), Racing (3 laps), Derby, Survival (last 1 minute 30 seconds)&lt;/li&gt;
&lt;li&gt;13 languages&lt;/li&gt;
&lt;li&gt;Controls: drag, arrow keys/WASD, gamepad&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Multiplayer with friends (up to 10 players in Derby, 4 in the other modes) came later, in an update after launch.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what did it cost? $45
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;What I used&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Code&lt;/td&gt;
&lt;td&gt;Claude Pro subscription (Sonnet)&lt;/td&gt;
&lt;td&gt;$20 (one month)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Graphics&lt;/td&gt;
&lt;td&gt;Drawn in code&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Background music&lt;/td&gt;
&lt;td&gt;Free tracks from Pixabay&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sound effects&lt;/td&gt;
&lt;td&gt;Made in code&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Launch&lt;/td&gt;
&lt;td&gt;Google Play developer registration&lt;/td&gt;
&lt;td&gt;$25 (one-time)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$45&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I use Claude Max 20× now, but Sperm Race was built &lt;strong&gt;entirely on the Pro plan.&lt;/strong&gt; It was done within a month, so I only paid for one month.&lt;/p&gt;

&lt;p&gt;For comparison, if you run a small game through the &lt;a href="https://tokensave.app/ai-game-cost-calculator" rel="noopener noreferrer"&gt;AI game cost calculator&lt;/a&gt;, the standard stack (image and music AI plus Claude Max) comes out at $128–181. Sperm Race came in at about 25–35% of that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next time
&lt;/h2&gt;

&lt;p&gt;So far it looks like "AI made a game in 3 days, easy." But the real struggle came after. Building took 3 days; fixing bugs took 14.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next: 3 days to build, 14 days to fix bugs&lt;/strong&gt; — why AI pokes at the wrong code when you ask it to "find the bug," and how I ended up finding bugs myself and having the AI fix them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://tokensave.app/blog/vibe-coding-a-game-beginner" rel="noopener noreferrer"&gt;tokensave.app&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>gamedev</category>
      <category>ai</category>
      <category>beginners</category>
      <category>claude</category>
    </item>
    <item>
      <title>Using Opus 5.5 Claude Code with Microsoft Foundry in VS Code</title>
      <dc:creator>Taswar Bhatti</dc:creator>
      <pubDate>Mon, 05 Oct 2026 12:17:07 +0000</pubDate>
      <link>https://dev.to/taswar_bhatti/using-opus-55-claude-code-with-microsoft-foundry-in-vs-code-7o8</link>
      <guid>https://dev.to/taswar_bhatti/using-opus-55-claude-code-with-microsoft-foundry-in-vs-code-7o8</guid>
      <description>&lt;p&gt;In my &lt;a href="https://taswar.zeytinsoft.com/claude-opus-5-5-csharp-developers-guide/" rel="noopener noreferrer"&gt;Claude Opus 5.5 for C# developers&lt;/a&gt; post I called Opus 5.5 from my own code: the &lt;code&gt;Anthropic.Foundry&lt;/code&gt; SDK, Entra ID, and the &lt;code&gt;effort&lt;/code&gt; parameter. This time the code calling Opus 5.5 isn't mine. It's &lt;strong&gt;Claude Code&lt;/strong&gt;, running in my terminal and in VS Code, pointed at my Foundry resource.&lt;/p&gt;

&lt;p&gt;Microsoft already has a good &lt;a href="https://techcommunity.microsoft.com/blog/azuredevcommunityblog/claude-code-on-microsoft-foundry-in-vs-code-%E2%80%94-a-practical-setup-guide-with-the-g/4524245" rel="noopener noreferrer"&gt;step-by-step setup guide for Claude Code on Microsoft Foundry in VS Code&lt;/a&gt;, so I'm not going to repeat it. If you've never connected Claude Code to Microsoft Foundry, start there. This post is about what happens &lt;em&gt;after&lt;/em&gt; it connects, when the model on the other end is Opus 5.5: how to make sure you're actually talking to it, which dial to turn, and where the tokens go when you're not looking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Opus 5.5 Changes the Claude Code Setup
&lt;/h2&gt;

&lt;p&gt;A quick recap of what shipped, and what each change means inside Claude Code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's 40% cheaper than Opus 5.&lt;/strong&gt; $4/M input, $20/M output, and $0.20/M cache reads. Claude Code sessions are mostly re-read context, so the cache price matters more than the headline price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Medium effort is the default.&lt;/strong&gt; Claude Code respects that: Opus 5.5 starts at &lt;code&gt;medium&lt;/code&gt;, while most other models start at &lt;code&gt;high&lt;/code&gt;. You can raise it per session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thinking is always on.&lt;/strong&gt; There's no thinking toggle to hunt for. Effort is the only dial, and &lt;code&gt;MAX_THINKING_TOKENS&lt;/code&gt; does nothing on Opus 5.5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output is 30%+ faster.&lt;/strong&gt; You'll feel this in the VS Code panel on long diffs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;More refusals.&lt;/strong&gt; The expanded safety classifiers apply in Claude Code too. If your repo is security tooling, expect some declines.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The catch: none of this matters if Claude Code isn't using Opus 5.5. And by default, on Foundry, it isn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Short Setup (and the Line Most Guides Miss)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmngvfj6rvgl7itrn7rm9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmngvfj6rvgl7itrn7rm9.png" alt="Install Claude Code CLI" width="571" height="369"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Install Claude Code CLI&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Deploy &lt;strong&gt;Claude Opus 5.5&lt;/strong&gt; from the Foundry Model Catalog to your Foundry resource, same as any other model. Give yourself the &lt;strong&gt;Azure AI User&lt;/strong&gt; or &lt;strong&gt;Cognitive Services User&lt;/strong&gt; role on that resource (either one is enough to call the model), and run &lt;code&gt;az login&lt;/code&gt;. No API key: Claude Code falls back to the Azure credential chain when no key is set, so your &lt;code&gt;az login&lt;/code&gt; session is the credential. If you plan to use the API Key then you will need to set your env key or the json file to have the &lt;code&gt;ANTHROPIC_FOUNDRY_API_KEY&lt;/code&gt; rather.&lt;/p&gt;
&lt;h2&gt;
  
  
  Remember to install the Extension in VSCode
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fntcdbkis6igwqhd0rbr8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fntcdbkis6igwqhd0rbr8.png" alt="VSCode Claude" width="799" height="281"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;VSCode Claude&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Now the part I care about. Instead of &lt;code&gt;setx&lt;/code&gt;-ing environment variables or pasting them into every shell, put them in &lt;strong&gt;&lt;code&gt;~/.claude/settings.json&lt;/code&gt;&lt;/strong&gt;. Both the CLI and the VS Code extension read that file, so you configure Foundry once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"env"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"CLAUDE_CODE_USE_FOUNDRY"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_FOUNDRY_RESOURCE"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your-foundry-resource-name"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-opus-5-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_DEFAULT_OPUS_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-opus-5-5"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two model lines do different jobs, and you want both:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ANTHROPIC_MODEL&lt;/code&gt;&lt;/strong&gt; makes Opus 5.5 the model for the session. &lt;strong&gt;This is the line most guides miss.&lt;/strong&gt; On Foundry, Claude Code's default model is Sonnet 4.5, not Opus. Setting only &lt;code&gt;ANTHROPIC_DEFAULT_OPUS_MODEL&lt;/code&gt; remaps the &lt;code&gt;opus&lt;/code&gt; alias, but you keep chatting with Sonnet until you switch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ANTHROPIC_DEFAULT_OPUS_MODEL&lt;/code&gt;&lt;/strong&gt; makes the &lt;code&gt;opus&lt;/code&gt; alias (in &lt;code&gt;/model&lt;/code&gt;, subagent definitions, and so on) resolve to your Opus 5.5 deployment instead of an older Opus you may not have deployed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use your actual deployment name if it isn't &lt;code&gt;claude-opus-5-5&lt;/code&gt;. &lt;code&gt;ANTHROPIC_FOUNDRY_RESOURCE&lt;/code&gt; takes the resource &lt;strong&gt;name&lt;/strong&gt; only, not a URL. If you need a private endpoint or custom domain, use &lt;code&gt;ANTHROPIC_FOUNDRY_BASE_URL&lt;/code&gt; instead. Don't set both.&lt;/p&gt;

&lt;p&gt;In VS Code, install the Claude Code extension and add one line to your VS Code &lt;code&gt;settings.json&lt;/code&gt; so it doesn't push you toward an Anthropic login:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"claudeCode.disableLoginPrompt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then check it, in the terminal or the VS Code panel: &lt;code&gt;/status&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;You're looking for the API provider set to &lt;strong&gt;Microsoft Foundry&lt;/strong&gt;, your resource name, and &lt;strong&gt;your Opus 5.5 deployment as the model&lt;/strong&gt;. If the model line says Sonnet, the &lt;code&gt;ANTHROPIC_MODEL&lt;/code&gt; line isn't being picked up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dev takeaway:&lt;/strong&gt; "Deployed in Foundry" and "used by Claude Code" are two different things. &lt;code&gt;/status&lt;/code&gt; is the five-second check that tells you which one you've got.&lt;/p&gt;

&lt;h2&gt;
  
  
  Effort: The Dial You'll Actually Touch
&lt;/h2&gt;

&lt;p&gt;In the SDK post, effort was a property on the request. In Claude Code it's a session setting, and there are three ways to set it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/effort&lt;/code&gt;&lt;/strong&gt; sets it for the current session: &lt;code&gt;/effort high&lt;/code&gt;, &lt;code&gt;/effort low&lt;/code&gt;, or &lt;code&gt;/effort auto&lt;/code&gt; to go back to the model default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The &lt;code&gt;/model&lt;/code&gt; picker&lt;/strong&gt; has an effort slider (left/right arrows). Whatever you pick there is remembered per model, so Opus 5.5 can keep its own setting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;CLAUDE_CODE_EFFORT_LEVEL&lt;/code&gt;&lt;/strong&gt; in the environment or the &lt;code&gt;env&lt;/code&gt; block wins over everything else, including &lt;code&gt;/effort&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point is easy to trip over: put &lt;code&gt;CLAUDE_CODE_EFFORT_LEVEL&lt;/code&gt; in your settings while testing, forget about it, and later &lt;code&gt;/effort high&lt;/code&gt; appears to do nothing. &lt;strong&gt;For interactive work, leave the env var out&lt;/strong&gt; and use &lt;code&gt;/effort&lt;/code&gt; or the &lt;code&gt;/model&lt;/code&gt; slider. Save the env var for scripted or CI runs where you want one fixed level. (The top-level &lt;code&gt;effortLevel&lt;/code&gt; user setting also won't apply to Opus 5.5, which is one more reason to use the per-model slider.)&lt;/p&gt;

&lt;p&gt;Here's how I map the levels to Claude Code work:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Effort&lt;/th&gt;
&lt;th&gt;What I use it for in Claude Code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;low&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Explain this file", rename a symbol, write a commit message&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;medium&lt;/code&gt; (default)&lt;/td&gt;
&lt;td&gt;Everyday feature work, bug fixes, writing tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;high&lt;/code&gt; / &lt;code&gt;xhigh&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Multi-file refactors, framework migrations, tricky concurrency bugs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Rarely, and only for the session: &lt;code&gt;/effort max&lt;/code&gt; when &lt;code&gt;xhigh&lt;/code&gt; clearly isn't getting there&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Thinking tokens are billed as output, at $20/M. Raising effort for the whole day costs you; raising it for one hard problem and dropping back is cheap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Opus Bill Hides
&lt;/h2&gt;

&lt;p&gt;Pinning everything to Opus 5.5 is easy. The surprise is how much &lt;em&gt;else&lt;/em&gt; then runs on Opus 5.5 too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background tasks.&lt;/strong&gt; Claude Code does small jobs behind the scenes, such as generating session titles. On the Anthropic API those go to Haiku. &lt;strong&gt;On Foundry they run on your primary model&lt;/strong&gt;, which is now Opus 5.5, unless you deploy a Haiku model and point &lt;code&gt;ANTHROPIC_DEFAULT_HAIKU_MODEL&lt;/code&gt; at it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Subagents.&lt;/strong&gt; When Claude Code fans work out to subagents (the Explore agent searching your repo, for example), each one uses the main conversation's model unless something says otherwise. That's Opus 5.5 for every file search. If you deploy a cheaper model, route subagents to it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"env"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"CLAUDE_CODE_USE_FOUNDRY"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_FOUNDRY_RESOURCE"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your-foundry-resource-name"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-opus-5-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_DEFAULT_OPUS_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-opus-5-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_DEFAULT_SONNET_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-sonnet-4-6"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_DEFAULT_HAIKU_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-haiku-4-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"CLAUDE_CODE_SUBAGENT_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-sonnet-4-6"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Again, those are deployment names, so use yours. Opus 5.5 does the planning and the edits in the main conversation; a cheaper model does the searching and reading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Only reference models you actually deployed.&lt;/strong&gt; Foundry has no startup model check, so Claude Code won't warn you about a typo or a missing deployment when it launches. You find out mid-session. While drafting this post, a subagent in my own session died with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The model claude-sonnet-4-5 is not available on your foundry deployment.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It had tried a model my resource didn't have. Everything else kept working, which is exactly why it's easy to miss. If you only deployed Opus 5.5, leave the Sonnet and Haiku lines out and accept that everything runs on Opus.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dev takeaway:&lt;/strong&gt; with one deployment, every token is an Opus token. That can be fine, since Opus 5.5 is cheaper than Opus 5, but make it a decision rather than a surprise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the $0.20 Cache Reads Count
&lt;/h2&gt;

&lt;p&gt;Cache reads at $0.20/M are the best part of the Opus 5.5 price sheet, and Claude Code is a cache-heavy workload: every turn re-sends your instructions, &lt;code&gt;CLAUDE.md&lt;/code&gt;, and the conversation so far. Caching is on automatically. Three things decide whether you actually get those cheap reads:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The default cache lifetime on Foundry is 5 minutes.&lt;/strong&gt; Step away for coffee, come back, and the next turn re-writes the whole context at full price. For long sessions with gaps, set &lt;code&gt;ENABLE_PROMPT_CACHING_1H&lt;/code&gt; to &lt;code&gt;1&lt;/code&gt; in the &lt;code&gt;env&lt;/code&gt; block. One-hour cache &lt;em&gt;writes&lt;/em&gt; are billed at a higher rate than 5-minute writes, so this pays off for long, stop-and-start sessions, not quick ones. (Recent Claude Code versions also have &lt;code&gt;CLAUDE_CODE_PROMPT_CACHE_TTL=1h&lt;/code&gt;, which applies only to the main conversation.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick your model at the start and stay on it.&lt;/strong&gt; Switching from Opus 5.5 to Sonnet and back mid-session throws away the cache each time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set up MCP servers before you start.&lt;/strong&gt; Some Azure-hosted deployments reject Claude Code's tool search, so it loads all MCP tools up front. Then connecting or removing an MCP server mid-session resets the cache.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Refusals Show Up in Your Editor Now
&lt;/h2&gt;

&lt;p&gt;In the SDK post I made refusal handling a required pattern, because a refusal is a successful HTTP response with no answer in it. In Claude Code you don't write that handler, but you'll still see the result: Opus 5.5 declines more requests than Opus 5 in the biology, cybersecurity, and reasoning-extraction categories.&lt;/p&gt;

&lt;p&gt;If you work on security tooling (scanners, fuzzers, detection rules), you'll occasionally hit a decline on a request that looks reasonable to you. Rephrase the request with the defensive context, or do that piece by hand. Don't try to engineer around the classifier; on a work resource, that's a conversation for your security team, not a prompt trick.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking What You Actually Spent
&lt;/h2&gt;

&lt;p&gt;Inside Claude Code, run &lt;strong&gt;&lt;code&gt;/usage&lt;/code&gt;&lt;/strong&gt; (&lt;code&gt;/cost&lt;/code&gt; is an alias). On Foundry you get the session's token counts, a &lt;strong&gt;prompt cache&lt;/strong&gt; line, and an &lt;strong&gt;estimated&lt;/strong&gt; dollar cost at list price. That estimate is great for "did turning effort up just double my session?", and the cache line tells you whether the 5-minute or 1-hour setting is actually working.&lt;/p&gt;

&lt;p&gt;It is not your bill. The real numbers live in &lt;strong&gt;Azure Cost Management&lt;/strong&gt; for the Foundry resource, at whatever price your agreement gives you. Anthropic's usage dashboards don't see Foundry traffic at all. For per-developer numbers across a team, Claude Code can export usage through OpenTelemetry. Tagging the Foundry resource (&lt;code&gt;team=...&lt;/code&gt;, &lt;code&gt;env=dev&lt;/code&gt;) makes chargeback easier.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Something's Off
&lt;/h2&gt;

&lt;p&gt;These are the Opus 5.5-specific problems I'd check first:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Likely cause&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;/status&lt;/code&gt; shows Sonnet, not Opus 5.5&lt;/td&gt;
&lt;td&gt;Only &lt;code&gt;ANTHROPIC_DEFAULT_OPUS_MODEL&lt;/code&gt; is set&lt;/td&gt;
&lt;td&gt;Add &lt;code&gt;ANTHROPIC_MODEL&lt;/code&gt; with your Opus 5.5 deployment name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;model ... is not available on your foundry deployment&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;An alias or subagent points at a model you didn't deploy&lt;/td&gt;
&lt;td&gt;Deploy it, or remove that line from &lt;code&gt;env&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;/effort&lt;/code&gt; seems to do nothing&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;CLAUDE_CODE_EFFORT_LEVEL&lt;/code&gt; is set and overrides it&lt;/td&gt;
&lt;td&gt;Remove the env var for interactive use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bill higher than &lt;code&gt;/usage&lt;/code&gt; suggests per session&lt;/td&gt;
&lt;td&gt;Background tasks and subagents running on Opus 5.5&lt;/td&gt;
&lt;td&gt;Deploy Haiku/Sonnet and set the Haiku and subagent model lines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Every turn after a break is expensive&lt;/td&gt;
&lt;td&gt;5-minute cache lifetime expired&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ENABLE_PROMPT_CACHING_1H=1&lt;/code&gt; for long sessions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;VS Code panel asks you to sign in to Anthropic&lt;/td&gt;
&lt;td&gt;Login prompt not disabled&lt;/td&gt;
&lt;td&gt;&lt;code&gt;"claudeCode.disableLoginPrompt": true&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;401&lt;/code&gt; / &lt;code&gt;403&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Missing role, or &lt;code&gt;az login&lt;/code&gt; in the wrong tenant&lt;/td&gt;
&lt;td&gt;Azure AI User or Cognitive Services User on the resource; &lt;code&gt;az login --tenant &amp;lt;id&amp;gt;&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One more: there's no &lt;code&gt;/logout&lt;/code&gt; on Foundry. To switch accounts or tenants, change your &lt;code&gt;az login&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom Line
&lt;/h2&gt;

&lt;p&gt;Opus 5.5 is the first Opus worth testing as your all-day default in Claude Code. It's cheaper, faster, and &lt;code&gt;medium&lt;/code&gt; effort is enough for most work. But "default" has to be deliberate on Foundry. Set &lt;code&gt;ANTHROPIC_MODEL&lt;/code&gt;, confirm it with &lt;code&gt;/status&lt;/code&gt;, decide whether background tasks and subagents should really run on Opus, and turn effort up only for the problem that needs it.&lt;/p&gt;

&lt;p&gt;Configure it once in &lt;code&gt;~/.claude/settings.json&lt;/code&gt;, and the CLI and VS Code both pick it up. Then keep an eye on &lt;code&gt;/usage&lt;/code&gt; for a week before you roll it out to the team.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Building AI features in C#? I write about practical, no-hype prompt engineering and Azure AI patterns for .NET developers. Check out &lt;a href="https://leanpub.com/promptengineeringfornetdevelopers" rel="noopener noreferrer"&gt;Prompt Engineering for .NET Developers&lt;/a&gt; — free, no Python required. You can also &lt;a href="https://taswar.zeytinsoft.com/subscribe/" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt; for more posts.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>vscode</category>
      <category>ai</category>
      <category>development</category>
    </item>
    <item>
      <title>AGENTS.md vs CLAUDE.md: Claude Code reads both now, just not at once.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 12:00:31 +0000</pubDate>
      <link>https://dev.to/manpreet171/agentsmd-vs-claudemd-claude-code-reads-both-now-just-not-at-once-hd</link>
      <guid>https://dev.to/manpreet171/agentsmd-vs-claudemd-claude-code-reads-both-now-just-not-at-once-hd</guid>
      <description>&lt;p&gt;Two files, one job: tell a coding agent about your repo. For a year the answer to "which one?" was simple. Claude Code read &lt;code&gt;CLAUDE.md&lt;/code&gt;. Almost everything else read &lt;code&gt;AGENTS.md&lt;/code&gt;. If you used both kinds of tool, you kept two files and watched them drift apart.&lt;/p&gt;

&lt;p&gt;On 18 September that changed. Claude Code 2.1.277 started reading &lt;code&gt;AGENTS.md&lt;/code&gt; natively. Some of the top results for this question still say it doesn't.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Claude Code reads AGENTS.md now. But only when there's no CLAUDE.md, and a single CLAUDE.local.md switches it off.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here's the current picture, the one setup that works on every version, and the part none of the comparison posts measure: what the file costs you on every turn, and what's missing from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short answer
&lt;/h2&gt;

&lt;p&gt;From Anthropic's &lt;a href="https://code.claude.com/docs/en/memory#agents-md" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;, as of today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Claude Code reads &lt;code&gt;AGENTS.md&lt;/code&gt; only when there's no &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/strong&gt; That means no &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;.claude/CLAUDE.md&lt;/code&gt; or &lt;code&gt;CLAUDE.local.md&lt;/code&gt; in the folder you're working in or any folder above it. Your personal &lt;code&gt;~/.claude/CLAUDE.md&lt;/code&gt; doesn't count, so that still loads alongside.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The trap.&lt;/strong&gt; Add a &lt;code&gt;CLAUDE.local.md&lt;/code&gt; for your own notes and your team's &lt;code&gt;AGENTS.md&lt;/code&gt; quietly stops loading. Nothing tells you.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;You can change the rule.&lt;/strong&gt; In &lt;code&gt;/config&lt;/code&gt;, "Project instructions" can be set to read either file (the default), both, or only &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Nearly everything else reads &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/strong&gt; Codex, Cursor, Copilot, Jules, Amp, Windsurf, Zed and more, per the list at &lt;a href="https://agents.md/" rel="noopener noreferrer"&gt;agents.md&lt;/a&gt;. Gemini CLI is the odd one out: it reads &lt;code&gt;GEMINI.md&lt;/code&gt; unless you point it elsewhere.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One more thing to check before any of that applies to you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ claude --version
2.1.257 (Claude Code)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the CLI on my own machine. It's older than 2.1.277, so here an &lt;code&gt;AGENTS.md&lt;/code&gt; on its own would be invisible to Claude Code, whatever the docs say. If your team updates at different speeds, some of you are reading the file and some aren't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which one should you use?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Only Claude Code touches the repo?&lt;/strong&gt; Keep &lt;code&gt;CLAUDE.md&lt;/code&gt;. There's no &lt;code&gt;AGENTS.md&lt;/code&gt; in this website's repo, because Claude Code is the only agent that works on it. A second file would be a second thing to keep true.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Several tools, or a team on mixed versions?&lt;/strong&gt; Make &lt;code&gt;AGENTS.md&lt;/code&gt; the one real file, and make &lt;code&gt;CLAUDE.md&lt;/code&gt; a single line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@AGENTS.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the setup Anthropic's docs recommend. Claude Code pulls the file in and never loads it twice, and it works on old versions too, because it doesn't depend on the new support at all. Two details that trip people up: an import inside a code fence is silently ignored, and on Windows a symlink instead of an import needs admin rights or Developer Mode, and Claude's own edit tools refuse to write through one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What not to do&lt;/strong&gt; is keep two hand-maintained copies. On the long-running &lt;a href="https://github.com/anthropics/claude-code/issues/6235" rel="noopener noreferrer"&gt;GitHub issue&lt;/a&gt; asking for &lt;code&gt;AGENTS.md&lt;/code&gt; support, it's a recurring complaint: the copies drift, and each agent ends up following different rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the file costs you
&lt;/h2&gt;

&lt;p&gt;Whichever file you pick, it's read at the start of every session and then sent again on every turn, like everything else in the conversation. So its size isn't a one-off. Here's this repo's &lt;code&gt;CLAUDE.md&lt;/code&gt;, checked by a small linter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; github:manpreet171/bridle agents CLAUDE.md
Linting CLAUDE.md

  ✔ 2 runnable &lt;span class="nb"&gt;command &lt;/span&gt;block&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt;
  ✘ does not say how to &lt;span class="nb"&gt;install&lt;/span&gt; / &lt;span class="nb"&gt;set &lt;/span&gt;up
  ✘ does not say how to run the tests
  ✔ how to build or run it
  &lt;span class="o"&gt;!&lt;/span&gt; 3 angle-bracket field&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt;, e.g. &amp;lt;slug&amp;gt; — check these are argument syntax, not unfilled template text
  ✔ nothing &lt;span class="k"&gt;in &lt;/span&gt;here tells the agent to &lt;span class="k"&gt;do &lt;/span&gt;something dangerous
  ✔ 12 prohibition&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt; — the agent knows where the edges are
  ✔ size: 916 words ≈ 1474 tokens, re-sent every turn

FAIL — 2 problem&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt; Your agent is reading this file on every run.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two failures don't apply here: it's a static site with no install step and no test suite. A general linter checks for a general project, so read it as a list of questions, not a verdict.&lt;/p&gt;

&lt;p&gt;The size line is the one that matters. About 1,500 tokens, on every turn. This project has run about 9,300 turns of Claude Code, so this one file accounts for roughly 14 million tokens, sent at the cheaper cached rate. That's an estimate, but the shape isn't. I measured where the rest of a session's tokens go in &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;a separate post&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Does a longer file make the agent worse at following it? Anthropic's &lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt; say to aim for under 200 lines, and its &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;best-practices guide&lt;/a&gt; warns that bloated files get ignored. The research is more mixed. An &lt;a href="https://arxiv.org/abs/2602.11988" rel="noopener noreferrer"&gt;ETH Zurich study&lt;/a&gt; found files written by people raised success rates slightly, files generated by an AI lowered them, and both added roughly 20% to cost. A &lt;a href="https://arxiv.org/abs/2605.10039" rel="noopener noreferrer"&gt;study of 1,650 Claude Code sessions&lt;/a&gt; found no measurable effect of length on following one simple rule. Length isn't proven to hurt. It is proven to cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's missing from it
&lt;/h2&gt;

&lt;p&gt;A file can be long and still not say the thing you actually need it to. Here's the same project from the other direction: not what the file says, but what I keep typing to the AI because the file doesn't say it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ npx -y toldya --dry
toldya · 13 sessions (1 Aug – 5 Oct) · 975 of your messages · 58 corrections

You keep telling your AI:
   1.   7×  keep it simple   (6 sessions)
   2.   3×  Don't confuse me   (3 sessions)

And 8 times you told it to try or check again: its first go missed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;916 words of rules, and neither line is in there. The two corrections I gave most often in this repo had never been written down, so every new session started without them and I typed them again.&lt;/p&gt;

&lt;p&gt;That's the real answer to "which file?". The name decides which tools read it. What's in it decides whether it was worth reading.&lt;/p&gt;

&lt;h2&gt;
  
  
  What belongs in it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What it can't work out on its own.&lt;/strong&gt; The odd build command, the folder that looks unused but isn't, the deploy that happens on push.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The edges.&lt;/strong&gt; What it must never touch, and what needs a person. Those are the lines that prevent damage.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What you keep correcting.&lt;/strong&gt; If you've typed it three times, it's a rule. Write it in your own words.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Not documentation.&lt;/strong&gt; Explaining the codebase is what the code is for. Every paragraph the agent could have read in the repo is a paragraph you pay for on every turn.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Not generated filler.&lt;/strong&gt; The ETH study's sharpest finding was that AI-written context files made results worse. Write it yourself, briefly.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Check your own
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;claude &lt;span class="nt"&gt;--version&lt;/span&gt;                             &lt;span class="c"&gt;# 2.1.277 or later to read AGENTS.md&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; github:manpreet171/bridle agents CLAUDE.md   &lt;span class="c"&gt;# or AGENTS.md&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; toldya &lt;span class="nt"&gt;--dry&lt;/span&gt;                          &lt;span class="c"&gt;# what you keep repeating&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both tools are free, read only your own files, and change nothing unless you say yes. &lt;a href="https://singhlabs.dev/bridle/" rel="noopener noreferrer"&gt;bridle&lt;/a&gt; reads the instruction file; &lt;a href="https://singhlabs.dev/toldya/" rel="noopener noreferrer"&gt;toldya&lt;/a&gt; reads your Claude Code history.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;This is how we set agents up for clients.&lt;/strong&gt; One short file that says what the agent can't guess and where the edges are, checked against what people actually keep correcting. Not a template, and never a generated one.&lt;/p&gt;

&lt;p&gt;Sources: every terminal block is a real run in this website's repo on 5 Oct 2026, the toldya one trimmed by one line (a link). Loading rules are from Anthropic's &lt;a href="https://code.claude.com/docs/en/memory#agents-md" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt; and the Claude Code changelog for 2.1.277. Tool support is from &lt;a href="https://agents.md/" rel="noopener noreferrer"&gt;agents.md&lt;/a&gt;. The studies are &lt;a href="https://arxiv.org/abs/2602.11988" rel="noopener noreferrer"&gt;arXiv 2602.11988&lt;/a&gt; and &lt;a href="https://arxiv.org/abs/2605.10039" rel="noopener noreferrer"&gt;arXiv 2605.10039&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/claude-md-rules-chat-history/" rel="noopener noreferrer"&gt;The best CLAUDE.md rules are hiding in your chat history&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;Where your Claude Code tokens actually go. Output is 0.2%.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/agents-md-vs-claude-md/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Code says it's done. Check the diff, not the paragraph.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 11:58:50 +0000</pubDate>
      <link>https://dev.to/manpreet171/claude-code-says-its-done-check-the-diff-not-the-paragraph-3g55</link>
      <guid>https://dev.to/manpreet171/claude-code-says-its-done-check-the-diff-not-the-paragraph-3g55</guid>
      <description>&lt;p&gt;"Done. All tests pass, and nothing else was touched." You read it, it's late, you merge.&lt;/p&gt;

&lt;p&gt;Claude Code's own issue tracker is full of what happens next. A test suite whose total quietly changed from 4,992 to 4,966 before the &lt;a href="https://github.com/anthropics/claude-code/issues/46940" rel="noopener noreferrer"&gt;"ALL PASSED"&lt;/a&gt;. Six tests &lt;a href="https://github.com/anthropics/claude-code/issues/45550" rel="noopener noreferrer"&gt;marked skip&lt;/a&gt;. A project reported 100% complete that an audit &lt;a href="https://github.com/anthropics/claude-code/issues/53983" rel="noopener noreferrer"&gt;put nearer 60%&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The summary isn't a lie. It's written to explain the work, and nobody explains the parts they'd rather you didn't see.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Asking the agent to try harder doesn't fix that. Checking does. Here's what goes wrong, what doesn't work, and what happened when I held this site's own Claude Code summaries up against their real diffs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four ways "done" isn't done
&lt;/h2&gt;

&lt;p&gt;Read enough of those issues and the same four shapes keep coming back:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The tests got easier, not the code better.&lt;/strong&gt; A test skipped, an assertion loosened, a test command narrowed to the ones that pass. The suite is green because there's less of it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The count got reframed.&lt;/strong&gt; A failure becomes "a flaky timeout". A smaller total gets reported as everything passing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Intent got ticked off as work.&lt;/strong&gt; A to-do marked done because it was planned, or a stub counted as a feature.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A file changed and never got mentioned.&lt;/strong&gt; The quietest one, and the one that bites three days later.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's common enough to measure. A &lt;a href="https://arxiv.org/abs/2605.29442" rel="noopener noreferrer"&gt;study of 20,574 real agent sessions&lt;/a&gt; found inaccurate self-reporting was about 23% of all the misbehaviour it caught, and a growing share over time. Anthropic's own &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;best-practices guide&lt;/a&gt; says it plainly: Claude stops when the work looks done.&lt;/p&gt;

&lt;h2&gt;
  
  
  What doesn't work
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Asking "did you actually do it?"&lt;/strong&gt; The same model that wrote the summary checks it with the same blind spot. In one report it &lt;a href="https://github.com/anthropics/claude-code/issues/39907" rel="noopener noreferrer"&gt;said yes three times&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A rule in &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/strong&gt; One reporter had a 250-line file, memory notes and hooks, and &lt;a href="https://github.com/anthropics/claude-code/issues/37818" rel="noopener noreferrer"&gt;still got false "done"s&lt;/a&gt;. In another, the rule against skipping tests was right there in context and got broken anyway. A rule asks. A check doesn't need to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I found in my own sessions
&lt;/h2&gt;

&lt;p&gt;Most writing on this describes the problem. Almost none of it shows a real summary next to its real diff, so I did that with two pieces of work Claude Code finished on this website today. For each one I took the summary it gave me at the end, word for word, and ran it through &lt;a href="https://singhlabs.dev/plumb/" rel="noopener noreferrer"&gt;plumb&lt;/a&gt;, a small tool that holds a summary against &lt;code&gt;git diff&lt;/code&gt; and prints only what doesn't match.&lt;/p&gt;

&lt;p&gt;The first was publishing a blog post, along with a fix to a script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ plumb check summary.md --base a12ccca
plumb — 9 files changed, 1 named in the summary

changed but never mentioned  (read these first)
  · .assetsignore modified
  · LOG.md modified
  · assets/blog/og-claude-code-token-usage.png added
  · baggage/index.html modified
  · blog/claude-code-token-usage/index.html added
  · blog/feed.xml modified
  · blog/index.html modified
  · sitemap.xml modified

8 things the summary did not tell you.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eight flags. I went through them one by one. Seven were things the summary did describe, just not by file name: it gave the post's URL, called the image "the social card", and said "that folder is now excluded from publishing" instead of naming &lt;code&gt;.assetsignore&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One wasn't. &lt;strong&gt;&lt;code&gt;baggage/index.html&lt;/code&gt;&lt;/strong&gt;: Claude added a link to a product page while publishing the post, and the summary never said so anywhere. Harmless, as it happens. But I only know it's harmless because I looked.&lt;/p&gt;

&lt;p&gt;The second was a batch of SEO fixes: 21 files changed, and again only one named by path. This time the summary described all 20 others in categories, "19 pages and the sitemap", and every one checked out. plumb also flagged two files as "mentioned but not changed", which the summary had only referred to. Nothing was missing.&lt;/p&gt;

&lt;p&gt;So, across two real summaries: 30 files changed, 2 named by path, one change genuinely left out. No skipped tests, no removed guards. Neither summary lied. One still hid a change, and I'd have merged it without blinking.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually works
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Read the diff, not the paragraph.&lt;/strong&gt; &lt;code&gt;/diff&lt;/code&gt; in Claude Code, or &lt;code&gt;git diff --stat&lt;/code&gt;, shows every file that changed. The summary is a guide to the diff, never a replacement for it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Make it end with a file list.&lt;/strong&gt; Ask for every file it touched, by path, at the end of each task. My two summaries were noisy to check because they talked in categories. A list turns "did it mention everything?" into a mechanical question.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Check on every stop, automatically.&lt;/strong&gt; Claude Code's &lt;a href="https://code.claude.com/docs/en/hooks" rel="noopener noreferrer"&gt;Stop hook&lt;/a&gt; receives the agent's last message, so a check can run the moment it says it's finished. &lt;a href="https://singhlabs.dev/trust-issues/" rel="noopener noreferrer"&gt;trust issues&lt;/a&gt; is a free plugin that does exactly that: the same four checks, every time, with nobody remembering to run them.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Count the tests, before and after.&lt;/strong&gt; If the number went down, ask why before you read anything else.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Get a second pair of eyes that didn't do the work.&lt;/strong&gt; Anthropic's guide suggests a separate verification step that checks nothing outside the task changed. A fresh session doesn't share the first one's reasons for skipping a test.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it on yours
&lt;/h2&gt;

&lt;p&gt;Save the agent's last message to a file, then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; github:manpreet171/plumb
&lt;span class="nv"&gt;$ &lt;/span&gt;plumb check summary.md                &lt;span class="c"&gt;# against your uncommitted changes&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;plumb check summary.md &lt;span class="nt"&gt;--base&lt;/span&gt; main    &lt;span class="c"&gt;# against a branch or commit&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Exit code 1 means it found something, so it works as a CI gate too. What it won't do, so you don't over-trust it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It matches names literally.&lt;/strong&gt; A URL or a category doesn't count as a mention, which is why my first run flagged seven things that weren't really missing. Ask for a file list and that noise goes away.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It can't tell whether the code works.&lt;/strong&gt; It finds what was left out of the story, not bugs. Run the tests for that.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Its "quiet cuts" are patterns.&lt;/strong&gt; It spots a removed assertion or an added &lt;code&gt;.skip&lt;/code&gt;, not every possible way to weaken a check.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;This is how we build agents for clients.&lt;/strong&gt; Nothing gets called done because the agent said so. Every run ends with what changed, and something other than the agent checks it.&lt;/p&gt;

&lt;p&gt;Sources: the plumb output is a real run (plumb 1.1.0) on 5 Oct 2026, against the exact end-of-task summary Claude Code gave for each piece of work, with the summary file's path shortened. The second run's output is described rather than shown because it lists 20 files. Incidents are from the linked GitHub issues on anthropics/claude-code; the session study is &lt;a href="https://arxiv.org/abs/2605.29442" rel="noopener noreferrer"&gt;arXiv 2605.29442&lt;/a&gt;; hook and review guidance is from Anthropic's &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;best practices&lt;/a&gt; and &lt;a href="https://code.claude.com/docs/en/hooks" rel="noopener noreferrer"&gt;hooks&lt;/a&gt; docs.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/a-green-tick-over-nothing/" rel="noopener noreferrer"&gt;My mailer printed “campaign sent”. It had sent to nobody.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;Where your Claude Code tokens actually go. Output is 0.2%.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/claude-code-says-done/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>testing</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Where your Claude Code tokens actually go. Output is 0.2%.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 11:38:18 +0000</pubDate>
      <link>https://dev.to/manpreet171/where-your-claude-code-tokens-actually-go-output-is-02-n8e</link>
      <guid>https://dev.to/manpreet171/where-your-claude-code-tokens-actually-go-output-is-02-n8e</guid>
      <description>&lt;p&gt;You hit the usage limit before lunch. So you do what everyone online says: make the AI talk less. Ban the "Great question!", cut the summaries, install one of the skills that trims its replies.&lt;/p&gt;

&lt;p&gt;I wanted to know where the tokens actually go, so I counted. Not estimated. Counted, from the transcripts Claude Code already keeps on your machine: every session still on mine, 173 of them, about 60,000 turns.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Everything Claude wrote back to me was 0.2% of the tokens. The rest was my own session, fed back in.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That number is real, but it isn't the whole story either, because most of those re-sent tokens are cheap ones. This post is both halves: where the tokens go, what they really cost, and the few things that actually move the number.&lt;/p&gt;

&lt;h2&gt;
  
  
  The count
&lt;/h2&gt;

&lt;p&gt;Here's the summary across every project, straight from the terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ baggage --all
baggage — 173 sessions, all projects, 60,454 turns

  before you typed anything              75k  tokens
    system prompt, every tool schema, every skill, CLAUDE.md.
    paid again on all 60,454 turns = 4.5B tokens, 17% of the bill,
    whether you called any of it or not.

  conversation you can see             27.8M  tokens
  what the API billed                  26.0B  tokens

  a 934x gap. Some of it is the fixed cost above; the rest
  is everything you picked up being re-sent on every later turn.

  output tokens           56.6M   0.2%
  re-sent context         25.3B   97.5%   &amp;lt;- the bill
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two lines matter. &lt;strong&gt;Output&lt;/strong&gt; (every reply, every line of code Claude wrote, its thinking included) is 56.6 million tokens. &lt;strong&gt;Re-sent context&lt;/strong&gt; is 25.3 billion. That's roughly 450 to 1.&lt;/p&gt;

&lt;p&gt;I'm not the first to see this shape. A &lt;a href="https://github.com/anthropics/claude-code/issues/24147" rel="noopener noreferrer"&gt;GitHub issue&lt;/a&gt; on Claude Code found the same thing in 30 days of someone else's transcripts, and a &lt;a href="https://dev.to/ploofnexa/i-measured-where-claude-code-actually-spends-tokens-968-is-re-reading-history-my-typing-was-16gm"&gt;dev.to post&lt;/a&gt; measured 0.5% output across 32 sessions. Different people, different work, same picture. What none of them show is &lt;em&gt;which things&lt;/em&gt; in a session cost the most, and what any of it means for the bill. That's the rest of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it works like that
&lt;/h2&gt;

&lt;p&gt;The model has no memory between turns. Every time you press enter, Claude Code sends the whole conversation again: the system prompt, the tool definitions, every file it read, every command output, every reply so far, and then your new message.&lt;/p&gt;

&lt;p&gt;Think of a suitcase you have to carry up every flight of stairs. Pack a brick on the second floor of a 200-floor building and you carry it up 198 more flights. A 4,000-token build log that lands at turn 12 of a 200-turn session isn't 4,000 tokens. It's 4,000 sent 188 more times.&lt;/p&gt;

&lt;p&gt;So the cost of anything in a session isn't its size. It's its size times the number of turns it stays.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching makes it cheaper, not smaller
&lt;/h2&gt;

&lt;p&gt;This is the part most guides get half right. Claude Code caches the conversation, and Anthropic &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;prices a cache read&lt;/a&gt; at a tenth of normal input (a twentieth on Opus 5.5, a fortieth on Fable 5.1). Output costs five times input. So 97.5% of the tokens is not 97.5% of the money.&lt;/p&gt;

&lt;p&gt;I priced my own split at those published rates, with the one-hour cache a subscription uses. Output comes to &lt;strong&gt;7 to 14% of the cost&lt;/strong&gt;, depending on the model. Re-reading and re-caching the session is the other &lt;strong&gt;86 to 93%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Less dramatic than 0.2%. Still the opposite of what "make it talk less" assumes.&lt;/p&gt;

&lt;p&gt;On a Pro or Max plan you never see dollars, you see a limit. Anthropic's &lt;a href="https://code.claude.com/docs/en/costs" rel="noopener noreferrer"&gt;cost docs&lt;/a&gt; say re-read history still draws on your usage, at the cached rate, so a one-line question late in a long session draws usage for the whole session. How heavily a cached token counts against the limit isn't published. I've seen people online insist it's full weight, and others insist it's free. Nobody I found has shown either, so I'm not going to guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  The floor you pay before you type
&lt;/h2&gt;

&lt;p&gt;That first block of the output is the one nobody shows you. Before you've typed a word, a session already carries the system prompt, the built-in tools, the descriptions of every skill you've installed, your &lt;code&gt;CLAUDE.md&lt;/code&gt; files and anything else that loads at start. A typical session on my machine starts at about 75,000 tokens.&lt;/p&gt;

&lt;p&gt;It rides along on every turn. Across all my sessions, that floor alone is &lt;strong&gt;17% of everything billed&lt;/strong&gt;, whether I used any of it or not. It's also the one line you can fix in thirty seconds: every paragraph of &lt;code&gt;CLAUDE.md&lt;/code&gt; you don't need is paid again on every turn of every session, and so is the description of every skill you never call.&lt;/p&gt;

&lt;h2&gt;
  
  
  The heaviest things I carry
&lt;/h2&gt;

&lt;p&gt;Here's one project, this website, with the list of what's costing the most rent, trimmed for length:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ baggage
baggage — 13 sessions in singhlabs, 9,307 turns

  before you typed anything              84k  tokens
    system prompt, every tool schema, every skill, CLAUDE.md.
    paid again on all 9,307 turns = 779.0M tokens, 18% of the bill,
    whether you called any of it or not.

  output tokens            7.1M   0.2%
  re-sent context          4.2B   98.0%   &amp;lt;- the bill

  HEAVIEST THINGS YOU ARE STILL CARRYING
  (rent = its size x the turns it stayed in context)

   304.4M   19.4%  3769x  assistant reply
   282.4M   18.0%  1376x  your message
   244.2M   15.6%  1415x  claude-in-chrome · browser_batch
    54.6M    3.5%   145x  WebSearch
    25.6M    1.6%   275x  claude-in-chrome · javascript_tool
    21.5M    1.4%    17x  claude-in-chrome · get_page_text
    14.9M    0.9%   159x  Agent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things surprised me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude's own replies are the biggest line: 19.4%.&lt;/strong&gt; Not because they were expensive to write. Writing all of them was part of the 0.2%. They're expensive because each one stays in the session and is sent again on every turn after it. So the "talk less" skills aren't wrong. They're right for the wrong reason: a short reply saves you far more in re-sends than it ever cost to write.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Browser automation is close behind: 15.6%.&lt;/strong&gt; And that's the text alone: page contents, element lists, logs of each click. baggage doesn't count images, so every screenshot rides along on top of that figure, uncounted. A session that clicks through a website carries a stack of pictures of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually moves the number
&lt;/h2&gt;

&lt;p&gt;Ranked by what the counts above say, not by what's easiest to write about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Start fresh between unrelated tasks.&lt;/strong&gt; &lt;code&gt;/clear&lt;/code&gt; drops everything you've been carrying, and Anthropic's docs say it costs nothing. The brick stays on floor two.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Lower the floor.&lt;/strong&gt; Uninstall skills you don't use, keep &lt;code&gt;CLAUDE.md&lt;/code&gt; to what Claude can't work out on its own, and run &lt;code&gt;/context&lt;/code&gt; once to see what loads before you type.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Keep big output out of the main session.&lt;/strong&gt; Send a long log to a file and search it, instead of printing it into the chat where it's carried to the end. Hand noisy exploration and browser work to a subagent, so only its answer comes back.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Don't break the cache mid-session.&lt;/strong&gt; Per the &lt;a href="https://code.claude.com/docs/en/prompt-caching" rel="noopener noreferrer"&gt;caching docs&lt;/a&gt;, switching model, changing effort, turning on fast mode or changing MCP servers can throw it away, and then the whole session is written to cache again at the higher rate. Decide those at the start.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Shorter replies, for the right reason.&lt;/strong&gt; Ask for the answer without the essay. It helps, because the essay gets carried.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On a paid plan, &lt;code&gt;/usage&lt;/code&gt; now shows which skills, subagents and MCP servers used your allowance. Worth a look before changing anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Count your own
&lt;/h2&gt;

&lt;p&gt;The tool that printed everything above is called &lt;a href="https://singhlabs.dev/baggage/" rel="noopener noreferrer"&gt;baggage&lt;/a&gt;. It's free, it reads the transcripts already on your machine, and nothing leaves it. The name on npm belongs to someone else, so install it from the repo:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; github:manpreet171/baggage
&lt;span class="nv"&gt;$ &lt;/span&gt;baggage          &lt;span class="c"&gt;# this project&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;baggage &lt;span class="nt"&gt;--all&lt;/span&gt;    &lt;span class="c"&gt;# every project&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What it doesn't tell you, so you don't over-read it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tokens, not dollars.&lt;/strong&gt; The totals are the exact counts the API reported. Pricing them is your model and your plan.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Item sizes are estimates.&lt;/strong&gt; The list ranks things with a rough four-characters-a-token ruler, the same ruler for everything. The totals above it are exact.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Only what's still on your machine.&lt;/strong&gt; Claude Code keeps about 30 days of history by default.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;One heavy user.&lt;/strong&gt; These are my numbers. Yours will differ, which is the point of running it.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;This is how we work.&lt;/strong&gt; Measure what's really happening before changing anything, then change the thing the numbers point at. It's the same rule we use when an AI system we build for a business starts costing more than it should.&lt;/p&gt;

&lt;p&gt;Sources: both terminal blocks are real baggage runs on my own machine on 5 Oct 2026, the second trimmed to whole lines for length. Prices and cache multipliers are from &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Anthropic's pricing page&lt;/a&gt;; the cost split weights my exact token counts by them. Cache and usage behaviour is from Claude Code's &lt;a href="https://code.claude.com/docs/en/costs" rel="noopener noreferrer"&gt;costs&lt;/a&gt; and &lt;a href="https://code.claude.com/docs/en/prompt-caching" rel="noopener noreferrer"&gt;prompt caching&lt;/a&gt; docs.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/claude-md-rules-chat-history/" rel="noopener noreferrer"&gt;The best CLAUDE.md rules are hiding in your chat history&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/loop-engineering-map/" rel="noopener noreferrer"&gt;I researched loop engineering to build a product. I built nothing.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>productivity</category>
      <category>llm</category>
    </item>
    <item>
      <title>MCP Ecosystem Week 41: When Developer Choice Outpaces Your Allowlist</title>
      <dc:creator>curatedmcp</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:54:39 +0000</pubDate>
      <link>https://dev.to/curatedmcp/mcp-ecosystem-week-41-when-developer-choice-outpaces-your-allowlist-1cg8</link>
      <guid>https://dev.to/curatedmcp/mcp-ecosystem-week-41-when-developer-choice-outpaces-your-allowlist-1cg8</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.curatedmcp.com/blog/week-2026-41" rel="noopener noreferrer"&gt;curatedmcp.com/blog/week-2026-41&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h1&gt;
  
  
  MCP Ecosystem Week 41: When Developer Choice Outpaces Your Allowlist
&lt;/h1&gt;

&lt;p&gt;The MCP ecosystem continues its steady climb toward ubiquity in AI-assisted development. This week brought no new catalog entries, but the view counts tell a story platform teams need to hear: developers are gravitating toward integrations that connect their existing tool chains—GitHub, OpenAI, Figma, Anthropic—directly into their coding agents. The question isn't whether your team will use these servers. It's whether you've decided to govern them or discovered them in an audit six months from now.&lt;/p&gt;

&lt;h2&gt;
  
  
  This Week in MCP
&lt;/h2&gt;

&lt;p&gt;No new servers were added to the CuratedMCP catalog this week. That pause is worth noting. It suggests we're entering a consolidation phase, where platform teams are working through the governance implications of the 80 risk-classified servers already available rather than chasing new additions.&lt;/p&gt;

&lt;p&gt;Use this breathing room. If you haven't audited which MCP servers your developers are actually running across Claude Code, Cursor, Windsurf, and GitHub Copilot, now is the time. Shadow usage—servers spinning up without your knowledge—is the governance blind spot we see most often in platform teams shipping AI coding tools at scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  On the Radar
&lt;/h2&gt;

&lt;p&gt;The five most-viewed servers this week reveal clear patterns in how developers want to extend their AI agents:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.curatedmcp.com/marketplace/github-copilot-mcp" rel="noopener noreferrer"&gt;GitHub Copilot MCP&lt;/a&gt;&lt;/strong&gt; (98k views) and &lt;strong&gt;&lt;a href="https://www.curatedmcp.com/marketplace/github-mcp" rel="noopener noreferrer"&gt;GitHub MCP&lt;/a&gt;&lt;/strong&gt; (76k views) dominate the list. The first wraps Copilot's own intelligence; the second manages repositories, issues, and workflows directly. Governance consideration: these servers integrate with your source control and CI/CD. Ensure your SSO and RBAC policies enforce per-repository or per-org limits. Token logging matters here—agent-driven GitHub operations can generate audit trails you'll need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.curatedmcp.com/marketplace/openai-mcp" rel="noopener noreferrer"&gt;OpenAI MCP&lt;/a&gt;&lt;/strong&gt; (87k views) gives agents direct access to GPT-4o, DALL-E, Whisper, and embeddings. Before allowlisting, clarify your org's multi-model policy. Are developers expected to route through Claude, or can they call competing LLMs? Token spend visibility becomes critical when agents can invoke multiple model providers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.curatedmcp.com/marketplace/figma-mcp" rel="noopener noreferrer"&gt;Figma MCP&lt;/a&gt;&lt;/strong&gt; (82k views) opens design files and component tokens to agents. Supply-chain risk consideration: your design system and brand assets flow through this connection. Confirm Figma's auth model (OAuth 2.0, service tokens) aligns with your identity provider. Audit who can access what.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.curatedmcp.com/marketplace/anthropic-claude-mcp" rel="noopener noreferrer"&gt;Anthropic Claude MCP&lt;/a&gt;&lt;/strong&gt; (76k views) nests Claude within Claude—useful for specialized reasoning tasks, but introduces recursive LLM calls and compounding token cost. If you're not tracking sub-agent behavior, this server will be a surprise on your monthly bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  Governance Take
&lt;/h2&gt;

&lt;p&gt;Here's what we're hearing from platform teams in the field: &lt;strong&gt;allowlist drift&lt;/strong&gt;. A security team approves MCP servers for Cursor in Q4. By Q1, developers are spinning up the same servers in Claude Code, and the approval matrix hasn't kept pace. Different IDEs, different enforcement points, no single source of truth.&lt;/p&gt;

&lt;p&gt;Add to that the TokenShield gap: most organizations have spend visibility for their primary model provider (usually Anthropic), but the moment agents start invoking GitHub APIs, Figma endpoints, and sub-agent LLM calls, the cost picture fragments. You're paying for tokens. You're also paying for API calls downstream that don't show up in your Claude bill.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The play:&lt;/strong&gt; build your allowlist once, enforce it everywhere. Map your MCP servers to teams, risk classifications, and cost centers. Log every server invocation—not just token count, but which user, which IDE, which repository access, which external API was called. TokenShield gives you spend visibility and measured, opt-in optimization across your deployed servers. But visibility only matters if your governance layer can act on it.&lt;/p&gt;




&lt;p&gt;Govern MCP usage across your team with &lt;a href="https://www.curatedmcp.com" rel="noopener noreferrer"&gt;CuratedMCP&lt;/a&gt; — or scan your own stack free at &lt;a href="https://www.curatedmcp.com/auditor" rel="noopener noreferrer"&gt;https://www.curatedmcp.com/auditor&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>claude</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Move a Claude Code task to a teammate without sharing a login</title>
      <dc:creator>Almast </dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:47:20 +0000</pubDate>
      <link>https://dev.to/wagglet/move-a-claude-code-task-to-a-teammate-without-sharing-a-login-3m6g</link>
      <guid>https://dev.to/wagglet/move-a-claude-code-task-to-a-teammate-without-sharing-a-login-3m6g</guid>
      <description>&lt;p&gt;A coding task often stalls because the person who wrote the brief cannot run it right now. Passing a provider login to someone else blurs who executed the work and what access they had. A better handoff transfers the task specification, then lets the next person work with their own authorized tools.&lt;/p&gt;

&lt;p&gt;Here is a small template for doing that. It works as a Markdown issue even if your team never adopts a task platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The handoff file
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Task: Add an empty-state message to the projects list&lt;/span&gt;

&lt;span class="gu"&gt;## Outcome&lt;/span&gt;
When the API returns zero projects, show:
“No projects yet. Create one to get started.”

&lt;span class="gu"&gt;## Scope&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Change the projects-list component and its focused tests.
&lt;span class="p"&gt;-&lt;/span&gt; Keep loading and API-error states distinct from the empty state.
&lt;span class="p"&gt;-&lt;/span&gt; Do not change the API or create a project automatically.

&lt;span class="gu"&gt;## Starting point&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Repository: &lt;span class="nt"&gt;&amp;lt;repo&lt;/span&gt; &lt;span class="na"&gt;URL&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Base branch or commit: &lt;span class="nt"&gt;&amp;lt;exact&lt;/span&gt; &lt;span class="na"&gt;ref&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Relevant paths: &lt;span class="nt"&gt;&amp;lt;component&amp;gt;&lt;/span&gt;, &lt;span class="nt"&gt;&amp;lt;tests&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Command to run: &lt;span class="nt"&gt;&amp;lt;project&lt;/span&gt;&lt;span class="err"&gt;'&lt;/span&gt;&lt;span class="na"&gt;s&lt;/span&gt; &lt;span class="na"&gt;documented&lt;/span&gt; &lt;span class="na"&gt;test&lt;/span&gt; &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;

&lt;span class="gu"&gt;## Access&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Author may write the brief and review the result.
&lt;span class="p"&gt;-&lt;/span&gt; Executor uses their own repository and Claude Code access.
&lt;span class="p"&gt;-&lt;/span&gt; Executor stops if either permission is missing.
&lt;span class="p"&gt;-&lt;/span&gt; No provider credentials, session cookies, or personal tokens are included.

&lt;span class="gu"&gt;## Evidence to return&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Commit or pull-request URL
&lt;span class="p"&gt;-&lt;/span&gt; Files changed
&lt;span class="p"&gt;-&lt;/span&gt; Test command and actual output
&lt;span class="p"&gt;-&lt;/span&gt; Screenshot only if captured from the running app
&lt;span class="p"&gt;-&lt;/span&gt; Any untested behavior or remaining risk

&lt;span class="gu"&gt;## Acceptance&lt;/span&gt;
A reviewer checks the empty, loading, and error states,
then accepts or returns the task with specific changes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact repository paths and command belong in the real task. Leaving them vague forces the executor to rediscover scope and makes review harder.&lt;/p&gt;

&lt;h2&gt;
  
  
  A worked handoff
&lt;/h2&gt;

&lt;p&gt;Suppose Maya prepares this brief but has no time to implement it. Arun has authorized access to the repository and uses his own Claude Code account. Maya gives Arun the task, not her login. Arun checks out the stated base ref, inspects the three UI states, edits the component, and runs the repository's documented tests.&lt;/p&gt;

&lt;p&gt;His delivery should identify the changed files and real commit, name the exact test command and result, and say whether he performed a visual check. Those details must come from the actual run. “Tests passed” is not acceptable unless the executor ran them and can identify the command that produced the result.&lt;/p&gt;

&lt;p&gt;Maya then reviews the diff and evidence. If the empty message appears while the API is still loading, she returns the task with that finding. If it meets the acceptance criteria, she accepts it. Delivery and acceptance are different events.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the separation buys you
&lt;/h2&gt;

&lt;p&gt;This pattern gives a team &lt;strong&gt;shared AI execution capacity&lt;/strong&gt; through transferable work while keeping each person's account and permissions separate. The author owns scope and acceptance criteria. The executor owns the work performed with their own authorized environment. The reviewer owns the verdict.&lt;/p&gt;

&lt;p&gt;A task system can preserve those boundaries and make the handoff easier to inspect; &lt;a href="https://wagglet.com/blog/close-team-skill-gaps-with-ai-task-handoffs" rel="noopener noreferrer"&gt;Wagglet's discussion of AI task handoffs&lt;/a&gt; is one example. The same discipline still matters in a plain issue tracker.&lt;/p&gt;

&lt;p&gt;It does not solve missing access, unclear requirements, or unreliable tests. When those are absent, improve the brief or stop the run. Borrowing another person's credentials is not a fix.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This article was drafted with AI assistance.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>coding</category>
      <category>softwaredevelopment</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>The best CLAUDE.md rules are hiding in your chat history.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:33:16 +0000</pubDate>
      <link>https://dev.to/manpreet171/the-best-claudemd-rules-are-hiding-in-your-chat-history-30ld</link>
      <guid>https://dev.to/manpreet171/the-best-claudemd-rules-are-hiding-in-your-chat-history-30ld</guid>
      <description>&lt;p&gt;Quick question. What's the one sentence you've typed to your AI more than any other?&lt;/p&gt;

&lt;p&gt;Not the one you &lt;em&gt;think&lt;/em&gt; you type. The one you actually type.&lt;/p&gt;

&lt;p&gt;I didn't know mine either. So I went back through three months of my own Claude Code history and counted. It was "keep it simple." &lt;strong&gt;42 times&lt;/strong&gt;, across 38 different sessions. And that rule was written nowhere my AI could read it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every correction you type and don't write down dies with the session.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you use Claude Code, Cursor or Codex every day, you have your own "keep it simple." This post is about finding it, writing it down, and checking whether it stuck. At the end there's a one-minute way to see your own top line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 42 times
&lt;/h2&gt;

&lt;p&gt;This is a real run on my real history, 143 Claude Code sessions between July and September:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;toldya · 143 sessions (2 Jul – 25 Sept) · 6379 of your messages · 528 corrections

You keep telling your AI:
   1.  42×  keep it simple   (38 sessions)
   2.  29×  dont complex this   (18 sessions)
   3.  15×  dont assume   (15 sessions)
   4.  13×  i dont want later on   (12 sessions)
   5.   9×  dont think too much   (9 sessions)

And 71 times you told it to try or check again: its first go missed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Numbers one and two are the same complaint, one with my typo. Together that's 71 times I told my AI to stop overbuilding. The last line is a different 71: "try again" and "check again", which isn't a rule, just an honest measure of how often the first attempt missed.&lt;/p&gt;

&lt;p&gt;This isn't a complaint about Claude. It's genuinely good at what I use it for. It's that every one of those corrections was a rule I never wrote down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why your AI forgets everything
&lt;/h2&gt;

&lt;p&gt;A coding agent doesn't remember yesterday. Each session starts from zero. What you corrected on Tuesday afternoon is gone by Wednesday morning.&lt;/p&gt;

&lt;p&gt;The one thing it reliably reads at the start of every session is a small text file in your project called &lt;code&gt;CLAUDE.md&lt;/code&gt;. Other tools use &lt;code&gt;AGENTS.md&lt;/code&gt;, same idea.&lt;/p&gt;

&lt;p&gt;Think of it as a sticky note on the monitor for a brilliant new colleague with amnesia. They wake up every day with no memory of you. The note is the only thing that carries over. If it's not on the note, it's gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everyone says "just add a rule"
&lt;/h2&gt;

&lt;p&gt;Every CLAUDE.md guide gives the same advice: when your AI repeats a mistake, add a rule for it. It's good advice. I agree with it.&lt;/p&gt;

&lt;p&gt;The gap is that &lt;strong&gt;you don't notice you're repeating yourself.&lt;/strong&gt; My "keep it simple" came up in 38 sessions, across different projects, weeks apart. Each time it felt like a one-off. I never once thought "this is the thirtieth time." Repetition across days is invisible from inside the day.&lt;/p&gt;

&lt;p&gt;That's why none of my top five were in my CLAUDE.md. Not laziness. I just had no idea.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why grep lies to you
&lt;/h2&gt;

&lt;p&gt;Claude Code keeps your history on your own machine, as text files in &lt;code&gt;~/.claude/projects&lt;/code&gt;. So the obvious move is to search it. I searched for "keep it simple." It said &lt;strong&gt;747&lt;/strong&gt;. The real number was 42.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F35jefn72fd7o2zvmpohn.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F35jefn72fd7o2zvmpohn.webp" alt="Bar chart: a plain search of the history counts 747 matches for keep it simple; counting only what I actually typed gives 42" width="799" height="359"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Same history, same phrase. 18 times too big.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The files hold everything, not just what you typed: the AI repeating your words back, summaries written when a chat gets long, tool output, logs and briefs you pasted. Counting properly means keeping only the short messages you typed, dropping questions and pastes, and grouping the ways you say the same thing. "Don't complicate it", "dont complex this" and "dont comlicate that" are one habit, not three.&lt;/p&gt;
&lt;h2&gt;
  
  
  So I built a tiny thing
&lt;/h2&gt;

&lt;p&gt;It's called &lt;a href="https://singhlabs.dev/toldya/" rel="noopener noreferrer"&gt;toldya&lt;/a&gt;, as in "I told you so." A free command-line tool that does the counting above, then offers to write the results onto the sticky note for you. One command, Node 18 or newer, no account:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx toldya --all
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It reads your Claude Code history on your machine, shows the list above, and asks about each one: &lt;strong&gt;y&lt;/strong&gt; adds it as a rule, &lt;strong&gt;n&lt;/strong&gt; skips it, &lt;strong&gt;e&lt;/strong&gt; lets you reword it first. Nothing is written without a yes. Or pick by number:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx toldya --all --add 1,2,3 --to CLAUDE.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Added 3 rules to CLAUDE.md. Run toldya again in a week to see if they stuck.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;## Things I kept repeating
- Keep it simple.
- Dont complex this.
- Dont assume.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Yes, it keeps my typo. They're your words. Fix them with &lt;strong&gt;e&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did it stick?
&lt;/h2&gt;

&lt;p&gt;Adding a rule doesn't mean the AI follows it. The file can be read perfectly and the rule can still not land.&lt;/p&gt;

&lt;p&gt;So toldya remembers which rules it added, and when. Run it again a week later and each rule shows how often you said it &lt;strong&gt;before&lt;/strong&gt; it went in, and how often &lt;strong&gt;since&lt;/strong&gt;. If "keep it simple" drops from 42 to near zero, the rule works. If it keeps climbing, the wording isn't landing.&lt;/p&gt;

&lt;p&gt;I can't show you a "since" number yet. The tool is a few days old, so there is no "since" to count. That's the honest state of it.&lt;/p&gt;

&lt;p&gt;And because it made me laugh: &lt;code&gt;npx toldya --all --card&lt;/code&gt; draws your top repeats as a picture, on your machine, with a "Save as PNG" button. Nothing is uploaded.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj49m14lpsjsc9dy9q1rp.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj49m14lpsjsc9dy9q1rp.webp" alt="toldya card titled Things I keep telling my AI, listing keep it simple 41 times, dont complex this 29 times, dont assume 15 times, next to the green character at a tally-mark wall" width="800" height="409"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Mine. Drawn earlier the same day, one "keep it simple" ago.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  The part I got wrong
&lt;/h2&gt;

&lt;p&gt;I had a bigger idea first: a team version that reads a project's pull-request review comments, finds what reviewers keep writing, and turns that into rules.&lt;/p&gt;

&lt;p&gt;Before building it, I pulled the last 1,000 review comments from each of Next.js, React, Astro, Supabase and Cal.com, and ran the same counting on them. Almost nothing repeated as a rule. The top "repeats" were "This is changed now" (13 times), "Fixed in the latest commit" (9) and "Thanks for the test" (7). That's conversation. Reviewers write about the code in front of them.&lt;/p&gt;

&lt;p&gt;People correcting their own AI are the opposite: the same handful of sentences, in the same words, for months. So the team version got dropped.&lt;/p&gt;
&lt;h2&gt;
  
  
  What it doesn't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Claude Code only, for now.&lt;/strong&gt; Codex stores history elsewhere and I haven't tested it on real files, so I'm not claiming it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It matches words, not meaning.&lt;/strong&gt; It won't connect "use the logger" with "why is there a console.log here?". That would need a model, and I wanted it small, free and local.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It sees only the history you have.&lt;/strong&gt; Claude Code keeps about 30 days by default.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It needs a pattern.&lt;/strong&gt; A phrase must appear three times, in two different sessions. New to Claude Code? You may get "nothing repeated yet." Good result.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It sends nothing anywhere.&lt;/strong&gt; Only your own messages are read. No account, no telemetry, no network calls.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Try it in a minute
&lt;/h2&gt;

&lt;p&gt;The safe version only shows the report and changes nothing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx toldya --all --dry
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The code is at &lt;a href="https://github.com/singhlabsdev/toldya" rel="noopener noreferrer"&gt;github.com/singhlabsdev/toldya&lt;/a&gt;, MIT, zero dependencies. If it finds something useful, a star helps other people find it. And I'd genuinely like to know your top line. Mine is "keep it simple." 42 times.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;This is how we build.&lt;/strong&gt; Small tools that count what really happened instead of guessing, run on your machine, and ask before they change anything. The same rule goes into the agents we build for businesses.&lt;/p&gt;

&lt;p&gt;Sources: every terminal block is a real toldya run on my own Claude Code history on 25 Sep 2026. The 747 is a plain grep of the same files that day. The review-comment figures come from the last 1,000 comments per repo, fetched from the GitHub API the same day.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/loop-engineering-map/" rel="noopener noreferrer"&gt;I researched loop engineering to build a product. I built nothing.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/a-green-tick-over-nothing/" rel="noopener noreferrer"&gt;My mailer printed “campaign sent”. It had sent to nobody.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/claude-md-rules-chat-history/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Audited My AI's Briefing File. It Had Turned Into a Diary.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:13:25 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/i-audited-my-ais-briefing-file-it-had-turned-into-a-diary-1hhc</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/i-audited-my-ais-briefing-file-it-had-turned-into-a-diary-1hhc</guid>
      <description>&lt;p&gt;Coding agents like Claude Code read a plain-text briefing file before they do anything else. Mine is called &lt;code&gt;CLAUDE.md&lt;/code&gt;, and it sits in the home directory of the small server I run my automation on. It tells the agent who I am, what's installed, which services run where, what the cron jobs do, and the handful of rules I don't want broken. It loads into every session, whatever I'm working on.&lt;/p&gt;

&lt;p&gt;I'd been adding to it for six months. After two model releases in ten days, I ran Claude Code's new prompt audit on it (&lt;code&gt;/checkup prompt-audit&lt;/code&gt;). The audit reads your instruction files and flags text written for older models: ALL-CAPS warnings, "think step by step", rigid scripts. It writes a report and a patch and changes nothing until you say so.&lt;/p&gt;

&lt;p&gt;I expected it to find shouting. It found almost none. What it found instead was more useful, and it changed how I think about the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I expected
&lt;/h2&gt;

&lt;p&gt;The audit's main checklist is about prompting habits that used to help and now hurt. Older models needed forceful instructions to follow anything reliably, so people wrote &lt;code&gt;IMPORTANT:&lt;/code&gt; and &lt;code&gt;NEVER&lt;/code&gt; everywhere. Current models follow instructions closely and literally, so the same shouting makes them over-apply a rule and behave rigidly in situations it was never meant for.&lt;/p&gt;

&lt;p&gt;My file had exactly one of those: a rule in capitals about using a site-specific command-line tool only for the site it was built for. The fix was to say it at normal volume and keep the reason: "the other sites have no command center." Two other spots used capitals for emphasis, but they stated facts, not rules, so the audit left them alone.&lt;/p&gt;

&lt;p&gt;That was the whole outdated-prompting section of the report. One finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually found
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A quarter of the file was history.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The cron section was meant to list my scheduled jobs. Over the months its header had grown into an 880-word paragraph about what I'd &lt;em&gt;removed&lt;/em&gt;: which jobs were cut on which date, why, where the backups went, the traffic numbers that made me drop three sites, which repository was left untouched. Every entry had been accurate when I wrote it. None of it was something the agent needed to do anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Several facts had quietly gone wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The file listed the Gemini CLI at a specific path. It wasn't installed there, or anywhere.&lt;/li&gt;
&lt;li&gt;It gave a version for one of my agent gateways that was two months out of date.&lt;/li&gt;
&lt;li&gt;It listed four API keys in my environment file. There were more than a dozen.&lt;/li&gt;
&lt;li&gt;It listed the tools wired into my voice assistant and was missing one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these would cause a crash. They'd cause a confident wrong answer: the agent telling me "use the Gemini CLI at this path" and spending a turn finding out it isn't there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One rule was a security habit dressed as documentation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Under my blog's repository it said: "Git push requires the token in the remote URL", followed by the exact command to paste a GitHub token into the repo's config. That was true once, because nothing else was set up to handle the login. But it's an instruction, so every time the agent pushed, it would write a live credential into a plain-text file. I'd already cleaned that exact pattern out of two other repos a few weeks earlier, and the instruction file was still teaching it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The flip
&lt;/h2&gt;

&lt;p&gt;I'd been treating &lt;code&gt;CLAUDE.md&lt;/code&gt; as documentation, a notebook about my server that the agent happens to read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It isn't documentation. It's a prompt that runs every session, and every sentence in it is a sentence the model acts on.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once you see it that way, the findings stop looking like housekeeping:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;History in a prompt isn't a record, it's noise the model has to work around.&lt;/strong&gt; "Removed on 2026-08-25, backup at this path, 82 days of rank history pruned" reads as context to me. To a model that treats everything in its instructions as relevant, it's a pile of half-relevant facts competing with the ones that matter. The one useful rule hidden in that paragraph ("these jobs were removed on purpose, so check before adding them back") was the part hardest to find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A stale fact in a prompt isn't out-of-date documentation, it's a wrong instruction.&lt;/strong&gt; A wiki page with an old version number just sits there. The same line in a prompt gets acted on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A how-to in a prompt isn't a note, it's a standing order.&lt;/strong&gt; "Put the token in the URL" stopped being a description of how I once got a push to work. It became the default.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The audit's own guide puts it in one line: the job is to find "specific instructions that no longer fit", not to make prompts shorter. My file didn't have an old-model problem. It had the problem every long-lived config file gets: it kept accumulating and nothing ever removed anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;The patch had nine hunks, each tied to one finding, so I could take or skip them one at a time. I took all of them.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Moved the history out.&lt;/strong&gt; The 880-word paragraph went verbatim into a separate &lt;code&gt;cron-history.md&lt;/code&gt; that isn't loaded automatically. The cron header is now one line with the job count and a pointer: "read it before re-adding a removed job." The schedule itself stayed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixed the facts.&lt;/strong&gt; Dropped the Gemini line, corrected the version, listed the real key names, and added the missing voice-agent tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kept rules, dropped stories.&lt;/strong&gt; Several sections had a current rule with an incident wrapped around it, like "Seen on 2026-09-22: a session started at 14:48 picked up the new package but still ran the old code…" The rule (restart the process after a config change, because clearing the session isn't enough) stayed, with its reason. The timestamped story went.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixed the token rule properly, not just the wording.&lt;/strong&gt; I set git to use the GitHub CLI's existing login for pushes (&lt;code&gt;gh auth setup-git&lt;/code&gt;), confirmed a dry-run push worked, and only then replaced the line with "plain &lt;code&gt;git push&lt;/code&gt;, never put a token in the remote URL." Then I checked every local repo for a token in its remote URL. There were none.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two smaller things went too. I removed account-balance figures from a list of fallback models, since balances drift and nothing on the machine could confirm them. I also moved four skills for a framework I'd stopped using out of the global skills folder: their descriptions were loading into every project.&lt;/p&gt;

&lt;p&gt;The file went from about 3,700 words to 2,760. That wasn't the goal, just a side effect of removing what didn't belong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the audit left alone, and why that matters
&lt;/h2&gt;

&lt;p&gt;The audit has an explicit list of things it must not delete, and it followed it. It kept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the rules that carry a reason, like "don't re-optimize this page's title, the low click rate comes from the search results layout, not the title";&lt;/li&gt;
&lt;li&gt;two identical copies of a scaffold file in two project folders, because they agree;&lt;/li&gt;
&lt;li&gt;my one-line note that I prefer short, direct answers, because that's context about me, not an outdated trick.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That restraint is the part I'd have got wrong doing this by hand. My instinct with a bloated file is to cut it hard, and a hard cut removes exactly the lines that only I could have written: the reasons. The model can work out how to be thorough by itself. It can't work out why I don't trust a particular page's metrics.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you keep one of these files
&lt;/h2&gt;

&lt;p&gt;Some things I'm doing differently now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Write the rule, not the story.&lt;/strong&gt; If a line describes something that happened, ask what the agent should &lt;em&gt;do&lt;/em&gt; differently because of it, and write that sentence instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give history its own file.&lt;/strong&gt; A changelog is valuable. It just shouldn't load into every session. Point to it from the briefing file, and the agent can read it when it's relevant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check facts against the machine, not your memory.&lt;/strong&gt; Paths, versions and key names drift silently. The audit caught four because it checked them, not because they looked wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read every how-to as a standing order.&lt;/strong&gt; If you wouldn't want the agent to do it every time without asking, it doesn't belong in the file as an instruction.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The file is shorter now, but that isn't the improvement. The improvement is that everything left in it is something I actually want the agent to act on.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
    <item>
      <title>How To Write Playwright tests in minutes with Playwright MCP and Claude Code</title>
      <dc:creator>Jakob Norlin</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:44:22 +0000</pubDate>
      <link>https://dev.to/jakobnorlin/how-to-write-playwright-tests-in-minutes-with-playwright-mcp-and-claude-code-1o0d</link>
      <guid>https://dev.to/jakobnorlin/how-to-write-playwright-tests-in-minutes-with-playwright-mcp-and-claude-code-1o0d</guid>
      <description>&lt;p&gt;&lt;em&gt;This article was originally published on the&lt;/em&gt; &lt;a href="https://endform.dev/blog/playwright-mcp-claude-code" rel="noopener noreferrer"&gt;&lt;em&gt;Endform blog&lt;/em&gt;&lt;/a&gt;&lt;em&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3egdamuwg279l903vgqt.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3egdamuwg279l903vgqt.webp" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;By the end of this walkthrough you'll have a single Playwright test that Claude Code wrote against your live app, that you've reviewed in two passes, run repeatedly to check for flakiness, and committed next to the plain-English scenario it came from. The trick that makes it trustworthy: the agent reads locators off the running page through the Playwright MCP server instead of guessing them from your source code.&lt;/p&gt;

&lt;p&gt;You need four things installed before step 1, so start there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you start
&lt;/h2&gt;

&lt;p&gt;Confirm all four of these first. Skipping one is the most common reason step 1 fails.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Node.js 20 or newer.&lt;/strong&gt; Check with &lt;code&gt;node --version&lt;/code&gt;. The MCP server launches through &lt;code&gt;npx&lt;/code&gt;, so Node is required no matter how you installed Claude Code.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Claude Code, installed and signed in.&lt;/strong&gt; Check with &lt;code&gt;claude --version&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A Playwright project.&lt;/strong&gt; You need &lt;code&gt;playwright.config.ts&lt;/code&gt; at the repo root. No project yet? Run &lt;code&gt;npm init playwright@latest&lt;/code&gt; to scaffold one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;An app to test.&lt;/strong&gt; Local dev server or staging URL, either works. Playwright can start it for you when the config says so.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Got all four? Connect Claude Code to a browser.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem this setup solves
&lt;/h2&gt;

&lt;p&gt;Here's a failure pattern you've probably lived through: Claude Code writes an end-to-end test, it passes locally, you merge, and CI goes red a day later on what looks like flakiness. The test was wrong from the start, it just had no way to show it.&lt;/p&gt;

&lt;p&gt;The reason is where the locators came from. Reading your components without ever opening the app, an agent spots a button with a class like &lt;code&gt;.btn-primary&lt;/code&gt; and writes a selector against it. Then a component library or some runtime logic rewrites that class on its way to the DOM, and by the time the browser renders, the selector points at a hashed string, a restructured node, or nothing. Nobody caught it because nobody ran the test against a real page. CI is the first thing that does.&lt;/p&gt;

&lt;p&gt;Compare the two ways the same button gets targeted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Hallucinated: guessed from training data, does not exist on this page&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;locator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;#submit-btn&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// Grounded: read from the accessibility tree Claude Code can see&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Place order&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Anthropic's Model Context Protocol is what closes the gap. The Playwright MCP server hands Claude Code a live browser: it can open the app, walk the accessibility tree, and build locators out of the roles and names the page actually exposes. It's working from a structured snapshot rather than a screenshot, which keeps the locators stable and the token cost low, because the model reasons over text it can already read. Same model, better inputs. Nothing here is a capability upgrade, only an access one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Register the Playwright MCP server
&lt;/h2&gt;

&lt;p&gt;A single command wires the server into Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add playwright &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @playwright/mcp@latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That entry gets saved to your local config, and the server process spins up the next time you open a session. &lt;code&gt;-y&lt;/code&gt; skips the install confirmation &lt;code&gt;npx&lt;/code&gt; would otherwise wait on, and &lt;code&gt;@latest&lt;/code&gt; grabs the newest published build so you're not stuck on a stale cache.&lt;/p&gt;

&lt;p&gt;By default this is a personal, single-project entry. Working on a team? Add &lt;code&gt;--scope project&lt;/code&gt;, which writes the same config to a &lt;code&gt;.mcp.json&lt;/code&gt; at the repo root so everyone shares one server without redoing setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"playwright"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stdio"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@playwright/mcp@latest"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't trust it until you've checked it. Run &lt;code&gt;claude mcp list&lt;/code&gt; and look for &lt;code&gt;✔ Connected&lt;/code&gt;. Immediately after adding, you may catch a &lt;code&gt;✘ Failed to connect&lt;/code&gt; while &lt;code&gt;npx&lt;/code&gt; is still pulling the package in the background; run the command again a few seconds later and it usually clears.&lt;/p&gt;

&lt;p&gt;If retrying doesn't clear it, the problem is elsewhere. Claude Code runs the server as a subprocess in its own environment, and that subprocess can't always resolve &lt;code&gt;npx&lt;/code&gt; the way your interactive shell does. Homebrew-installed Node on macOS is the usual culprit. Give it the absolute binary path instead of trusting PATH:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp remove playwright
claude mcp add playwright &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;which npx&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; @playwright/mcp@latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run &lt;code&gt;claude mcp list&lt;/code&gt; once more and wait for &lt;code&gt;✔ Connected&lt;/code&gt; before continuing.&lt;/p&gt;

&lt;p&gt;Connected only means the process is alive. It doesn't prove Claude Code can actually drive a browser yet. Open a session and call the tool by name, otherwise Claude Code might reach for a Bash command instead of the MCP server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using Playwright MCP, open [your app's local URL] and tell me the page title and the first heading you see.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it answers with something concrete off the page, not a hedge and not a generic description, the tool works. That answer is coming from a live snapshot rather than memory, which is the entire reason you connected it.&lt;/p&gt;

&lt;p&gt;These commands are current as of August 2026. MCP tooling changes quickly, so if a command here stops matching what you see, check the official Playwright MCP repo. Server verified and driving a browser, the next hurdle is login.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Get past the login wall with storage state
&lt;/h2&gt;

&lt;p&gt;Nearly every test worth writing sits behind authentication. A checkout, a settings screen, an admin panel, all of them dead ends if the &lt;a href="https://endform.dev/blog/agentic-ai-testing" rel="noopener noreferrer"&gt;agent&lt;/a&gt; can't clear the sign-in page.&lt;/p&gt;

&lt;p&gt;The MCP server has no idea about your suite's existing auth. It either carries its own browser profile across sessions or, in isolated mode, boots logged out every time. Either way it's disconnected from however your tests currently authenticate.&lt;/p&gt;

&lt;p&gt;Fix it by launching the server with a session already loaded:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @playwright/mcp@latest &lt;span class="nt"&gt;--isolated&lt;/span&gt; &lt;span class="nt"&gt;--storage-state&lt;/span&gt; .auth/user.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--isolated&lt;/code&gt; holds the profile in memory rather than writing it to disk, so every session starts fresh. &lt;code&gt;--storage-state&lt;/code&gt; reads a saved authenticated session from a file, and anything that changes mid-session gets thrown away at the end. But that file has to exist first.&lt;/p&gt;

&lt;p&gt;Playwright's auth docs suggest a dedicated setup test that signs in and saves the session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;setup&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;authFile&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;.auth/user.json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;password&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;TEST_USER_PASSWORD&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;password&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;TEST_USER_PASSWORD is not set&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;setup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;authenticate&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://your-app.example.com/login&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Email&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;test-user@example.com&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Password&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;password&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Log in&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;waitForURL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;**/dashboard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;context&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;storageState&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;authFile&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wire it to run automatically by adding a setup project in &lt;code&gt;playwright.config.ts&lt;/code&gt; that the browser projects depend on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;projects&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;testMatch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sr"&gt;/auth&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="sr"&gt;setup&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="sr"&gt;ts/&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;chromium&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;storageState&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;.auth/user.json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;],&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx playwright &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;setup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When it finishes, &lt;code&gt;.auth/user.json&lt;/code&gt; holds a live authenticated session, and &lt;code&gt;auth.setup.ts&lt;/code&gt; is now part of the suite, so login logic lives in exactly one place.&lt;/p&gt;

&lt;p&gt;Worth doing if your app allows it: create a throwaway test user through an API or seed script before login and tear it down after. Claude Code pokes at the app while it explores, and a disposable account keeps it from mutating a shared login someone else is on.&lt;/p&gt;

&lt;p&gt;With the file in place, re-register the server so it loads that session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp remove playwright
claude mcp add playwright &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @playwright/mcp@latest &lt;span class="nt"&gt;--isolated&lt;/span&gt; &lt;span class="nt"&gt;--storage-state&lt;/span&gt; .auth/user.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For teams, mirror it in &lt;code&gt;.mcp.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"playwright"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stdio"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"@playwright/mcp@latest"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--isolated"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--storage-state"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;".auth/user.json"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For messier auth flows, &lt;a href="https://endform.dev/blog/playwright-mcp" rel="noopener noreferrer"&gt;Endform's Playwright MCP guide&lt;/a&gt; goes deeper on this.&lt;/p&gt;

&lt;p&gt;Now that Claude Code browses as a logged-in user, it's time to write the prompt that turns a described flow into a committable test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Build the prompt in three parts
&lt;/h2&gt;

&lt;p&gt;This step is where a shippable test and a throwaway one diverge. Rather than one rambling paragraph that asks for a test and hopes, split the prompt into three parts that each do one job.&lt;/p&gt;

&lt;p&gt;Part one is environment setup: the app-specific facts the agent can't deduce on its own.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The website under test is at https://staging.yourapp.com.
You're already logged in as a test user through the storage state configured earlier.
The test user's cart currently holds one item, a placeholder t-shirt priced at $24.99.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Part two is the scenario, phrased the way you'd brief a teammate, not as pseudocode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Guest checkout with a saved card
1. Open the cart page.
2. Proceed to checkout.
3. Confirm the shipping address shown is the default one.
4. Select the saved Visa card ending in 4242.
5. Place the order.
6. Confirm the order confirmation page shows an order number and the correct total.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep this as a markdown file in the repo next to the tests it drives, not buried in a chat log. Playwright Test Agents work the same way: a planner writes the scenario as markdown, and a later step compiles it into code. Splitting the two into separate, reviewable artifacts is worth carrying over here.&lt;/p&gt;

&lt;p&gt;This part is also the one nobody on the team can write better than you. The environment facts are just facts, and the system prompt below is boilerplate you reuse everywhere. The scenario is the only part carrying judgment: does the address get confirmed before payment, does the saved card matter more than a fresh one, is the confirmation total worth asserting or is a loaded page enough? An agent can't rank those. Someone who knows the product has to.&lt;/p&gt;

&lt;p&gt;Part three is the system prompt, the standing rules that ride along on every test and describe what "good" looks like, not just what to do:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a Playwright test generator.
Explore the app using the Playwright MCP tools before writing any code, don't generate steps from assumption alone.
Prefer getByRole, getByLabel, and getByTestId locators over CSS selectors or XPath.
Don't add manual waitForTimeout calls, rely on Playwright's built-in auto-waiting and retrying assertions instead.
Group related steps with test.step for readability in the trace viewer.
Assert on outcomes a user would actually see, not on incidental implementation details.
Save the finished test to the tests directory, run it, and keep iterating until it passes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You don't have to write this cold. Debbie O'Brien, a longtime Playwright advocate and ex-member of Microsoft's Playwright team, keeps a public set of prompt files for exactly this, a better base than reinventing it.&lt;/p&gt;

&lt;p&gt;Stacked together, this is the full message to Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a Playwright test generator.
Explore the app using the Playwright MCP tools before writing any code, don't generate steps from assumption alone.
Prefer getByRole, getByLabel, and getByTestId locators over CSS selectors or XPath.
Don't add manual waitForTimeout calls, rely on Playwright's built-in auto-waiting and retrying assertions instead.
Group related steps with test.step for readability in the trace viewer.
Assert on outcomes a user would actually see, not on incidental implementation details.
Save the finished test to the tests directory, run it, and keep iterating until it passes.

The website under test is https://staging.yourapp.com.
You're already logged in as a test user through the storage state configured earlier.
The test user's cart currently holds one item, a placeholder t-shirt priced at $24.99.

# Guest checkout with a saved card
1. Open the cart page.
2. Proceed to checkout.
3. Confirm the shipping address shown is the default one.
4. Select the saved Visa card ending in 4242.
5. Place the order.
6. Confirm the order confirmation page shows an order number and the correct total.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The proof this prompt works is watching the agent call the MCP tools to explore before it writes any test code. A "explore first" instruction only counts if the agent obeys it. Three parts, three jobs, one message. Next, the generation itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Let it generate the test
&lt;/h2&gt;

&lt;p&gt;One look at the page isn't where Claude Code stops. It keeps navigating and reading through the MCP server until it has actually seen every element the scenario names. That's what separates this from feeding a model a plain-English description and hoping.&lt;/p&gt;

&lt;p&gt;Watch a run and the sequence is plain: open the page, click through the scenario's steps, read back what each returned, and only then start writing assertions. It writes the spec last, runs it, and keeps tweaking until it passes instead of handing you untested code.&lt;/p&gt;

&lt;p&gt;A real example: on a logout scenario, two headings both matched "Secure Area," so Playwright threw a strict mode violation, because a locator meant to act on one element matched several. Claude Code read the error, diagnosed it, and appended &lt;code&gt;exact: true&lt;/code&gt; so the locator demanded the full heading text rather than a substring. That correction landed in the same pass, no nudge from the developer.&lt;/p&gt;

&lt;p&gt;A checkout with a cart, saved card, and confirmation page runs the same loop, step by step, until green. Here's a second flow to show it's not a one-off.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://endform.dev/blog/playwright-mcp" rel="noopener noreferrer"&gt;Endform's guide to shipping quality end-to-end tests with Playwright MCP&lt;/a&gt; has a scenario worth reusing: check that a newly created team shows up in an activity log. Sign in, confirm a signup event is already logged, create a team, confirm the new event lands. Endform walks through the scenario and the review, not the code, so here's a plausible test around that flow, in the shape Claude Code would generate it against a live app:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;new team activity appears in the activity log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;open the activity log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://staging.yourapp.com/dashboard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;link&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Activity&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;confirm the signup event is already logged&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;you signed up&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;create a new team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Create a new team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Team name&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;QA Playground&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Create team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;confirm the new team event appears in the log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;you created a new team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice how deliberate the file is. Every link and button name is text Claude Code read off the page. The &lt;code&gt;test.step&lt;/code&gt; blocks track the scenario one to one, so a failure drops you straight on the broken step. No manual waits anywhere, because the retrying assertions cover timing.&lt;/p&gt;

&lt;p&gt;You end up with a runnable file that clears its first run more often than not, precisely because each locator was verified against the page before a line got written. That still isn't the same as a good test, which is what step 5 is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Review it in two passes
&lt;/h2&gt;

&lt;p&gt;Passing once doesn't earn a test a place in the suite. Run two review passes, because they catch different things: one on the code, one on the meaning.&lt;/p&gt;

&lt;p&gt;Pass one is the code. Read it and ask if you'd have written it this way. Generated tests tend toward the baroque, redundant checks, extra steps, logic that folds down to less, and trimming that is usually the first win. Hold the locators to the priority you set (&lt;code&gt;getByRole&lt;/code&gt; and &lt;code&gt;getByTestId&lt;/code&gt; over CSS or XPath), swap out anything brittle, and make sure each assertion proves something a user would notice rather than just confirming an action didn't throw. The &lt;code&gt;test.step&lt;/code&gt; groupings should read the way a person would narrate the flow.&lt;/p&gt;

&lt;p&gt;Pass two is the one no linter will ever do for you, and it's the one that matters more. Put the scenario next to the finished test and ask whether the thing being verified still matches what the scenario meant, not just whether it's green. Tests drift here quietly. On that logout test, Claude Code noticed the login banner showed once and got eaten by the stored session, so it retargeted the assertion to the permanent heading before wrapping up. Useful that it caught it, but it happened during generation, not review, and that's the point: next time it might not, and this pass is your backstop.&lt;/p&gt;

&lt;p&gt;The first draft gets much faster; the judgment moves into refinement. And that judgment, what the product does, what risk you're covering, what the scenario is really asserting, is yours to bring. Both passes done, one last check remains.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Verify, then commit
&lt;/h2&gt;

&lt;p&gt;The commands below use this tutorial's logout test as the example. A test that passed while being generated still hasn't earned your trust. Run it again, clean, outside the generation loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx playwright &lt;span class="nb"&gt;test &lt;/span&gt;tests/logout.spec.ts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A single green run tells you almost nothing about reliability. Fire it several times in a row:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx playwright &lt;span class="nb"&gt;test &lt;/span&gt;tests/logout.spec.ts &lt;span class="nt"&gt;--repeat-each&lt;/span&gt; 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--repeat-each&lt;/code&gt; reruns the same test N times in one invocation. It's among the fastest ways to smoke out a test that's green most runs and red occasionally, the exact flakiness that tends to surface only once it hits CI.&lt;/p&gt;

&lt;p&gt;Repeats clean? Commit the test together with the markdown scenario behind it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git add tests/logout.spec.ts tests/scenarios/logout.md
git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"Add logout test with scenario spec"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Committing them together keeps intent attached to implementation. Whoever reads the diff later sees what the test was meant to prove, not just the assertions.&lt;/p&gt;

&lt;p&gt;One thing to check before that first commit: your &lt;code&gt;.gitignore&lt;/code&gt;. A freshly scaffolded Playwright project ignores &lt;code&gt;playwright/.auth/&lt;/code&gt;, its default spot for storage state. But this setup writes to &lt;code&gt;.auth/&lt;/code&gt; at the repo root, a different path the default rule doesn't cover, which means a live authenticated session can slip into history unnoticed.&lt;/p&gt;

&lt;p&gt;Add the right path first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;".auth/"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; .gitignore
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one line keeps the session out of your repo history by intent rather than luck.&lt;/p&gt;

&lt;p&gt;Verified, stable over repeats, and committed alongside its scenario, the test is a dependable addition. What gets harder from here: longer flows, data that shifts between runs, and auth beyond a plain login form.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this workflow falls down
&lt;/h2&gt;

&lt;p&gt;Trust comes from being straight about where Claude Code struggles, not just where it shines. Four things break it fairly predictably.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Long multi-step flows.&lt;/strong&gt; The longer the flow, the more the agent loses the thread, since each step starts from a fresh snapshot rather than the whole sequence before it, and long MCP sessions dropping browser context is a known issue. Splitting the flow into smaller scenarios and stitching them later holds up better.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data that changes between runs.&lt;/strong&gt; The agent tends to assert against whatever it saw at generation time, a timestamp, an order ID, a count, and those move on the next run. Assertions last longer when they target something stable: a confirmation state, a completed action, the presence of a result, not the exact value from one run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deep conditionals.&lt;/strong&gt; While generating, the agent only travels one branch, one role, one account state, one flag. Branches it never walks stay invisible to it. Prompt each branch on its own, or write the conditional logic yourself, rather than expecting one pass to find every path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OAuth and third-party auth.&lt;/strong&gt; Usually the first wall you hit. Redirects out to Google, Microsoft, or any external identity provider leave your app, and the agent's control loop isn't built to follow through the handoff. One developer automating &lt;a href="https://endform.dev/blog/playwright-github-actions" rel="noopener noreferrer"&gt;GitHub's&lt;/a&gt; OAuth earned a temporary IP ban for it, and other providers can react to automated logins the same way, which makes this genuinely fragile. The move is to authenticate in a separate setup step, save the session with &lt;code&gt;storageState&lt;/code&gt;, and generate against an app that's already logged in.&lt;/p&gt;

&lt;p&gt;None of these shrink the workflow's value. They just mark where a scenario needs prep before you hand it over.&lt;/p&gt;

&lt;h2&gt;
  
  
  What comes next: keeping the suite fast
&lt;/h2&gt;

&lt;p&gt;We opened with a test that looked done and broke in CI regardless. Claude Code explored the real app through Playwright MCP, checked what was on the page, and produced a test you reviewed, verified, and committed.&lt;/p&gt;

&lt;p&gt;That changes the economics of test creation. Once a trustworthy test costs minutes instead of hours, teams write more of them, and a bigger suite brings its own problem. Those tests still run in CI, and as the count climbs, the bottleneck moves from writing tests to running them fast without flakiness.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://endform.dev" rel="noopener noreferrer"&gt;Endform&lt;/a&gt; runs every Playwright test on its own isolated machine in parallel, which keeps suite duration predictable no matter how many you add.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>mcp</category>
      <category>testing</category>
    </item>
    <item>
      <title>I Gave My AI Coding Agent 5 Skills and It Stopped Rewriting the Same Files</title>
      <dc:creator>The Android AI Architect</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:22:38 +0000</pubDate>
      <link>https://dev.to/theandroidaiarchitectdev/i-gave-my-ai-coding-agent-5-skills-and-it-stopped-rewriting-the-same-files-56mc</link>
      <guid>https://dev.to/theandroidaiarchitectdev/i-gave-my-ai-coding-agent-5-skills-and-it-stopped-rewriting-the-same-files-56mc</guid>
      <description>&lt;p&gt;Most AI coding setup advice stops at "install Claude Code." That's the easy part.&lt;/p&gt;

&lt;p&gt;The hard part is what happens three weeks in: the agent keeps rewriting the same helper file, forgets the project conventions, writes code without tests, and you spend more time reverting than shipping.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills are just instructions your agent reads
&lt;/h2&gt;

&lt;p&gt;An agent skill is a markdown file with a name, a description, and a procedure. No plugin system, no framework, no build step. Drop it in a directory the agent watches and it applies that procedure when the task matches.&lt;/p&gt;

&lt;p&gt;The difference is between hoping the agent does the right thing and telling it exactly what right means for your codebase.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five that changed my workflow
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;feature-plan&lt;/code&gt;&lt;/strong&gt; - before writing any code, produce a plan: what changes, which files, what could break. Skipping this is how you get 400 lines of code nobody wanted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;test-write&lt;/code&gt;&lt;/strong&gt; - write the test first, then make it pass. The skill encodes the pattern and makes the agent follow it instead of retrofitting tests onto code it already wrote.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;code-review&lt;/code&gt;&lt;/strong&gt; - a review pass with a specific checklist. Catches the things you skim past at 1am: unhandled null, a missing await, an N+1 query that won't matter until you have 10,000 users.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;bug-fix&lt;/code&gt;&lt;/strong&gt; - reproduce, isolate, fix, verify. In that order. The most common waste in agent-assisted debugging is changing code before you can reproduce the bug, then changing it again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;refactor&lt;/code&gt;&lt;/strong&gt; - safe refactoring with verification at each step, so cleanup doesn't become an outage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wiring it up
&lt;/h2&gt;

&lt;p&gt;Drop the folder in. Point your agent at the skills directory. For Claude Code the config is a JSON file; for Cursor there's a &lt;code&gt;.cursorrules&lt;/code&gt; file. Both are in the pack.&lt;/p&gt;

&lt;p&gt;The real payoff isn't the individual procedures. It's that they're consistent, so your agent stops improvising a different approach every session.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't replace
&lt;/h2&gt;

&lt;p&gt;Type hints, tests, and code review. Skills make your agent more consistent, not more competent. A skill that says "write good code" is useless; a skill that says "before editing any file, read CLAUDE.md and follow the existing naming convention" changes behavior.&lt;/p&gt;

&lt;p&gt;Be specific or be ignored.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one that gets skipped
&lt;/h2&gt;

&lt;p&gt;Nobody wires up their MCP servers and skills on day one. They install the editor, write a prompt, feel productive, and never revisit it.&lt;/p&gt;

&lt;p&gt;The setup takes twenty minutes. Not doing it costs you the same three weeks of drift every three weeks, forever.&lt;/p&gt;




&lt;p&gt;I packaged these five skills plus a Cursor config, project rules, and starter templates into the MCP Vibe Coder Pack:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MCP Vibe Coder Pack&lt;/strong&gt; - &lt;a href="https://baradigitaltools.gumroad.com/l/mcp-vibe-coder-pack" rel="noopener noreferrer"&gt;https://baradigitaltools.gumroad.com/l/mcp-vibe-coder-pack&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There's also a bundle that adds 20 freelance prompts and a remote jobs dataset, if you're trying to turn the shipping into income:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vibe Coder Power Pack&lt;/strong&gt; - &lt;a href="https://baradigitaltools.gumroad.com/l/vibe-coder-power-pack" rel="noopener noreferrer"&gt;https://baradigitaltools.gumroad.com/l/vibe-coder-power-pack&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>mcp</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
