DEV Community

jamilxt
jamilxt

Posted on

Stop Saying 'Think Step by Step': What Addy Osmani's Opus 5.5 Prompting Guide Actually Says

Open your saved prompts right now and search for the words "think carefully" or "step by step." If you find them, you are carrying habits from an older generation of models, and according to Anthropic's newest official guide, those habits are now making your replies slower for no gain.

Last week Addy Osmani published "Getting the most out of Opus 5.5 in Claude and Claude Code" on the Anthropic engineering blog, and it made the Hacker News front page with real discussion. The theme of the whole guide is uncomfortable for anyone who spent 2025 collecting prompt tricks: the model now decides how much to think on its own, so your job is no longer to push it toward reasoning. Your job is to define "done," tell it when to stop and ask, and then get out of its way.

One disclosure before the breakdown: everything below is built from Osmani's guide and Anthropic's own documentation, all linked inline. I have been running my own multi-week experiments in live repos, and those confirmed the guide's core claim before I read it, but the numbers and specific recommendations here are the guide's, not mine.

The core shift: the model thinks, so stop telling it to

Delete "think carefully" lines. In Anthropic's own testing in a chat product, removing a "think carefully" line made replies start sooner with no clear drop in quality. Opus 5.5 thinks before every reply and decides how much thinking the task deserves. The instruction is now dead weight, and in saved instructions it applies to every single message you send.

Give the whole task in one message. This is the biggest change from how most of us work. Osmani's example prompt has three parts, and every one of them matters:

  • The whole task: "Migrate the payment endpoints from the old client to the new one."
  • The finish line: "Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes."
  • The stop condition: "Stop and ask me only if a test fails for a reason you can't explain."

Why the change? Early testers ran Opus 5.5 on long coding tasks for hours with little oversight, and its biggest gains over Opus 5 are exactly there, carrying a change through a large repository until the tests pass. A finish line tells it when it is done. A stop condition tells it the one case where you want to be woken up. Everything else should be a status note, not a question.

Name the styles you do not want. For design work, "avoid a generic look" mostly swaps one generic look for another. Osmani's example is a list of specific habits to ban: no cream or off-white background, no italic accent words in headings, no numbered "01 / 02 / 03" section labels, no monospace labels, no pill-shaped buttons. Specific negatives beat vague positives, and if you dislike what it picks next, add that to the ban list and rerun.

Steering long runs: CLAUDE.md is now your steering wheel

The second section of the guide is the most valuable part for anyone running agent sessions, because it turns CLAUDE.md from a style guide into run control. Osmani's suggested rule is short enough to paste:

  • "When a step doesn't need my input, keep going. Put status notes in the same message as your next action."
  • "Stop and ask only when you can't continue without me, or before anything destructive: deleting data, force-pushing, or changing anything outside this repository."

Two details worth pausing on. First, the guide warns that Opus 5.5 sometimes stops mid-task to report instead of continuing, giving you a summary that names the next step without taking it, or an offer to continue. If your runs keep ending with "Want me to continue?", the rule above is the fix, and the fix for a single occurrence is to just reply "continue." Second, the destructive-command carve-out matters: a keep-going rule means fewer stops, so you keep your own checkpoint before anything risky or hard to undo, and you leave permission prompts on for destructive commands.

Subagents for audits and migrations. For work spread across a large codebase, Osmani suggests asking the model to fan out: "Give each service to its own subagent. When a subagent reports back, check its evidence before you accept it." Note that second half. Anthropic is not telling you to trust the parallel reports; they are telling you to have the model itself verify each subagent's evidence before it accepts it, then finish with a single table: service, affected yes or no, and the evidence.

Keep the task list in a file. Long runs fill the context window, and Claude Code then summarizes older turns. A checklist in TASKS.md survives that summarization and shows you at a glance what is done and what is left. Read the file, not the scrollback.

Checking the result: read "blocked on me" first

When a long run ends, Osmani says to look first for anything Claude is waiting on you for: a decision it left open or a change it wants you to approve. Only then read the rest of the summary. His trick for making this repeatable is to standardize the ending in CLAUDE.md: "End every run with three headings: Blocked on me, Changed, Found."

Three more checking habits from the guide:

  • Run a model review before the human review. One early tester reported Opus 5.5 at its lowest effort setting caught more bugs than Opus 5 at high effort, with fewer false alarms.
  • For research, add "Mark anything you couldn't confirm, and say where you looked." A model that tells you what it could not find is a model you can audit.
  • In the apps, attach the chart or screenshot instead of retyping the numbers. Osmani says Opus 5.5 reads charts and screenshots more accurately than Opus 5, including spatial meaning like which boxes an arrow connects.

The part nobody is talking about: silent model switching

Here is the section of the guide that deserves its own post. Opus 5.5 is the first Opus model to launch with Fable-level bio and cyber safeguards, and in Claude apps and Claude Code, most flagged messages do not get refused. They get silently routed to an older model, and your work just continues there, usually without you noticing.

Finding security vulnerabilities in source code is still allowed, and everyday questions still work. But Anthropic admits these safeguards can flag legitimate work, and the check covers everything in the conversation, including files and search results. That means a flag can come from earlier content in the session, not just your last message, and if you do not know the recovery steps you may be getting Opus 5-quality answers while believing you are on Opus 5.5.

The recovery steps from the guide:

  • In Claude Code: run /model to switch back, or press Esc twice to edit your last message and retry. To be asked first instead of switched automatically, run /config and change "Switch models when a message is flagged."
  • In the apps: pick Opus 5.5 in the model picker, and if the flagged message is still in the chat it may trigger again, so a new chat is safer. Settings, then Capabilities, has the same "Switch models when a message is flagged" toggle.

If you do security review or pen-test adjacent work, set that toggle to ask-first today. Getting silently downgraded mid-session is the kind of failure you can spend hours not noticing.

One related prompting note: do not ask the model to reproduce its internal reasoning in the reply. That request is one of the flag categories and can be declined outright. Ask for the useful output instead: "Explain why you chose this approach in three sentences."

Fast mode, and when it is worth it

Back-and-forth work, where you read each reply before sending the next message, is where speed actually matters, and that is exactly what fast mode targets: type /fast in Claude Code. You get the same model with text arriving sooner, but it is a research preview, needs extra usage turned on, and costs more per token than standard mode. For long autonomous runs, the opposite logic applies: the run is not interactive, so pay standard rates and let it cook.

The pre-flight checklist

Run through this before your next long task. This is the part worth bookmarking.

Before you hit send:

  • The task says what "done" looks like in one sentence
  • No "think hard" lines anywhere in your prompts or saved instructions
  • Design requests name the specific styles to leave out
  • Charts and screenshots are attached, not retyped

Before a long Claude Code run:

  • CLAUDE.md says when to keep going and when to stop and ask
  • The stop rule carves out destructive commands, and permission prompts are still on for them
  • Big audits and migrations are split across subagents, with evidence checks on each report
  • The task list lives in a file, not in the scrollback

Before you trust the result:

  • The "blocked on me" section is the first thing you read
  • A model review pass runs before the human review
  • Research answers mark what could not be confirmed, and where it looked

Once, today:

  • Check your flag settings: silent switch or ask-first, pick deliberately
  • Grep your saved prompts for "step by step" and delete what you find

That last one is the meta-lesson of the whole guide. Prompting advice has a shelf life, and the tricks that made Opus 5 feel controllable are now friction. The winning move in 2026 is fewer words in the prompt and more structure in the config: a defined finish line, a written stop rule, and settings you chose on purpose.

I write about AI tooling, agentic workflows, and what actually works in practice every week. Subscribe, it is free, and it keeps the deep dives coming.

Have you run Opus 5.5 on a long autonomous task yet? Did it stop too often to ask permission, or did it run clean? I want to hear how the "define done and let go" style is working for other people, because it is a real behavior change from how we all worked last year.

Top comments (0)