Opus 5.5 shipped on September 22, and its replies read differently: answer first, less throat-clearing. Box, testing it before launch, found its an...
For further actions, you may consider blocking this person and/or reporting abuse
Self-declaring the turn kind with tags like [status] or [analysis] fixes the prompt regex problem, but the edge that bites in longer agent loops is state drift mid-turn. An agent starts a task assuming a routine three-line fix under [status], executes a bash tool that fails an unexpected integration test, and suddenly needs an architectural pivot. If the response budget locks on that opening tag, the model truncates the explanation to squeeze under the status check. Allowing the model to emit a revised intent token after tool execution if the turn shifts from execution to debugging saved us a lot of retried turns.
agree, locking the budget on the opening intent would bite exactly there. in our setup the tag goes on the final reply rather than the task, so it's declared after the tool calls ran: the routine fix that hit a failing integration test just ends as [status][analysis], and a few checks (length among them) bind whatever the tag says. do you re-run the checks after each revised token or only on the last one?
Heads-up on
reply-shape.sh: theelifline renders as<<<;"$prompt". The stray;is a bash syntax error, and because bash parses the wholeifcompound before it runs any of it, the design branch never fires either. A syntax error exits with status 2, which UserPromptSubmit treats as "block this prompt", so a verbatim copy rejects every prompt instead of just losing the hint. The Stop hook has the same problem:&&and>&2render as&&and>&2, which looks like the editor escaping them. I ran the rendered script through bash: it stops on the guard line with exit 2 beforestop_hook_activeis ever read, so Claude keeps getting sent back after every reply until Claude Code's continuation cap ends the turn. Both scripts are written to be pasted, so a quick fix would save the next reader a confusing first run.@skillselion oh my, it seems there were some encoding problems during syndication. Thank you for letting me know, much appreciated! Fixed now.
The turn-kind tag idea is neat, but I wonder how it holds up once you have a dozen tags — do you end up with a taxonomy nobody remembers to use? The auto-memory handle seems like the sneakier win here, since it learns the shape from corrections instead of asking you to declare everything upfront.
it really depends on your system. The one that i'm maintaining is encompassing 9 different sub-projects, with some custom sub-sub cli toolsets as well. The tagging stayed thin.
I have 2 domains: reply kind ([status], [answer], [analysis], [narration], [decision], [handover]) and work domain (what the current work is about, that's a bit longer list).
A reply may lead with one to three tags.
The reason why I settled to this configuration because it enables a much higher variability in ruling and catching edge cases, while it's easy to maintain (memory can be simply directed to incorporate the ruling here, so that becomes reinforcement instead of SoT).
The two-lines-for-a-rename, full-walkthrough-for-architecture split is the right instinct. I have been doing something similar with CLAUDE.md rules that force a plan before any multi-file change, and the biggest surprise was how much of the verbosity problem was actually my prompts being vague, not the model being chatty. Have you tried scoring responses against the original ask, or is it still eyeball-and-adjust?
as far as responses go I'd say eyeball-and-adjust is perfectly fine, people tend to have very different tastes and thresholds for reading.
For scoring output against the task/ask at hand ... well ... that's a whole lot different game altogether. Eventually the series will get there too, but in order to have something like that, you'll have to lock down the event bus mechanism (as that is what gives you the management layer of any meaningful scoring) and then some structures/procedures and only after then we can talk about comparing the two.
Progressive disclosure is the right antidote to dumping an entire repository into every prompt. The strongest pattern here is letting file context and task intent decide what becomes visible, which should reduce both noise and accidental instruction conflicts. I would love to see a small trace view that shows which rule shaped each response.
We'll get there too together with the event bus in the next article. The pre-requisite was the self-classification and what it entails.