Anthropic shipped Claude Opus 5.5 on 22 September, and I've spent most of the week since with it open in one terminal or another. It's the first Opus in a while that I've just stopped thinking about. It does the work, it tells me what it did in plain English, and it doesn't burn through my allowance doing it.
I'm clearly not the only one. But alongside the praise there's a quieter thread of people saying the same thing I've been saying about every model for two years: it's brilliant, and it still doesn't remember anything. So let's look at both, and then at what I've been doing about the second one.
What people are saying
The headline numbers are good. Opus 5.5 is priced at $4 in and $20 out per million tokens, 20% cheaper than Opus 5, and it sits at the top of the Artificial Analysis Intelligence Index with a score of 58. eesel's review (https://www.eesel.ai/blog/claude-opus-5-5-review) called it, at medium effort, "the best tool I've used" for genuinely hard, long, agentic work. Subscribers get a slice of the saving too: Memeburn worked out (https://memeburn.com/anthropic-says-opus-5-5-costs-40-less-to-run-claude-subscribers-get-about-25-more-room) that Pro, Max and Team allowances stretch roughly 25% further than they did on Opus 5.
But the thing people keep coming back to isn't a benchmark. It's how it feels to work with. Opus 5 was verbose to the point of being tiring, and 5.5 has fixed most of that. The team at Every ran a week-long vibe check (https://every.to/vibe-check/vibe-check-opus-5-5-is-pulling-our-codex-converts-back-to-claude) and it pulled their writers back from Codex; their shortest summary was "Anthropic fixed Opus's personality." Over on Hacker News (https://news.ycombinator.com/item?id=49804316), one commenter called the change in output style "a very big improvement over Opus 5", and in the main launch thread (https://news.ycombinator.com/item?id=49803892) another simply said "Opus 5.5 is a game changer for me."
The code quality data backs up the vibe. SonarSource ran it through their Java benchmark (https://www.sonarsource.com/blog/claude-opus-5-5-an-evaluation/) and found it wrote 27.5% less code and used 40% fewer output tokens than Opus 5, with a pass rate within a point and total findings down 42%. Less code to review is the most underrated feature a coding model can have.
A couple of practical patterns are also emerging:
- Medium effort is the sweet spot. Moe Lueker's testing (https://moelueker.com/blog/claude-opus-5-5-review) on FrontierCode had medium scoring 54.6% at about $0.80 a task, while max scored slightly lower at nearly eight times the cost.- Plan with one, build with the other. Plenty of Claude Code users are pairing Fable 5.1 for planning and orchestration with Opus 5.5 for execution. One HN commenter, quoted by explainx (https://explainx.ai/blog/fable-5-1-vs-opus-5-5-comparison-2026), boiled it down to "plan with fable, implement with opus."That matches my own week. Opus 5.5 on medium is now my default for anything that isn't trivial. ## It's not all roses In fairness, the reception isn't unanimous. Some HN users found it just as wordy as Opus 5 on day one, and others weren't impressed until they bumped the effort up. CodeRabbit (https://www.coderabbit.ai/blog/opus-5-5-model-review) liked its capability but flagged that its code reviews used more tokens and left more comments, so the cheaper sticker price doesn't automatically mean a cheaper bill. Sonar's numbers also showed injection-style findings going up, which is a good reminder that less code isn't the same as safer code. And if you work in bioinformatics or security research, the new safeguards (https://roo.beehiiv.com/p/claude-opus-5-5-review) mean many of those requests get routed to an older model instead. None of that changes my overall view, but you should go in with your eyes open. ## The memory complaints This is the bit that caught my attention, because it's the problem I've been working on. A GitHub issue on the Claude Code repo (https://github.com/anthropics/claude-code/issues/96527), originally about Opus 4.6, lists Opus 5.5 among the affected models and describes the model losing track of repository context within a single conversation. The reporter sums up the symptom as the "model loses track of established facts within a session", with real consequences like files being committed to the wrong repository. Another issue (https://github.com/anthropics/claude-code/issues/97387) reports a Max plan user seeing Opus 5.5's context window capped at 150k tokens in Claude Code. And in the HN launch thread, one frustrated developer asked whether 5.5 still skips over CLAUDE.md and reads a couple of lines of a file when asked to read the whole thing. It's worth pulling these apart, because they're really two different problems. In-session problems (context caps, compaction losing detail, a model not reading the file you told it to) are Anthropic's to fix, and some of these look like bugs or config issues rather than the model itself. Nothing I'm about to describe will lift a 150k cap. Cross-session problems are different, and they're not a bug at all. They're simply how these models work. Every new session starts from zero. Your agent doesn't know your stack, doesn't know you tried that library last Tuesday and it segfaulted on Alpine, and doesn't know you prefer Forgejo over GitHub. A 1M-token context window doesn't change that. Buda put it better than I could (https://buda.im/blog/claude-opus-5-5-context-window-agent-memory): "A large window delays forgetting; it does not solve persistence." The first problem makes the second one worse. If you're relying on a giant CLAUDE.md or on stuffing everything into the context window so the model might remember it, any hiccup in how that context is read or compacted hits you hard. ## Where OmniMem fits OmniMem (https://omnimem.org) is the self-hosted memory server I built because I got fed up re-explaining myself every morning. It's an MCP server, so it works with anything that speaks MCP: Claude Code, Claude on the desktop and web, Codex, Cursor, OpenCode, Copilot, Kiro and others. It runs on your own hardware, uses local ONNX embeddings with Valkey vector search, and it's MIT licensed and free. Rather than a flat file the model may or may not read, OmniMem gives your agent a proper memory it can query:
- A one-call briefing. At the start of a session, briefing() hands the agent your project context, recent experience, stale memories, new articles and any contradictions. No warm-up, no hoping a file gets read.- Recall across four namespaces. Episodic memories (decisions, bugs, what you tried), project context, a knowledge base fed from RSS, and your preferences, all searched together and ranked by similarity, recency, lifecycle state and how hard the lesson was to learn.- The graveyard. Every abandoned approach is kept on purpose, along with why it failed and how much effort it burned. Before the agent reaches for that library again, it gets a warning and the thing that worked instead.- Contradiction detection and semantic dedup, so your memory doesn't quietly fill up with conflicting or duplicate facts.- A real lifecycle. "Forget about X" usually means "stop bringing it up", so memories can be deprioritised or archived rather than just deleted, and reinstated if they become relevant again.- Compiled skills. compile_skill() turns accumulated experience into a SKILL.md for a domain, proposed as a diff you review before anything is written.So how does that map to the complaints?
- "It forgets established facts": those facts live in OmniMem and come back via briefing and recall, rather than depending on surviving compaction.- "It ignores CLAUDE.md": OmniMem's instructions are injected when the MCP server connects, and memory is pulled in by query when it's relevant.- "Context gets capped": retrieving only what's relevant keeps the context lean, which matters more when the window is smaller than you expected.- "It made the same mistake again": that's exactly what the graveyard is for.To be clear about the limits: OmniMem doesn't change how Opus 5.5 reasons within a session, and it won't fix a platform bug. What it does is make sure the things that matter survive between sessions, between projects and between machines. Pair it with Opus 5.5's much better instruction following and the combination is genuinely lovely to work with. It's how I've been running all week. ## Use it today: OmniMem v6 OmniMem v6 is available now (the latest release is 6.7.1). You'll need Docker and Docker Compose, and that's about it. The quickest route is the installer: curl -fsSL https://code.squarecows.com/ric/omnimem/raw/branch/main/install.sh | bash Or, if you'd rather see what you're running first (and you should, it's a curl-to-bash), do it by hand:
- git clone https://code.squarecows.com/ric/omnimem.git- cd omnimem and cp .env.example .env- Edit .env to set VALKEY_PASSWORD, and optionally ANTHROPIC_API_KEY- docker compose up -d (the web UI is then at http://localhost:8080)Then point your agent at the MCP endpoint, http://localhost:8765/sse. There's a setup guide for each supported client, including the exact Claude Code config, at https://code.squarecows.com/ric/omnimem/src/branch/main/guides That gives you four containers (Valkey with vector search, the MCP server, the RSS worker and the web dashboard), and nothing leaves your machine. ## Coming soon: OmniMem v7 Docker Compose is great if you live on a Linux server like I do, but it's a barrier for a lot of people. v7 is a rebuild as a single Rust binary, and it comes with proper desktop installers: an MSI for Windows, a DMG for macOS and a Flatpak for Linux. Once it's running you'll get a system tray icon that takes you straight to the settings, with no YAML or .env editing required. The Docker image isn't going anywhere, so if you're running headless on a server nothing changes for you. In short, if v6 is "clone, configure, compose", v7 is "download, double-click, done". ## Give your agent a memory Opus 5.5 is the best model Anthropic has shipped for day-to-day work, and I'd switch to it today. Just don't expect a bigger context window to do the job of memory. Give it somewhere to keep what it learns. Head over to https://omnimem.org to get started, and if you hit a snag or have an idea, the issue tracker is open at https://code.squarecows.com/ric/omnimem/issues This article was researched and drafted with help from Claude Opus 5.5, which felt fitting. ### Sources
- Anthropic, Introducing Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5- eesel, Claude Opus 5.5 review: https://www.eesel.ai/blog/claude-opus-5-5-review- Every, Vibe Check: https://every.to/vibe-check/vibe-check-opus-5-5-is-pulling-our-codex-converts-back-to-claude- SonarSource evaluation: https://www.sonarsource.com/blog/claude-opus-5-5-an-evaluation/- CodeRabbit, Opus 5.5 for code review: https://www.coderabbit.ai/blog/opus-5-5-model-review- Moe Lueker, Claude Opus 5.5 review: https://moelueker.com/blog/claude-opus-5-5-review- Memeburn, subscriber limits: https://memeburn.com/anthropic-says-opus-5-5-costs-40-less-to-run-claude-subscribers-get-about-25-more-room- explainx, Fable 5.1 vs Opus 5.5: https://explainx.ai/blog/fable-5-1-vs-opus-5-5-comparison-2026- Roo, what breaks when you switch: https://roo.beehiiv.com/p/claude-opus-5-5-review- Hacker News launch thread: https://news.ycombinator.com/item?id=49803892- Hacker News, Artificial Analysis thread: https://news.ycombinator.com/item?id=49804316- GitHub anthropics/claude-code #96527: https://github.com/anthropics/claude-code/issues/96527- GitHub anthropics/claude-code #97387: https://github.com/anthropics/claude-code/issues/97387- Buda, 1M context is not agent memory: https://buda.im/blog/claude-opus-5-5-context-window-agent-memory
Top comments (0)