Every few weeks after a model launch, the same thread appears: "is it just me, or is Claude dumber today?" Claude Opus 5.5 shipped on September 22, and this time one developer started measuring before the complaints did. livenerf runs a frozen panel of hard questions against Opus 5.5 once a day for 30 days, with pinned tooling and public logs, so the next "Claude nerf" argument can be settled with a baseline from launch week. If you build on Claude Code or the API, the method is worth copying for your own agents.
TL;DR
- livenerf tests Claude Opus 5.5 once a day for 30 days, starting 2.5 days after launch. Six of 30 days were in on September 29, none missed.
- The panel is 78 questions Opus 5.5 gets right sometimes, picked from 2,336 screened. Questions it always gets right can't reveal a drop.
- It can detect a drop of about 7.5 accuracy points per 10-day window. The first possible call lands around October 24.
- In a deliberate test, lower "effort" cut output tokens by 26–62 % but accuracy by only 4–8 points. Token count is the early warning.
- Anthropic's own position, from its 2025 postmortem: "We never reduce model quality due to demand, time of day, or server load."
What is a Claude nerf, and why was it impossible to prove?
"Nerf" is gamer slang for a patch that makes something weaker. Applied to language models, it means the vendor quietly serves a worse model under the same name some days or weeks after launch. The livenerf README lists the usual suspects: quantization, a smaller model behind the same name, lower effort, or routing changes. It also lists the boring option: "It could also mean nothing happened and people are pattern-matching on noise."
The problem was always the missing starting point. In the README's words: "Nobody has had a clean day-0 baseline to check against, so every argument ends up as vibes versus vibes." Your memory of how good the model felt on launch day is the least reliable instrument in the room. You were excited, your prompts were different, and you remember the wins.
The Hacker News thread about the repo (786 points, over 300 comments) shows both camps. One commenter wrote: "\"Nerf\"ing models isn't real in the vast majority of reported cases." Another: "people are just getting used to the new level of intelligence." And one Claude Code user explained that they now dismiss the in-app feedback pop-up every time, because "the quality is more consistent", before adding, to their credit, "Complete adhoc and personal experience." That is the state of the evidence livenerf is trying to replace.
How livenerf measures Claude Opus 5.5
The design is a list of decisions that each remove one way the measurement could lie.
A panel that can move. The author screened 2,336 questions from GPQA Diamond, MMLU-Pro, competition math and AIME 2025–26, four samples each. Opus 5.5 got about 93 % right on the first try, and 97 % of the questions were always right or always wrong. A question the model always aces can't show a decline, and one it always fails can't either. The 78 questions it only sometimes gets right became the panel, because that is where a weaker model would show up first.
No AI judge. Grading is exact match, "no LLM judge, ever, since the judge would drift too." If the grader is another model, a nerf to the grader looks like a change in the subject.
Hermetic calls. Frozen system prompt, no tools, no MCP servers, no CLAUDE.md.
A pinned harness. It runs on a Claude Max subscription through headless Claude Code (claude -p), with the CLI pinned at version 2.1.280. The README is blunt about why: "Pin the CLI. This is not optional: a Claude Code update changes the harness, and a changed harness looks exactly like a changed model."
A subscription instead of the API. PLAN.md estimates the API version of the full daily run at about $55 a day, roughly $1,600 a month. The subscription route also measures what most complainers actually use.
A control arm. The older claude-opus-5 answers the GPQA questions every day too. "If both models move together, the harness or platform changed." If only Opus 5.5 moves, the change is in the model.
A pre-registered decision rule. A nerf is called only if the 99 % interval excludes zero in two consecutive 10-day windows, the drop is at least 3 points, and the control arm did not move with it. Writing the rule down before the data arrives is what keeps the author honest when day 20 looks exciting.
The tooling is built on Inspect from the UK AI Security Institute, and the statistics follow "Adding Error Bars to Evals" by Evan Miller, published by Anthropic. So Anthropic's own statistics paper is the method auditing Anthropic's model.
What this LLM benchmark can and can't see
Before trusting a detector, you break something on purpose and check that it notices. The author did that with Claude's effort setting, and the results are in VALIDATION.md:
| Setting vs. high effort | Output tokens | Accuracy |
|---|---|---|
| medium effort | −26 % | −4.2 ± 3.9 points |
| low effort | −62 % | −8.3 ± 4.5 points |
Median output tokens per sample went from 638 at high effort to 474 at medium and 293 at low. The README's summary: "Lower effort shows up much more clearly in tokens than in accuracy." Cut the thinking by almost two thirds and the score drops by eight points, which is barely above the detection threshold.
The blind spot is stated up front. "Swapping in Opus 5 was not distinguishable from Opus 5.5 at 99%" — the measured gap was −3.8 ± 6.3 points, with 23 % fewer tokens. The most plausible quiet downgrade, a same-family swap to the previous model, is exactly the one the rig can't call yet. I respect a benchmark that prints its own limits on the front page.
The audit also turned up problems in the benchmarks themselves: among the 78 panel questions, 8 answer keys look wrong and 30 are ambiguous. They were flagged and kept, so the panel stays constant. The nerf detector found bugs in the test set before it found a nerf.
One more serving detail: "The safety classifier sometimes answers with Opus 5 or refuses biology and some math questions." Those samples are rejected from the score.
Is Claude getting dumber? What Anthropic has said
Anthropic addressed the question directly in September 2025, after a wave of complaints about Claude quality. Its engineering post, "A postmortem of three recent issues", says: "To state it plainly: We never reduce model quality due to demand, time of day, or server load. The problems our users reported were due to infrastructure bugs alone."
The same post confirms users were seeing something real: "Between August and early September, three infrastructure bugs intermittently degraded Claude's response quality." And: "Approximately 30% of Claude Code users who made requests during this period had at least one message routed to the wrong server type."
So the 2025 complaints were right about the symptom and wrong about the motive. livenerf takes that history seriously. Its README notes that "The 2025 quality incidents turned out to be infrastructure bugs, not deliberate downgrades," and warns that "Launch week could easily be the worst week." Load, new serving paths and fresh bugs all peak right after launch, so a baseline taken in week one might be the low point.
As of September 29, six of 30 days are in. Days 1–10 are the baseline, then come two 10-day windows. The honest answer to "has Opus 5.5 been nerfed?" today is that nobody knows yet, including everyone who was sure.
How to detect a model downgrade in your own Claude Code agents
The part worth stealing is small and cheap:
- Pin the model version in API calls instead of an alias, and pin your Claude Code CLI version in CI. A harness update is indistinguishable from a model change.
- Log output tokens per task. The validation shows reduced effort moves tokens by a quarter to two thirds before accuracy moves much. If your agent suddenly writes shorter answers for the same work, that is your first signal.
- Keep a handful of frozen tasks with exact, machine-checkable answers, and run them on a schedule. Pick ones your model gets right only some of the time.
- Grade with code. Plain functions, exact match, no model in the grading loop.
- Write the decision rule down first. How big a drop, over how many days, before you call it.
A minimal sketch of the logging half (illustrative, not from the repo):
# simplified sketch: append one line per task run
import json, time
def log_run(task_id, model, cli_version, output_tokens, correct):
with open("model_health.jsonl", "a") as f:
f.write(json.dumps({
"ts": time.time(), "task": task_id, "model": model,
"cli": cli_version, "out_tokens": output_tokens, "correct": correct,
}) + "\n")
Check that file before you check Reddit.
Also in this episode
European data centers keep their numbers secret. NL Times reports that Lighthouse Reports, Trouw and other European outlets spent a year on freedom-of-information requests about data center energy and water use, and "Their attempts have yielded no results." The EU Energy Efficiency Directive has required data centers of 500 kW and up to report for three years. In the Netherlands, public electricity figures exist for 44 of 186 such sites. Microsoft's largest Dutch data center alone uses about 1 % of the country's electricity, according to RVO data. (HN thread)
Backblaze did the boring thing again. Its Q2 2026 Drive Stats cover 354,415 drives. The quarterly annualized failure rate rose to 1.73 %, "the highest it's been in quite a while", against a lifetime rate of 1.41 %. Three Seagate models (ST8000NM000A, ST12000NM000J, ST14000NM000J) had zero failures. Thirteen years of publishing the same table every quarter is what a public baseline looks like, and it is the discipline livenerf is copying. (HN thread)
Verdict: SHIP IT
I stamped livenerf SHIP IT. It is the first nerf argument I've seen with a timestamp, a public rulebook and a stated blind spot. My favourite line is in the credits: "a lot of this repo is written with the help of Claude, which is the model being measured." Claude helped build the meter that checks whether Claude got dumber, which is exactly why the graders are plain functions. In the author's words, you shouldn't have to trust the author, human or otherwise.
Just don't quote it before October 24.
FAQ
Has Claude Opus 5.5 been nerfed?
There is no evidence either way yet. livenerf has 6 of 30 daily runs, and its pre-registered rule needs two 10-day windows after the baseline. The first possible call is around October 24, 2026.
How does livenerf detect a nerf?
It runs 78 "sometimes right" questions through a pinned, headless Claude Code every day, grades them by exact match, and compares 10-day windows against the launch-week baseline with 99 % intervals, while a control model checks for harness changes.
Does Anthropic reduce model quality under load?
Anthropic says no: "We never reduce model quality due to demand, time of day, or server load." Its 2025 postmortem attributes that period's quality problems to three infrastructure bugs.
What is the cheapest way to watch my own agent for a downgrade?
Pin model and CLI versions, log output tokens per task, and run a few frozen, exactly graded tasks on a schedule. Falling token counts are the earliest signal.
Sources
- livenerf repository · README · PLAN.md · VALIDATION.md
- Hacker News: "Livenerf: Has Opus 5.5 been nerfed yet?"
- Evan Miller, "Adding Error Bars to Evals" (arXiv)
- Inspect, UK AI Security Institute
- Anthropic, "A postmortem of three recent issues"
- NL Times: data centers refusing to say how much water and electricity they use
- Backblaze Drive Stats for Q2 2026
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.


Top comments (0)