Claude Fable 5.1 Benchmarks: The Numbers That Matter
Anthropic launched Claude Fable 5.1 on September 1, 2026. Across nine published benchmarks, its largest gain is Terminal-Bench-Science: 52.6% versus 24.7% for Fable 5. Its smallest lead is CursorBench, where it beats Opus 5 by only 3.4 points despite costing twice as much. This guide collects the published results, explains what they measure, and shows how to validate them against your own workload.
The primary source is Anthropic’s launch post. For background, see what Claude Fable 5.1 is.
The full benchmark table
| Benchmark | What it measures | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | Agentic scientific research in a terminal | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 | Agentic coding in a terminal | 55.8% | 42.0% | 52.3% | 37.3% |
| GDPval-AA v2 | Economically valuable knowledge work (Elo) | 1853 | 1723 | 1824 | 1711 |
| OSWorld 2.0 (partial credit) | Computer use on real desktop tasks | 77.9% | 72.9% | 75.4% | Not reported |
| OSWorld 2.0 (strict) | Same, counting only full task completion | 41.7% | 36.1% | 39.6% | Not reported |
| Humanity’s Last Exam (no tools) | Expert-level reasoning questions | 60.9% | 57.8% | 56.6% | Not reported |
| Humanity’s Last Exam (with tools) | Same, with search and code | 65.0% | 63.8% | 63.6% | Not reported |
| AutomationBench | End-to-end business workflows | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 | IDE-style coding tasks | 73.4% | 70.5% | 70.0% | 67.2% |
Anthropic also published one Mythos 5.1 result: 60.9% on Terminal-Bench 4.0.
VentureBeat reported an additional Browserbase result: Fable 5.1 scored 82%, compared with 74% for Opus 5 and 57% for Fable 5 on the hardest browser-agent tasks. Because this figure comes from press coverage rather than Anthropic’s launch post, treat it as less authoritative.
The large gaps
Terminal-Bench-Science 0.1
Fable 5.1: 52.6% · Fable 5: 24.7% · Opus 5: 29.0%
This is the standout result. Fable 5.1 more than doubles its predecessor and nearly doubles Opus 5 on a benchmark that gives the model a terminal, scientific tools, and multistep research tasks.
It is the clearest evidence that Anthropic’s focus on “long-horizon problem-solving” produced a measurable improvement. However, this is a version 0.1 benchmark, so the task set and scores may change as it matures.
AutomationBench
Fable 5.1: 31.4% · Fable 5: 17.1% · Opus 5: 26.9%
AutomationBench evaluates end-to-end business workflows across applications. Absolute scores remain low for every model, but Fable 5.1 nearly doubles Fable 5 and leads Opus 5 by 4.5 points.
If your product automates multi-application business processes, this row is especially relevant.
Terminal-Bench 4.0
Fable 5.1: 55.8% · Fable 5: 42.0% · Opus 5: 52.3%
Fable 5.1 gains 13.8 points over Fable 5 on agentic coding and leads Opus 5 by 3.5 points.
Opus 5 had surpassed Fable 5 on this benchmark in July. Fable 5.1 restores the Fable tier’s advantage, although the lead is narrower than it was in June. See the Opus 5 benchmarks breakdown for the July results.
GDPval-AA v2
Fable 5.1: 1853 · Fable 5: 1723 · Opus 5: 1824
Fable 5.1 gains 130 Elo over Fable 5 on knowledge-work tasks. This is the benchmark Anthropic cites for its document, spreadsheet, and presentation claims.
Against Opus 5, the difference is only 29 Elo—a relatively small gap.
The small gaps
CursorBench 3.2.0
Fable 5.1: 73.4% · Fable 5: 70.5% · Opus 5: 70.0%
CursorBench measures IDE-style coding tasks. Fable 5.1 leads both models by about three points, making this its narrowest advantage.
This benchmark may matter most to developers. If your workload is primarily “help me in my editor,” these results alone do not justify paying twice as much.
OSWorld 2.0
Partial credit: Fable 5.1 77.9% vs Opus 5 75.4%
Strict completion: Fable 5.1 41.7% vs Opus 5 39.6%
Fable 5.1 leads Opus 5 by roughly two points under both scoring methods and leads Fable 5 by five points.
The strict score is the more sobering metric: fewer than half of tasks were fully completed by any model. Computer use remains one of the frontier’s weakest capabilities.
Humanity’s Last Exam
Without tools: Fable 5.1 60.9% vs Opus 5 56.6%
With tools: Fable 5.1 65.0% vs Opus 5 63.6%
Without tools, Fable 5.1 leads Opus 5 by 4.3 points—the largest of the smaller gaps. With search and code execution, the advantage falls to 1.4 points.
The no-tools score better represents raw reasoning. The with-tools score is closer to typical production usage.
Fable 5.1 vs GPT-5.6 Sol
Anthropic published GPT-5.6 Sol results for five rows. Fable 5.1 leads on each:
- Terminal-Bench-Science: 52.6% vs 22.4%, a 30.2-point lead
- Terminal-Bench 4.0: 55.8% vs 37.3%, an 18.5-point lead
- AutomationBench: 31.4% vs 19.6%, an 11.8-point lead
- CursorBench: 73.4% vs 67.2%, a 6.2-point lead
- GDPval-AA v2: 1853 vs 1711
Two caveats are important:
- Anthropic ran these competitor comparisons itself.
- GPT-5.6 Sol results are missing from five of the nine rows, and Anthropic did not explain why.
The earlier GPT-5.6 Sol vs Fable 5 comparison covered the previous matchup. On the published numbers, Fable 5.1 widens every gap reported for Fable 5.
What launch coverage skipped
Every result is vendor-run
Anthropic ran all benchmark tests, including the GPT-5.6 Sol comparisons. No independent lab had reproduced the results at launch.
Anthropic’s benchmark history has generally held up, but small gaps—especially on CursorBench and OSWorld—could change with a different harness, prompt set, or scoring method.
Mythos 5.1 is the higher-performing sibling
Anthropic’s system card and launch post describe Fable 5.1 and Mythos 5.1 as “the same model but with different levels of safeguards.”
On Terminal-Bench 4.0:
- Mythos 5.1: 60.9%
- Fable 5.1: 55.8%
The five-point difference shows the capability cost of the safeguards on agentic coding. It also means Anthropic’s most capable widely released model has a more capable sibling. The Mythos 5.1 vs Fable 5.1 comparison covers access differences.
Anthropic did not state the effort level
Anthropic’s effort documentation says Fable 5.1’s gains are largest at xhigh and max.
The launch post does not identify the effort level used for each result or state which settings were used for the comparison models. Because effort is a major quality and cost lever, this omission makes exact reproduction difficult.
Multilingual performance is flat
Anthropic says multilingual performance is on par with Fable 5 rather than improved. If your workload is not in English, the benchmark table may overstate the practical gain.
Safeguard metrics measure a different outcome
Anthropic reports:
- 85% fewer biology false positives on benign requests
- About 60% fewer cyber interventions per Claude Code session versus Fable 5
These figures come from Anthropic’s own traffic and are not externally reproducible. For users who frequently encounter refusals, however, they may matter more than capability scores.
What to test yourself
A launch table tells you where to investigate. Your own evals determine whether migration is worthwhile.
1. Run a long-horizon agent task
Use a representative task from your product and compare:
- Fable 5.1 at high effort
- Opus 5 at
xhigh - Fable 5 at high effort
This approximates Terminal-Bench-Science and AutomationBench. It will show whether the higher price pays off for your workflow.
2. Compare your hardest coding prompts
Run your ten most difficult coding prompts at the same effort level on Fable 5.1 and Opus 5. If the results are tied, you have reproduced the CursorBench story—and Opus 5 may be the better value.
3. Sweep effort levels
Run one routine task on Fable 5.1 at every setting from low through max. Anthropic claims that:
- Medium effort matches Fable 5
- Low effort can compete with Opus 5 on cost per task
- Gains are largest at
xhighandmax
Measure both quality and cost rather than relying on a single score.
4. Count benign refusals
Run benign security and life-sciences prompts on Fable 5 and Fable 5.1. Track refusal rates to test Anthropic’s 60% and 85% claims.
You can build all four evals as requests in Apidog. Define model and effort as environment variables, then add assertions for:
stop_reason- An expected string in the response
- Request success or failure
- Latency and token usage, if available
Run the collection once per model and compare the resulting history. This gives you a reproducible evaluation table without building new infrastructure. Download Apidog to get started; the Claude Fable 5.1 API walkthrough includes the request shapes.
How the benchmarks map to a decision
- Long-horizon agents, research, and automation: The gaps are large and consistent. Fable 5.1 is the strongest choice in Anthropic’s table. The pricing breakdown explains how cheaper cache reads narrow the cost difference on these workloads.
- IDE coding, knowledge Q&A, and tool-assisted reasoning: The advantage over Opus 5 is only one to four points. Stay on Opus 5 unless Fable 5.1 wins your higher-effort evals. See the Fable 5.1 vs Opus 5 comparison.
- Migration from Fable 5: Every published capability row improves, while the price does not. Anthropic also reports fewer false positives. The Fable 5.1 vs Fable 5 comparison covers migration considerations.
FAQ
What is Claude Fable 5.1’s best benchmark result?
Terminal-Bench-Science 0.1 at 52.6%. That is more than double Fable 5’s 24.7% and ahead of Opus 5’s 29.0% and GPT-5.6 Sol’s 22.4%. All figures come from Anthropic.
How does Fable 5.1 compare with Opus 5 on coding?
Fable 5.1 scores 55.8% vs 52.3% on Terminal-Bench 4.0 and 73.4% vs 70.0% on CursorBench 3.2.0. The advantage is real but narrow—roughly three points.
Is Fable 5.1 better than GPT-5.6 Sol?
On the five benchmarks where Anthropic published both results, yes. Fable 5.1 leads by 6 to 30 points. Anthropic ran the GPT-5.6 Sol tests itself, and four other benchmark rows have no GPT-5.6 Sol result.
Are the Fable 5.1 benchmarks independently verified?
Not at launch. Every number is Anthropic-run. Treat the results as claims to validate against your own workload.
How does Mythos 5.1 score?
Anthropic published one result: 60.9% on Terminal-Bench 4.0, five points above Fable 5.1’s 55.8%. Anthropic describes both as the same model with different safeguards.
Which effort level produced these scores?
Anthropic did not say. Its guidance is that Fable 5.1 performs best at xhigh and max, so reproduce the tests across effort levels before drawing conclusions.





Top comments (0)