Claude Opus 5.5 fixed all 18 critical defects in GoML's LatticeBench evaluation and resolved six of seven very_hard issues. The evaluation also revealed a surprising weakness: despite its performance on complex security defects, the model struggled with straightforward, one-line tickets.
At GoML, we found two cross-file mass-assignment vulnerabilities that had survived every previous test run. Opus 5.5 fixed both, setting a new benchmark for security-focused coding agents.
The model scored 87.3 on the AI Matic Bench Score, exceeding the previous highest score by 7.5 points. It also offers lower token pricing than its predecessor, Opus 5.
However, our findings suggest that strong benchmark performance does not eliminate the need for human review. Opus 5.5 modified 107 files, including security-sensitive settings and model fields that were outside the original task scope.
Our recommendation: use Opus 5.5 for complex code audits and agentic coding, but require diff reviews before merging changes into production.
Key Results at a Glance
| Metric | Claude Opus 5.5 |
|---|---|
| AI Matic Bench Score (eight-axis) | 87.3/100 |
| AI Matic Bench Score (nine-axis v0.2) | 87.6/100 |
| API input pricing per 1M tokens | $4 |
| API output pricing per 1M tokens | $20 |
| Cache reads per 1M tokens | $0.20 |
| Reported cost reduction on typical workloads | 40% |
| LatticeBench total coverage | 44 of 57 (77.2%), plus one partial |
| Critical defects fixed | 18 of 18 (100%) |
very_hard defects fixed |
6 of 7 (85.7%) |
| Bug-mode fixes without hints | 40 of 49 (81.6%) |
| Task-mode completion | 4 of 8 (50%) |
| Fabricated claims across 119 findings | 0 detected |
The results position Opus 5.5 as a strong candidate for security auditing and complex code changes. Its performance on smaller, explicitly described tasks remains a concern.
Why Our Opus 5.5 LLM Testing Does Not Match the Run Report
GoML's original LatticeBench scoring report gave Opus 5.5 a score of 79.3. Four evaluation areas—pricing, latency, deployment, and regulated-domain depth—lacked sufficient data, so we assigned neutral placeholder scores of 7.0.
The report also included an evidence-only score of 88.5.
Following the release of additional Anthropic material, we obtained evidence for all four areas and recalculated the overall score to 87.3.
Both earlier scores were mathematically correct based on their respective inputs. We independently recalculated the evidence-only score of 88.5 and reproduced it exactly.
The difference came from the available evidence, not a calculation error.
Results from LatticeBench
LatticeBench is a multi-service enterprise codebase containing 57 deliberately injected defects, including 18 critical-severity issues, alongside eight task prompts.
The benchmark uses diff-based scoring. Each patched file is compared with the original injected state and the expected solution recorded in the manifest.
The model's self-reported results are used to check whether it makes unsupported claims, rather than serving as the primary measure of whether a defect was fixed.
Benchmark Comparison
| Measure | Muse 1.3 | GPT-6 Astra | Grok 4.7 | Opus 5.5 |
|---|---|---|---|---|
| Bug-mode, no hints (49) | 21 | 33 | 25 | 40 (81.6%) |
| Task-mode completed (8) | 5 | 0 | 6 | 4 (50%) |
| Critical severity (18) | 9 | 15 | 14 | 18 (100%) |
| High severity (17) | Not split | 12 | 8 | 15 (88.2%) |
very_hard tier (7) |
3 | 4 | 3 | 6 (85.7%) |
| Total coverage (57) | 26 | 33 | 31 | 44 (77.2%) |
Note: Runs 1 and 2 were prose-scored; subsequent runs used diff-based scoring and are directly comparable.
The most important result is the model's ability to resolve complex, multi-file security defects that other models repeatedly missed.
The Two Chains No Model Had Solved
Two particularly important defects involved cross-file mass-assignment vulnerabilities.
SALES-001: Organization-Level Access Control
SALES-001 had survived five runs across five models.
The organization_id field remained editable in the opportunity-deal update schema, potentially allowing users to access another organization's data.
Fixing the issue required tracing the change through the schema, facade, and service layers.
Opus 5.5 removed the field exactly as the benchmark's expected solution specified.
SLA-002: SLA Policy Serializer
Opus 5.5 fixed SLA-002, the corresponding issue in the SLA policy serializer, using the same approach.
It also resolved four additional issues:
BUGS-002ENGAGEMENT-005BUGS-004COMMON-003
In total, Opus 5.5 resolved six of the seven very_hard benchmark entries.
What This Changes for the Benchmark
LatticeBench was designed to challenge models with multi-file mass-assignment defects. Across five previous runs, every tested model missed this pattern.
Opus 5.5 has now demonstrated that it can detect and resolve these issues. The only remaining very_hard miss is WM-006, involving a soft-delete manager swap.
This performance suggests that the benchmark's hardest tier needs more challenging tasks to continue distinguishing frontier models.
The Inverted Failure Mode
One of the most interesting findings is the difference between performance on complex defects and routine ticket-based work.
Opus 5.5 fixed 100% of critical defects but only 47.4% of medium-severity defects and 50% of explicit tickets.
Four of its eight task-mode misses involved straightforward, one-line index or wiring changes.
At the same time, the model successfully resolved complex issues such as SALES-003 and ORG-001.
This suggests that the problem may not be the inherent difficulty of the task. Instead, ticket framing, prompt structure, or task prioritization could be contributing factors.
Running task mode separately would help determine the cause.
When the Model Fixes the Wrong Bug
Some failures involved the model finding and fixing a legitimate bug in the correct part of the code while leaving the injected defect unresolved.
For WM-005, the model fixed a problem at line 303 but missed the injected defect at line 388. Grok 4.7 made a similar partial fix in another run, suggesting the code itself may contribute to the failure.
Two other cases stood out: ENGAGEMENT-003 and WM-003.
In ENGAGEMENT-003, Opus 5.5 added attempt_count persistence but failed to restore the idempotency guard. As a result, duplicate sends could still occur.
This is a significant review risk. A developer inspecting the diff might see a reasonable change in the correct file and assume the original issue had been resolved.
The lesson is straightforward: a plausible code change is not proof that the reported defect has been fixed. Every patch should be checked against the original failure condition and expected behavior.
The Out-of-Scope Problem in Opus 5.5 LLM Testing
The audit produced 119 findings across 107 files. Of these, 44 were linked to injected defects.
The remaining findings covered 62 files outside the benchmark manifest and four new files. Every cited path had a corresponding change, and all 103 modified Python files passed syntax checks.
No fabricated claims were detected. However, several changes went beyond the original task scope and could affect client-facing behavior.
Changes Requiring Human Review
| Change | Why It Matters |
|---|---|
| Access-token lifetime changed from 60 to 15 minutes; refresh-token lifetime changed from 30 to 7 days | The changes were justified using the README and .env.example, but they differed from the expected solution and could cause users to log in more frequently. |
Django SECRET_KEY default changed to a placeholder |
This could invalidate existing sessions and CSRF tokens wherever the default value was being used. |
BugLabelAssignment.label_id renamed to label without a migration |
The reasoning appeared correct, but the actual database schema must be verified before merging. |
| Kong Admin API port bound to loopback in the development Compose file | This is a reasonable security hardening measure, but it changes access behavior for clients connecting through another host. |
These changes highlight an important distinction between security improvements and authorized changes.
A model can identify real issues and make technically reasonable improvements while still introducing changes that were never requested.
The model also documented issues it chose not to change, including unreachable routers and design decisions requiring human approval.
That behavior deserves recognition. Clearly explaining what was left untouched makes the output easier to review and reduces the risk of unapproved changes being mistaken for part of the original fix.
Cost: A More Affordable Coding Agent
Opus 5.5 reduces token pricing compared with Opus 5.
| Pricing Metric | Opus 5.5 |
|---|---|
| Input tokens per 1M | $4 |
| Output tokens per 1M | $20 |
| Cache reads per 1M | $0.20 |
| Reported reduction on typical workloads | 40% |
| Reported output-generation speed improvement | More than 30% |
The reduction in cache-read pricing is particularly relevant for agentic coding tasks that repeatedly reuse context.
Anthropic reports that the combined changes reduce costs by 40% on typical workloads at default settings.
Cost and Performance in Real-World Tests
The reported comparisons also suggest improvements in practical coding workloads:
- FrontierCode and GDPval-AA v2.1: Opus 5.5 costs less than Astra.
- CursorBench: It costs approximately one-third as much as GPT-5.6 Sol while scoring 11 points higher.
- HAProxy C-to-Rust translation: It completed the task 2.5 hours faster than Fable 5.1 while costing 51% less.
- Large-codebase repair: It fixed a 200,000-line codebase in under three hours, compared with more than 20 hours for Opus 5.
- Code review: Deloitte reported that Opus 5.5 caught 72% of known bugs at its lowest effort setting, compared with 56% for Opus 5 at high effort.
These results indicate that Opus 5.5 may offer a compelling combination of cost, speed, and defect detection. Actual savings will depend on workload characteristics, context reuse, and review requirements.
Safety and Access Limits
Opus 5.5 achieved Anthropic's highest scores to date on its automated behavioral audit, which evaluates thousands of simulated scenarios.
External organizations, including Frontier Design and METR, also evaluated the model before release.
According to the reported findings, Opus 5.5 resisted prompt injection better than Opus 5 in every tested setting and matched Fable 5.1 for the lowest prompt-injection success rate on the Gray Swan benchmark.
The model also includes an action-screening classifier and an open-source sandbox that security teams can inspect.
In GoML's evaluation:
- Every secret-path attempt was blocked.
- No fabricated claims were detected across 119 findings.
- The modified Python files passed syntax checks.
These results are encouraging, but they do not remove the need for access controls, isolated testing environments, and human approval of sensitive changes.
Regulated-Domain Access
Access restrictions remain important for certain workloads.
- Biology research: Access requires approval through the Life Sciences Verification Program.
- Cybersecurity: Access requires approval through the Cyber Verification Program.
Both programs have safeguards similar to those used for Fable 5.1.
Teams planning security-audit engagements should apply for the necessary access before committing to delivery dates.
What the Results Show
Across the evaluation, six models fell within a 15-point score range.
The results also highlighted differences in reporting quality. Some models provided limited safety information, while the highest-scoring model published standard errors and acknowledged that its lead could be exaggerated.
This reinforces a broader point: benchmark performance alone is not enough to establish whether a model is suitable for enterprise deployment.
Transparent reporting, reproducible evaluations, accurate claims, and reviewable code changes are equally important.
The GoML Position: When to Use Claude Opus 5.5
Use Opus 5.5 When
1. Auditing or changing a large codebase
The model identified all 18 critical issues in our evaluation and completed a reported 200,000-line codebase audit in under three hours.
2. Replacing an existing Opus 5 workflow
Lower token prices and cheaper cache reads may reduce costs for tasks that reuse substantial amounts of context.
3. Human review is part of the workflow
The model's ability to explain which issues it changed and which it left untouched can help reviewers assess its work.
4. Delivering code changes for clients
Clear explanations and traceable diffs make it easier for reviewers to understand and validate the output.
Do Not Use Opus 5.5 When
1. Code is merged automatically without review
The 107-file audit included changes to token lifetimes, a Django SECRET_KEY, and a model field name. These changes require human review before merging.
2. The workload consists mainly of small, explicit fixes
The model completed only four of eight task-mode prompts. A less expensive model may be sufficient for routine tickets, provided it meets the required quality threshold.
3. The work involves biology or offensive security without the necessary access
Obtain the required verification-program approval before confirming delivery commitments.
4. Procurement requires confirmed data-policy details
The launch material did not establish data-retention or cloud-availability details. These must be confirmed directly before a client pilot.
Recommended Next Steps
Based on the findings, GoML recommends four actions.
-
Expand the benchmark: Add more challenging
very_hardtasks, given that Opus 5.5 resolved six of seven entries. - Separate task-mode evaluation: Determine whether missed one-line tickets result from prompt structure, task prioritization, or another factor.
-
Review out-of-scope changes: Specifically inspect token lifetimes, the Django
SECRET_KEY, and theBugLabelAssignmentfield rename before approving the patch. - Validate the scoring changes: Four additional fixes would raise coverage from 44 to 48 of 57 defects. Re-run the benchmark after verifying the fixes and their impact.
These steps will help establish whether the model's improved security performance translates into consistently reliable enterprise coding workflows.
Conclusion
Claude Opus 5.5 represents a significant improvement in GoML's security-focused coding evaluation. Fixing all 18 critical defects and six of seven very_hard issues demonstrates its ability to handle complex vulnerabilities that previous models repeatedly missed.
Its limitations are equally important. Missed routine tickets and changes outside the requested scope show why strong coding-agent performance must be paired with disciplined review.
For enterprise teams, the practical approach is to use Opus 5.5 for complex audits and agentic coding, require diff-based human review, and validate every fix against the original defect.
Follow the GoML blog for more LLM testing and updates on the AI Matic Bench Score, a practical standard for enterprise AI evaluation.
Top comments (1)
Checking the table against the sample sizes: Opus 5.5's 40/49 bug-mode fixes has a 95% Wilson interval of about 69-90%, and GPT-6 Astra's 33/49 is about 53-79%. Treated as independent groups that gap is Fisher p of about 0.16, so "ahead" is the observed order, not yet a separated one. The 18/18 critical result is 82-100% (Astra's 15/18 against it is p of about 0.23). The "inverted failure mode" rests on 4/8 task-mode completions (22-79%) and, if 47.4% medium is 9/19, a 27-68% interval. Those are the numbers I would hold loosely before concluding ticket framing is the cause.
Since every model fixed or missed the same 49 defects, the right comparison is paired: count the defects only Opus fixed against the ones only Astra fixed (a McNemar-style test on the discordant pairs). You have that per-defect data already, and it is usually much sharper than the aggregate. A second run per model would also show how much of the gap is run-to-run noise, which matters for the "7.5 points above the previous best" claim.
On the 79.3 to 87.3 rescore: the arithmetic is fine, but were the weights and the evidence for the four placeholder axes fixed before anyone saw the defect results? Publishing them in advance is what would make the composite comparable across models.