Claude Sonnet 5.5: benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1985 ± 16, rank #7 of 1607 rated models, from 162 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA-Briefcase | 63.23 | Rubric Pass Rate (%) | 100 |
| AA-Briefcase - XLSX - Rubric Pass Rate | 64.77 | Rubric Pass Rate (%) | 100 |
| AutomationBench-AA | 71.33 | Guardrail-Adjusted Objective Completion (%) | 100 |
| AutomationBench-AA - HR - Raw Objective Completion | 79.6 | Raw Objective Completion (%) | 100 |
| AutomationBench-AA - Operations - Raw Objective Completion | 97.05 | Raw Objective Completion (%) | 100 |
| AutomationBench-AA - Strict | 44.75 | Strict Task Success (%) | 100 |
| Bug Hunt Bench | 57 | Planted Bugs Fixed (out of 105) | 100 |
| Bug Hunt Bench - LMS | 33 | Planted Bugs Fixed (out of 60) | 100 |
| Bug Hunt Bench - VS Code Extension | 24 | Planted Bugs Fixed (out of 45) | 100 |
| Harvey LAB-AA - Trusts, Estates & Private Client | 95.8 | Criterion Pass Rate (%) | 100 |
| LLM Stats (BenchCAD (with Python tool)) | 96.3 | Score (%) | 100 |
| LLM Stats (BenchCAD) | 74.7 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.