Claude Sonnet 5.5: benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1985 ± 16, rank #7 of 1607 rated models, from 162 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA-Briefcase63.23Rubric Pass Rate (%)100
AA-Briefcase - XLSX - Rubric Pass Rate64.77Rubric Pass Rate (%)100
AutomationBench-AA71.33Guardrail-Adjusted Objective Completion (%)100
AutomationBench-AA - HR - Raw Objective Completion79.6Raw Objective Completion (%)100
AutomationBench-AA - Operations - Raw Objective Completion97.05Raw Objective Completion (%)100
AutomationBench-AA - Strict44.75Strict Task Success (%)100
Bug Hunt Bench57Planted Bugs Fixed (out of 105)100
Bug Hunt Bench - LMS33Planted Bugs Fixed (out of 60)100
Bug Hunt Bench - VS Code Extension24Planted Bugs Fixed (out of 45)100
Harvey LAB-AA - Trusts, Estates & Private Client95.8Criterion Pass Rate (%)100
LLM Stats (BenchCAD (with Python tool))96.3Score (%)100
LLM Stats (BenchCAD)74.7Score (%)100

Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.