Claude Opus 5.5: benchmark results

Provider: Anthropic. Access: API.

Unified ELO 2158 ± 15, rank #3 of 2928 rated models, from 51 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CursorBench 4.057.8Score (%)100
KernelBench Hub - Mega35.46Best Speedup vs Reference (x)100
LLM Stats (BenchCAD (with Python tool))96.2Score (%)100
LLM Stats (BenchCAD)73Score (%)100
LLM Stats (FrontierCode 1.1)54.4Score (%)100
LLM Stats (HealthBench)60.6Score (%)100
LLM Stats (Humanity's Last Exam (no tools, text-only))64.4Score (%)100
LLM Stats (Humanity's Last Exam (with tools, text-only))67.7Score (%)100
LLM Stats (OSWorld 2.0)81.8Score (%)100
LLM Stats (Program Bench)91.2Score (%)100
LLM Stats (SWE-Bench Multimodal)61.4Score (%)100
LLM Stats (SWE-bench Multilingual)93.9Score (%)100

Interactive version: theaggregate.ai/model?slug=claude-opus-5-5 · How It Works · Data refreshed daily, snapshot 2026-09-23.