Claude Opus 5.5: benchmark results
Provider: Anthropic. Access: API.
Unified ELO 2158 ± 15, rank #3 of 2928 rated models, from 51 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CursorBench 4.0 | 57.8 | Score (%) | 100 |
| KernelBench Hub - Mega | 35.46 | Best Speedup vs Reference (x) | 100 |
| LLM Stats (BenchCAD (with Python tool)) | 96.2 | Score (%) | 100 |
| LLM Stats (BenchCAD) | 73 | Score (%) | 100 |
| LLM Stats (FrontierCode 1.1) | 54.4 | Score (%) | 100 |
| LLM Stats (HealthBench) | 60.6 | Score (%) | 100 |
| LLM Stats (Humanity's Last Exam (no tools, text-only)) | 64.4 | Score (%) | 100 |
| LLM Stats (Humanity's Last Exam (with tools, text-only)) | 67.7 | Score (%) | 100 |
| LLM Stats (OSWorld 2.0) | 81.8 | Score (%) | 100 |
| LLM Stats (Program Bench) | 91.2 | Score (%) | 100 |
| LLM Stats (SWE-Bench Multimodal) | 61.4 | Score (%) | 100 |
| LLM Stats (SWE-bench Multilingual) | 93.9 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-opus-5-5 · How It Works · Data refreshed daily, snapshot 2026-09-23.