Claude Sonnet 5.5 (Medium): benchmark results
Provider: Anthropic. Released 2026-09-28. Access: API.
Unified ELO 1861 ± 25, rank #77 of 2066 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Epoch AI - Critpt | 16.86 | Score | 67.4 |
| Bug Hunt Bench - LMS | 14.7 | Planted Bugs Fixed (out of 60) | 65.7 |
| ChessBench GitHub - Elo | 2350 | Benchmark Elo | 64.1 |
| Bug Hunt Bench | 23.3 | Planted Bugs Fixed (out of 105) | 59.1 |
| CursorBench 4.0 | 39.2 | Score (%) | 58.3 |
| Argo-Bench | 25.42 | Mean task score (0-100) over 210 enterprise data-science tas | 54.8 |
| Argo-Bench - Solved | 11.9 | Share of the 210 tasks solved (%), a task counting as solved | 46.4 |
| Bug Hunt Bench - VS Code Extension | 8.7 | Planted Bugs Fixed (out of 45) | 41.4 |
| AI Coding Daily (Claude Code) - React-TS Code Quality | 18.33 | React-TS Code Quality (max 20) points, LLM-judged rubric sco | 40 |
| Mercor APEX | 44.6 | Best Pass@1 (%) | 32.7 |
| AI Coding Daily (Claude Code) - Laravel Code Quality | 17.88 | Laravel Code Quality (max 20) points, LLM-judged rubric scor | 20 |
| NonoBench | 20 | Overall Accuracy (%) | 14.7 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-5-medium · How It Works · Data refreshed daily, snapshot 2026-10-05.