Claude Sonnet 5.5 (Medium): benchmark results

Provider: Anthropic. Released 2026-09-28. Access: API.

Unified ELO 1861 ± 25, rank #77 of 2066 rated models, from 13 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Epoch AI - Critpt16.86Score67.4
Bug Hunt Bench - LMS14.7Planted Bugs Fixed (out of 60)65.7
ChessBench GitHub - Elo2350Benchmark Elo64.1
Bug Hunt Bench23.3Planted Bugs Fixed (out of 105)59.1
CursorBench 4.039.2Score (%)58.3
Argo-Bench25.42Mean task score (0-100) over 210 enterprise data-science tas54.8
Argo-Bench - Solved11.9Share of the 210 tasks solved (%), a task counting as solved46.4
Bug Hunt Bench - VS Code Extension8.7Planted Bugs Fixed (out of 45)41.4
AI Coding Daily (Claude Code) - React-TS Code Quality18.33React-TS Code Quality (max 20) points, LLM-judged rubric sco40
Mercor APEX44.6Best Pass@1 (%)32.7
AI Coding Daily (Claude Code) - Laravel Code Quality17.88Laravel Code Quality (max 20) points, LLM-judged rubric scor20
NonoBench20Overall Accuracy (%)14.7

Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-5-medium · How It Works · Data refreshed daily, snapshot 2026-10-05.