Claude Opus 5.5 (Medium): benchmark results
Provider: Anthropic. Released 2026-09-22. Access: API.
Unified ELO 1744 ± 1, rank #8 of 2032 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Coding Daily (Claude Code) - Laravel Code Quality | 18.95 | Laravel Code Quality (max 20) points, LLM-judged rubric scor | 100 |
| AI Coding Daily (Claude Code) - React-TS Code Quality | 19.67 | React-TS Code Quality (max 20) points, LLM-judged rubric sco | 100 |
| Maze-Bench | 85.17 | Mean Score | 97.2 |
| DuelLab Overall | 74.5 | DuelLab Score | 95.9 |
| DataBench | 62.5 | Score (%) | 95.2 |
| ARC-AGI-1 | 97.5 | Accuracy (%) | 95.1 |
| ARC-AGI-2 | 87.5 | Accuracy (%) | 92.2 |
| Agents on Rails | 33.3 | Successful Runs (%) | 91.7 |
| Bug Hunt Bench - LMS | 17.7 | Planted Bugs Fixed (out of 60) | 87.6 |
| Bug Hunt Bench | 30.3 | Planted Bugs Fixed (out of 105) | 82 |
| AI Coding Daily (Claude Code) - Total | 57.37 | Total points (max 60) | 80 |
| Chess Bench LLM | 1197 | Lichess Rating | 78.3 |
Interactive version: theaggregate.ai/model?slug=claude-opus-5-5-medium · How It Works · Data refreshed daily, snapshot 2026-09-26.