Claude Opus 5.5 (Medium): benchmark results

Provider: Anthropic. Released 2026-09-22. Access: API.

Unified ELO 1744 ± 1, rank #8 of 2032 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Coding Daily (Claude Code) - Laravel Code Quality18.95Laravel Code Quality (max 20) points, LLM-judged rubric scor100
AI Coding Daily (Claude Code) - React-TS Code Quality19.67React-TS Code Quality (max 20) points, LLM-judged rubric sco100
Maze-Bench85.17Mean Score97.2
DuelLab Overall74.5DuelLab Score95.9
DataBench62.5Score (%)95.2
ARC-AGI-197.5Accuracy (%)95.1
ARC-AGI-287.5Accuracy (%)92.2
Agents on Rails33.3Successful Runs (%)91.7
Bug Hunt Bench - LMS17.7Planted Bugs Fixed (out of 60)87.6
Bug Hunt Bench30.3Planted Bugs Fixed (out of 105)82
AI Coding Daily (Claude Code) - Total57.37Total points (max 60)80
Chess Bench LLM1197Lichess Rating78.3

Interactive version: theaggregate.ai/model?slug=claude-opus-5-5-medium · How It Works · Data refreshed daily, snapshot 2026-09-26.