Claude Opus 4 (Thinking 16K) — benchmark results
Claude Opus 4 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1723 ± 34, rank #235 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Generalization V1 (Lechmazur) | 1.69 | Avg Rank (lower is better) | 98.8 |
| Confabulation Leaderboard (Lechmazur) | 2.48 | Confabulation rate % (lower is better) | 98.4 |
| Chatbot Arena (Text) | 1424 | Elo | 75.4 |
| NYT Connections Older Models | 49.7 | Score (%) | 73.1 |
| ARC-AGI-2 | 8.61 | Accuracy (%) | 57.7 |
| Chatbot Arena (Vision) | 1206 | Arena Score | 56.3 |
| Step Game (Lechmazur) | 2.4 | TrueSkill μ | 52.7 |
| WeirdML | 43.72 | Average Score | 51.1 |
| OTIS Mock AIME 2024-25 | 60 | Accuracy (%) | 48.7 |
| Elimination Game (Lechmazur) | 3.86 | TrueSkill μ | 45.8 |
| ARC-AGI-1 | 35.67 | Accuracy (%) | 37.3 |
| VPCT | 38 | Accuracy (%) | 35.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-thinking-16k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.