Claude Opus 4 (Thinking) — benchmark results
Claude Opus 4 evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1763 ± 27, rank #175 of 1776 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchTable | 82.8 | Total Score (%) | 99.1 |
| AA MMLU-Pro | 87.32 | Accuracy (%) | 97.1 |
| Ducky Bench (Saxo Frog) | 1280 | ELO | 95.8 |
| TrackingAI IQ Test (Vision) | 87.5 | IQ Test Score (%) | 95 |
| AA MATH-500 | 98.2 | Accuracy (%) | 93.1 |
| Wolfram LLM Benchmarking Project | 62.4 | Correct Functionality (%) | 92.3 |
| Kagi LLM Benchmark | 74.3 | Accuracy (%) | 90 |
| SEAL - MASK | 87.87 | Score | 86.4 |
| TrackingAI IQ Test | 86.27 | IQ Test Score (%) | 80 |
| TrackingAI IQ Test (Offline) | 75 | Offline IQ Score (%) | 77.9 |
| Artificial Analysis Intelligence Index | 30.97 | Intelligence Index | 77.1 |
| AA Terminal-Bench Hard | 31.06 | Accuracy (%) | 73.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.