Claude Opus 4: benchmark results
Anthropic Claude Opus 4 model, the flagship Opus-tier Claude 4 row. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1647 ± 1, rank #134 of 1392 rated models, from 165 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FutureSearch DRB - Gather Evidence | 0.4 | Average Score | 100 |
| MIRAGE | 8 | Violent completion rate, $C_1$ (%) (self-reported) | 100 |
| Nejumi 4 - ALT - Truthfulness | 88.06 | Score (%) | 99.2 |
| BenchTable | 81.5 | Total Score (%) | 98 |
| Wordle Arena | 100 | Win Rate (%) | 97.9 |
| RAI-Bench - Refusal Rate (JD) | 100 | Rate (%) | 97.4 |
| Nejumi 4 - GLP - Function Calling | 70.06 | Score (%) | 95.9 |
| RubberDuckBench | 68.53 | Performance (%) | 94.7 |
| Nejumi 4 - GLP - Translation | 91.37 | Score (%) | 91.4 |
| CAIA - Pass@1 (With Tools) | 59.6 | Pass@1 (%) | 90.6 |
| MMMU Benchmark | 76.5 | Validation Score | 90.5 |
| Wolfram LLM Benchmarking Project | 61.8 | Correct Functionality (%) | 90.5 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4 · How It Works · Data refreshed daily, snapshot 2026-09-05.