Claude Opus 4 — benchmark results
Anthropic Claude Opus 4 model, the flagship Opus-tier Claude 4 row. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1681 ± 16, rank #306 of 1776 rated models, from 134 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FutureSearch DRB - Gather Evidence | 0.4 | Average Score | 100 |
| BenchTable | 81.5 | Total Score (%) | 98 |
| Wordle Arena | 100 | Win Rate (%) | 97.9 |
| RubberDuckBench | 68.53 | Performance (%) | 94.7 |
| Wolfram LLM Benchmarking Project | 61.8 | Correct Functionality (%) | 91.6 |
| CAIA - Pass@1 (With Tools) | 59.6 | Pass@1 (%) | 90.6 |
| MMMU Benchmark | 76.5 | Validation Score | 90.5 |
| Vals AI MGSM | 93.78 | Accuracy (%) | 89.2 |
| GosuEvals | 71.8 | Score (%) | 88.9 |
| AI for Education Pedagogy - Primary | 92.02 | Accuracy (%) | 88 |
| GSMA Open-Telco - TeleLogs | 65 | Score (%) | 87.2 |
| GSMA Open-Telco - 3GPP | 58 | Score (%) | 86.6 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.