Claude Opus 4 (2025-05-14) (Thinking): benchmark results
Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1637 ± 1, rank #310 of 1919 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Capabilities - MMLU-Pro | 87.5 | COT correct | 100 |
| HELM Capabilities - Omni-MATH | 61.6 | Acc | 94 |
| HELM Capabilities - GPQA | 70.85 | COT correct | 92 |
| HELM Capabilities - WildBench | 85.19 | WB Score | 90 |
| HELM Safety - Bbq | 97.1 | BBQ accuracy | 88.9 |
| HELM Safety | 96.94 | Mean score (self-reported) | 87 |
| HELM Safety - Safety Scenarios | 96.94 | Mean score | 86 |
| HELM Safety - Simple Safety Tests | 100 | LM Evaluated Safety score | 84.6 |
| HELM Safety - Harm Bench | 91.69 | LM Evaluated Safety score | 75.5 |
| HELM Capabilities - IFEval | 84.87 | IFEval Strict Acc | 74 |
| HELM Safety - Xstest | 97 | LM Evaluated Safety score | 69.2 |
| HELM Safety - Anthropic Red Team | 98.9 | LM Evaluated Safety score | 49.5 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-2025-05-14-thinking · How It Works · Data refreshed daily, snapshot 2026-09-08.