Claude Opus 4 (20250514) (Thinking) — benchmark results
Claude Opus 4 (20250514) evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-14. Access: API.
Unified ELO 1711 ± 26, rank #256 of 1776 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Web-Bench | 25.6 | Pass@1 (%) | 97.9 |
| UGI - Writing | 66.83 | Writing Score | 97.3 |
| UGI - Natural Intelligence | 56.7 | NatInt Score | 93.2 |
| WritingBench | 75.12 | Score (self-reported) | 67.3 |
| SEAL - MultiNRC | 33.93 | Score | 51.2 |
| Vals AI LiveCodeBench | 70.19 | Accuracy (%) | 40.9 |
| UGI Leaderboard | 30.12 | UGI Score | 31.7 |
| SEAL - TutorBench | 49.71 | Score | 30.8 |
| UGI - Willingness (W/10) | 1.2 | W/10 Score | 4 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-20250514-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.