Claude Opus 4 (20250514) (Thinking): benchmark results

Claude Opus 4 (20250514) evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-14. Access: API.

Unified ELO 1614 ± 1, rank #372 of 1761 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Web-Bench25.6Pass@1 (%)97.9
UGI - Writing66.83Writing Score96.6
UGI - Natural Intelligence56.7NatInt Score92.5
WritingBench75.12Score (self-reported)66.7
SEAL - MultiNRC33.93Score51.2
Vals AI LiveCodeBench70.19Accuracy (%)36.6
UGI Leaderboard30.12UGI Score31.4
SEAL - TutorBench49.71Score30.8
UGI - Willingness (W/10)1.2W/10 Score3.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-20250514-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.