Claude Opus 4 (20250514) (Thinking) — benchmark results

Claude Opus 4 (20250514) evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-14. Access: API.

Unified ELO 1711 ± 26, rank #256 of 1776 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Web-Bench25.6Pass@1 (%)97.9
UGI - Writing66.83Writing Score97.3
UGI - Natural Intelligence56.7NatInt Score93.2
WritingBench75.12Score (self-reported)67.3
SEAL - MultiNRC33.93Score51.2
Vals AI LiveCodeBench70.19Accuracy (%)40.9
UGI Leaderboard30.12UGI Score31.7
SEAL - TutorBench49.71Score30.8
UGI - Willingness (W/10)1.2W/10 Score4

Interactive version: theaggregate.ai/model?slug=claude-opus-4-20250514-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.