Claude Opus 4 (Thinking) — benchmark results

Claude Opus 4 evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1763 ± 27, rank #175 of 1776 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BenchTable82.8Total Score (%)99.1
AA MMLU-Pro87.32Accuracy (%)97.1
Ducky Bench (Saxo Frog)1280ELO95.8
TrackingAI IQ Test (Vision)87.5IQ Test Score (%)95
AA MATH-50098.2Accuracy (%)93.1
Wolfram LLM Benchmarking Project62.4Correct Functionality (%)92.3
Kagi LLM Benchmark74.3Accuracy (%)90
SEAL - MASK87.87Score86.4
TrackingAI IQ Test86.27IQ Test Score (%)80
TrackingAI IQ Test (Offline)75Offline IQ Score (%)77.9
Artificial Analysis Intelligence Index30.97Intelligence Index77.1
AA Terminal-Bench Hard31.06Accuracy (%)73.7

Interactive version: theaggregate.ai/model?slug=claude-opus-4-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.