Claude Opus 4 (Thinking): benchmark results

Claude Opus 4 evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1630 ± 1, rank #318 of 1761 rated models, from 31 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BenchTable82.8Total Score (%)99.1
Ducky Bench (Saxo Frog)1280ELO95.8
Wolfram LLM Benchmarking Project62.4Correct Functionality (%)91.2
Kagi LLM Benchmark74.3Accuracy (%)90.1
SEAL - MASK87.87Score86.4
ZEROBench-Sub25.06Score (self-reported)75
AA Terminal-Bench Hard31.06Accuracy (%)73.7
TrackingAI IQ Test (Offline)75Offline IQ Score (%)72.4
Artificial Analysis Intelligence Index24.28Intelligence Index71.5
Ducky Bench (Stabby Quack)1197ELO71.4
REAL Evals19.7Task Completion (%)67.9
SEAL - VISTA46.96Score67.7

Interactive version: theaggregate.ai/model?slug=claude-opus-4-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.