Claude Opus 4: benchmark results

Anthropic Claude Opus 4 model, the flagship Opus-tier Claude 4 row. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1647 ± 1, rank #134 of 1392 rated models, from 165 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FutureSearch DRB - Gather Evidence0.4Average Score100
MIRAGE8Violent completion rate, $C_1$ (%) (self-reported)100
Nejumi 4 - ALT - Truthfulness88.06Score (%)99.2
BenchTable81.5Total Score (%)98
Wordle Arena100Win Rate (%)97.9
RAI-Bench - Refusal Rate (JD)100Rate (%)97.4
Nejumi 4 - GLP - Function Calling70.06Score (%)95.9
RubberDuckBench68.53Performance (%)94.7
Nejumi 4 - GLP - Translation91.37Score (%)91.4
CAIA - Pass@1 (With Tools)59.6Pass@1 (%)90.6
MMMU Benchmark76.5Validation Score90.5
Wolfram LLM Benchmarking Project61.8Correct Functionality (%)90.5

Interactive version: theaggregate.ai/model?slug=claude-opus-4 · How It Works · Data refreshed daily, snapshot 2026-09-05.