Claude Opus 4 — benchmark results

Anthropic Claude Opus 4 model, the flagship Opus-tier Claude 4 row. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1681 ± 16, rank #306 of 1776 rated models, from 134 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FutureSearch DRB - Gather Evidence0.4Average Score100
BenchTable81.5Total Score (%)98
Wordle Arena100Win Rate (%)97.9
RubberDuckBench68.53Performance (%)94.7
Wolfram LLM Benchmarking Project61.8Correct Functionality (%)91.6
CAIA - Pass@1 (With Tools)59.6Pass@1 (%)90.6
MMMU Benchmark76.5Validation Score90.5
Vals AI MGSM93.78Accuracy (%)89.2
GosuEvals71.8Score (%)88.9
AI for Education Pedagogy - Primary92.02Accuracy (%)88
GSMA Open-Telco - TeleLogs65Score (%)87.2
GSMA Open-Telco - 3GPP58Score (%)86.6

Interactive version: theaggregate.ai/model?slug=claude-opus-4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.