Claude Opus 4.1 (Thinking) — benchmark results

Claude Opus 4.1 evaluated with thinking enabled. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1746 ± 44, rank #203 of 1776 rated models, from 23 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BenchTable82.7Total Score (%)98.8
AA MMLU-Pro87.99Accuracy (%)98.3
Wolfram LLM Benchmarking Project64.7Correct Functionality (%)94.5
Ducky Bench (Stabby Quack)1399ELO89.3
AA Long Context Reasoning66.33Accuracy (%)85.4
Artificial Analysis Intelligence Index33.71Intelligence Index81.5
AA Terminal-Bench Hard34.34Accuracy (%)79.4
MathVision66Overall Accuracy (%)79.2
AA SciCode40.9Accuracy (%)77.3
AA AIME 202580.33Accuracy (%)75.9
ZeroBench5Score (%)75.4
AA GPQA Diamond80.9Accuracy (%)74.1

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.