Claude Opus 4.1 (20250805) (Thinking) — benchmark results

Claude Opus 4.1 (20250805) evaluated with thinking enabled. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1744 ± 17, rank #207 of 1776 rated models, from 34 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Writing66.51Writing Score97
Vals AI MGSM94.44Accuracy (%)97
SEAL - MASK94.2Score95.5
Vals AI MATH 50095.4Accuracy (%)94.1
UGI - Natural Intelligence58.17NatInt Score93.7
Vals AI MMLU-Pro87.92Accuracy (%)88.4
Vals AI MedQA93.59Accuracy (%)77.7
SEAL - VISTA48.44Score74.2
Vals AI MedCode47.23Accuracy (%)72.6
SEAL Showdown1099.6Arena Score71.7
Vals AI TaxEval v273.67Accuracy (%)71.1
SEAL - EnigmaEval7.18Score69.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-20250805-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.