Claude Opus 4.1 (20250805) (Thinking): benchmark results

Claude Opus 4.1 (20250805) evaluated with thinking enabled. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1631 ± 1, rank #312 of 1761 rated models, from 36 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Vals AI MGSM94.44Accuracy (%)97.7
UGI - Writing66.51Writing Score96.4
Vals AI MATH 50095.4Accuracy (%)95.8
SEAL - MASK94.2Score95.5
UGI - Natural Intelligence58.17NatInt Score93
Vals AI MMLU-Pro87.92Accuracy (%)83.9
Vals AI MedQA93.59Accuracy (%)77.7
SEAL - VISTA48.44Score74.2
MedCode47.23Score (self-reported)71.6
Vals AI MedCode47.23Accuracy (%)69.7
SEAL Showdown1093Arena Score69.6
Vals AI TaxEval v273.67Accuracy (%)68.1

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-20250805-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.