Claude Opus 4.1 (Thinking): benchmark results

Claude Opus 4.1 evaluated with thinking enabled. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1626 ± 1, rank #333 of 1761 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BenchTable82.7Total Score (%)98.8
Wolfram LLM Benchmarking Project64.7Correct Functionality (%)93.3
Ducky Bench (Stabby Quack)1408ELO92.9
AA Long Context Reasoning76Accuracy (%)79.7
MathVision66Overall Accuracy (%)79.7
AA Terminal-Bench Hard34.34Accuracy (%)79.4
Artificial Analysis Intelligence Index26.89Intelligence Index76.3
AA-LCR73.33Accuracy (self-reported)74.2
Ducky Bench (Africa)1150ELO69.2
AA GPQA Diamond80.9Accuracy (%)68.1
AA IFBench55.44Accuracy (%)66.8
AA TAU-2 Bench71.37Accuracy (%)63.8

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.