Claude Opus 4.6 (Non-reasoning): benchmark results

Claude Opus 4.6 evaluated with reasoning disabled. Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1671 ± 1, rank #174 of 1761 rated models, from 24 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SEAL - MASK96.28Score100
SEAL - Professional Reasoning Benchmark - Finance53.28Score93.5
SEAL - Professional Reasoning Benchmark - Legal52.27Score90.3
FrontierMath - Tiers 1-338.28Accuracy (%, 290 problems)84.8
LLMEval-Logic Formalization Free43.5Accuracy (%)84.6
LLMEval-Logic Hard36Accuracy (%)84.6
LLMEval-Logic Hard Sub-Q74.4Accuracy (%)84.6
WeirdML65.87Average Score77.7
SEAL - Humanity's Last Exam (Text Only)19.37Score73.3
SEAL - MultiNRC48.34Score72.1
FrontierMath - Tier 414.58Accuracy (%, 48 problems)71.8
Generalization V2 (Lechmazur)68.8Inverse-Rank Score69

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.