Claude Opus 4.6 (Non-reasoning) — benchmark results

Claude Opus 4.6 evaluated with reasoning disabled. Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1832 ± 24, rank #113 of 1776 rated models, from 25 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SEAL - MASK96.28Score100
SEAL - Professional Reasoning Benchmark - Finance53.28Score96.4
SEAL - Professional Reasoning Benchmark - Legal52.27Score92.9
Epoch AI - ECI155.38ECI Score86.7
FrontierMath - Tiers 1-338.28Accuracy (%, 290 problems)84.8
LLMEval-Logic Formalization Free43.5Accuracy (%)84.6
LLMEval-Logic Hard36Accuracy (%)84.6
LLMEval-Logic Hard Sub-Q74.4Accuracy (%)84.6
WeirdML65.87Average Score83.2
Sycophancy (Lechmazur)2.5Sycophancy rate % (lower is better)75.8
SEAL - MultiNRC48.34Score72.1
FrontierMath - Tier 414.58Accuracy (%, 48 problems)71.8

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.