Claude Opus 4.6 (Non-reasoning) — benchmark results
Claude Opus 4.6 evaluated with reasoning disabled. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1832 ± 24, rank #113 of 1776 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SEAL - MASK | 96.28 | Score | 100 |
| SEAL - Professional Reasoning Benchmark - Finance | 53.28 | Score | 96.4 |
| SEAL - Professional Reasoning Benchmark - Legal | 52.27 | Score | 92.9 |
| Epoch AI - ECI | 155.38 | ECI Score | 86.7 |
| FrontierMath - Tiers 1-3 | 38.28 | Accuracy (%, 290 problems) | 84.8 |
| LLMEval-Logic Formalization Free | 43.5 | Accuracy (%) | 84.6 |
| LLMEval-Logic Hard | 36 | Accuracy (%) | 84.6 |
| LLMEval-Logic Hard Sub-Q | 74.4 | Accuracy (%) | 84.6 |
| WeirdML | 65.87 | Average Score | 83.2 |
| Sycophancy (Lechmazur) | 2.5 | Sycophancy rate % (lower is better) | 75.8 |
| SEAL - MultiNRC | 48.34 | Score | 72.1 |
| FrontierMath - Tier 4 | 14.58 | Accuracy (%, 48 problems) | 71.8 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.