Claude Opus 4.6 (Non-reasoning): benchmark results
Claude Opus 4.6 evaluated with reasoning disabled. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1671 ± 1, rank #174 of 1761 rated models, from 24 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SEAL - MASK | 96.28 | Score | 100 |
| SEAL - Professional Reasoning Benchmark - Finance | 53.28 | Score | 93.5 |
| SEAL - Professional Reasoning Benchmark - Legal | 52.27 | Score | 90.3 |
| FrontierMath - Tiers 1-3 | 38.28 | Accuracy (%, 290 problems) | 84.8 |
| LLMEval-Logic Formalization Free | 43.5 | Accuracy (%) | 84.6 |
| LLMEval-Logic Hard | 36 | Accuracy (%) | 84.6 |
| LLMEval-Logic Hard Sub-Q | 74.4 | Accuracy (%) | 84.6 |
| WeirdML | 65.87 | Average Score | 77.7 |
| SEAL - Humanity's Last Exam (Text Only) | 19.37 | Score | 73.3 |
| SEAL - MultiNRC | 48.34 | Score | 72.1 |
| FrontierMath - Tier 4 | 14.58 | Accuracy (%, 48 problems) | 71.8 |
| Generalization V2 (Lechmazur) | 68.8 | Inverse-Rank Score | 69 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.