Claude Opus 4.6: benchmark results
Anthropic Opus-tier Claude model focused on high-end reasoning, coding, and writing. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1699 ± 1, rank #53 of 1392 rated models, from 616 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - poetmt | 1.55 | Dataset z-score | 100 |
| AGC-Bench - sdat | 0.68 | Dataset z-score | 100 |
| AcademiClaw | 71.9 | Pass Rate (self-reported) | 100 |
| AttuneBench | 54.3 | Composite (self-reported) | 100 |
| AutoLab | 68 | Overall Score (self-reported) | 100 |
| AutoMedBench | 69.69 | Average Overall Score (self-reported) | 100 |
| CAR-bench | 58 | Avg Pass^3 (%) | 100 |
| Claw-Eval-Live | 66.7 | Pass Rate (self-reported) | 100 |
| ClawBench v1 | 61.4 | Reward Rate (%) | 100 |
| ClawForge | 45.3 | Strict Acc. (self-reported) | 100 |
| CodeRouterBench | 43.83 | Mean Task Score (%) | 100 |
| ContractBench | 77.8 | SR% (self-reported) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6 · How It Works · Data refreshed daily, snapshot 2026-09-05.