Claude Sonnet 4.6 (Non-reasoning): benchmark results
Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1652 ± 1, rank #333 of 3078 rated models, from 28 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BBQ Disambiguated Accuracy (Sonnet 5 & Opus 5 System Cards) | 88.1 | Accuracy (%) | 100 |
| RuleWeaver - Cross-Source - Rule Recall | 59.58 | Gold-rule recall (%) | 90 |
| RuleWeaver - Same-Source - Rule Recall | 71.67 | Gold-rule recall (%) | 90 |
| ComboShoppingBench - Coupon Legality | 96.6 | Pass rate (%) | 78.6 |
| ComboShoppingBench - Coupon-ID Validity | 99.3 | Pass rate (%) | 73.8 |
| RuleWeaver - Cross-Source - Rubric Score | 36.56 | Judge rubric score (0-100) | 70 |
| RuleWeaver - Same-Source - Rubric Score | 43.96 | Judge rubric score (0-100) | 70 |
| ComboShoppingBench - Coupon Optimality | 80.8 | Pass rate (%) | 66.7 |
| o11y-bench - Pass^3 | 58.73 | Tasks passed on all three attempts, Pass^3 (%) | 66.7 |
| Generalization V2 (Lechmazur) | 68.5 | Inverse-Rank Score | 65.5 |
| SteerBench-Work - Pass^5 | 82.1 | Scenarios Correct in All 5 Trials (%) | 63.8 |
| Multi-turn Debate (Lechmazur) | 1551.7 | Bradley-Terry Rating | 62.7 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.