Claude Sonnet 4.6 (Non-reasoning): benchmark results

Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1652 ± 1, rank #333 of 3078 rated models, from 28 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BBQ Disambiguated Accuracy (Sonnet 5 & Opus 5 System Cards)88.1Accuracy (%)100
RuleWeaver - Cross-Source - Rule Recall59.58Gold-rule recall (%)90
RuleWeaver - Same-Source - Rule Recall71.67Gold-rule recall (%)90
ComboShoppingBench - Coupon Legality96.6Pass rate (%)78.6
ComboShoppingBench - Coupon-ID Validity99.3Pass rate (%)73.8
RuleWeaver - Cross-Source - Rubric Score36.56Judge rubric score (0-100)70
RuleWeaver - Same-Source - Rubric Score43.96Judge rubric score (0-100)70
ComboShoppingBench - Coupon Optimality80.8Pass rate (%)66.7
o11y-bench - Pass^358.73Tasks passed on all three attempts, Pass^3 (%)66.7
Generalization V2 (Lechmazur)68.5Inverse-Rank Score65.5
SteerBench-Work - Pass^582.1Scenarios Correct in All 5 Trials (%)63.8
Multi-turn Debate (Lechmazur)1551.7Bradley-Terry Rating62.7

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.