Claude Sonnet 4.6 (Low): benchmark results

Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1642 ± 1, rank #381 of 3078 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FinLifeBench - Financial State - Evidence Recall86.6Recall of gold evidence sessions (%)90
FinLifeBench - Life-Event History - Event-Anchor F172F1 (%; event type with first-establishing session)90
Context Arena70.38Average Score (%)81.4
DeepResearchBench50.4Average Score80
FinLifeBench - Life-Event History - Exact History Match8.3Checkpoints with the exact history (%)80
GIM0.39IRT ability (theta)57.8
o11y-bench - Pass^350.79Tasks passed on all three attempts, Pass^3 (%)49
o11y-bench - Pass@376.19Tasks passed on at least one of three attempts, Pass@3 (%)45.1
Epoch AI - Mystery Game Puzzles16Score41.3
FinLifeBench - Financial State - Checkpoint State Accuracy71.1Cell-level state accuracy (%)40
ObviousBench79.17Answer pass³ (%)33.4
FinLifeBench - Financial State - Granular Change Accuracy37.9GCA@15 (%)20

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-low · How It Works · Data refreshed daily, snapshot 2026-09-19.