Claude Sonnet 4.6 (Low): benchmark results
Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1642 ± 1, rank #381 of 3078 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FinLifeBench - Financial State - Evidence Recall | 86.6 | Recall of gold evidence sessions (%) | 90 |
| FinLifeBench - Life-Event History - Event-Anchor F1 | 72 | F1 (%; event type with first-establishing session) | 90 |
| Context Arena | 70.38 | Average Score (%) | 81.4 |
| DeepResearchBench | 50.4 | Average Score | 80 |
| FinLifeBench - Life-Event History - Exact History Match | 8.3 | Checkpoints with the exact history (%) | 80 |
| GIM | 0.39 | IRT ability (theta) | 57.8 |
| o11y-bench - Pass^3 | 50.79 | Tasks passed on all three attempts, Pass^3 (%) | 49 |
| o11y-bench - Pass@3 | 76.19 | Tasks passed on at least one of three attempts, Pass@3 (%) | 45.1 |
| Epoch AI - Mystery Game Puzzles | 16 | Score | 41.3 |
| FinLifeBench - Financial State - Checkpoint State Accuracy | 71.1 | Cell-level state accuracy (%) | 40 |
| ObviousBench | 79.17 | Answer pass³ (%) | 33.4 |
| FinLifeBench - Financial State - Granular Change Accuracy | 37.9 | GCA@15 (%) | 20 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-low · How It Works · Data refreshed daily, snapshot 2026-09-19.