Claude Sonnet 4.6 (Thinking) — benchmark results
Claude Sonnet 4.6 evaluated with thinking enabled. Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1884 ± 51, rank #73 of 1776 rated models, from 28 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Pencil Puzzle Bench - Nurikabe | 13.3 | Direct-ask Success Rate (%) | 96 |
| Pencil Puzzle Bench - Norinori | 60 | Direct-ask Success Rate (%) | 94 |
| Pencil Puzzle Bench - Hitori | 33.3 | Direct-ask Success Rate (%) | 93 |
| Pencil Puzzle Bench - Nurimisaki | 20 | Direct-ask Success Rate (%) | 91 |
| Wolfram LLM Benchmarking Project | 61.5 | Correct Functionality (%) | 91 |
| Pencil Puzzle Bench - LITS | 13.3 | Direct-ask Success Rate (%) | 90 |
| ProfBench | 55.6 | Overall Rubric Score (%) | 89.4 |
| Pencil Puzzle Bench - Tapa | 20 | Direct-ask Success Rate (%) | 89 |
| Pencil Puzzle Bench - Slitherlink | 13.3 | Direct-ask Success Rate (%) | 88 |
| Pencil Puzzle Bench - Light Up | 13.3 | Direct-ask Success Rate (%) | 83 |
| Pencil Puzzle Bench - Mashu | 6.7 | Direct-ask Success Rate (%) | 82 |
| Pencil Puzzle Bench - Shikaku | 13.3 | Direct-ask Success Rate (%) | 82 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.