Claude Sonnet 4.6 (Thinking): benchmark results

Claude Sonnet 4.6 evaluated with thinking enabled. Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1661 ± 1, rank #211 of 1761 rated models, from 29 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Pencil Puzzle Bench - Nurikabe13.3Direct-ask Success Rate (%)96
Pencil Puzzle Bench - Norinori60Direct-ask Success Rate (%)94
Pencil Puzzle Bench - Hitori33.3Direct-ask Success Rate (%)93
Pencil Puzzle Bench - Nurimisaki20Direct-ask Success Rate (%)91
Pencil Puzzle Bench - LITS13.3Direct-ask Success Rate (%)90
Wolfram LLM Benchmarking Project61.5Correct Functionality (%)90
Pencil Puzzle Bench - Tapa20Direct-ask Success Rate (%)89
Pencil Puzzle Bench - Slitherlink13.3Direct-ask Success Rate (%)88
ProfBench55.6Overall Rubric Score (%)86.1
Pencil Puzzle Bench - Light Up13.3Direct-ask Success Rate (%)83
Pencil Puzzle Bench - Mashu6.7Direct-ask Success Rate (%)82
Pencil Puzzle Bench - Shikaku13.3Direct-ask Success Rate (%)82

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.