Claude Sonnet 4.6 (Thinking) — benchmark results

Claude Sonnet 4.6 evaluated with thinking enabled. Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1884 ± 51, rank #73 of 1776 rated models, from 28 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Pencil Puzzle Bench - Nurikabe13.3Direct-ask Success Rate (%)96
Pencil Puzzle Bench - Norinori60Direct-ask Success Rate (%)94
Pencil Puzzle Bench - Hitori33.3Direct-ask Success Rate (%)93
Pencil Puzzle Bench - Nurimisaki20Direct-ask Success Rate (%)91
Wolfram LLM Benchmarking Project61.5Correct Functionality (%)91
Pencil Puzzle Bench - LITS13.3Direct-ask Success Rate (%)90
ProfBench55.6Overall Rubric Score (%)89.4
Pencil Puzzle Bench - Tapa20Direct-ask Success Rate (%)89
Pencil Puzzle Bench - Slitherlink13.3Direct-ask Success Rate (%)88
Pencil Puzzle Bench - Light Up13.3Direct-ask Success Rate (%)83
Pencil Puzzle Bench - Mashu6.7Direct-ask Success Rate (%)82
Pencil Puzzle Bench - Shikaku13.3Direct-ask Success Rate (%)82

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.