Claude Sonnet 4.6 (Medium): benchmark results

Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1660 ± 1, rank #284 of 3078 rated models, from 20 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ALE-Bench1327.3Performance (Self-Refine x1) (self-reported)94.3
BeQu - Experiment 1 - Entailment F143.2Entailment F1 (%)84.2
BeQu - Experiment 1 - Entailment Recall32.4Entailment Recall (%)84.2
NVIDIA ComputeEval - Math Libs89.1Pass@1 (%, zero-shot, 2026.1 release, 101 problems)84.2
BeQu - Experiment 1 - Entailment Precision64.6Entailment Precision (%)81.6
Context Arena69.61Average Score (%)79.1
WeirdML66.07Average Score76.9
NVIDIA ComputeEval - cuDNN17.7Pass@1 (%, zero-shot, 2026.1 release, 113 problems)73.7
NVIDIA ComputeEval - cuBLAS95.1Pass@1 (%, zero-shot, 2026.1 release, 81 problems)71.1
Epoch AI - GPQA Diamond83.33Accuracy (%)68.1
OTIS Mock AIME 2024-2582.22Accuracy (%)65.8
GIM0.84IRT ability (theta)64.4

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-medium · How It Works · Data refreshed daily, snapshot 2026-09-19.