Claude Sonnet 4.6 (Medium): benchmark results
Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1660 ± 1, rank #284 of 3078 rated models, from 20 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ALE-Bench | 1327.3 | Performance (Self-Refine x1) (self-reported) | 94.3 |
| BeQu - Experiment 1 - Entailment F1 | 43.2 | Entailment F1 (%) | 84.2 |
| BeQu - Experiment 1 - Entailment Recall | 32.4 | Entailment Recall (%) | 84.2 |
| NVIDIA ComputeEval - Math Libs | 89.1 | Pass@1 (%, zero-shot, 2026.1 release, 101 problems) | 84.2 |
| BeQu - Experiment 1 - Entailment Precision | 64.6 | Entailment Precision (%) | 81.6 |
| Context Arena | 69.61 | Average Score (%) | 79.1 |
| WeirdML | 66.07 | Average Score | 76.9 |
| NVIDIA ComputeEval - cuDNN | 17.7 | Pass@1 (%, zero-shot, 2026.1 release, 113 problems) | 73.7 |
| NVIDIA ComputeEval - cuBLAS | 95.1 | Pass@1 (%, zero-shot, 2026.1 release, 81 problems) | 71.1 |
| Epoch AI - GPQA Diamond | 83.33 | Accuracy (%) | 68.1 |
| OTIS Mock AIME 2024-25 | 82.22 | Accuracy (%) | 65.8 |
| GIM | 0.84 | IRT ability (theta) | 64.4 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-medium · How It Works · Data refreshed daily, snapshot 2026-09-19.