Claude 3 Sonnet — benchmark results
Provider: Anthropic. Released 2024-03-04. Access: API.
Unified ELO 1039 ± 24, rank #1770 of 1776 rated models, from 406 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CRMArena - PVI | 41.5 | PVI Score (%) | 100 |
| InfiBench | 58.2 | Score (%) | 91.4 |
| OpenVLM MMMU - Computer Science | 70 | Accuracy (%) | 89.4 |
| OpenVLM MMT-Bench - Interactive Segmentation | 57.1 | Score (%) | 87.6 |
| OpenVLM MMMU - Finance | 66.7 | Accuracy (%) | 84.3 |
| OpenVLM MMMU - Music | 40 | Accuracy (%) | 84.1 |
| CZ-EVAL - Critical Thinking | 68.89 | Accuracy (%) | 81.8 |
| CZ-EVAL - Klokan (Math) | 44.28 | Accuracy (%) | 81.8 |
| OpenVLM MMT-Bench - Counting By Reasoning | 85 | Score (%) | 81.3 |
| OpenVLM LLaVA-Bench - Detail | 79.4 | Score | 78.8 |
| OpenVLM MMMU - Economics | 63.3 | Accuracy (%) | 77.6 |
| SynthPAI | 70.9 | Average accuracy in % | 76.5 |
Interactive version: theaggregate.ai/model?slug=claude-3-sonnet · How the rankings work · Data refreshed daily, snapshot 2026-07-22.