Claude 3 Sonnet: benchmark results
Provider: Anthropic. Released 2024-03-04. Access: API.
Unified ELO 1429 ± 1, rank #1031 of 1392 rated models, from 403 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CRMArena - PVI | 41.5 | PVI Score (%) | 100 |
| OpenVLM MMMU - Computer Science | 70 | Accuracy (%) | 89.4 |
| OpenVLM MMT-Bench - Interactive Segmentation | 57.1 | Score (%) | 87.6 |
| OpenVLM MMMU - Finance | 66.7 | Accuracy (%) | 84.3 |
| OpenVLM MMMU - Music | 40 | Accuracy (%) | 84.1 |
| CZ-EVAL - Critical Thinking | 68.89 | Accuracy (%) | 81.8 |
| CZ-EVAL - Klokan (Math) | 44.28 | Accuracy (%) | 81.8 |
| OpenVLM MMT-Bench - Counting By Reasoning | 85 | Score (%) | 81.3 |
| OpenVLM LLaVA-Bench - Detail | 79.4 | Score | 78.8 |
| OpenVLM MMMU - Economics | 63.3 | Accuracy (%) | 77.6 |
| SynthPAI | 70.9 | Average accuracy in % | 76.5 |
| OpenVLM MMMU - Clinical Medicine | 63.3 | Accuracy (%) | 76.1 |
Interactive version: theaggregate.ai/model?slug=claude-3-sonnet · How It Works · Data refreshed daily, snapshot 2026-09-05.