Claude 3 Opus: benchmark results
Anthropic's flagship Claude 3 model from March 2024, API-only with a 200K context. Provider: Anthropic. Released 2024-03-04. Access: API.
Unified ELO 1505 ± 1, rank #666 of 1392 rated models, from 263 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CRM LLM Leaderboard | 67.5 | CRM Accuracy (0-100) | 100 |
| CZ-EVAL - Critical Thinking | 76.45 | Accuracy (%) | 100 |
| CZ-EVAL - Culture | 92.45 | Accuracy (%) | 100 |
| CZ-EVAL - Verbal | 80.18 | Accuracy (%) | 100 |
| Deception Resistance (Lechmazur) | 0.28 | Vulnerability Score (lower is better) | 100 |
| LingOly | 46.3 | Exact Match Accuracy | 100 |
| VNTL Leaderboard | 74.59 | Accuracy (%) | 97.7 |
| Vellum - GPQA | 95.4 | Accuracy (%) | 97.1 |
| MathBench | 63 | Average Score | 93.9 |
| Compl-AI Board | 84.9 | Average Compliance Score (%) | 92.9 |
| Judge Arena | 1312 | ELO Score | 92.6 |
| MixEval | 63.5 | Score | 92.2 |
Interactive version: theaggregate.ai/model?slug=claude-3-opus · How It Works · Data refreshed daily, snapshot 2026-09-05.