Claude 3 Opus — benchmark results
Anthropic's flagship Claude 3 model from March 2024, API-only with a 200K context. Provider: Anthropic. Released 2024-03-04. Access: API.
Unified ELO 1425 ± 11, rank #1105 of 1776 rated models, from 278 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CRM LLM Leaderboard | 67.5 | CRM Accuracy (0-100) | 100 |
| CZ-EVAL - Critical Thinking | 76.45 | Accuracy (%) | 100 |
| CZ-EVAL - Culture | 92.45 | Accuracy (%) | 100 |
| CZ-EVAL - Verbal | 80.18 | Accuracy (%) | 100 |
| Deception Resistance (Lechmazur) | 0.28 | Vulnerability Score (lower is better) | 100 |
| LingOly | 46.3 | Exact Match Accuracy | 100 |
| Vellum - GPQA | 95.4 | Accuracy (%) | 98.4 |
| VNTL Leaderboard | 74.59 | Accuracy (%) | 97.7 |
| InfiBench | 63.89 | Score (%) | 97.1 |
| RepoQA | 90.6 | Score (self-reported) | 96.9 |
| MathBench | 63 | Average Score | 93.9 |
| Compl-AI Board | 84.9 | Average Compliance Score (%) | 92.9 |
Interactive version: theaggregate.ai/model?slug=claude-3-opus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.