Claude Opus 4.7: benchmark results
Anthropic's flagship Opus model for complex reasoning, coding, and writing tasks. Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1709 ± 1, rank #44 of 1392 rated models, from 759 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench | 0.8 | Mean JRT z-score | 100 |
| AGC-Bench - Humor | 1.09 | JRT z-score | 100 |
| AGC-Bench - c3_crosstalk | 2.02 | Dataset z-score | 100 |
| AGC-Bench - crowd_vote | 1.52 | Dataset z-score | 100 |
| AGC-Bench - fig_qa | 2.5 | Dataset z-score | 100 |
| AGC-Bench - grapheval_review_advisor | 1.77 | Dataset z-score | 100 |
| AGC-Bench - metaphoric_analogies | 3.98 | Dataset z-score | 100 |
| AcuityBench | 85.3 | QA Exact (self-reported) | 100 |
| AgentCIBench | 13.7 | Leakage (self-reported) | 100 |
| Biomni-Bench (Humanlaya) | 73.34 | Score (0-100) | 100 |
| CVerifBench | 98.3 | Total (self-reported) | 100 |
| CharXiv Reasoning (Anthropic No Tools) | 81.3 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7 · How It Works · Data refreshed daily, snapshot 2026-09-05.