Claude Opus 4.7 (Non-reasoning): benchmark results
Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1770 ± 27, rank #222 of 2075 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| o11y-bench - Pass^3 | 79.37 | Tasks passed on all three attempts, Pass^3 (%) | 99 |
| ALE-Bench | 1323.05 | Performance (Self-Refine x1) (self-reported) | 92.2 |
| WeirdML | 76.4 | Average Score | 85.7 |
| o11y-bench - Pass@3 | 87.3 | Tasks passed on at least one of three attempts, Pass@3 (%) | 85.3 |
| AstroAlertBench | 48.87 | End-to-end five-class accuracy (%; 1,500 ZTF alerts, 300 eac | 72.7 |
| TUA-Bench | 49.7 | Success rate (%; mean of the verifier's scalar task reward o | 66.7 |
| DGEval - Regulatory Recall | 17.4 | Accuracy (%) | 61.8 |
| DGEval - IMDG Code MCQ | 65.8 | Accuracy (%) | 38.2 |
| Generalization V2 (Lechmazur) | 52.6 | Inverse-Rank Score | 34.5 |
| Chess Bench LLM | -48 | Lichess Rating | 23.3 |
| NYT Connections Extended | 10.8 | Score (%) | 8.3 |
| ChessBench GitHub - Elo | 1455 | Benchmark Elo | 0 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-04.