Claude Opus 4.7 (Non-reasoning): benchmark results

Provider: Anthropic. Released 2026-04-16. Access: API.

Unified ELO 1770 ± 27, rank #222 of 2075 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
o11y-bench - Pass^379.37Tasks passed on all three attempts, Pass^3 (%)99
ALE-Bench1323.05Performance (Self-Refine x1) (self-reported)92.2
WeirdML76.4Average Score85.7
o11y-bench - Pass@387.3Tasks passed on at least one of three attempts, Pass@3 (%)85.3
AstroAlertBench48.87End-to-end five-class accuracy (%; 1,500 ZTF alerts, 300 eac72.7
TUA-Bench49.7Success rate (%; mean of the verifier's scalar task reward o66.7
DGEval - Regulatory Recall17.4Accuracy (%)61.8
DGEval - IMDG Code MCQ65.8Accuracy (%)38.2
Generalization V2 (Lechmazur)52.6Inverse-Rank Score34.5
Chess Bench LLM-48Lichess Rating23.3
NYT Connections Extended10.8Score (%)8.3
ChessBench GitHub - Elo1455Benchmark Elo0

Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-04.