Claude Opus 4.7 (Low): benchmark results

Provider: Anthropic. Released 2026-04-16. Access: API.

Unified ELO 1743 ± 25, rank #259 of 2066 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
TeachObs - Scene Coding (Transcript)27.7Macro F1 (%; 39 observation codes per 15-second classroom sc100
o11y-bench - Pass@385.71Tasks passed on at least one of three attempts, Pass@3 (%)79.4
o11y-bench - Pass^363.49Tasks passed on all three attempts, Pass^3 (%)79.4
TeachObs - Lesson Narrative Coverage29.5Reference coverage (%; share of the experts' atomic claims a75
TeachObs - Scene Coding (Transcript + Frame)43.8Macro F1 (%; 39 observation codes per 15-second classroom sc75
Chess Puzzles (Epoch AI)20Accuracy (%)59.5
DGEval - IMDG Code MCQ68.3Accuracy (%)58.8
DGEval - Regulatory Recall16.2Accuracy (%)55.9
Equation-Suffix Prediction (Kimi K2.6 Scorer)0.15Likelihood lift (mean clipLL2 per target token over the same50
Equation-Suffix Prediction (Qwen3-8B Scorer)0.18Likelihood lift (mean clipLL2 per target token over the same50
ObviousBench84.03Answer pass³ (%)37.8
Context Arena28.63Average Score (%)19.2

Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-low · How It Works · Data refreshed daily, snapshot 2026-10-05.