Claude Opus 4.7 (Low): benchmark results
Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1743 ± 25, rank #259 of 2066 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| TeachObs - Scene Coding (Transcript) | 27.7 | Macro F1 (%; 39 observation codes per 15-second classroom sc | 100 |
| o11y-bench - Pass@3 | 85.71 | Tasks passed on at least one of three attempts, Pass@3 (%) | 79.4 |
| o11y-bench - Pass^3 | 63.49 | Tasks passed on all three attempts, Pass^3 (%) | 79.4 |
| TeachObs - Lesson Narrative Coverage | 29.5 | Reference coverage (%; share of the experts' atomic claims a | 75 |
| TeachObs - Scene Coding (Transcript + Frame) | 43.8 | Macro F1 (%; 39 observation codes per 15-second classroom sc | 75 |
| Chess Puzzles (Epoch AI) | 20 | Accuracy (%) | 59.5 |
| DGEval - IMDG Code MCQ | 68.3 | Accuracy (%) | 58.8 |
| DGEval - Regulatory Recall | 16.2 | Accuracy (%) | 55.9 |
| Equation-Suffix Prediction (Kimi K2.6 Scorer) | 0.15 | Likelihood lift (mean clipLL2 per target token over the same | 50 |
| Equation-Suffix Prediction (Qwen3-8B Scorer) | 0.18 | Likelihood lift (mean clipLL2 per target token over the same | 50 |
| ObviousBench | 84.03 | Answer pass³ (%) | 37.8 |
| Context Arena | 28.63 | Average Score (%) | 19.2 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-low · How It Works · Data refreshed daily, snapshot 2026-10-05.