Claude Opus 4.7 (Medium): benchmark results
Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1760 ± 21, rank #227 of 2055 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DexHoldem Agentic Perception | 34.3 | Strict problem accuracy (%; every applicable field of the pa | 100 |
| DexHoldem Agentic Perception - Opponent Chip Inventory | 43.8 | Field accuracy (%; exact per-denomination count of the oppon | 100 |
| DexHoldem Agentic Perception - Community Cards | 43.6 | Field accuracy (%; visible community cards as an order-insen | 85.7 |
| RepoRef | 63.25 | Success rate (%; exact match of the submitted GitHub issue, | 85.7 |
| DexHoldem Agentic Perception - Current Bet Chips | 31.2 | Field accuracy (%; exact per-denomination chip counts of the | 78.6 |
| DexHoldem Agentic Perception - Turn Ownership | 93.5 | Field accuracy (%; whose turn it is, all 36 problems; 36 tab | 78.6 |
| DexHoldem Agentic Perception - Robot Chip Inventory | 37.5 | Field accuracy (%; exact per-denomination count of the robot | 71.4 |
| ChainSWE (Oracle) | 64.5 | Per-bug resolution rate (%; share of the 304 bugs in 100 chr | 66.7 |
| Equation-Suffix Prediction (Kimi K2.6 Scorer) | 0.15 | Likelihood lift (mean clipLL2 per target token over the same | 62.5 |
| Equation-Suffix Prediction (Qwen3-8B Scorer) | 0.18 | Likelihood lift (mean clipLL2 per target token over the same | 62.5 |
| ChainSWE (Sequential with Memory) | 39.5 | Per-bug resolution rate (%; share of the 304 bugs in 100 chr | 50 |
| DexHoldem Agentic Perception - Field Average | 49.1 | Field accuracy (%; unweighted mean of the eight field accura | 42.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-medium · How It Works · Data refreshed daily, snapshot 2026-09-29.