Claude Opus 4.7 (Medium): benchmark results

Provider: Anthropic. Released 2026-04-16. Access: API.

Unified ELO 1760 ± 21, rank #227 of 2055 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
DexHoldem Agentic Perception34.3Strict problem accuracy (%; every applicable field of the pa100
DexHoldem Agentic Perception - Opponent Chip Inventory43.8Field accuracy (%; exact per-denomination count of the oppon100
DexHoldem Agentic Perception - Community Cards43.6Field accuracy (%; visible community cards as an order-insen85.7
RepoRef63.25Success rate (%; exact match of the submitted GitHub issue, 85.7
DexHoldem Agentic Perception - Current Bet Chips31.2Field accuracy (%; exact per-denomination chip counts of the78.6
DexHoldem Agentic Perception - Turn Ownership93.5Field accuracy (%; whose turn it is, all 36 problems; 36 tab78.6
DexHoldem Agentic Perception - Robot Chip Inventory37.5Field accuracy (%; exact per-denomination count of the robot71.4
ChainSWE (Oracle)64.5Per-bug resolution rate (%; share of the 304 bugs in 100 chr66.7
Equation-Suffix Prediction (Kimi K2.6 Scorer)0.15Likelihood lift (mean clipLL2 per target token over the same62.5
Equation-Suffix Prediction (Qwen3-8B Scorer)0.18Likelihood lift (mean clipLL2 per target token over the same62.5
ChainSWE (Sequential with Memory)39.5Per-bug resolution rate (%; share of the 304 bugs in 100 chr50
DexHoldem Agentic Perception - Field Average49.1Field accuracy (%; unweighted mean of the eight field accura42.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-medium · How It Works · Data refreshed daily, snapshot 2026-09-29.