Claude Opus 4.7 (Claude Code): benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1802 ± 16, rank #89 of 1605 rated models, from 39 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BioSecBench-Refusal62.3Red-Team total refusal (%)100
CLI-Advantage Suite (CLI Agents)68.9Success rate (%; 45 task templates x 3 seeds = 135 instances100
LoopsBench - Claude Code Loop - Dependency Depth0.61Dependency depth reached before the first residual handoff (100
LoopsBench - Claude Code Loop - Resolve Rate25Resolve rate (%; with outer-loop continuation)100
LoopsBench - Claude Code Loop - Resolve Rate (Single Segment)16.96Resolve rate (%; without outer-loop continuation)100
LoopsBench - Claude Code Loop - Test Pass Rate53.05Released-test pass rate (%; with outer-loop continuation)100
LoopsBench - Claude Code Loop - Test Pass Rate (Single Segment)41.18Released-test pass rate (%; without outer-loop continuation)100
LoopsBench - Vendor Loops - Dependency Depth0.61Dependency depth reached before the first residual handoff (100
LoopsBench - Vendor Loops - Resolve Rate25Resolve rate (%; with outer-loop continuation)100
LoopsBench - Vendor Loops - Resolve Rate (Single Segment)16.96Resolve rate (%; without outer-loop continuation)100
LoopsBench - Vendor Loops - Test Pass Rate53.05Released-test pass rate (%; with outer-loop continuation)100
LoopsBench - Vendor Loops - Test Pass Rate (Single Segment)41.18Released-test pass rate (%; without outer-loop continuation)100

Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.