Claude Opus 4.6 (Claude Code): benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1779 ± 18, rank #113 of 1605 rated models, from 40 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AgentKernelArena - HIP-to-HIP6.69Mean speedup (x; arithmetic mean over the 24 HIP-to-HIP task100
HarnessAudit-Bench - Task Completion0.82Task completion rate (0-1; weighted hidden completion checkp100
ResearchClawBench21.5Overall (self-reported)100
SkillSafetyBench15.5Attack success rate (%; 155 skill-facing attack cases across100
SkillSafetyBench - Agency Scope and Authorization16Attack success rate (%; 25 cases of actions beyond the inten100
SkillSafetyBench - Context Trust and Prompt Manipulation32Attack success rate (%; 25 cases of misleading contextual si100
SkillSafetyBench - Data Boundary, Output and Externalization11.1Attack success rate (%; 27 cases of corrupted or abused fina100
SkillSafetyBench - Execution, Runtime and Protocol7.7Attack success rate (%; 26 cases of execution redirected thr100
SkillSafetyBench - Knowledge, Model and Supply Chain11.5Attack success rate (%; 26 cases of compromised knowledge so100
SkillSafetyBench - Memory, Recovery, Audit and Persistence15.4Attack success rate (%; 26 cases of unsafe state that persis100
SkillEvolBench (Curated-Revision) - Deployment Success38.9Frozen deployment success rate (%; ESR, share of the 90 cont88.9
SkillEvolBench (No-Skill) - Deployment Success37.8Frozen deployment success rate (%; ESR, share of the 90 cont88.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.