Claude Opus 4.6 (Claude Code): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1779 ± 18, rank #113 of 1605 rated models, from 40 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AgentKernelArena - HIP-to-HIP | 6.69 | Mean speedup (x; arithmetic mean over the 24 HIP-to-HIP task | 100 |
| HarnessAudit-Bench - Task Completion | 0.82 | Task completion rate (0-1; weighted hidden completion checkp | 100 |
| ResearchClawBench | 21.5 | Overall (self-reported) | 100 |
| SkillSafetyBench | 15.5 | Attack success rate (%; 155 skill-facing attack cases across | 100 |
| SkillSafetyBench - Agency Scope and Authorization | 16 | Attack success rate (%; 25 cases of actions beyond the inten | 100 |
| SkillSafetyBench - Context Trust and Prompt Manipulation | 32 | Attack success rate (%; 25 cases of misleading contextual si | 100 |
| SkillSafetyBench - Data Boundary, Output and Externalization | 11.1 | Attack success rate (%; 27 cases of corrupted or abused fina | 100 |
| SkillSafetyBench - Execution, Runtime and Protocol | 7.7 | Attack success rate (%; 26 cases of execution redirected thr | 100 |
| SkillSafetyBench - Knowledge, Model and Supply Chain | 11.5 | Attack success rate (%; 26 cases of compromised knowledge so | 100 |
| SkillSafetyBench - Memory, Recovery, Audit and Persistence | 15.4 | Attack success rate (%; 26 cases of unsafe state that persis | 100 |
| SkillEvolBench (Curated-Revision) - Deployment Success | 38.9 | Frozen deployment success rate (%; ESR, share of the 90 cont | 88.9 |
| SkillEvolBench (No-Skill) - Deployment Success | 37.8 | Frozen deployment success rate (%; ESR, share of the 90 cont | 88.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.