Claude Opus 4.8 (Claude Code): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1891 ± 16, rank #30 of 1605 rated models, from 40 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DrawAI-Bench - Editability | 91.6 | Editability score (0-100; rules and VLM rubric, strongest ob | 100 |
| DrawAI-Bench - Editability - Formula | 100 | Editability score (0-100; rules and VLM rubric, strongest ob | 100 |
| DrawAI-Bench - Editability - Text | 88.7 | Editability score (0-100; rules and VLM rubric, strongest ob | 100 |
| GameReplica | 71.6 | Overall replication score (%; equal-weight mean of rule cons | 100 |
| GameReplica - Implementation Consistency | 67.4 | Implementation consistency (%; share of reference rules the | 100 |
| GameReplica - Rule Consistency | 54.2 | Rule consistency (%; item-by-item agreement of the rule docu | 100 |
| GameReplica - Visual Fidelity | 93.1 | Visual fidelity (%; screenshot comparison of original and re | 100 |
| Legal Hallucination Detection Benchmark (no ex | 68.8 | F1 (%) - Agentic (self-reported) | 100 |
| BioSecBench-Refusal | 53.6 | Red-Team total refusal (%) | 93.3 |
| Vibe Code Bench v1.1 | 77.49 | Score (%) | 91.9 |
| DrawAI-Bench - Editability - Connector | 97.1 | Editability score (0-100; rules and VLM rubric, strongest ob | 91.7 |
| DrawAI-Bench - Editability - Shape | 91.6 | Editability score (0-100; rules and VLM rubric, strongest ob | 91.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.