Claude Opus 4.8 (Claude Code): benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1891 ± 16, rank #30 of 1605 rated models, from 40 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
DrawAI-Bench - Editability91.6Editability score (0-100; rules and VLM rubric, strongest ob100
DrawAI-Bench - Editability - Formula100Editability score (0-100; rules and VLM rubric, strongest ob100
DrawAI-Bench - Editability - Text88.7Editability score (0-100; rules and VLM rubric, strongest ob100
GameReplica71.6Overall replication score (%; equal-weight mean of rule cons100
GameReplica - Implementation Consistency67.4Implementation consistency (%; share of reference rules the 100
GameReplica - Rule Consistency54.2Rule consistency (%; item-by-item agreement of the rule docu100
GameReplica - Visual Fidelity93.1Visual fidelity (%; screenshot comparison of original and re100
Legal Hallucination Detection Benchmark (no ex68.8F1 (%) - Agentic (self-reported)100
BioSecBench-Refusal53.6Red-Team total refusal (%)93.3
Vibe Code Bench v1.177.49Score (%)91.9
DrawAI-Bench - Editability - Connector97.1Editability score (0-100; rules and VLM rubric, strongest ob91.7
DrawAI-Bench - Editability - Shape91.6Editability score (0-100; rules and VLM rubric, strongest ob91.7

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.