Claude Mythos Preview: benchmark results

Anthropic preview Claude model for high-end reasoning, coding, and general tasks. Provider: Anthropic. Released 2026-04-07. Access: Private.

Unified ELO 1767 ± 1, rank #13 of 1392 rated models, from 79 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Anthropic ECI (AECI)159.2AECI Score100
BioMysteryBench Verified - Human Solvable (Anthropic)82.6Score (%)100
BioPipelineBench Verified88.1Accuracy (self-reported)100
BioPipelineBench Verified (Anthropic)88.1Score (%)100
ChartQAPro73.6With tools score (self-reported)100
ExploitBench v8-bench78Mean Capability (%)100
LAB-Bench FigQA (Mythos Preview No Tools)79.7Accuracy (%)100
LAB-Bench FigQA (Mythos Preview Tools)89Accuracy (%)100
LABBench2 - Clinical Trial Questions (Anthropic)86.3Score (%)100
LLM Stats (CharXiv-R)93.2Score (%)100
LLM Stats (MMMLU)92.7Score (%)100
LatchBio SpatialBench (Anthropic)53.8Score (%)100

Interactive version: theaggregate.ai/model?slug=claude-mythos-preview · How It Works · Data refreshed daily, snapshot 2026-09-05.