Claude Mythos Preview — benchmark results

Anthropic preview Claude model for high-end reasoning, coding, and general tasks. Provider: Anthropic. Released 2026-04-07. Access: API.

Unified ELO 2103 ± 21, rank #9 of 1776 rated models, from 81 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Anthropic ECI (AECI)159.2AECI Score100
BioMysteryBench Verified - Human Solvable (Anthropic)82.6Score (%)100
BioPipelineBench Verified88.1Accuracy (self-reported)100
BioPipelineBench Verified (Anthropic)88.1Score (%)100
ChartQAPro73.6With tools score (self-reported)100
DeepSearchQA94.4Score (self-reported)100
ExploitBench v8-bench78Mean Capability (%)100
Humanity's Last Exam (Fable/Mythos Tools)64.7Score (%)100
LAB-Bench FigQA (Mythos Preview No Tools)79.7Accuracy (%)100
LAB-Bench FigQA (Mythos Preview Tools)89Accuracy (%)100
LABBench2 - Clinical Trial Questions (Anthropic)86.3Score (%)100
LLM Stats (CharXiv-R)93.2Score (%)100

Interactive version: theaggregate.ai/model?slug=claude-mythos-preview · How the rankings work · Data refreshed daily, snapshot 2026-07-22.