Claude Mythos Preview — benchmark results
Anthropic preview Claude model for high-end reasoning, coding, and general tasks. Provider: Anthropic. Released 2026-04-07. Access: API.
Unified ELO 2103 ± 21, rank #9 of 1776 rated models, from 81 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Anthropic ECI (AECI) | 159.2 | AECI Score | 100 |
| BioMysteryBench Verified - Human Solvable (Anthropic) | 82.6 | Score (%) | 100 |
| BioPipelineBench Verified | 88.1 | Accuracy (self-reported) | 100 |
| BioPipelineBench Verified (Anthropic) | 88.1 | Score (%) | 100 |
| ChartQAPro | 73.6 | With tools score (self-reported) | 100 |
| DeepSearchQA | 94.4 | Score (self-reported) | 100 |
| ExploitBench v8-bench | 78 | Mean Capability (%) | 100 |
| Humanity's Last Exam (Fable/Mythos Tools) | 64.7 | Score (%) | 100 |
| LAB-Bench FigQA (Mythos Preview No Tools) | 79.7 | Accuracy (%) | 100 |
| LAB-Bench FigQA (Mythos Preview Tools) | 89 | Accuracy (%) | 100 |
| LABBench2 - Clinical Trial Questions (Anthropic) | 86.3 | Score (%) | 100 |
| LLM Stats (CharXiv-R) | 93.2 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-mythos-preview · How the rankings work · Data refreshed daily, snapshot 2026-07-22.