Claude Mythos 5.1: benchmark results
Anthropic September 2026 Mythos model for vetted cybersecurity and life-sciences users. Shares underlying weights with Fable 5.1 but uses different safeguards; benchmark results remain separately attributed. Provider: Anthropic. Released 2026-09-01. Access: Private.
Unified ELO 2163 ± 29, rank #3 of 2656 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BBQ Disambiguated Accuracy (Mythos 5.1 System Card) | 89.9 | Score (%) | 100 |
| LLM Stats (Terminal-Bench 4.0) | 60.9 | Score (%) | 100 |
| LatchBio SingleCellBench (Mythos 5.1 System Card) | 61.9 | Score (%) | 100 |
| LatchBio SpatialBench Verified (Mythos 5.1 System Card) | 77.6 | Score (%) | 100 |
| Organic Chemistry V2 (Mythos 5.1 System Card) | 69.2 | Score (%) | 100 |
| Protein Design Library Ranking (Mythos 5.1 System Card) | 49.3 | Score (%) | 100 |
| Protein Design Sequence Generation (Mythos 5.1 System Card) | 46 | Score (%) | 100 |
| ProteinGym Hard | 49.3 | Rank correlation (self-reported) | 100 |
| Protocols Troubleshooting (Mythos 5.1 System Card) | 70.2 | Score (%) | 100 |
| Vellum - Humanity's Last Exam | 65 | Accuracy (%) | 98.5 |
| LLM Stats Score | 45.06 | LLM Stats Score (conservative rating) | 91.2 |
| AA-Omniscience Net Score | 57 | Net score (self-reported) | 88.9 |
Interactive version: theaggregate.ai/model?slug=claude-mythos-5-1 · How It Works · Data refreshed daily, snapshot 2026-09-19.