Claude Mythos 5.1: benchmark results

Anthropic September 2026 Mythos model for vetted cybersecurity and life-sciences users. Shares underlying weights with Fable 5.1 but uses different safeguards; benchmark results remain separately attributed. Provider: Anthropic. Released 2026-09-01. Access: Private.

Unified ELO 2163 ± 29, rank #3 of 2656 rated models, from 19 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BBQ Disambiguated Accuracy (Mythos 5.1 System Card)89.9Score (%)100
LLM Stats (Terminal-Bench 4.0)60.9Score (%)100
LatchBio SingleCellBench (Mythos 5.1 System Card)61.9Score (%)100
LatchBio SpatialBench Verified (Mythos 5.1 System Card)77.6Score (%)100
Organic Chemistry V2 (Mythos 5.1 System Card)69.2Score (%)100
Protein Design Library Ranking (Mythos 5.1 System Card)49.3Score (%)100
Protein Design Sequence Generation (Mythos 5.1 System Card)46Score (%)100
ProteinGym Hard49.3Rank correlation (self-reported)100
Protocols Troubleshooting (Mythos 5.1 System Card)70.2Score (%)100
Vellum - Humanity's Last Exam65Accuracy (%)98.5
LLM Stats Score45.06LLM Stats Score (conservative rating)91.2
AA-Omniscience Net Score57Net score (self-reported)88.9

Interactive version: theaggregate.ai/model?slug=claude-mythos-5-1 · How It Works · Data refreshed daily, snapshot 2026-09-19.