BioMysteryBench Human-Difficult — leaderboard

Anthropic BioMysteryBench slice covering 23 real-world bioinformatics tasks no human benchmarker solved after QC, evaluated by average accuracy across five trials per problem.

Metric: Accuracy (self-reported). Source: benchmarklist.com. Status: saturation imminent. 7 models tracked.

Top models

#ModelScore
1Claude Mythos 546.1
2Claude Opus 4.840
3Claude Mythos Preview29.6
4Claude Opus 4.724.7
5Claude Opus 4.623.5
6Claude Sonnet 4.619.1
7Claude Haiku 4.55.2

Interactive version: theaggregate.ai/benchmark?slug=biomysterybench-human-difficult · How the rankings work · Data refreshed daily, snapshot 2026-07-22.