BioMysteryBench Human-Solvable — leaderboard

Anthropic BioMysteryBench slice covering 76 real-world bioinformatics tasks solved by at least one human benchmarker, evaluated by average accuracy across five trials per problem.

Metric: Accuracy (self-reported). Source: benchmarklist.com. Status: saturation imminent. 7 models tracked.

Top models

#ModelScore
1Claude Mythos 583.9
2Claude Mythos Preview82.6
3Claude Opus 4.880.4
4Claude Opus 4.778.9
5Claude Opus 4.677.4
6Claude Sonnet 4.671.8
7Claude Haiku 4.536.8

Interactive version: theaggregate.ai/benchmark?slug=biomysterybench-human-solvable · How the rankings work · Data refreshed daily, snapshot 2026-07-22.