PhageBench: leaderboard

Metric: Accuracy (%) averaged over the five PhageBench tasks (phage contig identification, contamination detection, completeness estimation, lifestyle classification, host prediction), each item a raw-nucleotide multiple-choice question with shuffled options, zero-shot chain-of-thought prompting (reasoning models at reasoning effort medium with at most 2,048 reasoning tokens); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 8 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Medium)56.33
2GPT-5.2 (Medium)50.95
3Claude Sonnet 4.5 (Thinking)47.88
4Qwen 3 Max (Thinking)45.69
5GPT-OSS-120B42.27
6GPT-4o Mini40.15

Interactive version: theaggregate.ai/benchmark?slug=phagebench · How It Works · Data refreshed daily, snapshot 2026-10-07.