PhageBench: leaderboard
Metric: Accuracy (%) averaged over the five PhageBench tasks (phage contig identification, contamination detection, completeness estimation, lifestyle classification, host prediction), each item a raw-nucleotide multiple-choice question with shuffled options, zero-shot chain-of-thought prompting (reasoning models at reasoning effort medium with at most 2,048 reasoning tokens); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Medium) | 56.33 |
| 2 | GPT-5.2 (Medium) | 50.95 |
| 3 | Claude Sonnet 4.5 (Thinking) | 47.88 |
| 4 | Qwen 3 Max (Thinking) | 45.69 |
| 5 | GPT-OSS-120B | 42.27 |
| 6 | GPT-4o Mini | 40.15 |
Interactive version: theaggregate.ai/benchmark?slug=phagebench · How It Works · Data refreshed daily, snapshot 2026-10-07.