PhageBench - Completeness Estimation: leaderboard

Metric: Accuracy (%) on 1,000 four-option items grading a phage genome as complete, high, medium or low quality, each a raw-nucleotide multiple-choice question with shuffled options, zero-shot chain-of-thought prompting (reasoning models at reasoning effort medium with at most 2,048 reasoning tokens); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.

Top models

#ModelScore
1GPT-5.2 (Medium)43.1
2Claude Sonnet 4.5 (Thinking)35.5
3Qwen 3 Max (Thinking)33.8
4Gemini 3 Flash (Medium)32.1
5GPT-4o Mini24.5
6GPT-OSS-120B24.2

Interactive version: theaggregate.ai/benchmark?slug=phagebench-completeness-estimation · How It Works · Data refreshed daily, snapshot 2026-10-07.