PhageBench - Contamination Detection: leaderboard
Metric: Accuracy (%) on 1,200 two-option items deciding whether a phage contig carries host-derived sequence, each a raw-nucleotide multiple-choice question with shuffled options, zero-shot chain-of-thought prompting (reasoning models at reasoning effort medium with at most 2,048 reasoning tokens); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Medium) | 58.83 |
| 2 | GPT-5.2 (Medium) | 54.17 |
| 3 | Qwen 3 Max (Thinking) | 52.58 |
| 4 | Claude Sonnet 4.5 (Thinking) | 51.83 |
| 5 | GPT-4o Mini | 50.08 |
| 6 | GPT-OSS-120B | 46.58 |
Interactive version: theaggregate.ai/benchmark?slug=phagebench-contamination-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.