BABE (Biology Arena): leaderboard
Metric: Average score (%), the paper's fixed weighting of its strong- and weak-correlation scores (about 51% strong in every row, not the 45/55 question split), over the BABE questions (expert-written triplets per peer-reviewed paper or study across 12 biology subfields, reviewed by a second expert panel; about 45% strong-correlation and 55% weak-correlation questions); single inference trial; the paper does not state its grading rule; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.1 (High) | 52.31 | #131 (GPT-5.1) |
| 2 | Gemini 3 Pro (Preview) | 52.02 | #64 |
| 3 | O3 (High) | 51.62 | #121 (O3) |
| 4 | Gemini 2.5 Pro | 42.17 | #145 |
| 5 | Claude Sonnet 4.5 (Thinking) | 41.93 | #138 (Claude Sonnet 4.5) |
| 6 | Doubao-Seed-1.6-thinking-250715 | 39.54 | #156 |
| 7 | Claude Opus 4.1 (Thinking) | 37.35 | #132 (Claude Opus 4.1) |
| 8 | GPT-4.1 | 36.86 | #240 |
| 9 | Claude Sonnet 4.5 | 31.66 | #138 |
| 10 | Qwen 3 Max (2025-09-23) | 26.54 | #249 |
| 11 | GPT-4o (2024-11-20) | 23.93 | #369 |
| 12 | GLM-4.5V | 20.83 | #339 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=babe-biology-arena · How It Works · Data refreshed daily, snapshot 2026-10-11.