BABE (Biology Arena) - Weak Correlation: leaderboard
Metric: Score (%) on the weak-correlation triplets of independent questions about one source, among the BABE questions (expert-written triplets per peer-reviewed paper or study across 12 biology subfields, reviewed by a second expert panel; about 45% strong-correlation and 55% weak-correlation questions); single inference trial; the paper does not state its grading rule; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 55.16 | #64 |
| 2 | GPT-5.1 (High) | 52.86 | #131 (GPT-5.1) |
| 3 | O3 (High) | 52.05 | #121 (O3) |
| 4 | Gemini 2.5 Pro | 42.33 | #145 |
| 5 | Claude Sonnet 4.5 (Thinking) | 41.91 | #138 (Claude Sonnet 4.5) |
| 6 | GPT-4.1 | 41.34 | #240 |
| 7 | Doubao-Seed-1.6-thinking-250715 | 39.4 | #156 |
| 8 | Claude Opus 4.1 (Thinking) | 38.46 | #132 (Claude Opus 4.1) |
| 9 | Claude Sonnet 4.5 | 29.55 | #138 |
| 10 | GPT-4o (2024-11-20) | 27.22 | #369 |
| 11 | Qwen 3 Max (2025-09-23) | 25.32 | #249 |
| 12 | GLM-4.5V | 23.72 | #339 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=babe-biology-arena-weak-correlation · How It Works · Data refreshed daily, snapshot 2026-10-11.