BABE (Biology Arena) - Strong Correlation: leaderboard

Metric: Score (%) on the strong-correlation triplets, where each question's answer feeds the next, among the BABE questions (expert-written triplets per peer-reviewed paper or study across 12 biology subfields, reviewed by a second expert panel; about 45% strong-correlation and 55% weak-correlation questions); single inference trial; the paper does not state its grading rule; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.1 (High)51.79#131 (GPT-5.1)
2O3 (High)51.22#121 (O3)
3Gemini 3 Pro (Preview)49.05#64
4Gemini 2.5 Pro42.02#145
5Claude Sonnet 4.5 (Thinking)41.94#138 (Claude Sonnet 4.5)
6Doubao-Seed-1.6-thinking-25071539.67#156
7Claude Opus 4.1 (Thinking)36.3#132 (Claude Opus 4.1)
8Claude Sonnet 4.533.67#138
9GPT-4.132.61#240
10Qwen 3 Max (2025-09-23)27.7#249
11GPT-4o (2024-11-20)20.8#369
12GLM-4.5V18.09#339

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=babe-biology-arena-strong-correlation · How It Works · Data refreshed daily, snapshot 2026-10-11.