BABE (Biology Arena): leaderboard

Metric: Average score (%), the paper's fixed weighting of its strong- and weak-correlation scores (about 51% strong in every row, not the 45/55 question split), over the BABE questions (expert-written triplets per peer-reviewed paper or study across 12 biology subfields, reviewed by a second expert panel; about 45% strong-correlation and 55% weak-correlation questions); single inference trial; the paper does not state its grading rule; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.1 (High)52.31#131 (GPT-5.1)
2Gemini 3 Pro (Preview)52.02#64
3O3 (High)51.62#121 (O3)
4Gemini 2.5 Pro42.17#145
5Claude Sonnet 4.5 (Thinking)41.93#138 (Claude Sonnet 4.5)
6Doubao-Seed-1.6-thinking-25071539.54#156
7Claude Opus 4.1 (Thinking)37.35#132 (Claude Opus 4.1)
8GPT-4.136.86#240
9Claude Sonnet 4.531.66#138
10Qwen 3 Max (2025-09-23)26.54#249
11GPT-4o (2024-11-20)23.93#369
12GLM-4.5V20.83#339

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=babe-biology-arena · How It Works · Data refreshed daily, snapshot 2026-10-11.