BIABench: leaderboard
Metric: Outcome score (0-100): mean over 16 end-to-end bioimage-analysis tasks of the per-task mean of three runs, scored against each source study's ground truth (brief instruction, DeepSeek Harness, medium effort). Source: biabench.github.io. Saturation forecast: Around January 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (Medium) | 65.4 |
| 2 | GPT-5.6 Sol (Medium) | 58.3 |
Interactive version: theaggregate.ai/benchmark?slug=biabench · How It Works · Data refreshed daily, snapshot 2026-10-01.