BIABench: leaderboard

Metric: Outcome score (0-100): mean over 16 end-to-end bioimage-analysis tasks of the per-task mean of three runs, scored against each source study's ground truth (brief instruction, DeepSeek Harness, medium effort). Source: biabench.github.io. Saturation forecast: Around January 2028. 6 models tracked.

Top models

#ModelScore
1Claude Opus 5 (Medium)65.4
2GPT-5.6 Sol (Medium)58.3

Interactive version: theaggregate.ai/benchmark?slug=biabench · How It Works · Data refreshed daily, snapshot 2026-10-01.