Benchmarks.bio - SpatialBench-Long: leaderboard

Long-form Benchmarks.bio spatial transcriptomics tasks that require multi-step biological data analysis, tool use, and synthesis over larger assay contexts.

Metric: Pass Rate (%). Source: benchmarks.bio. Status: years away from saturation. 12 models tracked.

Top models

#ModelScore
1GPT-5.511.11
2Claude Opus 4.811.11
3Gemini 3.5 Flash11.11
4Claude Opus 4.69.72
5Claude Opus 4.78.33
6Grok 4.20 Beta (0309) (Reasoning)6.94
7GPT-5.45.56
8Kimi K2.65.56
9Claude Sonnet 4.64.17
10Gemini 3.1 Pro (Preview)4.17
11Grok 4.34.17
12Gemini 2.5 Pro1.39

Interactive version: theaggregate.ai/benchmark?slug=benchmarks-bio-spatialbench-long · How It Works · Data refreshed daily, snapshot 2026-09-05.