Benchmarks.bio - TxBench-PP — leaderboard

Benchmarks.bio TxBench-PP evaluates agentic perturbation/transcriptomics analysis workflows with deterministic grading over realistic biological data-analysis tasks.

Metric: Pass Rate (%). Source: benchmarks.bio. Status: saturation imminent. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.859.33
2GPT-5.555.33
3Gemini 3.5 Flash51.33
4GPT-5.449.67
5Claude Opus 4.749.33
6Claude Opus 4.644.67
7Gemini 3.1 Pro (Preview)40
8Claude Sonnet 4.636
9Kimi K2.629.67
10Grok 4.20 0309 (Reasoning)19.67
11Grok 4.318.33

Interactive version: theaggregate.ai/benchmark?slug=benchmarks-bio-txbench-pp · How the rankings work · Data refreshed daily, snapshot 2026-07-22.