BenchBench-Protocol: leaderboard
Benchling's wet-lab protocol-modification benchmark: 149 expert-reviewed tasks derived from the differences between a published protocol and the version a scientist actually ran, across 96 source protocols in nine domains of wet-lab biology. Responses are graded against weighted rubrics recovered from the real modification, so the score measures whether a model reasons about prior choices and downstream steps rather than reciting protocol text.
Metric: Normalized Rubric Score (%). Source: www.benchling.com. Status: years away from saturation. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 | 59.2 |
| 2 | GPT-5.6 Sol | 47.1 |
| 3 | Kimi K3 | 45.7 |
| 4 | GPT-5.6 Terra | 45.3 |
| 5 | GPT-5.6 Luna | 41.3 |
| 6 | Inkling | 39.4 |
| 7 | Grok 4.5 | 38.9 |
| 8 | GLM-5.2 | 36.9 |
| 9 | Gemini 3.6 Flash | 34.1 |
Interactive version: theaggregate.ai/benchmark?slug=benchbench-protocol · How It Works · Data refreshed daily, snapshot 2026-09-05.