BenchBench-Protocol: leaderboard

Benchling's wet-lab protocol-modification benchmark: 149 expert-reviewed tasks derived from the differences between a published protocol and the version a scientist actually ran, across 96 source protocols in nine domains of wet-lab biology. Responses are graded against weighted rubrics recovered from the real modification, so the score measures whether a model reasons about prior choices and downstream steps rather than reciting protocol text.

Metric: Normalized Rubric Score (%). Source: www.benchling.com. Status: years away from saturation. 9 models tracked.

Top models

#ModelScore
1Claude Opus 559.2
2GPT-5.6 Sol47.1
3Kimi K345.7
4GPT-5.6 Terra45.3
5GPT-5.6 Luna41.3
6Inkling39.4
7Grok 4.538.9
8GLM-5.236.9
9Gemini 3.6 Flash34.1

Interactive version: theaggregate.ai/benchmark?slug=benchbench-protocol · How It Works · Data refreshed daily, snapshot 2026-09-05.