BulkPR-Bench: leaderboard

Metric: Relational Delivery Score (%; 581 newly authored candidate pull requests on frozen snapshots of 18 real repositories, released in batches of K=32 under the buffered protocol (up to 4 candidates deferred for up to 16 batches); credit for safe delivery in executable order and correct rejection within each gold relation group of the realized merge trace, repository-weighted, three runs per repository; every model runs in the same Claude Code v2.1.138 scaffold with a 150-turn budget). Source: arxiv.org. Saturation forecast: Around May 2027. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.866.6
2GPT-5.462
3GLM-5.257.9
4Kimi K349.4
5Qwen 3.7 Max47.2
6DeepSeek V4 Pro40.9

Interactive version: theaggregate.ai/benchmark?slug=bulkpr-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.