BulkPR-Bench - Critical Relation Recall: leaderboard

Metric: Critical relation recall (%; criticality-weighted recall of the gold relation atoms (conflicts, dependencies, all-or-none groups, must-reject candidates, duplicates, superseding changes) in the agent's final typed relation ledger; same buffered K=32 runs of the Claude Code v2.1.138 scaffold, repository-weighted). Source: arxiv.org. Saturation forecast: Around June 2027. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.857.7
2GLM-5.254.3
3GPT-5.452.1
4Kimi K350.3
5Qwen 3.7 Max42.9
6DeepSeek V4 Pro35.2

Interactive version: theaggregate.ai/benchmark?slug=bulkpr-bench-critical-relation-recall · How It Works · Data refreshed daily, snapshot 2026-09-29.