SWE-PRBench — leaderboard

Pull-request review benchmark with a public paper-baseline leaderboard for model review quality and false-positive behavior.

Metric: Overall (sbar) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 8 models tracked.

Top models

#ModelScore
1Claude Haiku 4.515.3
2Claude Sonnet 4.615.2
3DeepSeek V315
4Mistral Large 314.7
5GPT-4o11.3
6GPT-4o Mini10.8
7Mistral Small10.6
8Llama 3.3 70B Instruct7.9

Interactive version: theaggregate.ai/benchmark?slug=swe-prbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.