Merge-Bench (All Languages) - Normalized Equivalence: leaderboard

Metric: Code-normalized equivalent (%): share of conflicts whose output matches the developer resolution after removing formatting and comments, all 7,938 merge-conflict hunks from 1,439 repositories in 11 languages, the model returns the resolved snippet in Markdown, compared with the resolution the developers committed; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro52.6
2O3 Pro45.1
3Claude Opus 444.8
4DeepSeek R1 052836.5
5Grok 431.7
6Qwen 3 235B A22B30.6

Interactive version: theaggregate.ai/benchmark?slug=merge-bench-all-languages-normalized-equivalence · How It Works · Data refreshed daily, snapshot 2026-10-07.