Merge-Bench (All Languages): leaderboard

Metric: Equivalent text (%): share of conflicts whose output matches the developer resolution exactly as a string, all 7,938 merge-conflict hunks from 1,439 repositories in 11 languages, the model returns the resolved snippet in Markdown, compared with the resolution the developers committed; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 6 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro47.1
2Claude Opus 440.3
3O3 Pro39.2
4DeepSeek R1 052832
5Grok 427.7
6Qwen 3 235B A22B25.8

Interactive version: theaggregate.ai/benchmark?slug=merge-bench-all-languages · How It Works · Data refreshed daily, snapshot 2026-10-07.