PAIR-Bench - Gap Closure: leaderboard

Metric: Gap closure (%; share of the remaining hidden-test pass-rate gap after the first attempt that the best program within the hint budget closes, averaged over initially failing instances; higher is better). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1DeepSeek V3.298.8
2Gemini 2.5 Flash Lite97.51
3Qwen 3 Coder 30B A3B Instruct94.25
4Ministral 3 14B91.73
5Llama 3.3 70B Instruct89.32
6GPT-4o Mini89.08
7Mistral Small 3.287.92
8Gemma 3 27B80.06

Interactive version: theaggregate.ai/benchmark?slug=pair-bench-gap-closure · How It Works · Data refreshed daily, snapshot 2026-09-29.