PAIR-Bench: leaderboard

Metric: Final fix rate (%; share of the 440 wrong-answer Codeforces Python programs repaired to pass every hidden test within a 10-turn budget of progressive hints generated by GPT-OSS 120B, at most 3 turns per failure scenario; higher is better). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1DeepSeek V3.299.31
2Gemini 2.5 Flash Lite95.45
3Qwen 3 Coder 30B A3B Instruct90.9
4Llama 3.3 70B Instruct85.9
5Ministral 3 14B85.9
6GPT-4o Mini85.45
7Mistral Small 3.279.77
8Gemma 3 27B72.04

Interactive version: theaggregate.ai/benchmark?slug=pair-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.