PAIR-Bench - Initial Fix: leaderboard

Metric: Initial fix rate (%; share of the 440 wrong-answer programs a model repairs to pass every hidden test in its first attempt, before any hint; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1DeepSeek V3.265.68
2Gemini 2.5 Flash Lite53.18
3Qwen 3 Coder 30B A3B Instruct35.9
4GPT-4o Mini29.77
5Llama 3.3 70B Instruct27.72
6Ministral 3 14B22.95
7Mistral Small 3.213.4
8Gemma 3 27B10.22

Interactive version: theaggregate.ai/benchmark?slug=pair-bench-initial-fix · How It Works · Data refreshed daily, snapshot 2026-09-29.