PAIR-Bench - Hint Efficiency: leaderboard

Metric: Hint efficiency score (out of 6): 7 minus the hint level, from 1 for coarse symptoms to 6 for implementation guidance, at which each attempted failure scenario is closed, 0 when it is never closed, averaged over scenarios and initially failing instances; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1DeepSeek V3.24.79
2GPT-4o Mini3.8
3Gemini 2.5 Flash Lite3.76
4Qwen 3 Coder 30B A3B Instruct3.61
5Ministral 3 14B3.29
6Llama 3.3 70B Instruct3.21
7Mistral Small 3.23.02
8Gemma 3 27B2.76

Interactive version: theaggregate.ai/benchmark?slug=pair-bench-hint-efficiency · How It Works · Data refreshed daily, snapshot 2026-09-29.