KMP-Bench (Skills) - Error Correction: leaderboard

Metric: Correction accuracy (%): share of turn-two corrected solutions reaching the right answer, on KMP-Skills (K-8 math problems with LLM-generated pedagogical components); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.0 Flash96.7#331
2Phi-492.1#701
3GPT-4o90.9#333
4Qwen 2.5 72B Instruct89.5#436
5Qwen 2.5 32B Instruct87.8#491
6GPT-4o Mini84.3#588
7Qwen 2.5 14B Instruct82#634
8Qwen2.5-Math-72B-Instruct81#708
9Qwen 2.5 7B Instruct73.2#846
10Llama 3.1 8B Instruct55.9#1018
11Qwen2.5-Math-7B-Instruct52.3#1422
12Mistral Nemo Instruct (2407)48.1#916

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=kmp-bench-skills-error-correction · How It Works · Data refreshed daily, snapshot 2026-10-11.