KMP-Bench (Skills) - Multi-Turn Problem Solving: leaderboard

Metric: Accuracy (%) on the final turn of a three-turn sequence (a seed problem and two progressively harder follow-up questions, each answered with the earlier answers as context), on KMP-Skills (K-8 math problems with LLM-generated pedagogical components); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.0 Flash91.2#331
2Phi-485.6#701
3GPT-4o84.1#333
4Qwen2.5-Math-72B-Instruct83.2#708
5Qwen 2.5 32B Instruct82.6#491
6Qwen 2.5 72B Instruct82.2#436
7Qwen 2.5 14B Instruct81.4#634
8GPT-4o Mini79.4#588
9Qwen2.5-Math-7B-Instruct76.1#1422
10Qwen 2.5 7B Instruct71#846
11Mistral Nemo Instruct (2407)50.2#916
12Llama 3.1 8B Instruct50.1#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=kmp-bench-skills-multi-turn-problem-solving · How It Works · Data refreshed daily, snapshot 2026-10-11.