KMP-Bench (Skills) - Error Identification F1: leaderboard

Metric: F1 (%) for turn one of the error task, judging whether a student solution is correct and locating its first erroneous step, on KMP-Skills (K-8 math problems with LLM-generated pedagogical components); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScoreOverall rank
1Phi-480.3#701
2Qwen 2.5 72B Instruct78.6#436
3GPT-4o78.3#333
4Qwen2.5-Math-72B-Instruct73.5#708
5Qwen 2.5 14B Instruct70.9#634
6Qwen 2.5 32B Instruct69.8#491
7GPT-4o Mini66.1#588
8Qwen 2.5 7B Instruct44.3#846
9Qwen2.5-Math-7B-Instruct44.1#1422
10Mistral Nemo Instruct (2407)22.5#916
11Llama 3.1 8B Instruct9#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=kmp-bench-skills-error-identification-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.