MalruleLib: leaderboard

Metric: Malrule Reasoning Accuracy (%; predict the student's next answer from one worked mistake; cross-template, answer only). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-OSS-120B56.9
2Qwen 3 Next 80B A3B (Thinking)56.4
3GPT-OSS-20B53.6
4Qwen 3 Next 80B A3B Instruct51.3
5Qwen 3 4B39.1
6Phi-436.7
7Llama 3.3 70B Instruct34.6
8Phi-4 Mini Instruct18.2

Interactive version: theaggregate.ai/benchmark?slug=malrulelib · How It Works · Data refreshed daily, snapshot 2026-09-25.