Open LMM Reasoning - WeMath - RoteMemorization (Strict) — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 83 models tracked.

Top models

#ModelScore
1GPT-4.1 Mini100
2InternVL2.5-2B81.9
3Llama 3.2 11B Instruct80
4Qwen 2 VL 2B75.4
5Aquila-VL-2B67.5
6InternVL2-8B65.1
7Qwen 2 VL 7B57.3
8Ovis2-8B55.6
9Gemma 3 4B53.7
10GPT-4.1 Nano50.6
11Gemma 3 12B42.5
12QVQ-72B-Preview40.7
13InternVL3-8B40.6
14Qwen 2 VL 72B40.5
15Grok 2 (1212)39.7

Interactive version: theaggregate.ai/benchmark?slug=open-lmm-reasoning-wemath-rotememorization-strict · How the rankings work · Data refreshed daily, snapshot 2026-07-22.