Open LMM Reasoning - DynaMath - Subject-arithmetic; Avg — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 100 models tracked.

Top models

#ModelScore
1GPT-4.1 Mini73.8
2GPT-4o ChatGPT72.7
3GPT-4.1 (2025-04-14)71.2
4Doubao-1.5-Pro61.2
5InternVL3-38B54.2
6Claude 3.7 Sonnet53.5
7Claude 3.5 Sonnet (20241022)52.3
8Gemini 2.0 Flash51.2
9Gemini 1.5 Pro (002)49.6
10GPT-4.1 Nano49.6
11QVQ-72B-Preview48.8
12InternVL3-14B47.7
13InternVL3-78B47.3
14Gemma 3 27B46.5
15InternVL3-8B45

Interactive version: theaggregate.ai/benchmark?slug=open-lmm-reasoning-dynamath-subject-arithmetic-avg · How the rankings work · Data refreshed daily, snapshot 2026-07-22.