Open LMM Reasoning - DynaMath - Subject-puzzle test — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 100 models tracked.

Top models

#ModelScore
1GPT-4.1 Mini47.1
2GPT-4.1 (2025-04-14)41.2
3InternVL3-38B35.3
4Claude 3.5 Sonnet (20241022)35.3
5Doubao-1.5-Pro35.3
6InternVL3-78B29.4
7Claude 3.7 Sonnet29.4
8Gemini 2.0 Flash29.4
9GPT-4o ChatGPT29.4
10Gemma 3 12B23.5
11Grok 2 (1212)17.6
12QVQ-72B-Preview17.6
13InternVL3-8B11.8
14GPT-4.1 Nano11.8
15InternVL2.5-78B11.8

Interactive version: theaggregate.ai/benchmark?slug=open-lmm-reasoning-dynamath-subject-puzzle-test · How the rankings work · Data refreshed daily, snapshot 2026-07-22.