HELM VHELM - Mmmu - Clinical Medicine: leaderboard

Metric: Prefix Quasi-Exact Match (%). Source: crfm.stanford.edu. 48 models tracked.

Top models

#ModelScore
1GPT-4o (2024-08-06)80
2Gemini 1.5 Pro (002)76.67
3Gemini 2.0 Flash (Preview)76.67
4O3 (2025-04-16)76.67
5GPT-4.576.67
6Gemini 2.0 Flash Lite76.67
7Claude 3.5 Sonnet (20241022)73.33
8Claude 3.5 Sonnet (20240620)73.33
9Gemini 2.0 Flash73.33
10Gemini 2.5 Pro (Preview 03-25)73.33
11GPT-4o (2024-05-13)73.33
12O4 Mini (2025-04-16)73.33
13GPT-4o (2024-11-20)73.33
14GPT-4o Mini (2024-07-18)70
15GPT-4.1 (2025-04-14)70

Interactive version: theaggregate.ai/benchmark?slug=helm-vhelm-mmmu-clinical-medicine · How It Works · Data refreshed daily, snapshot 2026-09-19.