HELM VHELM - Mmmu - Public Health: leaderboard

Metric: Prefix Quasi-Exact Match (%). Source: crfm.stanford.edu. 48 models tracked.

Top models

#ModelScore
1O4 Mini (2025-04-16)100
2Gemini 2.5 Pro (Preview 03-25)96.67
3O3 (2025-04-16)96.67
4O1 (2024-12-17)96.67
5Gemini 1.5 Pro (002)90
6Claude 3.5 Sonnet (20240620)90
7GPT-4.586.67
8Llama 4 Maverick Instruct FP883.33
9Gemini 2.0 Pro (Preview 02-05)83.33
10GPT-4.1 Mini80
11Gemini 2.0 Flash80
12Claude 3.5 Sonnet (20241022)76.67
13GPT-4o (2024-08-06)76.67
14GPT-4 Turbo76.67
15Gemini 2.0 Flash (Preview)76.67

Interactive version: theaggregate.ai/benchmark?slug=helm-vhelm-mmmu-public-health · How It Works · Data refreshed daily, snapshot 2026-09-19.