Medmarks - MedHallu Hard — leaderboard

Metric: Score (%). Source: medmarks.ai. 71 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)74.51
2GPT-5.1 (Medium)60.85
3Claude Sonnet 4.559.96
4Ling-flash-2.054.44
5Hermes 4 70B52
6GLM-4.7 FP851.99
7Qwen 3 235B A22B (Thinking)51.72
8Granite 4.0 H Small50.86
9Hermes-4-14B50.76
10Llama 3.1 8B Instruct49.42
11GPT-OSS-120B (High)48.88
12Phi-4 (Reasoning)48.36
13Qwen 3 4B (Reasoning)47.93
14GPT-OSS-20B (Low)47.28
15Llama 3.3 70B Instruct47.12

Interactive version: theaggregate.ai/benchmark?slug=medmarks-medhallu-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.