Medmarks - LongHealth Task 2 — leaderboard

Metric: Score (%). Source: medmarks.ai. 71 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)90.96
2Claude Sonnet 4.590.58
3Qwen 3 Next 80B A3B Instruct90.33
4GPT-5.1 (Medium)90.33
5Qwen 3 235B A22B (Thinking)90.25
6Ministral-3-8B-Instruct-251289.29
7Magistral Small88.96
8Ministral-3-14B-Instruct-251288.88
9GLM-4.7 FP888.75
10GPT-5.2 (Medium)88.71
11Llama 3.3 70B Instruct88.62
12Qwen 2.5 32B Instruct88.46
13INTELLECT-388.46
14GPT-OSS-120B (High)88.38
15MiniMax-M2.188.17

Interactive version: theaggregate.ai/benchmark?slug=medmarks-longhealth-task-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.