Medmarks - MedConceptsQA Hard — leaderboard

Metric: Score (%). Source: medmarks.ai. 71 models tracked.

Top models

#ModelScore
1Grok 499.53
2Gemini 3 Pro (Preview)99.25
3Claude Sonnet 4.595.85
4GPT-5.1 (Medium)94.37
5GPT-5.2 (Medium)93.25
6GLM-4.7 FP886.6
7Qwen 3 235B A22B (Thinking)79.27
8MiniMax-M2.176.72
9MiniMax-M276.72
10Qwen 3 Next 80B A3B (Thinking)71.2
11Qwen 3 Next 80B A3B Instruct71.15
12GPT-OSS-120B (High)69.07
13Llama 3.3 70B Instruct67.95
14GLM 4.5 Air67.58
15INTELLECT-367.15

Interactive version: theaggregate.ai/benchmark?slug=medmarks-medconceptsqa-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.