Medmarks - MedConceptsQA Medium — leaderboard

Metric: Score (%). Source: medmarks.ai. 71 models tracked.

Top models

#ModelScore
1Grok 499.82
2Gemini 3 Pro (Preview)99.02
3Claude Sonnet 4.597.35
4GPT-5.1 (Medium)97.07
5GPT-5.2 (Medium)95.93
6GLM-4.7 FP893.17
7Qwen 3 235B A22B (Thinking)85.95
8MiniMax-M284.02
9MiniMax-M2.183.83
10Llama 3.3 70B Instruct76.17
11Qwen 3 Next 80B A3B (Thinking)75.6
12GLM 4.5 Air75.38
13INTELLECT-374.1
14GPT-OSS-120B (High)73.8
15Qwen 3 Next 80B A3B Instruct72.62

Interactive version: theaggregate.ai/benchmark?slug=medmarks-medconceptsqa-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.