LiveClin: leaderboard

Metric: Case accuracy (%): share of the 1,407 first-half-2025 clinical cases (6,605 sequential ten-option questions generated from PubMed Central case reports and verified by physicians) in which every question is answered correctly, the conversation history kept across a case's questions; zero-shot, temperature 0 or the official setting for reasoning modes; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 26 models tracked.

Top models

#ModelScoreOverall rank
1O335.7#121
2GPT-535.5#91
3Gemini 2.5 Pro33.7#145
4GPT-4.131.8#240
5Claude 3.5 Sonnet (20241022)27.3#286
6GPT-4.1 Mini26.1#346
7Claude 3.7 Sonnet (20250219) (Thinking)25.6#196 (Claude 3.7 Sonnet (20250219))
8Gemini 2.0 Flash21.4#331
9Gemini 2.5 Flash19#237
10Qwen 2.5 VL 72B Instruct18.6#364
11Gemini 1.5 Pro17.7#442
12Qwen 2.5 VL 32B Instruct17.2#443
13Claude 3.5 Haiku (20241022)15.3#574
14GPT-4o Mini (2024-07-18)14.8#546
15GPT-4o (2024-11-20)14.7#369

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=liveclin · How It Works · Data refreshed daily, snapshot 2026-10-11.