GPAgentBench-2K - Diagnosis Accuracy: leaderboard

Metric: Diagnosis accuracy (%; final diagnosis judged correct by a GPT-5.4 judge against the ground truth; zero-shot GP agent with six clinical actions over 2,428 physician-validated primary-care cases, GPT-4o-mini patient simulator, at most 10 turns). Source: arxiv.org. Saturation forecast: Around April 2028. 21 models tracked.

Top models

#ModelScore
1Claude Opus 4.770
2Gemini 3.1 Pro (Preview)66.8
3MiMo-V2.5-Pro66
4GPT-5.464.1
5DeepSeek V4 Flash63.9
6Qwen 3.7 Max63.1
7Claude Sonnet 4.661.7
8MiniMax-M2.760
9DeepSeek V4 Pro59.5
10Kimi K2.656.7
11GLM-5.154.2
12Lingshu-32B50.9
13MedGemma-27B49.5
14Lingshu-7B47.1

Interactive version: theaggregate.ai/benchmark?slug=gpagentbench-2k-diagnosis-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.