GPAgentBench-2K - Treatment Score: leaderboard

Metric: Treatment score (%; prescribed treatment scored by a GPT-5.4 judge against the ground-truth treatment; zero-shot GP agent with six clinical actions over 2,428 physician-validated primary-care cases, GPT-4o-mini patient simulator, at most 10 turns). Source: arxiv.org. Saturation forecast: Around September 2028. 21 models tracked.

Top models

#ModelScore
1GPT-5.453.2
2Kimi K2.649.2
3Claude Opus 4.746
4Gemini 3.1 Pro (Preview)41.6
5DeepSeek V4 Pro41.2
6Qwen 3.7 Max41.1
7DeepSeek V4 Flash40.6
8MiMo-V2.5-Pro38.3
9Claude Sonnet 4.636
10GLM-5.134.3
11MiniMax-M2.731.3
12MedGemma-27B29.7
13Lingshu-32B25.2
14Lingshu-7B23

Interactive version: theaggregate.ai/benchmark?slug=gpagentbench-2k-treatment-score · How It Works · Data refreshed daily, snapshot 2026-09-26.