MedCase-Structured (5-Shot): leaderboard

Metric: Diagnostic accuracy (%; the model reads a case as a terminology-coded FHIR R4 bundle with the final diagnosis hidden and names the most likely diagnosis; an LLM judge (Claude Sonnet 4.6) decides clinical equivalence to the reference; mean of three runs; five examples from the training split in the prompt; medium reasoning, temperature 1.0, 800 generated tokens). Source: arxiv.org. Saturation forecast: Around March 2027. 3 models tracked.

Top models

#ModelScore
1GPT-5.4 (Medium)58.18
2Claude Opus 4.6 (Medium)56.3
3Gemini 3.1 Pro (Preview) (Medium)40.91

Interactive version: theaggregate.ai/benchmark?slug=medcase-structured-5-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.