MINT (Medical) - MedMCQA (Question First, Initial Answer): leaderboard

Metric: Initial-answer accuracy (%) on the 174 MedMCQA-derived diagnosis cases: the multiple-choice question comes first with an instruction to wait, evidence shards follow one per turn, and the first answer the model commits to is scored over the cases it answered (abstentions are left out of the denominator; Ask-Question-First); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 11 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.689.6
2O4 Mini86.1
3GPT-5 Mini84.4
4GPT-OSS-20B75.7
5Qwen 3 4B58.7
6MedGemma 1.5 4B37.8

Interactive version: theaggregate.ai/benchmark?slug=mint-medical-medmcqa-question-first-initial-answer · How It Works · Data refreshed daily, snapshot 2026-10-07.