MINT (Medical) - MedQA (Question First, Initial Answer): leaderboard
Metric: Initial-answer accuracy (%) on the 600 MedQA-derived diagnosis cases: the multiple-choice question comes first with an instruction to wait, evidence shards follow one per turn, and the first answer the model commits to is scored over the cases it answered (abstentions are left out of the denominator; Ask-Question-First); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 87.7 |
| 2 | GPT-5 Mini | 76.5 |
| 3 | O4 Mini | 76.3 |
| 4 | GPT-OSS-20B | 61.4 |
| 5 | Qwen 3 4B | 53 |
| 6 | MedGemma 1.5 4B | 31.8 |
Interactive version: theaggregate.ai/benchmark?slug=mint-medical-medqa-question-first-initial-answer · How It Works · Data refreshed daily, snapshot 2026-10-07.