MINT (Medical) - Derm-Private (Question First, Initial Answer): leaderboard
Metric: Initial-answer accuracy (%) on the 99 Derm-Private-derived diagnosis cases: the multiple-choice question comes first with an instruction to wait, evidence shards follow one per turn, and the first answer the model commits to is scored over the cases it answered (abstentions are left out of the denominator; Ask-Question-First); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 94.8 |
| 2 | GPT-5 Mini | 84.8 |
| 3 | O4 Mini | 84.7 |
| 4 | GPT-OSS-20B | 64.7 |
| 5 | Qwen 3 4B | 58.8 |
| 6 | MedGemma 1.5 4B | 38.4 |
Interactive version: theaggregate.ai/benchmark?slug=mint-medical-derm-private-question-first-initial-answer · How It Works · Data refreshed daily, snapshot 2026-10-07.