DermoBench - Reasoning: leaderboard
Metric: Reasoning Score (0-100; mean of two Gemini-2.5-Pro judge scores, free and morphology-grounded CoT). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 53.7 |
| 2 | QVQ-72B-Preview | 52.85 |
| 3 | Qwen 3 VL 8B Instruct | 50.48 |
| 4 | Llama 3.2 90B | 50.38 |
| 5 | Claude Sonnet 4.5 (Thinking) | 48.95 |
| 6 | GLM-4.5V | 48.73 |
| 7 | Lingshu-7B | 48.23 |
| 8 | GPT-4o Mini | 47.24 |
| 9 | Qwen 2.5 VL 72B Instruct | 45.05 |
| 10 | Nemotron Nano 12B v2 VL | 34.65 |
Interactive version: theaggregate.ai/benchmark?slug=dermobench-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-25.