DermoBench - Reasoning: leaderboard

Metric: Reasoning Score (0-100; mean of two Gemini-2.5-Pro judge scores, free and morphology-grounded CoT). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash53.7
2QVQ-72B-Preview52.85
3Qwen 3 VL 8B Instruct50.48
4Llama 3.2 90B50.38
5Claude Sonnet 4.5 (Thinking)48.95
6GLM-4.5V48.73
7Lingshu-7B48.23
8GPT-4o Mini47.24
9Qwen 2.5 VL 72B Instruct45.05
10Nemotron Nano 12B v2 VL34.65

Interactive version: theaggregate.ai/benchmark?slug=dermobench-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-25.