JMed48k (With Images): leaderboard

Metric: Accuracy (%) on the 2,579 scored JMed48k-Eval items that include images, unweighted macro average over the 11 professions, official MHLW gold answers, single-shot Japanese prompts at temperature 0 with reasoning disabled where the model has a switch, JSON-constrained answers; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro79.7
2GPT-578.8
3GPT-5 Mini70.6
4Gemini 2.5 Flash (Non-reasoning)66.7
5Qwen 3.5 397B A17B (Non-reasoning)63.6
6Gemma 4 31B (IT)61.6
7Claude Sonnet 457.7
8Qwen 3.5 27B (Non-reasoning)57.6
9Llama 4 Maverick53.8
10Grok 4.2053.6
11Lingshu-32B45
12Qwen 3.5 9B (Non-reasoning)41.8
13MedGemma-27B-IT34.2
14MedGemma-4B-IT23.6

Interactive version: theaggregate.ai/benchmark?slug=jmed48k-with-images · How It Works · Data refreshed daily, snapshot 2026-10-07.