JMed48k (Text-Only) - Physical Therapist: leaderboard

Metric: Accuracy (%) on the text-only scored items of the Japanese Physical Therapist national licensing examination (2021-2025) in JMed48k-Eval, official MHLW gold answers, single-shot Japanese prompts at temperature 0 with reasoning disabled where the model has a switch, JSON-constrained answers; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro96.7
2GPT-596.2
3GPT-5 Mini93.2
4Claude Sonnet 490.5
5Qwen 3.5 397B A17B (Non-reasoning)89.1
6Gemini 2.5 Flash (Non-reasoning)88.7
7Gemma 4 31B (IT)86.7
8DeepSeek R186.5
9Grok 4.2085.6
10Qwen 3.5 27B (Non-reasoning)83.2
11Llama 4 Maverick79.2
12Qwen 3.5 9B (Non-reasoning)69.7
13Lingshu-32B65.8
14MedGemma-27B-IT59.8
15MedGemma-4B-IT30.6

Interactive version: theaggregate.ai/benchmark?slug=jmed48k-text-only-physical-therapist · How It Works · Data refreshed daily, snapshot 2026-10-07.