MedMT-Bench (Text Subset): leaderboard

Metric: Accuracy (%) on the 285 text-only dialogues of MedMT-Bench's long multi-turn medical dialogues (health consultation, treatment recommendation and rehabilitation scenarios, text and image-text), where the final-turn reply is checked against fine-grained test points, judged by a Gemini-2.5-Pro evaluator at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScoreOverall rank
1Kimi K244.91#236
2Qwen 3 32B (Thinking)41.75#424 (Qwen 3 32B)
3GLM-4.5 (Thinking)40.85#265 (GLM-4.5)
4Qwen 3 32B (Non-reasoning)34.74#424 (Qwen 3 32B)
5Qwen 3 8B (Thinking)34.39#667 (Qwen 3 8B)
6GLM-4.5 (Non-reasoning)34.39#265 (GLM-4.5)
7Qwen 3 8B (Non-reasoning)33.33#667 (Qwen 3 8B)
8Llama 3.1 70B Instruct24.91#548
9Llama 3.1 8B Instruct18.25#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=medmt-bench-text-subset · How It Works · Data refreshed daily, snapshot 2026-10-11.