MedBench v5 - CCR-LLM: leaderboard

Metric: Macro-average score (0-100) over the 36 text tasks of the MedBench v5 Clinical Cognitive Responsiveness LLM track (medical knowledge QA, language understanding and generation, clinical reasoning and decision-making, safety and ethics; macro-recall and LLM-judge scoring), each task normalized to 0-100; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.769.16
2Gemini 3.1 Pro (Preview)68.61
3Kimi K2.668.19
4GLM-5.167.7
5GPT-5.567.37
6Grok 4.2066.87
7Seed 2.0 Pro66.56
8DeepSeek V4 Pro65.86

Interactive version: theaggregate.ai/benchmark?slug=medbench-v5-ccr-llm · How It Works · Data refreshed daily, snapshot 2026-09-29.