PhysAssistBench (Chinese): leaderboard

Metric: Mean rubric score (%, mRS) over the 1,296 Chinese turns (324 four-turn clinical sessions built from MIMIC-IV records: 4 scenarios by 3 data-richness tiers by 27 sessions); the model assists a physician with 18 FHIR EHR, patient-interview and control tools (at most 16 tool calls per turn) and each turn scores the fraction of binary rubric items a fixed GPT-5.4-mini judge marks passed; thinking enabled (GPT-5 series at high effort), temperature 0.2; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 14 models tracked.

Top models

#ModelScore
1GLM-571.5
2Kimi K2.669.9
3Claude Opus 4.7 (Thinking)69.9
4Seed 1.868.7
5Gemini 3.1 Pro (Preview)68
6MiniMax-M2.766.8
7GPT-5.4 (High)65.9
8Qwen 3.5 35B A3B65.8
9GPT-5.4 Mini (High)61.5
10Qwen 3.5 27B61.4
11Qwen 3.5 9B58
12Qwen 3.5 4B47.4

Interactive version: theaggregate.ai/benchmark?slug=physassistbench-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.