LingxiDiagBench (Static) - 12-Class Accuracy: leaderboard

Metric: Accuracy (times 100, 0-100) on twelve-category ICD-10 diagnosis (F20, F31, F32, F39, F41, F42, F43, F45, F51, F98, Z71, others; multi-label, exact match of the predicted label set); LingxiDiagBench-Static on the 1,000 test cases of LingxiDiag-16K (synthetic Chinese psychiatric consultation dialogues with EMRs matching the Shanghai Mental Health Center outpatient distribution): zero-shot diagnosis from the complete dialogue; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 13 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini40.9#176
2Claude Haiku 4.539.5#271
3Grok 4.1 Fast35.1#208
4Kimi K2 (Thinking)33.5#236 (Kimi K2)
5DeepSeek V3.232.3#198
6GPT-OSS-20B25.9#499
7Qwen 3 32B24.1#424
8Gemini 3 Flash17.2#93
9Qwen 3 1.7B14.5#1186
10Qwen 3 4B2.1#823
11Qwen 3 8B1.2#667

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=lingxidiagbench-static-12-class-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.