LingxiDiagBench (Static): leaderboard

Metric: Overall score (times 100, 0-100): mean of the eleven printed classification metrics (accuracy, macro-F1 and weighted F1 of the 2-class and 4-class tasks; accuracy, top-1 and top-3 accuracy, macro-F1 and weighted F1 of the 12-class task); LingxiDiagBench-Static on the 1,000 test cases of LingxiDiag-16K (synthetic Chinese psychiatric consultation dialogues with EMRs matching the Shanghai Mental Health Center outpatient distribution): zero-shot diagnosis from the complete dialogue; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 13 models tracked.

Top models

#ModelScoreOverall rank
1Grok 4.1 Fast52.1#208
2Claude Haiku 4.551.6#271
3Gemini 3 Flash51#93
4Qwen 3 32B50.6#424
5GPT-5 Mini50.4#176
6DeepSeek V3.250.1#198
7Kimi K2 (Thinking)48.4#236 (Kimi K2)
8GPT-OSS-20B47.9#499
9Qwen 3 4B47.4#823
10Qwen 3 8B47.3#667
11Qwen 3 1.7B44.8#1186

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=lingxidiagbench-static · How It Works · Data refreshed daily, snapshot 2026-10-11.