LingxiDiagBench (Static) - 2-Class Accuracy: leaderboard

Metric: Accuracy (times 100, 0-100) on binary depression-versus-anxiety classification (cases without comorbidity); LingxiDiagBench-Static on the 1,000 test cases of LingxiDiag-16K (synthetic Chinese psychiatric consultation dialogues with EMRs matching the Shanghai Mental Health Center outpatient distribution): zero-shot diagnosis from the complete dialogue; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash85.4#93
2Grok 4.1 Fast84.1#208
3Qwen 3 8B83.5#667
4Qwen 3 32B82.7#424
5Claude Haiku 4.582.5#271
6Qwen 3 4B82.5#823
7DeepSeek V3.282#198
8Kimi K2 (Thinking)81.8#236 (Kimi K2)
9GPT-5 Mini80.3#176
10Qwen 3 1.7B78.6#1186
11GPT-OSS-20B77.8#499

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=lingxidiagbench-static-2-class-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.