SuiChat-CN - Without Context: leaderboard

Metric: Macro F1 (%) of suicide-risk level classification of Chinese group-chat topics into six levels (L0 to L5) on the SuiChat-CN test set, macro F1 over the levels, with the surrounding group-chat context removed; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro73.4
2Kimi K2.673.3
3Kimi K2.573
4Kimi K272.4
5GLM-5.172.3
6Qwen 3.5 Plus72.2
7Seed-OSS-36B-Instruct71.3
8DeepSeek R171.1
9GLM-4 32B (0414)70.8
10DeepSeek V370.7
11GPT-569.2
12Qwen 3 235B A22B69.2
13GLM-4.669.1
14Llama 3.3 70B Instruct68.9
15Qwen 3.5 397B A17B68.6

Interactive version: theaggregate.ai/benchmark?slug=suichat-cn-without-context · How It Works · Data refreshed daily, snapshot 2026-10-07.