SuiChat-CN - Without Keywords: leaderboard

Metric: Macro F1 (%) of suicide-risk level classification of Chinese group-chat topics into six levels (L0 to L5) on the SuiChat-CN test set, macro F1 over the levels, with the risk keywords removed from the input; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 44 models tracked.

Top models

#ModelScore
1DeepSeek R176.3
2GLM-5.175.1
3Qwen 3.5 Plus75.1
4Kimi K274.7
5Gemini 2.5 Pro74.5
6Seed-OSS-36B-Instruct74.2
7Qwen Max73.1
8Qwen 3.5 397B A17B72.9
9GLM-4 32B (0414)71.6
10Qwen 3 235B A22B70.5
11GPT-570.1
12GLM-569.9
13Llama 3.3 70B Instruct69.8
14Kimi K2.569.3
15GLM-4.669

Interactive version: theaggregate.ai/benchmark?slug=suichat-cn-without-keywords · How It Works · Data refreshed daily, snapshot 2026-10-07.