SuiChat-CN - Few-Shot: leaderboard

Metric: Macro F1 (%) of suicide-risk level classification of Chinese group-chat topics into six levels (L0 to L5) on the SuiChat-CN test set, macro F1 over the levels, few-shot prompting with labeled examples; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 44 models tracked.

Top models

#ModelScore
1Seed-OSS-36B-Instruct77.3
2Gemini 2.5 Pro76.4
3Ring-flash-2.076
4DeepSeek V375.1
5DeepSeek R174.1
6Qwen Max73.6
7Kimi K2.573.4
8Qwen 3.5 Plus73.3
9Kimi K272.8
10Qwen 3.5 397B A17B72.4
11GPT-572
12Kimi K2.671.9
13Llama 3.3 70B Instruct71.2
14GLM-5.171.1
15Qwen 3 235B A22B70.8

Interactive version: theaggregate.ai/benchmark?slug=suichat-cn-few-shot · How It Works · Data refreshed daily, snapshot 2026-10-07.