ChineseSafe Benchmark — leaderboard
Chinese content safety evaluation: measures LLM accuracy in identifying unsafe vs safe content in Chinese across multiple safety categories.
Metric: Accuracy (%). Source: huggingface.co. Status: saturation imminent. 55 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | deepseek-llm-67B-chat | 76.76 |
| 2 | Qwen 3 32B | 75.26 |
| 3 | Llama 4 Maverick | 75.02 |
| 4 | Qwen 3 4B | 74.95 |
| 5 | GPT-4o | 73.78 |
| 6 | Phi-3-small-8k-instruct | 72.73 |
| 7 | Phi-4 | 72.24 |
| 8 | Qwen 2.5 3B Instruct | 71.81 |
| 9 | Gemma 1.1 7B (IT) | 71.7 |
| 10 | GPT-4 Turbo | 71.67 |
| 11 | deepseek-llm-7B-chat | 71.63 |
| 12 | Gemma 3 4B (IT) | 71.41 |
| 13 | Gemini 2.5 Flash (Preview 05-20) | 71.27 |
| 14 | GLM-4 9B Chat | 70.96 |
| 15 | Mistral 7B Instruct (v0.3) | 70.41 |
Interactive version: theaggregate.ai/benchmark?slug=chinesesafe-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.