CASTLE Student Safety (Chinese) - Content and Information Bias: leaderboard

Metric: Average Safety Score (1-5) on CASTLE's Chinese scenarios, the Content and Information Bias risk category (stereotypes, hallucination risks, inappropriate content and knowledge limitations): mean of Risk Sensitivity, Emotional Empathy and Student Alignment, each rated 1-5 by a Claude-Haiku-4.5 judge validated against ten expert annotators; Non-Personalized setting (the student's query only, no profile); a safety propensity rather than task accuracy; higher is better. Source: arxiv.org. 16 models tracked.

Top models

#ModelScoreOverall rank
1Claude Haiku 4.52.34#271
2Gemini 2.5 Flash2.12#237
3Qwen 3 235B A22B2.12#304
4QwQ-32B2.1#410
5Qwen 2.5 72B Instruct2.07#436
6Ministral-3-14B-Instruct-25122.07#590
7InternLM3-8B-Instruct2#847
8Qwen 2.5 32B Instruct1.98#491
9GPT-4o1.96#333
10GPT-5.21.94#105
11GLM-4 9B Chat1.89#904
12Qwen 2.5 7B Instruct1.88#846
13Llama 3 8B Instruct1.81#1115
14deepseek-llm-7B-chat1.8#1335
15Mistral 7B Instruct1.7#1345

Interactive version: theaggregate.ai/benchmark?slug=castle-student-safety-chinese-content-and-information-bias · How It Works · Data refreshed daily, snapshot 2026-10-11.