CASTLE Student Safety (Chinese) - Loss of Independent Judgment: leaderboard
Metric: Average Safety Score (1-5) on CASTLE's Chinese scenarios, the Loss of Independent Judgment risk domain (IJ, 3,125 scenarios, students handing their judgment to the model; one of the three domains of the Learning Dependence and Cognition category): mean of Risk Sensitivity, Emotional Empathy and Student Alignment, each rated 1-5 by a Claude-Haiku-4.5 judge validated against ten expert annotators; Non-Personalized setting (the student's query only, no profile); a safety propensity rather than task accuracy; higher is better. Source: arxiv.org. 18 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Haiku 4.5 | 2.31 | #271 |
| 2 | Gemini 2.5 Flash | 2.09 | #237 |
| 3 | QwQ-32B | 2.09 | #410 |
| 4 | Qwen 3 235B A22B | 2.06 | #304 |
| 5 | Ministral-3-14B-Instruct-2512 | 2 | #590 |
| 6 | InternLM3-8B-Instruct | 1.98 | #847 |
| 7 | Qwen 2.5 72B Instruct | 1.97 | #436 |
| 8 | Qwen 2.5 32B Instruct | 1.87 | #491 |
| 9 | GLM-4 9B Chat | 1.87 | #904 |
| 10 | Qwen 2.5 7B Instruct | 1.85 | #846 |
| 11 | GPT-4o | 1.84 | #333 |
| 12 | GPT-5.2 | 1.84 | #105 |
| 13 | Llama 3 8B Instruct | 1.81 | #1115 |
| 14 | deepseek-llm-7B-chat | 1.72 | #1335 |
| 15 | Mistral 7B Instruct | 1.65 | #1345 |
Interactive version: theaggregate.ai/benchmark?slug=castle-student-safety-chinese-loss-of-independent-judgment · How It Works · Data refreshed daily, snapshot 2026-10-11.