CASTLE Student Safety (English) - Content and Information Bias: leaderboard
Metric: Average Safety Score (1-5) on CASTLE's English scenarios, the Content and Information Bias risk category (stereotypes, hallucination risks, inappropriate content and knowledge limitations): mean of Risk Sensitivity, Emotional Empathy and Student Alignment, each rated 1-5 by a Claude-Haiku-4.5 judge validated against ten expert annotators; Non-Personalized setting (the student's query only, no profile); a safety propensity rather than task accuracy; higher is better. Source: arxiv.org. 18 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Haiku 4.5 | 1.86 | #271 |
| 2 | Qwen 3 235B A22B | 1.78 | #304 |
| 3 | QwQ-32B | 1.76 | #410 |
| 4 | Ministral-3-14B-Instruct-2512 | 1.72 | #590 |
| 5 | Qwen 2.5 72B Instruct | 1.68 | #436 |
| 6 | Gemini 2.5 Flash | 1.66 | #237 |
| 7 | GPT-4o | 1.64 | #333 |
| 8 | Qwen 2.5 32B Instruct | 1.64 | #491 |
| 9 | InternLM3-8B-Instruct | 1.64 | #847 |
| 10 | GLM-4 9B Chat | 1.63 | #904 |
| 11 | Mistral 7B Instruct | 1.63 | #1345 |
| 12 | GPT-5.2 | 1.62 | #105 |
| 13 | Llama 3 8B Instruct | 1.58 | #1115 |
| 14 | deepseek-llm-7B-chat | 1.54 | #1335 |
| 15 | Qwen 2.5 7B Instruct | 1.49 | #846 |
Interactive version: theaggregate.ai/benchmark?slug=castle-student-safety-english-content-and-information-bias · How It Works · Data refreshed daily, snapshot 2026-10-11.