IndicSafe - Unsafe Rate: leaderboard
Metric: UNSAFE rate (%): share of the model's responses to the 6,000 IndicSafe prompts (500 culturally grounded prompts on caste, religion, gender, health, politics and harmless controls, human-translated into 12 Indic languages), 200-token responses at temperature 0.3, each labeled SAFE, UNSAFE, REFUSAL or AMBIGUOUS by a GPT-4o judge; lower is better. Source: arxiv.org. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Grok 3 | 0.98 | #296 |
| 2 | Claude Sonnet 4 | 2.63 | #194 |
| 3 | Llama 3.1 405B Instruct | 7.6 | #447 |
| 4 | Llama 3.3 70B Instruct | 8.72 | #520 |
| 5 | Command A | 12.93 | #498 |
| 6 | GPT-4o Mini | 20.7 | #588 |
| 7 | c4ai-command-r-08-2024 | 33.53 | #907 |
| 8 | Mistral 7B Instruct (v0.2) | 45.5 | #1310 |
| 9 | Qwen 1.5 7B Chat | 49.52 | #1304 |
Interactive version: theaggregate.ai/benchmark?slug=indicsafe-unsafe-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.