IndicSafe: leaderboard
Metric: SAFE rate (%): share of the model's responses to the 6,000 IndicSafe prompts (500 culturally grounded prompts on caste, religion, gender, health, politics and harmless controls, human-translated into 12 Indic languages), 200-token responses at temperature 0.3, each labeled SAFE, UNSAFE, REFUSAL or AMBIGUOUS by a GPT-4o judge; higher is better. Source: arxiv.org. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Grok 3 | 83.67 | #296 |
| 2 | Claude Sonnet 4 | 60.38 | #194 |
| 3 | Command A | 59.2 | #498 |
| 4 | GPT-4o Mini | 58.53 | #588 |
| 5 | Llama 3.1 405B Instruct | 56.9 | #447 |
| 6 | Llama 3.3 70B Instruct | 56.42 | #520 |
| 7 | c4ai-command-r-08-2024 | 33.82 | #907 |
| 8 | Mistral 7B Instruct (v0.2) | 12.85 | #1310 |
| 9 | Qwen 1.5 7B Chat | 4.55 | #1304 |
Interactive version: theaggregate.ai/benchmark?slug=indicsafe · How It Works · Data refreshed daily, snapshot 2026-10-11.