IndicSafe - Unsafe Rate: leaderboard

Metric: UNSAFE rate (%): share of the model's responses to the 6,000 IndicSafe prompts (500 culturally grounded prompts on caste, religion, gender, health, politics and harmless controls, human-translated into 12 Indic languages), 200-token responses at temperature 0.3, each labeled SAFE, UNSAFE, REFUSAL or AMBIGUOUS by a GPT-4o judge; lower is better. Source: arxiv.org. 10 models tracked.

Top models

#ModelScoreOverall rank
1Grok 30.98#296
2Claude Sonnet 42.63#194
3Llama 3.1 405B Instruct7.6#447
4Llama 3.3 70B Instruct8.72#520
5Command A12.93#498
6GPT-4o Mini20.7#588
7c4ai-command-r-08-202433.53#907
8Mistral 7B Instruct (v0.2)45.5#1310
9Qwen 1.5 7B Chat49.52#1304

Interactive version: theaggregate.ai/benchmark?slug=indicsafe-unsafe-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.