IndicSafe: leaderboard

Metric: SAFE rate (%): share of the model's responses to the 6,000 IndicSafe prompts (500 culturally grounded prompts on caste, religion, gender, health, politics and harmless controls, human-translated into 12 Indic languages), 200-token responses at temperature 0.3, each labeled SAFE, UNSAFE, REFUSAL or AMBIGUOUS by a GPT-4o judge; higher is better. Source: arxiv.org. 10 models tracked.

Top models

#ModelScoreOverall rank
1Grok 383.67#296
2Claude Sonnet 460.38#194
3Command A59.2#498
4GPT-4o Mini58.53#588
5Llama 3.1 405B Instruct56.9#447
6Llama 3.3 70B Instruct56.42#520
7c4ai-command-r-08-202433.82#907
8Mistral 7B Instruct (v0.2)12.85#1310
9Qwen 1.5 7B Chat4.55#1304

Interactive version: theaggregate.ai/benchmark?slug=indicsafe · How It Works · Data refreshed daily, snapshot 2026-10-11.