HindiGen — leaderboard

Hindi generative benchmark evaluating LLMs on question answering, grammar, and safety in Hindi using the 3C3H scoring framework (Correctness, Completeness, Conciseness, Helpfulness, Honesty, Harmlessness).

Metric: 3C3H Score (%). Source: huggingface.co. Status: saturation imminent. 35 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)85.56
2O1 (2024-12-17)79.64
3Claude 3.5 Sonnet (20241022)77.47
4O4 Mini (2025-04-16)75.52
5Claude Opus 4 (20250514)74.49
6GPT-4o (2024-08-06)74.45
7GPT-4o (2024-05-13)73.98
8Gemini 2.5 Flash (Preview 05-20)73.65
9GPT-4.173.37
10GPT-4o (2024-11-20)72.44
11Gemini 2.5 Pro (Preview 05-06)71.77
12Claude 3.7 Sonnet (20250219)70.77
13Llama 3.1 70B Instruct70.45
14Claude Sonnet 4 (20250514)69.75
15GPT-4o Mini (2024-07-18)65.5

Interactive version: theaggregate.ai/benchmark?slug=hindigen · How the rankings work · Data refreshed daily, snapshot 2026-07-22.