IndicContextEval: leaderboard

Metric: Word error rate (%), averaged over the benchmark, with only the target language given in the prompt (L1); 56 hours of natural read and extempore speech from 555 speakers in 8 Indian languages (Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu) and 23 professional domains, native-script output; lower is better. Source: arxiv.org. Saturation forecast: Around September 2027. 5 models tracked.

Top models

#ModelScore
1Gemini 3 Flash18.9
2Gemma 3n E4B38.73

Interactive version: theaggregate.ai/benchmark?slug=indiccontexteval · How It Works · Data refreshed daily, snapshot 2026-09-29.