L3Cube-IndicQuest v2 (English): leaderboard
Metric: Weighted accuracy (%; all 3,471 English questions, domains weighted by size; closed-book short answers to curriculum-grounded India-specific factual questions in English, judged correct or incorrect by Gemma 3 12B against the gold answer (paraphrases and transliterations accepted)). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 86.9 |
| 2 | Gemma 4 31B | 62.2 |
| 3 | GPT-5.4 Mini | 60.1 |
| 4 | Sarvam 30B | 48.7 |
| 5 | Gemma 2 9B | 41.1 |
| 6 | Llama 3.1 8B | 35.4 |
Interactive version: theaggregate.ai/benchmark?slug=l3cube-indicquest-v2-english · How It Works · Data refreshed daily, snapshot 2026-09-29.