GAIA-v2-LILT - Hindi: leaderboard

Metric: Pass@1 accuracy (%) on the 165 GAIA validation tasks in Hindi after the GAIA-v2-LILT human audit (functional and cultural alignment of each machine-translated query and answer), solved by the smolagents Open Deep Research agent with web search (Exa), speech recognition, image captioning and file-reading tools, at most 12 manager steps and 20 search-subagent steps; exact match after removing spaces, lowercasing and stripping punctuation; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1Open Deep Research + Gemini 3.1 Pro63.6
2Open Deep Research + Claude Opus 4.662.4
3Open Deep Research + GPT-5.460

Interactive version: theaggregate.ai/benchmark?slug=gaia-v2-lilt-hindi · How It Works · Data refreshed daily, snapshot 2026-10-07.