GAIA-v2-LILT - German: leaderboard
Metric: Pass@1 accuracy (%) on the 165 GAIA validation tasks in German after the GAIA-v2-LILT human audit (functional and cultural alignment of each machine-translated query and answer), solved by the smolagents Open Deep Research agent with web search (Exa), speech recognition, image captioning and file-reading tools, at most 12 manager steps and 20 search-subagent steps; exact match after removing spaces, lowercasing and stripping punctuation; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Open Deep Research + Gemini 3.1 Pro | 66.7 |
| 2 | Open Deep Research + Claude Opus 4.6 | 66.7 |
| 3 | Open Deep Research + GPT-5.4 | 63.6 |
Interactive version: theaggregate.ai/benchmark?slug=gaia-v2-lilt-german · How It Works · Data refreshed daily, snapshot 2026-10-07.