GAIA-v2-LILT - German: leaderboard

Metric: Pass@1 accuracy (%) on the 165 GAIA validation tasks in German after the GAIA-v2-LILT human audit (functional and cultural alignment of each machine-translated query and answer), solved by the smolagents Open Deep Research agent with web search (Exa), speech recognition, image captioning and file-reading tools, at most 12 manager steps and 20 search-subagent steps; exact match after removing spaces, lowercasing and stripping punctuation; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1Open Deep Research + Gemini 3.1 Pro66.7
2Open Deep Research + Claude Opus 4.666.7
3Open Deep Research + GPT-5.463.6

Interactive version: theaggregate.ai/benchmark?slug=gaia-v2-lilt-german · How It Works · Data refreshed daily, snapshot 2026-10-07.