LakeQA: leaderboard

Metric: End-to-end exact match (%) on the 1,007 LakeQA-full tasks; the agent answers a multi-hop question over a data lake of about 40 million Data.gov and Wikipedia documents with one tool call per round (search, list, download, inspect, query) until it answers or hits the turn limit; exact match against the annotator-verified answer; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.532.87
2GPT-5.218.37
3DeepSeek R115.99
4Claude Haiku 4.511.42
5GPT-5 Mini5.16
6Llama 3.3 70B Instruct5.06

Interactive version: theaggregate.ai/benchmark?slug=lakeqa · How It Works · Data refreshed daily, snapshot 2026-09-29.