LakeQA - Mini: leaderboard

Metric: End-to-end exact match (%) on the 135 LakeQA-mini tasks, a stratified subset of LakeQA-full that preserves the reasoning-intensity distribution; the agent answers a multi-hop question over a data lake of about 40 million Data.gov and Wikipedia documents with one tool call per round (search, list, download, inspect, query) until it answers or hits the turn limit; exact match against the annotator-verified answer; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1Claude Opus 4.545.93
2Claude Sonnet 4.533.33
3GPT-5.217.04
4Claude Haiku 4.511.11
5DeepSeek R16.67
6GPT-5 Mini2.22
7Llama 3.3 70B Instruct1.48

Interactive version: theaggregate.ai/benchmark?slug=lakeqa-mini · How It Works · Data refreshed daily, snapshot 2026-09-29.