DataSpace: leaderboard

Metric: Task accuracy (%; share of the 410 DataSpace tasks, each a question over a task-local heterogeneous workspace of CSV, JSON, SQLite, Markdown, PDF and video artifacts, whose returned table matches the gold table under the deterministic evaluator: header-invariant column alignment, type- and precision-aware normalization and order-aware row comparison; DataSpace-Agent reference harness unless the row names another harness). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Grok 4.566.34
2GPT-5.6 Sol64.63
3Kimi K353.41
4MiMo-V2.539.27
5Claude Sonnet 532.93
6MiniMax-M328.54

Interactive version: theaggregate.ai/benchmark?slug=dataspace · How It Works · Data refreshed daily, snapshot 2026-09-29.