WorkSurface-Bench: leaderboard

Enterprise-agent benchmark that splits a workspace into three knowledge surfaces: retrievable documents, structured tables and a dependency graph: and asks whether an agent routes each of 1,151 atomic questions to the right one, acquires the evidence and answers correctly. The published score is the all-tools ReAct setting, where every surface is exposed and the model has to choose; the aggregate weights answer 0.35, evidence 0.30, routing 0.25 and efficiency 0.10.

Metric: Aggregate Score (0-1). Source: huggingface.co. Status: saturation imminent. 4 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)0.76
2DeepSeek V4 Pro0.69
3GPT-5.50.69
4GPT-4o Mini0.58

Interactive version: theaggregate.ai/benchmark?slug=worksurface-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.