WorkSurface-Bench: leaderboard
Enterprise-agent benchmark that splits a workspace into three knowledge surfaces: retrievable documents, structured tables and a dependency graph: and asks whether an agent routes each of 1,151 atomic questions to the right one, acquires the evidence and answers correctly. The published score is the all-tools ReAct setting, where every surface is exposed and the model has to choose; the aggregate weights answer 0.35, evidence 0.30, routing 0.25 and efficiency 0.10.
Metric: Aggregate Score (0-1). Source: huggingface.co. Status: saturation imminent. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 0.76 |
| 2 | DeepSeek V4 Pro | 0.69 |
| 3 | GPT-5.5 | 0.69 |
| 4 | GPT-4o Mini | 0.58 |
Interactive version: theaggregate.ai/benchmark?slug=worksurface-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.