PAUSE - Data and Log Tracking (Hard): leaderboard
Metric: Task completion (%; 57 hard tasks needing multi-step service execution and configuration changes that only the user can make, share of task targets met as judged by Gemini-3-Flash; multi-turn tasks in a simulated personal service environment with 50 assistant tools, user-controlled permissions and system configurations and a Gemini-3-Flash user simulator; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 59.12 |
| 2 | Gemini 3 Pro | 48.39 |
| 3 | GPT-5 | 47.34 |
| 4 | GPT-5 Mini | 43.84 |
| 5 | DeepSeek V3.2 (Thinking) | 35.11 |
| 6 | Gemini 2.5 Pro | 19.33 |
| 7 | GPT-4.1 | 17.56 |
| 8 | DeepSeek V3.2 (Non-reasoning) | 14.06 |
| 9 | GPT-4.1 Mini | 10.53 |
Interactive version: theaggregate.ai/benchmark?slug=pause-data-and-log-tracking-hard · How It Works · Data refreshed daily, snapshot 2026-09-29.