PAUSE - Shopping: leaderboard
Metric: Task score (%; 60 constrained shopping tasks, state-based check of the purchased product, quantity and size, voucher use and budget, aggregated per task; multi-turn tasks in a simulated personal service environment with 50 assistant tools, user-controlled permissions and system configurations and a Gemini-3-Flash user simulator; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 72.1 |
| 2 | GPT-5 | 69.1 |
| 3 | Gemini 3 Flash | 59 |
| 4 | DeepSeek V3.2 (Thinking) | 55 |
| 5 | GPT-5 Mini | 47.3 |
| 6 | DeepSeek V3.2 (Non-reasoning) | 41.7 |
| 7 | Gemini 2.5 Pro | 37.7 |
| 8 | GPT-4.1 | 19.7 |
| 9 | GPT-4.1 Mini | 18.3 |
Interactive version: theaggregate.ai/benchmark?slug=pause-shopping · How It Works · Data refreshed daily, snapshot 2026-09-29.