VitaBench — leaderboard
66-tool agent benchmark across food delivery, retail, and travel domains with 100 cross-scenario and 300 single-scenario tasks. Even the best model reaches only 62% on single-scenario and 32.5% on cross-scenario.
Metric: Cross-Scenario Avg@4 (%). Source: vitabench.github.io. Status: saturation imminent. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (High) | 32.5 |
| 2 | Gemini 3 Pro (High) | 31.5 |
| 3 | LongCat-Flash-Thinking-2601 | 29.3 |
| 4 | Claude Opus 4.5 | 28.5 |
| 5 | O3 (High) | 26.3 |
| 6 | GPT-5.2 (xHigh) | 24.3 |
| 7 | DeepSeek V3.2 | 24 |
| 8 | Claude Sonnet 4.5 | 23.5 |
| 9 | LongCat-Flash-Chat | 22.8 |
| 10 | O4 Mini (High) | 19.5 |
| 11 | GLM-4.7 | 18.3 |
| 12 | Qwen 3 235B A22B (Thinking) | 14.5 |
| 13 | Qwen 3 Max | 14.3 |
| 14 | Seed 1.8 | 13.8 |
| 15 | Kimi K2 (Thinking) | 12.8 |
Interactive version: theaggregate.ai/benchmark?slug=vitabench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.