VitaBench — leaderboard

66-tool agent benchmark across food delivery, retail, and travel domains with 100 cross-scenario and 300 single-scenario tasks. Even the best model reaches only 62% on single-scenario and 32.5% on cross-scenario.

Metric: Cross-Scenario Avg@4 (%). Source: vitabench.github.io. Status: saturation imminent. 21 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (High)32.5
2Gemini 3 Pro (High)31.5
3LongCat-Flash-Thinking-260129.3
4Claude Opus 4.528.5
5O3 (High)26.3
6GPT-5.2 (xHigh)24.3
7DeepSeek V3.224
8Claude Sonnet 4.523.5
9LongCat-Flash-Chat22.8
10O4 Mini (High)19.5
11GLM-4.718.3
12Qwen 3 235B A22B (Thinking)14.5
13Qwen 3 Max14.3
14Seed 1.813.8
15Kimi K2 (Thinking)12.8

Interactive version: theaggregate.ai/benchmark?slug=vitabench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.