VitaBench 2.0 (RAG Memory): leaderboard

Metric: Avg@4 (%, times 100): mean task success over four runs at temperature 0 of VitaBench 2.0 personalized agentic tasks (temporally ordered per-user task sequences across domains, tool-calling agent in an executable environment, preferences scattered over fragmented interaction histories, gpt-4.1-2025-04-14 as user simulator and rubric evaluator), with the interaction history kept in a vector-retrieval (RAG) memory bank queried per task instead of the full context; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 27 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro43
2Claude Opus 4.6 (Thinking)43
3DeepSeek V4 Pro (Non-reasoning)42.4
4GPT-541
5Seed 2.0 Pro (Non-reasoning)40.6
6DeepSeek R1 052839
7Kimi K2.6 (Non-reasoning)38.3
8GLM-5.1 (Non-reasoning)38.3
9Seed-1.637.5
10Claude Sonnet 4.5 (Thinking)37.4
11O336.2
12Seed 2.0 Pro33.9
13GLM-4.633.6
14GLM-4.533.6
15GLM-5.132.8

Interactive version: theaggregate.ai/benchmark?slug=vitabench-2-0-rag-memory · How It Works · Data refreshed daily, snapshot 2026-10-07.