VitaBench 2.0 (RAG Memory): leaderboard
Metric: Avg@4 (%, times 100): mean task success over four runs at temperature 0 of VitaBench 2.0 personalized agentic tasks (temporally ordered per-user task sequences across domains, tool-calling agent in an executable environment, preferences scattered over fragmented interaction histories, gpt-4.1-2025-04-14 as user simulator and rubric evaluator), with the interaction history kept in a vector-retrieval (RAG) memory bank queried per task instead of the full context; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 27 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4 Pro | 43 |
| 2 | Claude Opus 4.6 (Thinking) | 43 |
| 3 | DeepSeek V4 Pro (Non-reasoning) | 42.4 |
| 4 | GPT-5 | 41 |
| 5 | Seed 2.0 Pro (Non-reasoning) | 40.6 |
| 6 | DeepSeek R1 0528 | 39 |
| 7 | Kimi K2.6 (Non-reasoning) | 38.3 |
| 8 | GLM-5.1 (Non-reasoning) | 38.3 |
| 9 | Seed-1.6 | 37.5 |
| 10 | Claude Sonnet 4.5 (Thinking) | 37.4 |
| 11 | O3 | 36.2 |
| 12 | Seed 2.0 Pro | 33.9 |
| 13 | GLM-4.6 | 33.6 |
| 14 | GLM-4.5 | 33.6 |
| 15 | GLM-5.1 | 32.8 |
Interactive version: theaggregate.ai/benchmark?slug=vitabench-2-0-rag-memory · How It Works · Data refreshed daily, snapshot 2026-10-07.