VitaBench 2.0 (Agentic Memory): leaderboard
Metric: Avg@4 (%, times 100): mean task success over four runs at temperature 0 of VitaBench 2.0 personalized agentic tasks (temporally ordered per-user task sequences across domains, tool-calling agent in an executable environment, preferences scattered over fragmented interaction histories, gpt-4.1-2025-04-14 as user simulator and rubric evaluator), with the interaction history kept in an agentic memory (MemAgent) that the model itself updates instead of the full context; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 27 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 (Thinking) | 45.4 |
| 2 | DeepSeek V4 Pro | 44.9 |
| 3 | Seed 2.0 Pro | 42.8 |
| 4 | DeepSeek V4 Pro (Non-reasoning) | 42.7 |
| 5 | Seed 2.0 Pro (Non-reasoning) | 42.6 |
| 6 | GLM-5.1 (Non-reasoning) | 42.3 |
| 7 | GPT-5 | 42.1 |
| 8 | DeepSeek R1 0528 | 41.2 |
| 9 | O3 | 40.1 |
| 10 | Claude Sonnet 4.5 (Thinking) | 39.7 |
| 11 | Kimi K2.6 (Non-reasoning) | 39.7 |
| 12 | Seed-1.6 | 38.3 |
| 13 | Gemini 2.5 Pro | 37.8 |
| 14 | GLM-5.1 | 35.2 |
| 15 | MiniMax-M2.7 | 35.1 |
Interactive version: theaggregate.ai/benchmark?slug=vitabench-2-0-agentic-memory · How It Works · Data refreshed daily, snapshot 2026-10-07.