VitaBench 2.0 (Agentic Memory): leaderboard

Metric: Avg@4 (%, times 100): mean task success over four runs at temperature 0 of VitaBench 2.0 personalized agentic tasks (temporally ordered per-user task sequences across domains, tool-calling agent in an executable environment, preferences scattered over fragmented interaction histories, gpt-4.1-2025-04-14 as user simulator and rubric evaluator), with the interaction history kept in an agentic memory (MemAgent) that the model itself updates instead of the full context; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 27 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)45.4
2DeepSeek V4 Pro44.9
3Seed 2.0 Pro42.8
4DeepSeek V4 Pro (Non-reasoning)42.7
5Seed 2.0 Pro (Non-reasoning)42.6
6GLM-5.1 (Non-reasoning)42.3
7GPT-542.1
8DeepSeek R1 052841.2
9O340.1
10Claude Sonnet 4.5 (Thinking)39.7
11Kimi K2.6 (Non-reasoning)39.7
12Seed-1.638.3
13Gemini 2.5 Pro37.8
14GLM-5.135.2
15MiniMax-M2.735.1

Interactive version: theaggregate.ai/benchmark?slug=vitabench-2-0-agentic-memory · How It Works · Data refreshed daily, snapshot 2026-10-07.