AMemGym: leaderboard

Metric: Normalized memory score (times 100): overall question-answering accuracy rescaled between a random-guess baseline (0) and the same model's upper bound when given the ground-truth user states (100), averaged over evaluation periods, on the AMemGym base configuration: 20 synthetic user profiles, each with 10 multiple-choice personalization questions (4 to 7 options) asked after each of 11 interaction periods while an LLM-simulated user (GPT-4.1) reveals evolving user states over about 47 turns; the native LLM keeps the whole history in context; temperature 0; on-policy: the model's own conversation with the simulated user is the history; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 12 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 433.6#194
2Gemini 2.5 Flash32.7#237
3Gemini 2.5 Flash Lite26.9#413
4GLM-4.625.7#246
5GPT-4.124.4#240
6Gemini 2.0 Flash24.4#331
7Kimi K223.4#236
8GPT-4.1 Mini20.3#346
9DeepSeek V315.2#312
10DeepSeek V3.1 Terminus15.1#212
11GPT-4o Mini14.9#588
12Qwen 3 235B A22B 2507 Instruct14.8#291

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=amemgym · How It Works · Data refreshed daily, snapshot 2026-10-11.