AMemGym - Overall Accuracy: leaderboard

Metric: Overall score (times 100): mean multiple-choice question-answering accuracy over all evaluation periods on the AMemGym base configuration: 20 synthetic user profiles, each with 10 multiple-choice personalization questions (4 to 7 options) asked after each of 11 interaction periods while an LLM-simulated user (GPT-4.1) reveals evolving user states over about 47 turns; the native LLM keeps the whole history in context; temperature 0; on-policy; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash44.8#237
2GLM-4.640.4#246
3Kimi K238.9#236
4Qwen 3 235B A22B 2507 Instruct33.1#291
5DeepSeek V3.1 Terminus32.7#212

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=amemgym-overall-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.