AMemGym - Overall Accuracy: leaderboard
Metric: Overall score (times 100): mean multiple-choice question-answering accuracy over all evaluation periods on the AMemGym base configuration: 20 synthetic user profiles, each with 10 multiple-choice personalization questions (4 to 7 options) asked after each of 11 interaction periods while an LLM-simulated user (GPT-4.1) reveals evolving user states over about 47 turns; the native LLM keeps the whole history in context; temperature 0; on-policy; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Flash | 44.8 | #237 |
| 2 | GLM-4.6 | 40.4 | #246 |
| 3 | Kimi K2 | 38.9 | #236 |
| 4 | Qwen 3 235B A22B 2507 Instruct | 33.1 | #291 |
| 5 | DeepSeek V3.1 Terminus | 32.7 | #212 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=amemgym-overall-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.