Memora (Reasoning-Enabled Inference) - Recommending: leaderboard
Metric: Aggregated forgetting-aware memory accuracy (0 to 300) on recommending questions (personalized suggestions that respect current preferences), reasoning-enabled decoding with the provider default reasoning allocation, on Memora, simulated multi-session conversations of ten personas whose memories are added, updated and deleted over weeks to quarters, with the full history in context (oldest sessions truncated beyond the context window); Forgetting-Aware Memory Accuracy per question rewards valid memories and penalizes obsolete ones, judged by majority vote of GPT-4.1, Claude Haiku 4.5 and Gemini 2.5 Flash, normalized to 0-100 per duration and summed over the weekly, monthly and quarterly settings; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 (Thinking) | 165.08 |
| 2 | GPT-5.2 | 158.36 |
| 3 | Qwen 3 32B | 150.81 |
| 4 | Gemini 3 Pro (Preview) | 139.11 |
Interactive version: theaggregate.ai/benchmark?slug=memora-reasoning-enabled-inference-recommending · How It Works · Data refreshed daily, snapshot 2026-10-07.