Memora (Standard Inference) - Remembering: leaderboard

Metric: Aggregated forgetting-aware memory accuracy (0 to 300) on remembering questions (direct recall of stored preferences, activities and goals), standard decoding with no reasoning tokens, on Memora, simulated multi-session conversations of ten personas whose memories are added, updated and deleted over weeks to quarters, with the full history in context (oldest sessions truncated beyond the context window); Forgetting-Aware Memory Accuracy per question rewards valid memories and penalizes obsolete ones, judged by majority vote of GPT-4.1, Claude Haiku 4.5 and Gemini 2.5 Flash, normalized to 0-100 per duration and summed over the weekly, monthly and quarterly settings; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.

Top models

#ModelScore
1GPT-5.2 (Non-reasoning)68.63
2Claude Sonnet 4.568.17
3Qwen 3 32B (Non-reasoning)66.5

Interactive version: theaggregate.ai/benchmark?slug=memora-standard-inference-remembering · How It Works · Data refreshed daily, snapshot 2026-10-07.