Episodic Memory — leaderboard

Tests how well LLMs encode, store, and recall episodic events across long narratives (200 chapters, ~100K tokens, 686 Q&A pairs). Measures simple recall and chronological awareness.

Metric: Simple Recall (%). Source: github.com. Status: saturation imminent. 21 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro96.8
2Gemini 2.5 Flash96
3GPT-594.2
4GPT-5 Mini83
5Claude Sonnet 479
6Grok 4 Fast (Reasoning)72.6
7Gemini 2.0 Flash (Thinking)70.8
8GPT-4o67
9Grok 4 Fast (Non-reasoning)60.2
10DeepSeek V360
11Gemini 2.0 Flash59.6
12DeepSeek R157.2
13Llama 3.1 405B50.4
14GPT-4o Mini49.2
15Claude 3.5 Sonnet47

Interactive version: theaggregate.ai/benchmark?slug=episodic-memory · How the rankings work · Data refreshed daily, snapshot 2026-07-22.