MemLens (64K) - Multi-Session Reasoning: leaderboard

Metric: Accuracy (%) on the multi-session reasoning (MSR, 143 questions) at a 64K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 27 models tracked.

Top models

#ModelScore
1Kimi K2.535.66
2Gemini 3.1 Pro (Preview)31.47
3GPT-5.427.27
4Qwen 3.5 122B A10B27.27
5GLM-4.6V23.78
6Gemma 3 27B (IT)22.38
7Gemma 3 12B (IT)22.38
8Qwen 3.5 27B20.98
9Qwen 3 VL 235B A22B Instruct20.98
10Phi-4 Multimodal Instruct20.98
11Nemotron Nano 12B v2 VL20.98
12Qwen 3 VL 30B A3B Instruct20.28
13Qwen 3 VL 8B Instruct19.58
14Qwen 3 VL 4B Instruct19.58
15Gemma 3 4B (IT)18.88

Interactive version: theaggregate.ai/benchmark?slug=memlens-64k-multi-session-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.