MemLens (32K) - Multi-Session Reasoning: leaderboard

Metric: Accuracy (%) on the multi-session reasoning (MSR, 143 questions) at a 32K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 27 models tracked.

Top models

#ModelScore
1Kimi K2.544.06
2Gemini 3.1 Pro (Preview)32.17
3Qwen 3.5 122B A10B30.07
4Qwen 3.5 27B29.37
5Qwen 3 VL 235B A22B (Thinking)25.17
6Gemma 3 27B (IT)23.78
7Gemma 3 12B (IT)22.38
8Gemma 3 4B (IT)21.68
9Qwen 3 VL 4B (Thinking)20.98
10Phi-4 Multimodal Instruct20.28
11GLM-4.6V20.28
12Qwen 3 VL 235B A22B Instruct18.88
13Qwen 3 VL 30B A3B Instruct18.88
14Qwen 3 VL 8B Instruct17.48
15Qwen 3.5 2B17.48

Interactive version: theaggregate.ai/benchmark?slug=memlens-32k-multi-session-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.