MemLens (64K) - Temporal Reasoning: leaderboard

Metric: Accuracy (%) on the temporal reasoning (TR, 194 questions) at a 64K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 27 models tracked.

Top models

#ModelScore
1Qwen 3 VL 30B A3B Instruct55.15
2Qwen 3 VL 235B A22B Instruct54.12
3GLM-4.6V53.09
4Kimi K2.552.06
5Qwen 3.5 122B A10B51.55
6Qwen 3 VL 8B (Thinking)50.52
7Qwen 3 VL 8B Instruct47.42
8Gemma 3 27B (IT)44.33
9Gemma 3 12B (IT)44.33
10Qwen 3 VL 235B A22B (Thinking)42.78
11Nemotron Nano 12B v2 VL42.78
12Gemini 3.1 Pro (Preview)42.27
13GPT-5.442.27
14Phi-4 Multimodal Instruct39.18
15Qwen 3 VL 30B A3B (Thinking)37.11

Interactive version: theaggregate.ai/benchmark?slug=memlens-64k-temporal-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.