MemLens (32K) - Temporal Reasoning: leaderboard

Metric: Accuracy (%) on the temporal reasoning (TR, 194 questions) at a 32K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 27 models tracked.

Top models

#ModelScore
1Qwen 3 VL 30B A3B Instruct60.82
2Qwen 3 VL 8B Instruct58.25
3Qwen 3 VL 235B A22B Instruct55.67
4Kimi K2.553.09
5GLM-4.6V53.09
6Qwen 3 VL 8B (Thinking)52.58
7Qwen 3.5 122B A10B51.55
8Qwen 3 VL 235B A22B (Thinking)49.48
9Nemotron Nano 12B v2 VL44.85
10Gemma 3 12B (IT)43.81
11Phi-4 Multimodal Instruct43.81
12Qwen 3 VL 4B (Thinking)43.3
13Qwen 3 VL 4B Instruct42.78
14Gemma 3 4B (IT)41.75
15Gemma 3 27B (IT)41.24

Interactive version: theaggregate.ai/benchmark?slug=memlens-32k-temporal-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.