MemLens (32K) - Knowledge Update: leaderboard

Metric: Accuracy (%) on the knowledge update (KU, 116 questions) at a 32K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 27 models tracked.

Top models

#ModelScore
1Kimi K2.550.86
2Gemini 3.1 Pro (Preview)49.14
3Qwen 3.5 122B A10B49.14
4GPT-5.447.41
5GLM-4.6V43.97
6Qwen 3.5 9B43.1
7Qwen 3.5 27B42.24
8Qwen 3 VL 235B A22B Instruct40.52
9Nemotron Nano 12B v2 VL39.66
10Qwen 3 VL 30B A3B Instruct37.93
11Qwen 3 VL 235B A22B (Thinking)36.21
12Qwen 3.5 4B34.48
13Qwen 3 VL 8B Instruct33.62
14Qwen 3.5 2B33.62
15Qwen 3 VL 4B (Thinking)33.62

Interactive version: theaggregate.ai/benchmark?slug=memlens-32k-knowledge-update · How It Works · Data refreshed daily, snapshot 2026-10-07.