MemLens (128K) - Information Extraction: leaderboard

Metric: Accuracy (%) on the information extraction (IE, 246 questions) at a 128K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 25 models tracked.

Top models

#ModelScore
1GPT-5.460.16
2Qwen 3 VL 30B A3B Instruct56.5
3Gemini 3.1 Pro (Preview)55.79
4Qwen 3.5 27B54.07
5GLM-4.6V52.85
6Kimi K2.551.63
7Qwen 3 VL 235B A22B Instruct51.22
8Qwen 3.5 9B45.12
9Qwen 3.5 122B A10B43.09
10Qwen 3 VL 30B A3B (Thinking)33.74
11Qwen 3 VL 8B (Thinking)33.33
12Qwen 3 VL 8B Instruct32.52
13Qwen 3.5 4B29.67
14Qwen 3 VL 235B A22B (Thinking)29.27
15Qwen 3 VL 4B Instruct28.46

Interactive version: theaggregate.ai/benchmark?slug=memlens-128k-information-extraction · How It Works · Data refreshed daily, snapshot 2026-10-07.