MemLens (64K) - Information Extraction: leaderboard

Metric: Accuracy (%) on the information extraction (IE, 246 questions) at a 64K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 27 models tracked.

Top models

#ModelScore
1Qwen 3.5 122B A10B67.07
2GPT-5.463.01
3Qwen 3.5 27B62.2
4GLM-4.6V61.79
5Qwen 3.5 9B60.57
6Gemini 3.1 Pro (Preview)58.94
7Qwen 3 VL 30B A3B Instruct56.91
8Qwen 3 VL 235B A22B Instruct56.1
9Nemotron Nano 12B v2 VL56.1
10Kimi K2.552.44
11Qwen 3.5 4B49.19
12GLM-4.5V45.93
13Qwen 3 VL 8B Instruct44.72
14Qwen 3 VL 4B Instruct42.28
15Qwen 3 VL 8B (Thinking)42.28

Interactive version: theaggregate.ai/benchmark?slug=memlens-64k-information-extraction · How It Works · Data refreshed daily, snapshot 2026-10-07.