MemLens (32K) - Information Extraction: leaderboard

Metric: Accuracy (%) on the information extraction (IE, 246 questions) at a 32K-token history, interleaved multimodal multi-session conversation history with the evidence sessions hidden among haystack sessions, judged correct or incorrect by a Qwen3-VL-235B-A22B-Instruct judge; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 27 models tracked.

Top models

#ModelScore
1Qwen 3.5 122B A10B74.39
2Qwen 3.5 27B70.33
3GPT-5.469.51
4Qwen 3.5 9B67.48
5Qwen 3 VL 30B A3B Instruct65.45
6Qwen 3.5 2B64.23
7Nemotron Nano 12B v2 VL63.82
8GLM-4.5V61.79
9Qwen 3 VL 235B A22B Instruct60.98
10GLM-4.6V59.35
11Gemini 3.1 Pro (Preview)57.32
12Qwen 3.5 4B56.91
13Qwen 3 VL 8B Instruct52.85
14Qwen 3 VL 4B Instruct52.03
15Qwen 3 VL 235B A22B (Thinking)51.63

Interactive version: theaggregate.ai/benchmark?slug=memlens-32k-information-extraction · How It Works · Data refreshed daily, snapshot 2026-10-07.