RefMem-Bench - Direct-Answer: leaderboard
Metric: Answer accuracy (%; short free-form answers judged equivalent to a reference by a GPT-4o judge; mean of the eight reflective-memory dimensions over the 20,279-question test split; dialogue history with image captions in context, temperature 0). Source: arxiv.org. Saturation forecast: Around April 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 VL 8B | 21.1 |
| 2 | DeepSeek R1 Distill Qwen 14B | 19.1 |
Interactive version: theaggregate.ai/benchmark?slug=refmem-bench-direct-answer · How It Works · Data refreshed daily, snapshot 2026-09-26.