ImplicitMemBench: leaderboard

Metric: Overall score: mean of procedural-memory accuracy, classical-conditioning accuracy (%) and the GPT-4o-mini-judged priming influence score (0 to 95 bands) on the 300 ImplicitMemBench items (learning, interference and test phases, first-attempt scoring, mean of three runs at temperature 0); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 17 models tracked.

Top models

#ModelScore
1DeepSeek R165.3
2Qwen 3 32B64.13
3GPT-563
4Qwen 3 8B62.35
5O361.79
6O4 Mini (High)60.87
7GLM-4.558.59
8Gemini 2.5 Pro55.69
9Claude Opus 4.155.65
10Gemini 2.5 Flash55.43
11GPT-4o Mini50.88
12Qwen 2.5 72B Instruct50.78
13GPT-4o50.32
14Claude Sonnet 449.84
15Llama 3.3 70B Instruct49.44

Interactive version: theaggregate.ai/benchmark?slug=implicitmembench · How It Works · Data refreshed daily, snapshot 2026-10-07.