ImplicitMemBench - Classical Conditioning: leaderboard

Metric: First-action accuracy (%) on the 100 classical-conditioning items: after stimulus and consequence pairings and unrelated turns, the first action on the conditioned stimulus is judged by GPT-4o-mini for learned protective behavior (learning, interference and test phases, first-attempt scoring, mean of three runs at temperature 0); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 17 models tracked.

Top models

#ModelScore
1DeepSeek R169.67
2Qwen 3 32B67
3GPT-564
4Qwen 3 8B64
5O4 Mini (High)60
6O357.67
7GLM-4.553.33
8Claude Sonnet 451.67
9Gemini 2.5 Flash49
10Gemini 2.5 Pro47.33
11Llama 3.3 70B Instruct47.33
12Qwen 2.5 72B Instruct47
13GPT-4o Mini44
14GPT-4o43.67
15Claude Opus 4.141.67

Interactive version: theaggregate.ai/benchmark?slug=implicitmembench-classical-conditioning · How It Works · Data refreshed daily, snapshot 2026-10-07.