TriggerBench - Negative Clean PM Accuracy: leaderboard

Metric: PM accuracy (%) on the 289 negative controls, where the context makes the reminder unnecessary and restraint is correct; base dialogues of about 2.5K tokens; TriggerBench: 1,265 prospective-memory tasks from 488 dialogue blueprints in five dimensions (state tracking, temporal grounding, logical adherence, attention recovery, safe coding); a constraint is stated early in a dialogue and a later trigger turn calls for acting on it without being asked; GPT-4o (T=0) judges whether the reply fulfils the proactive intent; temperature 0.6; higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 9 models tracked.

Top models

#ModelScore
1GPT-4o82.7
2GPT-4.176.82
3Qwen 3 32B71.97
4Qwen 3 235B A22B 2507 Instruct69.9
5Gemma 3 27B (IT)63.67
6Qwen 3 235B A22B 2507 FP8 (Thinking)62.63
7GPT-5.2 (Medium)46.37
8GPT-5.2 (High)43.94
9GPT-5.2 (Non-reasoning)42.9

Interactive version: theaggregate.ai/benchmark?slug=triggerbench-negative-clean-pm-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.