AuthMem-Bench (Task Success Rate): leaderboard

Metric: Task success rate (%; share of 350 authorized counterpart cases in which the action model issues the required tool call; washed focal memory with no authority metadata). Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini56.6
2Gemini 3.1 Pro (Preview)52.9
3GPT-5.552
4Gemini 3.5 Flash51.7
5GLM-5.250.9
6Qwen 3.7 Max46
7DeepSeek V4 Pro35.7

Interactive version: theaggregate.ai/benchmark?slug=authmem-bench-task-success-rate · How It Works · Data refreshed daily, snapshot 2026-09-26.