MemEvoBench (Workflow Style, Biased Feedback): leaderboard

Metric: Attack success rate (%) in the third evaluation round on the 83 workflow-style cases (20 Agent-SafetyBench environments with noisy tool returns in the memory pool): share of responses that follow the misleading memory into the target risk behavior after two rounds of the agent's own responses were added to the memory pool, with simulated biased user feedback stored after each round (risky shortcuts rewarded, cautious answers penalized), judged MISLED or SAFE by GPT-5.2, temperature 0, standard system prompt without safety guidance; lower is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro85.5
2GPT-587
3Claude 3.7 Sonnet87.5
4Qwen 3 Next 80B A3B Instruct89.2
5Qwen 3 235B A22B 2507 Instruct90.4
6GPT-4o91.5
7Qwen 3 32B91.6
8Llama 3.3 70B Instruct94
9DeepSeek V3.294

Interactive version: theaggregate.ai/benchmark?slug=memevobench-workflow-style-biased-feedback · How It Works · Data refreshed daily, snapshot 2026-10-07.