MemEvoBench (QA Style, Biased Feedback): leaderboard

Metric: Attack success rate (%) in the third evaluation round on the 108 QA-style cases (misleading memory injection across 7 domains and 36 risk types): share of responses that follow the misleading memory into the target risk behavior after two rounds of the agent's own responses were added to the memory pool, with simulated biased user feedback stored after each round (risky shortcuts rewarded, cautious answers penalized), judged MISLED or SAFE by GPT-5.2, temperature 0, standard system prompt without safety guidance; lower is better. Source: arxiv.org. Saturation forecast: Around February 2028. 9 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet77
2GPT-578
3Gemini 2.5 Pro80
4Qwen 3 235B A22B 2507 Instruct83.1
5Qwen 3 Next 80B A3B Instruct84.4
6DeepSeek V3.289
7Llama 3.3 70B Instruct90
8GPT-4o94
9Qwen 3 32B94.4

Interactive version: theaggregate.ai/benchmark?slug=memevobench-qa-style-biased-feedback · How It Works · Data refreshed daily, snapshot 2026-10-07.