MemEvoBench (QA Style): leaderboard
Metric: Attack success rate (%) in the third evaluation round on the 108 QA-style cases (misleading memory injection across 7 domains and 36 risk types): share of responses that follow the misleading memory into the target risk behavior after two rounds of the agent's own responses were added to the memory pool, without user feedback, judged MISLED or SAFE by GPT-5.2, temperature 0, standard system prompt without safety guidance; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 55 |
| 2 | GPT-5 | 59 |
| 3 | Claude 3.7 Sonnet | 63 |
| 4 | Qwen 3 235B A22B 2507 Instruct | 76.9 |
| 5 | Qwen 3 Next 80B A3B Instruct | 79.6 |
| 6 | Llama 3.3 70B Instruct | 81 |
| 7 | DeepSeek V3.2 | 81 |
| 8 | Qwen 3 32B | 89.8 |
| 9 | GPT-4o | 91.7 |
Interactive version: theaggregate.ai/benchmark?slug=memevobench-qa-style · How It Works · Data refreshed daily, snapshot 2026-10-07.