PASB - Short-Term Memory Extraction: leaderboard
Metric: Extraction success rate (%): share of the 40 cases in which the attacker retrieves the specified short-term context fragment (PASB direct prompt injection against the deployed OpenClaw personal agent's memory, 40 cases per task in a sandboxed testbed with canary markers; no defense; mean over repeated trials); lower is better. Source: arxiv.org. 3 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Qwen 2.5 7B Instruct | 33.5 | #846 |
| 2 | GPT-4o Mini | 38.2 | #588 |
| 3 | Llama 3.1 70B Instruct | 41 | #548 |
Interactive version: theaggregate.ai/benchmark?slug=pasb-short-term-memory-extraction · How It Works · Data refreshed daily, snapshot 2026-10-11.