LongPIBench - Email Summary: leaderboard

Metric: Attack success rate (%; share of 100 long synthetic email threads to summarize with a draft reply; success when the summary contains the attacker's link, under the authority-spoofing prompt injection embedded in the context and the default attack goal, no defense; default inference settings, up to 20,000 output tokens (32,768 for Qwen3-8B)). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1GPT-4.168
2GPT-4o69
3Llama 3.1 8B Instruct70
4deepseek-llm-7B-chat79
5Llama 3.2 3B Instruct83
6Qwen 3 8B97

Interactive version: theaggregate.ai/benchmark?slug=longpibench-email-summary · How It Works · Data refreshed daily, snapshot 2026-09-29.