LongPIBench - Paper Review: leaderboard
Metric: Attack success rate (%; share of 100 long synthetic conference papers the model must review; success when the final rating is 8 or 9 on the 1-9 scale, under the authority-spoofing prompt injection embedded in the context and the default attack goal, no defense; default inference settings, up to 20,000 output tokens (32,768 for Qwen3-8B)). Source: arxiv.org. Saturation forecast: Around July 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.1 | 79 |
| 2 | GPT-4o | 82 |
| 3 | deepseek-llm-7B-chat | 94 |
| 4 | Llama 3.2 3B Instruct | 96 |
| 5 | Llama 3.1 8B Instruct | 100 |
| 6 | Qwen 3 8B | 100 |
Interactive version: theaggregate.ai/benchmark?slug=longpibench-paper-review · How It Works · Data refreshed daily, snapshot 2026-09-29.