LongProc — leaderboard

Long-context procedural generation benchmark testing models on generating long, structured outputs following complex procedural instructions.

Metric: Accuracy @ 0.5K (%). Source: princeton-pli.github.io. Status: saturated. 17 models tracked.

Top models

#ModelScore
1GPT-4o (2024-08-06)94.8
2Gemini 1.5 Pro (001)89.2
3Gemini 1.5 Flash (001)78.9
4Claude 3.5 Sonnet (20241022)78.4
5Llama 3.1 70B Instruct72.9
6Qwen 2.5 72B Instruct68.7
7Llama 3.1 8B Instruct26.9
8Llama 3.2 3B Instruct13.5
9Llama 3.2 1B Instruct4

Interactive version: theaggregate.ai/benchmark?slug=longproc · How the rankings work · Data refreshed daily, snapshot 2026-07-22.