PredicateLongBench - Unary For-All: leaderboard
Metric: Accuracy (%; share of the 100 synthetic word lists of about 128K tokens whose answer is exactly the one word sequence satisfying the predicate, retrieve every matching run when exactly one exists (universal quantifier); up to 16K output tokens including reasoning, an answer past the cap counts as wrong). Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 (Thinking) | 99 |
| 2 | GPT-5.4 (High) | 93 |
| 3 | GPT-5.4 (Non-reasoning) | 93 |
| 4 | Claude Opus 4.6 | 90 |
| 5 | GLM-5.1 | 73 |
| 6 | Gemini 3.1 Pro (Preview) (High) | 59 |
| 7 | Qwen 3.5 397B A17B | 43 |
| 8 | MiniMax-M2.7 | 30 |
| 9 | Gemini 3.1 Pro (Preview) (Low) | 20 |
Interactive version: theaggregate.ai/benchmark?slug=predicatelongbench-unary-for-all · How It Works · Data refreshed daily, snapshot 2026-09-29.