PredicateLongBench - Lexicographic Locate (LongBench): leaderboard

Metric: Accuracy (%; share of the 93 LongBench v2 documents under 170K tokens used as word lists whose answer is exactly the one word sequence satisfying the predicate, locate the lexicographically sorted run as long as the longest in the document, no decoys; up to 16K output tokens including reasoning, an answer past the cap counts as wrong). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)61
2GPT-5.4 (High)36
3Claude Opus 4.610
4Gemini 3.1 Pro (Preview) (Low)10
5Gemini 3.1 Pro (Preview) (High)6
6GLM-5.13
7GPT-5.4 (Non-reasoning)1
8Qwen 3.5 397B A17B0
9MiniMax-M2.70

Interactive version: theaggregate.ai/benchmark?slug=predicatelongbench-lexicographic-locate-longbench · How It Works · Data refreshed daily, snapshot 2026-09-29.