EntLORE - Explicit Lookup (Gold Documents): leaderboard

Metric: Answer accuracy (%; L1, 469 questions answered by a stated fact; gold-document agentic reference: the gold document paths read through the same agent tool loop; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1GLM-5.289.2
2GPT-5.488.9
3Claude Sonnet 4.6 (Thinking)88.5
4Kimi K2.688.2
5GPT-5.4 Mini87.5
6DeepSeek V4 Pro (Reasoning)86.4
7DeepSeek V4 Flash (Reasoning)85.6
8Qwen 3.5 397B A17B85.2

Interactive version: theaggregate.ai/benchmark?slug=entlore-explicit-lookup-gold-documents · How It Works · Data refreshed daily, snapshot 2026-09-29.