EntLORE - Latent Organizational Reasoning (Gold Documents): leaderboard

Metric: Answer accuracy (%; L3, 234 questions whose target relation is stated in no document; gold-document agentic reference: the gold document paths read through the same agent tool loop; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (Thinking)76.3
2DeepSeek V4 Pro (Reasoning)73.5
3DeepSeek V4 Flash (Reasoning)71.3
4GLM-5.270.9
5Kimi K2.670.6
6GPT-5.469.5
7Qwen 3.5 397B A17B62.5
8GPT-5.4 Mini62.1

Interactive version: theaggregate.ai/benchmark?slug=entlore-latent-organizational-reasoning-gold-documents · How It Works · Data refreshed daily, snapshot 2026-09-29.