EntLORE - Latent Organizational Reasoning (Agentic Retrieval): leaderboard

Metric: Answer accuracy (%; L3, 234 questions whose target relation is stated in no document; agentic retrieval: search and fetch tools over a dense index of the released corpus, at most 30 LLM iterations; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1GPT-5.417.5
2DeepSeek V4 Flash (Reasoning)15.1
3Claude Sonnet 4.6 (Thinking)14.9
4GPT-5.4 Mini14.8
5DeepSeek V4 Pro (Reasoning)14.3
6GLM-5.212.3
7Qwen 3.5 397B A17B9.5
8Kimi K2.69.4

Interactive version: theaggregate.ai/benchmark?slug=entlore-latent-organizational-reasoning-agentic-retrieval · How It Works · Data refreshed daily, snapshot 2026-09-29.