EntLORE - Latent Organizational Reasoning (LLM Wiki): leaderboard

Metric: Answer accuracy (%; L3, 234 questions whose target relation is stated in no document; LLM wiki: the corpus compiled offline into a navigable wiki read through the same agent loop; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Around April 2028. 8 models tracked.

Top models

#ModelScore
1GPT-5.435.5
2Claude Sonnet 4.6 (Thinking)34.6
3GLM-5.230.8
4DeepSeek V4 Flash (Reasoning)30.6
5DeepSeek V4 Pro (Reasoning)29.6
6GPT-5.4 Mini26.5
7Qwen 3.5 397B A17B20.4
8Kimi K2.614

Interactive version: theaggregate.ai/benchmark?slug=entlore-latent-organizational-reasoning-llm-wiki · How It Works · Data refreshed daily, snapshot 2026-09-29.