EntLORE - Cross-Source Composition (LLM Wiki): leaderboard

Metric: Answer accuracy (%; L2, 204 questions composing facts stated in several sources; LLM wiki: the corpus compiled offline into a navigable wiki read through the same agent loop; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Around June 2027. 8 models tracked.

Top models

#ModelScore
1DeepSeek V4 Flash (Reasoning)59.6
2DeepSeek V4 Pro (Reasoning)56.9
3GLM-5.253.2
4GPT-5.452.5
5Claude Sonnet 4.6 (Thinking)51.5
6Kimi K2.649.7
7Qwen 3.5 397B A17B49.3
8GPT-5.4 Mini47.1

Interactive version: theaggregate.ai/benchmark?slug=entlore-cross-source-composition-llm-wiki · How It Works · Data refreshed daily, snapshot 2026-09-29.