EntLORE - Cross-Source Composition (GraphRAG): leaderboard

Metric: Answer accuracy (%; L2, 204 questions composing facts stated in several sources; corpus-induced GraphRAG: an entity graph and community reports induced offline from the released corpus; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Around May 2027. 8 models tracked.

Top models

#ModelScore
1DeepSeek V4 Flash (Reasoning)51.7
2GPT-5.451.5
3DeepSeek V4 Pro (Reasoning)48.8
4Qwen 3.5 397B A17B46.6
5GLM-5.246.5
6Claude Sonnet 4.6 (Thinking)43.7
7Kimi K2.640.9
8GPT-5.4 Mini32.9

Interactive version: theaggregate.ai/benchmark?slug=entlore-cross-source-composition-graphrag · How It Works · Data refreshed daily, snapshot 2026-09-29.