EntLORE - Cross-Source Composition (Gold Documents): leaderboard

Metric: Answer accuracy (%; L2, 204 questions composing facts stated in several sources; gold-document agentic reference: the gold document paths read through the same agent tool loop; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B95.9
2GLM-5.295.3
3GPT-5.495.2
4DeepSeek V4 Pro (Reasoning)95.2
5Kimi K2.694.9
6DeepSeek V4 Flash (Reasoning)94.7
7Claude Sonnet 4.6 (Thinking)94.5
8GPT-5.4 Mini84.9

Interactive version: theaggregate.ai/benchmark?slug=entlore-cross-source-composition-gold-documents · How It Works · Data refreshed daily, snapshot 2026-09-29.