EntLORE - Explicit Lookup (Agentic Retrieval): leaderboard

Metric: Answer accuracy (%; L1, 469 questions answered by a stated fact; agentic retrieval: search and fetch tools over a dense index of the released corpus, at most 30 LLM iterations; 907 questions over 2,341 anonymized enterprise documents; programmatic scoring for entity, set, count and ordered answers, atomic-claim entailment judged by Claude Opus 4.8 for free-form answers). Source: arxiv.org. Saturation forecast: Around August 2027. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (Thinking)54.7
2GLM-5.251
3DeepSeek V4 Flash (Reasoning)49.6
4GPT-5.449.1
5DeepSeek V4 Pro (Reasoning)48.4
6Kimi K2.646.5
7Qwen 3.5 397B A17B46
8GPT-5.4 Mini42.5

Interactive version: theaggregate.ai/benchmark?slug=entlore-explicit-lookup-agentic-retrieval · How It Works · Data refreshed daily, snapshot 2026-09-29.