DoRA (Defense Documents) - Completeness (GTE Retrieval): leaderboard

Metric: RAGEval keypoint completeness (%) of grounded answers to the 1,259-question DoRA defense-document test split with the top-3 passages retrieved by GTE (gte-multilingual-base) as context: per-question share of the reference answer's keypoints the answer covers correctly, judged by GPT-4o-mini and averaged over questions; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Claude 3.5 Haiku70.61
2GPT-4o69.18
3Ministral-8B-Instruct-241067.99
4GPT-4o Mini67.81
5Llama 3.1 8B Instruct67.43
6Qwen 2.5 7B Instruct66.98
7GPT-3.5 Turbo65.87
8Llama 3 8B Instruct65.41

Interactive version: theaggregate.ai/benchmark?slug=dora-defense-documents-completeness-gte-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.