EpiQAL - Text-Grounded Recall: leaderboard

Metric: Exact Match (%; zero-shot, no CoT, full article). Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScore
1GPT-5 Mini90.5
2Mistral Large 2 (Nov) Instruct (2411)90.1
3Qwen 3 30B A3B 2507 Instruct88.2
4Qwen 3 32B87.2
5Qwen 3 8B81.1
6Llama 3.1 8B Instruct79.8
7GPT-4.1 Nano76.8
8GPT-4o Mini76.6
9Mistral 7B Instruct (v0.3)72.2
10Phi-4 Mini Instruct58.3
11Llama 3.2 3B Instruct36.6

Interactive version: theaggregate.ai/benchmark?slug=epiqal-text-grounded-recall · How It Works · Data refreshed daily, snapshot 2026-09-25.