FAB-Bench (16K Context): leaderboard

Metric: Mean of six GPT-4.1-mini G-Eval scores (0-1: factuality, technical depth, completeness, relevance, context utilization, support quality) over 200 questions, AnythingLLM retrieval with a 16K-token context window and 4K output tokens. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1DeepSeek V3.2 Exp0.78
2Gemini 2.5 Flash0.73
3Qwen 2.5 72B Instruct0.69

Interactive version: theaggregate.ai/benchmark?slug=fab-bench-16k-context · How It Works · Data refreshed daily, snapshot 2026-09-25.