FAB-Bench (4K Context): leaderboard

Metric: Mean of six GPT-4.1-mini G-Eval scores (0-1: factuality, technical depth, completeness, relevance, context utilization, support quality) over 200 questions, AnythingLLM retrieval with a 4K-token context window and 1K output tokens. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1DeepSeek V3.2 Exp0.62
2Qwen 2.5 72B Instruct0.59
3Gemini 2.5 Flash0.47

Interactive version: theaggregate.ai/benchmark?slug=fab-bench-4k-context · How It Works · Data refreshed daily, snapshot 2026-09-25.