FAB-Bench (32K Context): leaderboard

Metric: Mean of six GPT-4.1-mini G-Eval scores (0-1: factuality, technical depth, completeness, relevance, context utilization, support quality) over 200 questions, AnythingLLM retrieval with a 32K-token context window and 4K output tokens. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1DeepSeek V3.2 Exp0.88
2Gemini 2.5 Flash0.84
3Qwen 2.5 72B Instruct0.8

Interactive version: theaggregate.ai/benchmark?slug=fab-bench-32k-context · How It Works · Data refreshed daily, snapshot 2026-09-25.