FAB-Bench (4K Context): leaderboard
Metric: Mean of six GPT-4.1-mini G-Eval scores (0-1: factuality, technical depth, completeness, relevance, context utilization, support quality) over 200 questions, AnythingLLM retrieval with a 4K-token context window and 1K output tokens. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3.2 Exp | 0.62 |
| 2 | Qwen 2.5 72B Instruct | 0.59 |
| 3 | Gemini 2.5 Flash | 0.47 |
Interactive version: theaggregate.ai/benchmark?slug=fab-bench-4k-context · How It Works · Data refreshed daily, snapshot 2026-09-25.