FMNB Leaderboard — leaderboard

Fiction Multi-Needle Benchmark: long-context evaluation testing models on retrieving and synthesizing information from multiple locations in fiction texts at 8K-32K context lengths.

Metric: Score. Source: huggingface.co. Status: saturated. 49 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct100
2Gemma 4 31B (IT)100
3Qwen 3.5 27B100
4GLM 4.5 Air99
5Mistral Small 3.293
6Magistral Small88
7Llama 4 Scout Instruct83
8Mistral Nemo Instruct (2407)55
9MN-12B-Mag-Mell-R153

Interactive version: theaggregate.ai/benchmark?slug=fmnb-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.