FMNB Leaderboard — leaderboard
Fiction Multi-Needle Benchmark: long-context evaluation testing models on retrieving and synthesizing information from multiple locations in fiction texts at 8K-32K context lengths.
Metric: Score. Source: huggingface.co. Status: saturated. 49 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.3 70B Instruct | 100 |
| 2 | Gemma 4 31B (IT) | 100 |
| 3 | Qwen 3.5 27B | 100 |
| 4 | GLM 4.5 Air | 99 |
| 5 | Mistral Small 3.2 | 93 |
| 6 | Magistral Small | 88 |
| 7 | Llama 4 Scout Instruct | 83 |
| 8 | Mistral Nemo Instruct (2407) | 55 |
| 9 | MN-12B-Mag-Mell-R1 | 53 |
Interactive version: theaggregate.ai/benchmark?slug=fmnb-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.