FanOutQA — leaderboard

Metric: Accuracy (%, closedbook loose). Source: fanoutqa.com. 16 models tracked.

Top models

#ModelScore
1Mistral Large 2 (Jul)47.45
2Mixtral 8x7B Instruct46.96
3Llama 3 70B Instruct46.62
4GPT-OSS-120B46.6
5GPT-4 Turbo45.97
6Mistral-Small-Instruct-240945.01
7Claude 3 Opus44.77
8GPT-4o44.12
9Llama 2 70B Chat44.03
10Mistral 7B Instruct42.71
11GPT-OSS-20B39.96
12GPT-3.5 Turbo39.8
13GPT-435.5
14Claude 2.134.09
15Gemma 1.1 7B (IT)18.65

Interactive version: theaggregate.ai/benchmark?slug=fanoutqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.