6G-Bench: leaderboard

Metric: Accuracy (%) on the full set of 6G-Bench's 3,722 expert-validated four-option multiple-choice questions on network-level semantic reasoning for AI-native 6G networks, deterministic single-shot answers (temperature 0, one letter in a JSON object); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScoreOverall rank
1Llama 4 Maverick82.9#451
2Qwen3 Coder Next81.8#321
3Ministral 3 14B79.5#636
4GPT-5.2 Instant79.4#205
5GPT-5.2 Codex79#89
6DeepSeek V3.2 Exp78.9#227
7Olmo 3.1 32B Instruct78.3#754
8Ministral 3 8B78.3#676
9Claude Haiku 4.576.4#271
10Hermes 4 70B76.2#489
11GPT-4.1 Nano72.6#716
12Phi-469.3#701
13Llama 3.1 8B Instruct66.6#1018
14lfm-2.2-6B65.9
15Qwen 2.5 7B Instruct63.2#846

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=6g-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.