6G-Bench - AI-Native Networking and Agentic Control: leaderboard

Metric: Accuracy (%) on the AI-native networking and agentic control tasks (T11, T17-T20, T27) of 6G-Bench's 3,722 expert-validated four-option multiple-choice questions on network-level semantic reasoning for AI-native 6G networks, deterministic single-shot answers (temperature 0, one letter in a JSON object), group score is the unweighted mean of its task accuracies; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1Llama 4 Maverick85.5#451
2GPT-5.2 Codex84#89
3Qwen3 Coder Next83.8#321
4GPT-5.2 Instant83.8#205
5Olmo 3.1 32B Instruct81.1#754
6DeepSeek V3.2 Exp80.7#227
7Qwen 3 235B A22B 2507 Instruct79.9#291
8DeepSeek V3.279.8#198
9Phi-479.5#701
10GPT-4o Mini79.3#588
11Hermes 4 70B79.3#489
12Claude Haiku 4.579#271
13Ministral 3 14B78.7#636
14Qwen 3 VL 32B Instruct78.1#276
15GPT-5 Mini77.5#176

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=6g-bench-ai-native-networking-and-agentic-control · How It Works · Data refreshed daily, snapshot 2026-10-11.