6G-Bench - Intent and Policy Reasoning: leaderboard

Metric: Accuracy (%) on the intent and policy reasoning tasks (T1, T2, T3, T12, T15) of 6G-Bench's 3,722 expert-validated four-option multiple-choice questions on network-level semantic reasoning for AI-native 6G networks, deterministic single-shot answers (temperature 0, one letter in a JSON object), group score is the unweighted mean of its task accuracies; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1Qwen3 Coder Next88.6#321
2Llama 4 Maverick88.1#451
3GPT-5.2 Instant87.8#205
4DeepSeek V3.287.6#198
5Ministral 3 14B87#636
6GPT-5.2 Codex86.8#89
7DeepSeek V3.2 Exp86.3#227
8Olmo 3.1 32B Instruct86.1#754
9Hermes 4 70B86#489
10Claude Haiku 4.585.3#271
11Ministral 3 8B85.3#676
12Qwen 3 VL 32B Instruct85.2#276
13Phi-485.1#701
14Qwen 3 235B A22B 2507 Instruct84.5#291
15GPT-4o Mini84.4#588

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=6g-bench-intent-and-policy-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.