6G-Bench - Trust, Security and SLA Awareness: leaderboard

Metric: Accuracy (%) on the trust, security and SLA awareness tasks (T9, T10, T26, T30) of 6G-Bench's 3,722 expert-validated four-option multiple-choice questions on network-level semantic reasoning for AI-native 6G networks, deterministic single-shot answers (temperature 0, one letter in a JSON object), group score is the unweighted mean of its task accuracies; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 27 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek V3.283.8#198
2DeepSeek V3.2 Exp81.7#227
3Qwen3 Coder Next81.2#321
4Llama 4 Maverick80.8#451
5Ministral 3 14B78.6#636
6Hermes 4 70B77.5#489
7Qwen 3 VL 32B Instruct77.3#276
8GPT-5.2 Instant77.3#205
9Olmo 3.1 32B Instruct77#754
10Ministral 3 8B76.7#676
11GPT-4o Mini76.3#588
12Phi-476.2#701
13GPT-5.2 Codex76.1#89
14Claude Haiku 4.575.2#271
15GPT-5 Mini70.3#176

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=6g-bench-trust-security-and-sla-awareness · How It Works · Data refreshed daily, snapshot 2026-10-11.