PhysAI-Bench: leaderboard

Metric: Decision accuracy (%; four-option choice of the next action at decision points extracted from autonomous UAV mission traces, with mission context, physical and sensor state, MCP tool responses, A2A messages and 6G network conditions available before the decision; 500 human-verified held-out instances from episodes disjoint from the development set; strict option-label match, parse and inference failures scored wrong; mean of three runs with each model's prompting (0-, 3- or 5-shot) and temperature selected on a 35-instance development set). Source: arxiv.org. Saturation forecast: Around 2033. 30 models tracked.

Top models

#ModelScore
1GPT-5.352
2GPT-5.2 Instant49.4
3Grok 4.549.07
4Qwen 3.7 Plus47.73
5Qwen 3.7 Max46.4
6DeepSeek V3.246.07
7Hermes 4 70B45.87
8Claude Haiku 4.544.8
9Gemma 4 31B (IT)44.67
10Gemma 4 26B A4B (IT)43.87
11Hermes-3-Llama-3.1-405B43.8
12Hermes 4 405B43.67
13Gemma 3 12B (IT)42.47
14Kimi K2.541.87
15Mistral Small 440.8

Interactive version: theaggregate.ai/benchmark?slug=physai-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.