PhysAI-Bench: leaderboard
Metric: Decision accuracy (%; four-option choice of the next action at decision points extracted from autonomous UAV mission traces, with mission context, physical and sensor state, MCP tool responses, A2A messages and 6G network conditions available before the decision; 500 human-verified held-out instances from episodes disjoint from the development set; strict option-label match, parse and inference failures scored wrong; mean of three runs with each model's prompting (0-, 3- or 5-shot) and temperature selected on a 35-instance development set). Source: arxiv.org. Saturation forecast: Around 2033. 30 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.3 | 52 |
| 2 | GPT-5.2 Instant | 49.4 |
| 3 | Grok 4.5 | 49.07 |
| 4 | Qwen 3.7 Plus | 47.73 |
| 5 | Qwen 3.7 Max | 46.4 |
| 6 | DeepSeek V3.2 | 46.07 |
| 7 | Hermes 4 70B | 45.87 |
| 8 | Claude Haiku 4.5 | 44.8 |
| 9 | Gemma 4 31B (IT) | 44.67 |
| 10 | Gemma 4 26B A4B (IT) | 43.87 |
| 11 | Hermes-3-Llama-3.1-405B | 43.8 |
| 12 | Hermes 4 405B | 43.67 |
| 13 | Gemma 3 12B (IT) | 42.47 |
| 14 | Kimi K2.5 | 41.87 |
| 15 | Mistral Small 4 | 40.8 |
Interactive version: theaggregate.ai/benchmark?slug=physai-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.