PilotBench: leaderboard

Metric: Pilot-Score (5-100): 0.6 x regression score plus 0.4 x instruction-following score, on PilotBench (one-step trajectory and attitude prediction from 34-channel telemetry of 708 general-aviation trajectories in nine flight phases, fixed output schema, temperature 0); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 19 models tracked.

Top models

#ModelScore
1Qwen 3 32B90.23
2Qwen 2.5 72B Instruct88.58
3DeepSeek V388.31
4GPT-4o Mini88.09
5O3 Mini87.76
6Doubao-1.5-Pro-32k87.52
7Qwen 2.5 32B Instruct85.27
8QwQ-32B84.9
9Qwen 2.5 14B Instruct84.44
10Qwen 2.5 VL 72B Instruct77.35
11Qwen 2.5 7B Instruct71.11
12glm-4-9B69.03
13DeepSeek R1 Distill Llama 70B65
14DeepSeek R1 Distill Qwen 32B59.37
15Qwen 3 14B49.54

Interactive version: theaggregate.ai/benchmark?slug=pilotbench · How It Works · Data refreshed daily, snapshot 2026-10-07.