VEHBench - Specification Triage: leaderboard

Metric: P1 composite (0-100): weighted combination of macro-F1, action-conditioned, missing-information and infeasibility detection scores and subtype F1 for deciding whether to propose a design, abstain as infeasible or request missing information on a harvester brief, x100; LLM-assisted vibration energy harvester design against a physics oracle; one complete run per model. Source: arxiv.org. Saturation forecast: Around 2033. 12 models tracked.

Top models

#ModelScore
1Qwen 3 Max57.4
2Gemini 3.1 Pro (Preview)54.9
3O4 Mini51.8
4DeepSeek R150.4
5GPT-5.442.8
6Hy3-preview (Reasoning)42.5
7DeepSeek V336.9
8Llama 3.3 70B36.1
9MiMo-V2.5-Pro (Non-reasoning)31.7
10Qwen 3.6 Plus30.6
11DeepSeek V4 Pro (Non-reasoning)30.6
12Claude Sonnet 4.620.7

Interactive version: theaggregate.ai/benchmark?slug=vehbench-specification-triage · How It Works · Data refreshed daily, snapshot 2026-09-29.