VEHBench - Specification Triage: leaderboard
Metric: P1 composite (0-100): weighted combination of macro-F1, action-conditioned, missing-information and infeasibility detection scores and subtype F1 for deciding whether to propose a design, abstain as infeasible or request missing information on a harvester brief, x100; LLM-assisted vibration energy harvester design against a physics oracle; one complete run per model. Source: arxiv.org. Saturation forecast: Around 2033. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 Max | 57.4 |
| 2 | Gemini 3.1 Pro (Preview) | 54.9 |
| 3 | O4 Mini | 51.8 |
| 4 | DeepSeek R1 | 50.4 |
| 5 | GPT-5.4 | 42.8 |
| 6 | Hy3-preview (Reasoning) | 42.5 |
| 7 | DeepSeek V3 | 36.9 |
| 8 | Llama 3.3 70B | 36.1 |
| 9 | MiMo-V2.5-Pro (Non-reasoning) | 31.7 |
| 10 | Qwen 3.6 Plus | 30.6 |
| 11 | DeepSeek V4 Pro (Non-reasoning) | 30.6 |
| 12 | Claude Sonnet 4.6 | 20.7 |
Interactive version: theaggregate.ai/benchmark?slug=vehbench-specification-triage · How It Works · Data refreshed daily, snapshot 2026-09-29.