CADEngBench - Functional Editing (L2-E): leaderboard
Metric: Edit pass rate (%; share of 300 functional-edit tasks on supplied source CAD, native CadQuery or a replayed Fusion 360 JSON patch, where the requested change is measured and protected non-target geometry is preserved; one response per task). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) (Minimal) | 72.3 |
| 2 | GPT-5.2 (Non-reasoning) | 70.3 |
| 3 | Kimi K2.5 (Non-reasoning) | 67 |
| 4 | Claude Sonnet 4.5 | 66.7 |
| 5 | Qwen 3.5 35B A3B (Non-reasoning) | 61.7 |
| 6 | Llama 4 Maverick | 60.7 |
| 7 | GLM-4.6V (Non-reasoning) | 59.3 |
Interactive version: theaggregate.ai/benchmark?slug=cadengbench-functional-editing-l2-e · How It Works · Data refreshed daily, snapshot 2026-09-29.