Modelica Agent Workflow Benchmark - Model Repair: leaderboard
Metric: Tasks passed (%; 132 Model Repair tasks: correct a Modelica model with an injected fault (123 fail model checking, 9 fail simulation) so that the benchmark-owned evaluator's interface, model-check and simulation gates pass; one fresh isolated run per task and harness-backend pair, the agent's explicit submission scored outside the loop; OpenModelica 1.26.1). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 5 (Medium) | 94.7 |
| 2 | DeepSeek V4 Flash (0731) | 93.94 |
Interactive version: theaggregate.ai/benchmark?slug=modelica-agent-workflow-benchmark-model-repair · How It Works · Data refreshed daily, snapshot 2026-09-29.