Modelica Agent Workflow Benchmark - Model Tuning: leaderboard

Metric: Tasks passed (%; 50 Model Tuning tasks: choose bounded parameters of a fixed Modelica model so that the simulated response meets private target and behavioural conditions; one fresh isolated run per task and harness-backend pair, the agent's explicit submission scored outside the loop; OpenModelica 1.26.1). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1Claude Sonnet 5 (Medium)72
2DeepSeek V4 Flash (0731)68

Interactive version: theaggregate.ai/benchmark?slug=modelica-agent-workflow-benchmark-model-tuning · How It Works · Data refreshed daily, snapshot 2026-09-29.