ProOPF-B (Few-Shot) - Level 2: leaderboard

Metric: Objective-value accuracy (%) on the 30 Level 2 cases (parameter changes the model must infer from a qualitative operating scenario, scored on a predefined parameter instantiation); ProOPF-B's 121 expert-annotated optimal power flow modeling cases (natural-language operating requirements turned into executable MATPOWER code); a case is correct when the generated implementation executes and its optimal objective value matches the expert implementation's within a small tolerance; few-shot prompting; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek R116.67#245
2GPT-5.110#131
3Claude 3.5 Sonnet10#337
4Claude Sonnet 4.56.67#138
5GPT-5.26.67#105
6Gemini 3 Pro6.67#77
7DeepSeek V3.20#198
8Qwen 3 30B A3B Instruct0#460
9Qwen3 Coder0#285

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=proopf-b-few-shot-level-2 · How It Works · Data refreshed daily, snapshot 2026-10-11.