NLCO - Solver-Calling Code (Set-L) - Accuracy: leaderboard
Metric: Accuracy of generated Gurobi solver code: feasible and optimal (%). Source: arxiv.org. Saturation forecast: Around December 2026. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (High) | 83.77 |
| 2 | Claude Sonnet 4.5 (Thinking) | 80.9 |
| 3 | GPT-5.1 (Medium) | 76.88 |
| 4 | DeepSeek V3.2 (Thinking) | 76.19 |
| 5 | Qwen 3 235B A22B 2507 Instruct | 69.26 |
| 6 | Grok 4.1 Fast (Reasoning) | 67.26 |
| 7 | DeepSeek V3.2 (Non-reasoning) | 66.98 |
| 8 | O4 Mini (High) | 56.47 |
| 9 | QwQ-32B | 54.7 |
| 10 | Llama 4 Maverick Instruct | 53.07 |
| 11 | Qwen 3 14B (Reasoning) | 51.21 |
| 12 | MiMo-V2-Flash (Non-reasoning) | 41.72 |
| 13 | Ministral-3-14B-Instruct-2512 | 40.6 |
| 14 | Qwen 3 14B (Non-reasoning) | 39.58 |
Interactive version: theaggregate.ai/benchmark?slug=nlco-solver-calling-code-set-l-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-25.