Text2Opt-Bench - Resource Allocation: leaderboard

Metric: Pass@1 accuracy (%; direct-translation LP/MILP resource allocation, 248 evaluation instances with 2-20 variables; solver-verified optimal objective). Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Claude Opus 4.689.9
2GPT-587.9
3Claude Sonnet 4.684.7
4DeepSeek R180.6
5O4 Mini80.2
6DeepSeek V3.279
7Llama 3.3 70B Instruct49.6
8GPT-5 Nano49.2
9Qwen 2.5 7B Instruct13.3

Interactive version: theaggregate.ai/benchmark?slug=text2opt-bench-resource-allocation · How It Works · Data refreshed daily, snapshot 2026-09-26.