NLCO - Set-M - Accuracy: leaderboard

Metric: Accuracy: feasible and matching the solver-optimal objective (%). Source: arxiv.org. Saturation forecast: Around December 2026. 16 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (High)76.8
2GPT-5.1 (Medium)72.8
3Grok 4.1 Fast (Reasoning)70.9
4O4 Mini (High)70.7
5Claude Sonnet 4.5 (Thinking)70.4
6DeepSeek V3.2 (Thinking)69.1
7Qwen 3 235B A22B 2507 Instruct54.9
8DeepSeek V3.2 (Non-reasoning)47.7
9Qwen 3 14B (Reasoning)42.4
10QwQ-32B39.1
11MiMo-V2-Flash (Non-reasoning)37.4
12Ministral-3-14B-Instruct-251225.4
13Llama 4 Maverick Instruct19
14Qwen 3 14B (Non-reasoning)17.5

Interactive version: theaggregate.ai/benchmark?slug=nlco-set-m-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-25.