ProLLM - StackUnseen — leaderboard

ProLLM benchmark evaluating LLMs on unseen Stack Overflow coding questions not in training data. Measures acceptance rate on real developer problems.

Metric: Score (%). Source: www.prollm.ai. Status: saturated. 35 models tracked.

Top models

#ModelScore
1GPT-5.288.6
2Grok 488.6
3Gemini 3 Pro (Preview)86.2
4Gemini 3 Flash83
5GPT-5 Mini82.4
6Qwen 3.5 397B A17B76.3
7Kimi K2 (Thinking)76.1
8MiniMax-M2.566
9Kimi K2.564.9
10GPT-5 Nano60.4
11GLM-555.1
12DeepSeek R1 052852.4
13Mistral Large 351.6
14DeepSeek V3.148.1
15Qwen 3 32B45.7

Interactive version: theaggregate.ai/benchmark?slug=prollm-stackunseen · How the rankings work · Data refreshed daily, snapshot 2026-07-22.