ProLLM - StackEval — leaderboard

ProLLM benchmark evaluating LLMs on Stack Overflow coding questions. Measures acceptance rate on real developer problems across languages and frameworks.

Metric: Score (%). Source: www.prollm.ai. Status: saturated. 44 models tracked.

Top models

#ModelScore
1GPT-4.198.6
2GPT-4.1 Mini98.4
3DeepSeek V397.6
4O3 Mini (Medium)97.4
5GPT-4.597.3
6O3 Mini (High)97.1
7GPT-4o96.1
8DeepSeek R195.6
9GPT-4.1 Nano95.2
10Llama 3.1 Nemotron 70B Instruct94.6
11Gemini 1.5 Pro94.3
12GPT-4 Turbo94.3
13Gemini 2.0 Flash94.2
14MiniMax-Text-0192.8
15Llama 4 Maverick Instruct92.3

Interactive version: theaggregate.ai/benchmark?slug=prollm-stackeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.