ClassEval-Pro: leaderboard

Metric: Class-level full-success pass@1 (%): the generated class passes all of the task's method-level and class-level tests, on ClassEval-Pro's 300 class-level Python tasks (11 domains, single- and cross-domain class compositions mined from GitHub code created after January 2025), holistic generation (the whole class from its skeleton in one call), unbiased pass@1 from five samples at temperature 0.2, run against each task's test suite (over 90% line coverage) with a 60-second timeout; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 5 models tracked.

Top models

#ModelScore
1Qwen 3 Coder 480B A35B Instruct45.6
2Kimi K245.1
3Gemini 2.5 Pro44.7
4Qwen 3 30B A3B40.5
5GPT-5.127.9

Interactive version: theaggregate.ai/benchmark?slug=classeval-pro · How It Works · Data refreshed daily, snapshot 2026-10-07.