ClassEval-Pro: leaderboard
Metric: Class-level full-success pass@1 (%): the generated class passes all of the task's method-level and class-level tests, on ClassEval-Pro's 300 class-level Python tasks (11 domains, single- and cross-domain class compositions mined from GitHub code created after January 2025), holistic generation (the whole class from its skeleton in one call), unbiased pass@1 from five samples at temperature 0.2, run against each task's test suite (over 90% line coverage) with a 60-second timeout; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 Coder 480B A35B Instruct | 45.6 |
| 2 | Kimi K2 | 45.1 |
| 3 | Gemini 2.5 Pro | 44.7 |
| 4 | Qwen 3 30B A3B | 40.5 |
| 5 | GPT-5.1 | 27.9 |
Interactive version: theaggregate.ai/benchmark?slug=classeval-pro · How It Works · Data refreshed daily, snapshot 2026-10-07.