ClassEval-Pro - Method-Level: leaderboard
Metric: Method-level full-success pass@1 (%): share of the class's methods whose tests all pass, on ClassEval-Pro's 300 class-level Python tasks (11 domains, single- and cross-domain class compositions mined from GitHub code created after January 2025), holistic generation (the whole class from its skeleton in one call), unbiased pass@1 from five samples at temperature 0.2, run against each task's test suite (over 90% line coverage) with a 60-second timeout; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 Coder 480B A35B Instruct | 84.2 |
| 2 | Kimi K2 | 83 |
| 3 | Gemini 2.5 Pro | 80.5 |
| 4 | Qwen 3 30B A3B | 78.2 |
| 5 | GPT-5.1 | 61.3 |
Interactive version: theaggregate.ai/benchmark?slug=classeval-pro-method-level · How It Works · Data refreshed daily, snapshot 2026-10-07.