SuperCLUE-Code3 (May 2024) - Overall: leaderboard
Metric: Difficulty-Weighted Score. Source: www.superclueai.com. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 71.68 |
| 2 | GPT-4 Preview (0125) | 68 |
| 3 | GPT-4 | 63.74 |
| 4 | Llama 3 70B Instruct | 62.56 |
| 5 | DeepSeek V2 | 62.52 |
| 6 | GPT-3.5 Turbo (0125) | 55.51 |
| 7 | deepseek-coder-6.7B-instruct | 47.77 |
| 8 | Gemini 1.0 Pro | 46.5 |
| 9 | Qwen-14B-Chat | 24.67 |
| 10 | Llama 2 13B Chat | 6.06 |
Interactive version: theaggregate.ai/benchmark?slug=superclue-code3-may-2024-overall · How It Works · Data refreshed daily, snapshot 2026-09-19.