Codestral: benchmark results
Provider: Mistral. Released 2024-05-29. Access: Open.
Unified ELO 1581 ± 24, rank #707 of 2656 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PECC - Advent of Code | 50.77 | Solve rate (%) | 90 |
| PECC - Advent of Code (LeetCode-style) | 37.5 | Solve rate (%) | 90 |
| PECC - Project Euler | 6.2 | Solve rate (%) | 80 |
| PECC - Project Euler (Story) | 5.71 | Solve rate (%) | 80 |
| Wolfram LLM Benchmarking Project | 34.4 | Correct Functionality (%) | 38.8 |
| BaxBench - No Security Reminder - Correct | 27.63 | Correct solutions, pass@1 (%) | 29.7 |
| BaxBench - No Security Reminder - Correct & Secure | 11.53 | Correct and secure solutions, sec_pass@1 (%) | 29.7 |
| Aider Code Editing Leaderboard | 45.9 | Exercises completed correctly after one retry, pass_rate_2 ( | 28 |
| SnorkelGraph | 24.5 | Accuracy@1 (%, 200 graph reasoning questions) | 25 |
| SnorkelUnderwrite | 34 | Overall accuracy (%, LLM-as-a-judge) | 21.9 |
| SnorkelSequences | 38.4 | Accuracy@1 (%, 250 compositional sequence questions) | 19.4 |
| SnorkelSpatial | 13.64 | Accuracy@1 (%, 330 spatial reasoning questions) | 18.8 |
Interactive version: theaggregate.ai/model?slug=codestral · How It Works · Data refreshed daily, snapshot 2026-09-19.