Codestral: benchmark results

Provider: Mistral. Released 2024-05-29. Access: Open.

Unified ELO 1581 ± 24, rank #707 of 2656 rated models, from 13 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PECC - Advent of Code50.77Solve rate (%)90
PECC - Advent of Code (LeetCode-style)37.5Solve rate (%)90
PECC - Project Euler6.2Solve rate (%)80
PECC - Project Euler (Story)5.71Solve rate (%)80
Wolfram LLM Benchmarking Project34.4Correct Functionality (%)38.8
BaxBench - No Security Reminder - Correct27.63Correct solutions, pass@1 (%)29.7
BaxBench - No Security Reminder - Correct & Secure11.53Correct and secure solutions, sec_pass@1 (%)29.7
Aider Code Editing Leaderboard45.9Exercises completed correctly after one retry, pass_rate_2 (28
SnorkelGraph24.5Accuracy@1 (%, 200 graph reasoning questions)25
SnorkelUnderwrite34Overall accuracy (%, LLM-as-a-judge)21.9
SnorkelSequences38.4Accuracy@1 (%, 250 compositional sequence questions)19.4
SnorkelSpatial13.64Accuracy@1 (%, 330 spatial reasoning questions)18.8

Interactive version: theaggregate.ai/model?slug=codestral · How It Works · Data refreshed daily, snapshot 2026-09-19.