codestral-2508: benchmark results
Provider: Mistral. Released 2025-07-30. Access: API.
Unified ELO 1618 ± 33, rank #559 of 2656 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Tinybird AI SQL Benchmark - First-Attempt Success Rate | 100 | Questions answered with a valid query on the first attempt ( | 93.4 |
| Tinybird AI SQL Benchmark - Success Rate | 100 | Questions answered with a valid query within 3 retries (%) | 72.5 |
| Guesswork 2026-07 | 1 | MAE (z-score units) | 66.7 |
| AI Chess Leaderboard (Reasoning) | 731 | Elo | 56.6 |
| Tinybird AI SQL Benchmark - Exactness | 47.22 | Result exactness vs human reference queries (0-100) | 45.4 |
| Wolfram LLM Benchmarking Project | 37.8 | Correct Functionality (%) | 43 |
| Guesswork 2026-08 | 1.21 | MAE (z-score units) | 40.6 |
| LM Market Cap LMC Score | 40 | LMC Score (0-100) | 28 |
| Design Arena (UI Components) | 1032 | Elo | 16.1 |
| Design Arena (3D) | 1049 | Elo | 13.9 |
| Design Arena (Website) | 1025 | Elo | 13.7 |
| Design Arena (Data Viz) | 1032 | Elo | 13.2 |
Interactive version: theaggregate.ai/model?slug=codestral-2508 · How It Works · Data refreshed daily, snapshot 2026-09-19.