codestral-2508: benchmark results

Provider: Mistral. Released 2025-07-30. Access: API.

Unified ELO 1618 ± 33, rank #559 of 2656 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Tinybird AI SQL Benchmark - First-Attempt Success Rate100Questions answered with a valid query on the first attempt (93.4
Tinybird AI SQL Benchmark - Success Rate100Questions answered with a valid query within 3 retries (%)72.5
Guesswork 2026-071MAE (z-score units)66.7
AI Chess Leaderboard (Reasoning)731Elo56.6
Tinybird AI SQL Benchmark - Exactness47.22Result exactness vs human reference queries (0-100)45.4
Wolfram LLM Benchmarking Project37.8Correct Functionality (%)43
Guesswork 2026-081.21MAE (z-score units)40.6
LM Market Cap LMC Score40LMC Score (0-100)28
Design Arena (UI Components)1032Elo16.1
Design Arena (3D)1049Elo13.9
Design Arena (Website)1025Elo13.7
Design Arena (Data Viz)1032Elo13.2

Interactive version: theaggregate.ai/model?slug=codestral-2508 · How It Works · Data refreshed daily, snapshot 2026-09-19.