MLX Benchmark V2 - Coding — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 21 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.693.94
2GPT-5.4 Nano75.76
3Grok 4.1 Fast63.64
4Gemma 4 26B A4B (IT)60.61
5Gemini 3 Flash (Preview)54.55
6Gemini 2.5 Flash Lite (Preview 09-2025)51.52
7DeepSeek V4 Flash39.39
8Qwen 3.6 35B A3B36.36
9DeepSeek V4 Pro35.71
10GPT-5 Nano0
11Kimi K2.50
12GLM-5.10
13Kimi K2.60
14Nemotron 3 Ultra0

Interactive version: theaggregate.ai/benchmark?slug=mlx-benchmark-v2-coding · How the rankings work · Data refreshed daily, snapshot 2026-07-22.