MLX Benchmark V2 - Coding: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 21 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.693.94
2GPT-5.4 Nano75.76
3Grok 4.1 Fast63.64
4Gemma 4 26B A4B (IT)60.61
5Gemini 3 Flash (Preview)54.55
6Gemini 2.5 Flash Lite (Preview 09-2025)51.52
7DeepSeek V4 Flash39.39
8Qwen 3.6 35B A3B36.36
9DeepSeek V4 Pro35.71
10GPT-5 Nano0
11Kimi K2.50
12Kimi K2.60
13GLM-5.10
14Nemotron 3 Ultra0

Interactive version: theaggregate.ai/benchmark?slug=mlx-benchmark-v2-coding · How It Works · Data refreshed daily, snapshot 2026-09-05.