MLX Benchmark V2 — leaderboard

Benchmark for evaluating model knowledge of Apple's MLX machine-learning framework across implementation, debugging, and conceptual questions.

Metric: Accuracy (%). Source: huggingface.co. Status: saturation imminent. 22 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.689.62
2Gemini 3 Flash (Preview)82.39
3Qwen 3.6 Max Preview80.13
4Gemma 4 26B A4B (IT)75.19
5GPT-5.4 Nano75.19
6Grok 4.1 Fast72.69
7DeepSeek V4 Pro71.84
8DeepSeek V4 Flash71.54
9Gemini 2.5 Flash Lite (Preview 09-2025)67.31
10Qwen 3.6 35B A3B52.5
11GPT-5 Nano41.92
12GLM-5.119.23
13Nemotron 3 Ultra16.15
14Kimi K2.54.81
15Kimi K2.63.1

Interactive version: theaggregate.ai/benchmark?slug=mlx-benchmark-v2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.