DSAEval: leaderboard

Metric: Total score (0-10; mean of two LLM judges). Source: ama-cmfai.github.io. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.58.16
2MiMo-V2-Pro7.91
3GPT-5.27.71
4MiniMax-M2.77.7
5MiMo-V2-Flash7.64
6MiniMax-M27.64
7Gemini 3 Pro7.31
8Grok 4.1 Fast7.25
9GPT-5 Nano7.07
10DeepSeek V3.27.03
11GLM-4.6V6.87
12Qwen 3 VL 30B A3B (Thinking)5.32
13Ministral 3 14B5.18

Interactive version: theaggregate.ai/benchmark?slug=dsaeval · How It Works · Data refreshed daily, snapshot 2026-09-25.