DeepSeek V4 Flash (Max): benchmark results

DeepSeek V4 Flash evaluated at the max reasoning-effort setting. Provider: DeepSeek. Released 2026-04-23. Access: Open.

Unified ELO 1643 ± 1, rank #268 of 1761 rated models, from 43 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (CodeForces)100Score96.9
MathArena - APEX Shortlist 202589.36Accuracy (%)89.2
LLM Stats (MathArena Apex)85.7Score (%)87.5
LLM Stats Score39.48LLM Stats Score (conservative rating)85.2
ZeroEval GPQA Diamond88.1GPQA Diamond Score84.2
MathArena - ARXIV_FALSE March19.64Accuracy (%)78.6
SuperCLUE-Writing - Mainstream Genre Writing82.57Score76.9
LLM2014 Logic 2026-0550.97Median Score76.3
Chess Bench LLM919Lichess Rating75.5
LLM Stats (IMO-AnswerBench)88.4Score (%)73.7
Epoch AI - Dtbench86.4Score73.1
MathArena - APEX 202527.08Accuracy (%)72.7

Interactive version: theaggregate.ai/model?slug=deepseek-v4-flash-max · How It Works · Data refreshed daily, snapshot 2026-09-05.