DeepSeek V4 Flash (Max) — benchmark results

DeepSeek V4 Flash evaluated at the max reasoning-effort setting. Provider: DeepSeek. Released 2026-04-23. Access: Open.

Unified ELO 1871 ± 23, rank #87 of 1776 rated models, from 36 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (CodeForces)100Score96.7
CritPt7.1Accuracy (self-reported)91.7
MathArena - APEX Shortlist 202589.36Accuracy (%)88.9
ZeroEval GPQA Diamond88.1GPQA Diamond Score86.5
LLM Stats (MathArena Apex)85.7Score (%)83.3
AA-LCR63Score (self-reported)80.3
LLM Stats (HMMT Feb 26)94.8Score (%)80
MathArena - ARXIV_FALSE March19.64Accuracy (%)78.6
LLM2014 Logic 2026-0550.97Median Score76.3
OckBench83Accuracy (%)76
MathArena - APEX 202527.08Accuracy (%)75
LLM2014 Logic 2026-0456.54Median Score72.5

Interactive version: theaggregate.ai/model?slug=deepseek-v4-flash-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.