DeepSeek V4 Flash (Thinking): benchmark results

Provider: DeepSeek. Released 2026-04-24. Access: Open.

Unified ELO 1609, rank #357 of 2032 rated models, from 22 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EpiQAL - Text-Grounded Recall92.6Exact Match (%; zero-shot, no CoT, full article)100
PortBench-QA - Regime Detection84.3Item score (%; market regime plus the implied allocation dir100
PortBench-QA81.9Mean item score (%; mean of seven QA templates, 50 test ques94.4
NetConfArena - Task Pass Rate79.4Instances with every test case satisfied (%)85.7
NetConfArena - Test-Case Score93.3Mean fraction of task test cases satisfied (%)85.7
EpiQAL - Masked-Input Reasoning66.6Exact Match (%; zero-shot, no CoT, full article)84.6
EpiQAL - Multi-Step Inference70.3Exact Match (%; zero-shot, no CoT, full article)84.6
NYT Connections Older Models43.9Score (%)74.1
NetConfArena - Configuration F173.9Configuration line F1 against reference (%)71.4
WebGameBench62.7Usable Rate (%; validly scored artifacts labelled Excellent 69.2
PortBench-QA - Max-Sharpe Allocation93.2Item score (%; long-only maximum-Sharpe weights for three or66.7
QEncodeBench - 3-Coloring25Semantic pass@1 (%; L3 gate, graph 3-coloring (all edges bic66.7

Interactive version: theaggregate.ai/model?slug=deepseek-v4-flash-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.