DeepSeek V4 Flash (Thinking): benchmark results
Provider: DeepSeek. Released 2026-04-24. Access: Open.
Unified ELO 1609, rank #357 of 2032 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EpiQAL - Text-Grounded Recall | 92.6 | Exact Match (%; zero-shot, no CoT, full article) | 100 |
| PortBench-QA - Regime Detection | 84.3 | Item score (%; market regime plus the implied allocation dir | 100 |
| PortBench-QA | 81.9 | Mean item score (%; mean of seven QA templates, 50 test ques | 94.4 |
| NetConfArena - Task Pass Rate | 79.4 | Instances with every test case satisfied (%) | 85.7 |
| NetConfArena - Test-Case Score | 93.3 | Mean fraction of task test cases satisfied (%) | 85.7 |
| EpiQAL - Masked-Input Reasoning | 66.6 | Exact Match (%; zero-shot, no CoT, full article) | 84.6 |
| EpiQAL - Multi-Step Inference | 70.3 | Exact Match (%; zero-shot, no CoT, full article) | 84.6 |
| NYT Connections Older Models | 43.9 | Score (%) | 74.1 |
| NetConfArena - Configuration F1 | 73.9 | Configuration line F1 against reference (%) | 71.4 |
| WebGameBench | 62.7 | Usable Rate (%; validly scored artifacts labelled Excellent | 69.2 |
| PortBench-QA - Max-Sharpe Allocation | 93.2 | Item score (%; long-only maximum-Sharpe weights for three or | 66.7 |
| QEncodeBench - 3-Coloring | 25 | Semantic pass@1 (%; L3 gate, graph 3-coloring (all edges bic | 66.7 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-flash-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.