DeepSeek V4.1 Flash (High): benchmark results
Provider: DeepSeek. Released 2026-09-10. Access: Open.
Unified ELO 1716 ± 1, rank #96 of 3078 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Coding Daily (OpenCode) - React-TS Code Quality | 18.67 | React-TS Code Quality (max 20) points, LLM-judged rubric sco | 100 |
| Vals AI SkillsBench | 69.8 | Accuracy (%) | 100 |
| Chess Bench LLM | 1294 | Lichess Rating | 84 |
| Vals AI CyberBench | 78.82 | Accuracy (%) | 84 |
| Vals AI MedScribe | 85.5 | Accuracy (%) | 80.6 |
| AI Coding Daily (OpenCode) - Total | 48.52 | Total points (max 60) | 78.9 |
| LLM2014 Logic 2026-09 | 52.09 | Median Score | 68.8 |
| Vals AI SAGE | 47.88 | Accuracy (%) | 68.8 |
| Vals AI LegalBench | 83.28 | Accuracy (%) | 64.6 |
| Vals AI Public Benefits Bench | 64.28 | Accuracy (%) | 62.9 |
| Vals AI ProofBench | 54 | Accuracy (%) | 54.3 |
| Bug Hunt Bench - VS Code Extension | 9 | Planted Bugs Fixed (out of 45) | 52.2 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-1-flash-high · How It Works · Data refreshed daily, snapshot 2026-09-19.