DeepSeek V4.1 Flash: benchmark results
Provider: DeepSeek. Released 2026-09-10. Access: Open.
Unified ELO 1966 ± 10, rank #29 of 2656 rated models, from 190 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Humanity's Last Exam (Self-Reported, With Tools) | 63.9 | Accuracy (%) | 100 |
| LLM Stats (AutomationBench) | 54.8 | Score (%) | 100 |
| LLM Stats (CyberGym) | 88.1 | Score (%) | 100 |
| LLM Stats (NL2Repo) | 64 | Score (%) | 100 |
| LLM Stats (Terminal-Bench 2.1) | 90.6 | Score (%) | 100 |
| LLM Stats (ZEROBench) | 49 | Score (%) | 100 |
| Nejumi 4 - jaster (0-shot) - NIILC | 63.01 | Character F1 (x100) | 100 |
| PLCC - Grammar | 98 | Accuracy (%) | 100 |
| SecIT Bench (Pydantic AI) | 82.9 | Accuracy (%) | 100 |
| Nejumi 4 - jaster (2-shot) - JaMP | 84 | Exact match (%) | 98.5 |
| Featherbench | 100 | Pass Rate (%) | 97.6 |
| LLM Stats (DeepSWE 1.1) | 74.2 | Score (%) | 96.8 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-1-flash · How It Works · Data refreshed daily, snapshot 2026-09-19.