DeepSeek V4 Flash Vision (Reasoning, Max Effort): benchmark results

Provider: DeepSeek. Released 2026-08-21. Access: Open.

Unified ELO 1685 ± 1, rank #143 of 1761 rated models, from 19 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Long Context Reasoning81.33Accuracy (%)93.3
BenchmarkList ECI146.04Capability Index (ECI)93.3
AA GPQA Diamond91.31Accuracy (%)92.3
Artificial Analysis Intelligence Index41.54Intelligence Index91.1
AA-LCR78Accuracy (self-reported)90.8
AA GDPval1581.74ELO90.7
Tau3 Banking41.03Success Rate (%)86.9
AA CritPt10.86Accuracy (%)86.3
AA Humanity's Last Exam34.48Accuracy (%)86.2
AA Omniscience - Health38.9Accuracy (%)86
AA Omniscience - Science, Engineering & Mathematics43.6Accuracy (%)84.1
AA Omniscience - Humanities & Social Sciences38.5Accuracy (%)83.4

Interactive version: theaggregate.ai/model?slug=deepseek-v4-flash-vision-reasoning-max-effort · How It Works · Data refreshed daily, snapshot 2026-09-05.