Step3 VL 10B: benchmark results

StepFun's Apache-2.0 10B vision-language model pairing a perception encoder with a Qwen3-8B decoder, rivaling multimodal models 10-20x its size (January 2026). Provider: StepFun. Released 2026-01-20. Access: Open.

Unified ELO 1524 ± 1, rank #576 of 1392 rated models, from 42 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (MMBench)91.8Score (%)100
LLM Stats (MathVista)84Score (%)89.2
MathVision75.95Overall Accuracy (%)86.2
LLM Stats (Multi-Challenge)62.6Score (%)78.6
AA IFBench50.2Accuracy (%)59.9
LLM Stats Score26.07LLM Stats Score (conservative rating)59.5
AA Humanity's Last Exam10.84Accuracy (%)56.3
LLM Stats (MathVision)70.8Score (%)50
AA GPQA Diamond68.99Accuracy (%)46.9
MMGist43.3Macro ↑ (self-reported)46.2
DiffCap-Bench66.1F1* (self-reported)40
AA MMMU-Pro63.99Accuracy (%)38

Interactive version: theaggregate.ai/model?slug=step3-vl-10b · How It Works · Data refreshed daily, snapshot 2026-09-05.