Step3 VL 10B — benchmark results

StepFun's Apache-2.0 10B vision-language model pairing a perception encoder with a Qwen3-8B decoder, rivaling multimodal models 10-20x its size (January 2026). Provider: StepFun. Released 2026-01-20. Access: Open.

Unified ELO 1496 ± 31, rank #813 of 1776 rated models, from 43 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (MMBench)91.8Score (%)100
LLM Stats (MathVista)84Score (%)89.2
MathVision75.95Overall Accuracy (%)85.7
LLM Stats (Multi-Challenge)62.6Score (%)78.6
AA Humanity's Last Exam10.19Accuracy (%)62.6
AA IFBench50.2Accuracy (%)59.8
LLM Stats (MathVision)70.8Score (%)54.8
AA GPQA Diamond68.99Accuracy (%)51.2
MMGist43.3Macro ↑ (self-reported)46.2
AA SciCode31.13Accuracy (%)45.8
AA MMMU-Pro63.99Accuracy (%)43
DiffCap-Bench66.1F1* (self-reported)40

Interactive version: theaggregate.ai/model?slug=step3-vl-10b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.