Step3 VL 10B — benchmark results
StepFun's Apache-2.0 10B vision-language model pairing a perception encoder with a Qwen3-8B decoder, rivaling multimodal models 10-20x its size (January 2026). Provider: StepFun. Released 2026-01-20. Access: Open.
Unified ELO 1496 ± 31, rank #813 of 1776 rated models, from 43 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (MMBench) | 91.8 | Score (%) | 100 |
| LLM Stats (MathVista) | 84 | Score (%) | 89.2 |
| MathVision | 75.95 | Overall Accuracy (%) | 85.7 |
| LLM Stats (Multi-Challenge) | 62.6 | Score (%) | 78.6 |
| AA Humanity's Last Exam | 10.19 | Accuracy (%) | 62.6 |
| AA IFBench | 50.2 | Accuracy (%) | 59.8 |
| LLM Stats (MathVision) | 70.8 | Score (%) | 54.8 |
| AA GPQA Diamond | 68.99 | Accuracy (%) | 51.2 |
| MMGist | 43.3 | Macro â (self-reported) | 46.2 |
| AA SciCode | 31.13 | Accuracy (%) | 45.8 |
| AA MMMU-Pro | 63.99 | Accuracy (%) | 43 |
| DiffCap-Bench | 66.1 | F1* (self-reported) | 40 |
Interactive version: theaggregate.ai/model?slug=step3-vl-10b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.