Step3 VL 10B: benchmark results
StepFun's Apache-2.0 10B vision-language model pairing a perception encoder with a Qwen3-8B decoder, rivaling multimodal models 10-20x its size (January 2026). Provider: StepFun. Released 2026-01-20. Access: Open.
Unified ELO 1524 ± 1, rank #576 of 1392 rated models, from 42 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (MMBench) | 91.8 | Score (%) | 100 |
| LLM Stats (MathVista) | 84 | Score (%) | 89.2 |
| MathVision | 75.95 | Overall Accuracy (%) | 86.2 |
| LLM Stats (Multi-Challenge) | 62.6 | Score (%) | 78.6 |
| AA IFBench | 50.2 | Accuracy (%) | 59.9 |
| LLM Stats Score | 26.07 | LLM Stats Score (conservative rating) | 59.5 |
| AA Humanity's Last Exam | 10.84 | Accuracy (%) | 56.3 |
| LLM Stats (MathVision) | 70.8 | Score (%) | 50 |
| AA GPQA Diamond | 68.99 | Accuracy (%) | 46.9 |
| MMGist | 43.3 | Macro ↑ (self-reported) | 46.2 |
| DiffCap-Bench | 66.1 | F1* (self-reported) | 40 |
| AA MMMU-Pro | 63.99 | Accuracy (%) | 38 |
Interactive version: theaggregate.ai/model?slug=step3-vl-10b · How It Works · Data refreshed daily, snapshot 2026-09-05.