Step-3: benchmark results
Provider: StepFun. Access: API.
Unified ELO 1568 ± 1, rank #372 of 1392 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FlagEval VQA - Geographic Reasoning | 58.2 | Score | 69.6 |
| FlagEval VQA - Text Understanding (Chinese) | 73 | Score | 65.2 |
| FlagEval VQA - Visual Puzzles | 29.5 | Score | 60.9 |
| LLM2014 Logic 2025-08 | 40.26 | Median Score | 52.3 |
| FlagEval VQA - Long-tail Recognition | 35 | Score | 52.2 |
| FlagEval VQA - Overall | 43.57 | Score | 52.2 |
| FlagEval VQA - Scene Understanding | 53.8 | Score | 52.2 |
| LLM2014 Logic 2025-09 | 37.03 | Median Score | 48.9 |
| LLM2014 Logic 2025-10 | 37.92 | Median Score | 46.9 |
| FlagEval VQA - Multi-Image Analysis | 42.5 | Score | 43.5 |
| FlagEval VQA - Subject Knowledge | 55.9 | Score | 43.5 |
| LLM2014 Logic 2025-11 | 33.59 | Median Score | 40.4 |
Interactive version: theaggregate.ai/model?slug=step-3 · How It Works · Data refreshed daily, snapshot 2026-09-05.