Moonshot v1 8K: benchmark results
Provider: Moonshot. Access: API.
Unified ELO 1589 ± 14, rank #669 of 2656 rated models, from 681 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OpenCompass Code - TACO (CompassBench 2405) | 19.6 | Score (%) | 100 |
| OpenVLM OCRBench - Handwritten Mathematical Expression Recognition | 89 | Correct Answers (%) | 99.3 |
| OpenVLM Q-Bench - What (Others) | 87.5 | Accuracy (%) | 97.5 |
| OpenVLM MMT-Bench - Action Quality Assessment | 40 | Score (%) | 97.1 |
| OpenVLM MMT-Bench - Equation-to-LaTeX | 95 | Score (%) | 95.1 |
| Open LMM Reasoning - MathVerse - Analytic | 55.8 | Accuracy (%) | 94.9 |
| OpenVLM MMT-Bench - Sketch-to-Code | 50 | Score (%) | 94.4 |
| OpenCompass Math - High School - English (CompassBench 2405) | 79 | Score (%) | 93.4 |
| OpenVLM A-Bench - Aesthetic Quality | 68.5 | Accuracy (%) | 93.4 |
| OpenVLM AI2D - Eclipses | 96.8 | Accuracy (%) | 93.1 |
| OpenVLM MMT-Bench - Attribute Hallucination | 85 | Score (%) | 92.5 |
| OpenVLM MMT-Bench - Order Hallucination | 55 | Score (%) | 92.2 |
Interactive version: theaggregate.ai/model?slug=moonshot-v1-8k · How It Works · Data refreshed daily, snapshot 2026-09-19.