Moonshot v1 8K: benchmark results

Provider: Moonshot. Access: API.

Unified ELO 1589 ± 14, rank #669 of 2656 rated models, from 681 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
OpenCompass Code - TACO (CompassBench 2405)19.6Score (%)100
OpenVLM OCRBench - Handwritten Mathematical Expression Recognition89Correct Answers (%)99.3
OpenVLM Q-Bench - What (Others)87.5Accuracy (%)97.5
OpenVLM MMT-Bench - Action Quality Assessment40Score (%)97.1
OpenVLM MMT-Bench - Equation-to-LaTeX95Score (%)95.1
Open LMM Reasoning - MathVerse - Analytic55.8Accuracy (%)94.9
OpenVLM MMT-Bench - Sketch-to-Code50Score (%)94.4
OpenCompass Math - High School - English (CompassBench 2405)79Score (%)93.4
OpenVLM A-Bench - Aesthetic Quality68.5Accuracy (%)93.4
OpenVLM AI2D - Eclipses96.8Accuracy (%)93.1
OpenVLM MMT-Bench - Attribute Hallucination85Score (%)92.5
OpenVLM MMT-Bench - Order Hallucination55Score (%)92.2

Interactive version: theaggregate.ai/model?slug=moonshot-v1-8k · How It Works · Data refreshed daily, snapshot 2026-09-19.