LongCat Flash (Thinking): benchmark results
Provider: Meituan. Released 2025-09-21. Access: API.
Unified ELO 1599 ± 1, rank #631 of 3078 rated models, from 20 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ZeroEval MATH-500 | 99.2 | MATH-500 Score | 100 |
| LLM Stats (AIME 2024) | 93.3 | Score (%) | 93.4 |
| LLM Stats (ZebraLogic) | 95.5 | Score (%) | 87.5 |
| ZeroEval GPQA Diamond | 81.5 | GPQA Diamond Score | 65 |
| LLM Stats Score | 28.61 | LLM Stats Score (conservative rating) | 63.6 |
| SuperCLUE General (September 2025) - Precise Instruction Following | 41.98 | Score | 50 |
| LLM Stats (MMLU-Redux) | 89.3 | Score (%) | 43.4 |
| SuperCLUE General (November 2025) - Hallucination Control | 77.79 | Score | 32.3 |
| SuperCLUE General (November 2025) - Precise Instruction Following | 26.45 | Score | 27.4 |
| NYT Connections Extended | 17.3 | Score (%) | 18.1 |
| SuperCLUE General (September 2025) - Hallucination Control | 60.14 | Score | 16.7 |
| SuperCLUE General (September 2025) - Math Reasoning | 37.27 | Score | 16.7 |
Interactive version: theaggregate.ai/model?slug=longcat-flash-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.