LongCat Flash (Thinking): benchmark results

Provider: Meituan. Released 2025-09-21. Access: API.

Unified ELO 1599 ± 1, rank #631 of 3078 rated models, from 20 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ZeroEval MATH-50099.2MATH-500 Score100
LLM Stats (AIME 2024)93.3Score (%)93.4
LLM Stats (ZebraLogic)95.5Score (%)87.5
ZeroEval GPQA Diamond81.5GPQA Diamond Score65
LLM Stats Score28.61LLM Stats Score (conservative rating)63.6
SuperCLUE General (September 2025) - Precise Instruction Following41.98Score50
LLM Stats (MMLU-Redux)89.3Score (%)43.4
SuperCLUE General (November 2025) - Hallucination Control77.79Score32.3
SuperCLUE General (November 2025) - Precise Instruction Following26.45Score27.4
NYT Connections Extended17.3Score (%)18.1
SuperCLUE General (September 2025) - Hallucination Control60.14Score16.7
SuperCLUE General (September 2025) - Math Reasoning37.27Score16.7

Interactive version: theaggregate.ai/model?slug=longcat-flash-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.