LongCat-Flash-Thinking-2601 — benchmark results

January 2026 refresh of Meituan's open LongCat Flash Thinking reasoning MoE with large-scale agentic RL training. Provider: Meituan. Released 2026-01-29. Access: API.

Unified ELO 1725 ± 18, rank #228 of 1776 rated models, from 13 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
VitaBench29.3Cross-Scenario Avg@4 (%)85
LLM Stats (BrowseComp-zh)69Score (%)75
OpenEvals - Humanity's Last Exam25.2Accuracy (%)71.4
ZeroEval GPQA Diamond80.5GPQA Diamond Score63.3
ClawProBench57.48Final Score (self-reported)58.9
LLM2014 Logic 2026-0337.69Median Score56.1
OpenClawProBench64.5Overall Score (%)50
OpenEvals - GPQA Diamond80.5Accuracy (%)50
LLM2014 Logic 2026-0431.26Median Score37.5
LLM Stats (BrowseComp)56.6Score (%)35.7
LLM2014 Logic 2026-0521.78Median Score23.7
LoHoSearch9.74Accuracy (%)20

Interactive version: theaggregate.ai/model?slug=longcat-flash-thinking-2601 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.