LongCat-Flash-Thinking-2601 — benchmark results
January 2026 refresh of Meituan's open LongCat Flash Thinking reasoning MoE with large-scale agentic RL training. Provider: Meituan. Released 2026-01-29. Access: API.
Unified ELO 1725 ± 18, rank #228 of 1776 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| VitaBench | 29.3 | Cross-Scenario Avg@4 (%) | 85 |
| LLM Stats (BrowseComp-zh) | 69 | Score (%) | 75 |
| OpenEvals - Humanity's Last Exam | 25.2 | Accuracy (%) | 71.4 |
| ZeroEval GPQA Diamond | 80.5 | GPQA Diamond Score | 63.3 |
| ClawProBench | 57.48 | Final Score (self-reported) | 58.9 |
| LLM2014 Logic 2026-03 | 37.69 | Median Score | 56.1 |
| OpenClawProBench | 64.5 | Overall Score (%) | 50 |
| OpenEvals - GPQA Diamond | 80.5 | Accuracy (%) | 50 |
| LLM2014 Logic 2026-04 | 31.26 | Median Score | 37.5 |
| LLM Stats (BrowseComp) | 56.6 | Score (%) | 35.7 |
| LLM2014 Logic 2026-05 | 21.78 | Median Score | 23.7 |
| LoHoSearch | 9.74 | Accuracy (%) | 20 |
Interactive version: theaggregate.ai/model?slug=longcat-flash-thinking-2601 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.