LexKairos (Thinking Mode): leaderboard

Metric: Mean score (%; average of the nine sub-task scores, accuracy or F1, over 3,800 Chinese legal temporal items from real judicial cases and statutes; thinking mode, built-in reasoning enabled at high effort with the vanilla prompt; temperature 0, 8,000 output tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (High)84.57
2GPT-5.4 (High)82.75
3DeepSeek V4 Flash (High)79.94
4Qwen 3.5 9B (Thinking)69.79

Interactive version: theaggregate.ai/benchmark?slug=lexkairos-thinking-mode · How It Works · Data refreshed daily, snapshot 2026-09-29.