LexKairos (Chain-of-Thought): leaderboard

Metric: Mean score (%; average of the nine sub-task scores, accuracy or F1, over 3,800 Chinese legal temporal items from real judicial cases and statutes; chain-of-thought prompting, a generic step-by-step reasoning instruction appended to the prompt of models without a built-in thinking mode; temperature 0, 8,000 output tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 2 models tracked.

Top models

#ModelScore
1Llama 3.1 8B31.55

Interactive version: theaggregate.ai/benchmark?slug=lexkairos-chain-of-thought · How It Works · Data refreshed daily, snapshot 2026-09-29.