TSQueryBench - Explanation Generation: leaderboard

Metric: Explanation generation accuracy (%, times 100): share of the model's free-form explanations rated fully correct by a domain-knowledgeable annotator who verifies every numerical claim against the series on TSQueryBench, 500 synthetic time series in ten query types (spikes, drops, structural breaks, mean and volatility shifts, trends, ordering, shape), zero-shot with structured JSON output; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 6 models tracked.

Top models

#ModelScore
1Qwen 3 8B38
2Gemma 4 31B (IT)36
3Claude 3 Haiku31
4DeepSeek V3.230
5Llama 3.1 8B Instruct18
6Gemma 3 12B (IT)13

Interactive version: theaggregate.ai/benchmark?slug=tsquerybench-explanation-generation · How It Works · Data refreshed daily, snapshot 2026-10-07.