TSQueryBench - Explanation Generation: leaderboard
Metric: Explanation generation accuracy (%, times 100): share of the model's free-form explanations rated fully correct by a domain-knowledgeable annotator who verifies every numerical claim against the series on TSQueryBench, 500 synthetic time series in ten query types (spikes, drops, structural breaks, mean and volatility shifts, trends, ordering, shape), zero-shot with structured JSON output; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 8B | 38 |
| 2 | Gemma 4 31B (IT) | 36 |
| 3 | Claude 3 Haiku | 31 |
| 4 | DeepSeek V3.2 | 30 |
| 5 | Llama 3.1 8B Instruct | 18 |
| 6 | Gemma 3 12B (IT) | 13 |
Interactive version: theaggregate.ai/benchmark?slug=tsquerybench-explanation-generation · How It Works · Data refreshed daily, snapshot 2026-10-07.