TSQueryBench - Independent Scoring: leaderboard
Metric: Independent scoring accuracy (%, times 100): share of the 1,500 candidate explanations whose correctness label (correct, partially correct, incorrect) the model assigns correctly when it sees one explanation at a time on TSQueryBench, 500 synthetic time series in ten query types (spikes, drops, structural breaks, mean and volatility shifts, trends, ordering, shape), zero-shot with structured JSON output; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 8B | 72 |
| 2 | Gemma 4 31B (IT) | 67 |
| 3 | DeepSeek V3.2 | 56 |
| 4 | Gemma 3 12B (IT) | 47 |
| 5 | Claude 3 Haiku | 35 |
| 6 | Llama 3.1 8B Instruct | 29 |
Interactive version: theaggregate.ai/benchmark?slug=tsquerybench-independent-scoring · How It Works · Data refreshed daily, snapshot 2026-10-07.