TSQueryBench - Relative Ranking: leaderboard
Metric: Relative ranking accuracy (%, times 100): share of instances where the model picks the correct explanation among a correct, a partially correct and an incorrect candidate in randomized order on TSQueryBench, 500 synthetic time series in ten query types (spikes, drops, structural breaks, mean and volatility shifts, trends, ordering, shape), zero-shot with structured JSON output; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 8B | 84 |
| 2 | DeepSeek V3.2 | 74 |
| 3 | Gemma 4 31B (IT) | 67 |
| 4 | Gemma 3 12B (IT) | 56 |
| 5 | Llama 3.1 8B Instruct | 49 |
| 6 | Claude 3 Haiku | 41 |
Interactive version: theaggregate.ai/benchmark?slug=tsquerybench-relative-ranking · How It Works · Data refreshed daily, snapshot 2026-10-07.