TeleResilienceBench - TeleQnA: leaderboard

Metric: Correct flip rate (CFR, %): share of the 359 TeleQnA instances (telecom standards question answering, GSMA Open-Telco ot-lite) where the model, given the question, the options and the first half of a flawed reasoning trace from qwen3.5:2b that ended in a wrong answer, continues the reasoning in its thinking stream and returns the correct option; one continuation prompt, Ollama builds on one RTX 4090; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1nemotron-3-nano:4b (Ollama)34
2gemma4:31b (Ollama)32.9
3gemma4:e2b (Ollama)32.6
4gemma4:26b (Ollama)30.1
5gemma4:e4b (Ollama)26.7
6qwen3.5:4b (Ollama)23.7
7qwen3.5:27b (Ollama)23.4
8qwen3.5:9b (Ollama)21.2

Interactive version: theaggregate.ai/benchmark?slug=teleresiliencebench-teleqna · How It Works · Data refreshed daily, snapshot 2026-10-07.