Front-Door Routing Benchmark: leaderboard
Metric: Exact classification accuracy (0-1, times 100) on 60 prompts (six task families, ten each, labelled by one author) classified zero-shot into a six-label routing taxonomy as JSON; greedy decoding, 128 output tokens, 4-bit NF4 quantized checkpoints; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen2.5-3B-Instruct [4bit] | 78.3 |
| 2 | Phi-3.5-mini-instruct [4bit] | 71.67 |
| 3 | Qwen2.5-1.5B-Instruct [4bit] | 40 |
Interactive version: theaggregate.ai/benchmark?slug=front-door-routing-benchmark · How It Works · Data refreshed daily, snapshot 2026-10-07.