Front-Door Routing Benchmark: leaderboard

Metric: Exact classification accuracy (0-1, times 100) on 60 prompts (six task families, ten each, labelled by one author) classified zero-shot into a six-label routing taxonomy as JSON; greedy decoding, 128 output tokens, 4-bit NF4 quantized checkpoints; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1Qwen2.5-3B-Instruct [4bit]78.3
2Phi-3.5-mini-instruct [4bit]71.67
3Qwen2.5-1.5B-Instruct [4bit]40

Interactive version: theaggregate.ai/benchmark?slug=front-door-routing-benchmark · How It Works · Data refreshed daily, snapshot 2026-10-07.