OpenRouter Tau2-Bench Airline — leaderboard

OpenRouter's own τ²-Bench Airline run: 50 multi-turn airline customer-service tasks where the model must use tools and follow policy, executed against the production endpoints it serves.

Metric: Accuracy (%). Source: openrouter.ai. Status: years away from saturation. 108 models tracked.

Top models

#ModelScore
1Claude Fable 582
2Nova Micro80.7
3Gemini 3 Flash (Preview)79.3
4Qwen 3.5 397B A17B78.4
5GPT-5.578
6Claude Opus 4.878
7Qwen 3.5 122B A10B78
8GPT-5.6 Sol77.3
9Claude Opus 577.3
10Step 3.7 Flash77.3
11Nemotron 3 Ultra76.9
12Claude Sonnet 4.676.7
13Claude Opus 4.676.7
14Gemma 4 31B76.7
15Claude Opus 4.576.7

Interactive version: theaggregate.ai/benchmark?slug=openrouter-tau2-bench-airline · How It Works · Data refreshed daily, snapshot 2026-08-06.