TaxBench — leaderboard

TaxBench evaluates AI models on real-world tax tasks from Rivet's active tax workflows, spanning tax knowledge and judgment, tax calculations, and agentic data-retrieval question answering.

Metric: Mean pass^5 (computed) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1GPT-5.5 Pro29.27
2GPT-5.4 Pro (xHigh)27.03
3GPT-5.524.43
4Claude Opus 4.621.37
5Gemini 3.1 Pro (Preview)20.1
6Grok 4.1 Fast19.63
7GPT-5.2 Pro17.77
8Gemini 3.1 Flash15.93
9Grok 4.20 (Reasoning)15.2
10Claude Opus 4.714.37
11Grok 4.2012.23
12Claude Sonnet 4.611.2
13GPT-5.49.33
14Gemini 2.5 Pro9
15Claude Sonnet 4.58.03

Interactive version: theaggregate.ai/benchmark?slug=taxbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.