CHI-Bench: leaderboard

Healthcare workflow agent benchmark across prior authorization, utilization management, and care management, with agents operating simulated clinical apps through MCP-style tools.

Metric: Overall Pass@1 (%). Source: actava.ai. Status: years away from saturation. 26 models tracked.

Top models

#ModelScore
1Claude Opus 554.7
2Claude Opus 4.837.3
3Claude Opus 4.628
4Claude Sonnet 4.626.2
5GPT-5.6 Sol25.3
6Kimi K325.3
7Claude Opus 4.724.4
8Claude Fable 524
9GPT-5.520.9
10Claude Sonnet 520
11GLM-5.118.7
12GLM-5.218.7
13Qwen 3.6 Max16.4
14GPT-5.416
15Kimi K2.615.6

Interactive version: theaggregate.ai/benchmark?slug=chi-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.