CHI-Bench — leaderboard

Healthcare workflow agent benchmark across prior authorization, utilization management, and care management, with agents operating simulated clinical apps through MCP-style tools.

Metric: Overall Pass@1 (%). Source: actava.ai. Status: saturation imminent. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.837.3
2Claude Opus 4.628
3Claude Sonnet 4.626.2
4Claude Opus 4.724.4
5GPT-5.520.9
6Claude Sonnet 520
7GLM-5.118.7
8GLM-5.218.7
9GPT-5.416
10Kimi K2.615.6
11DeepSeek V4 Pro14.2
12Gemini 3 Flash12.5
13GPT-5.4 Mini8.4
14Gemini 3.1 Pro (Preview)7.1
15Claude Haiku 4.56.2

Interactive version: theaggregate.ai/benchmark?slug=chi-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.