CAIS Risk Index — leaderboard

Composite CAIS AI Dashboard risk index averaging VCT refusal risk, HLE miscalibration, MASK risk, Machiavelli, and TextQuests Harm for models with all component scores. Lower is better.

Metric: Risk Index (self-reported). Source: benchmarklist.com. 37 models tracked.

Top models

#ModelScore
1GPT-4o (2024-11-20)67
2Gemini 2.5 Flash Lite66.4
3Gemini 3.1 Flash Lite (Preview)61.7
4Gemini 2.5 Flash60.1
5O3 Mini (High)60.1
6Gemini 3 Flash (Preview)59.4
7Gemini 2.5 Pro59
8Gemini 3.5 Flash58.2
9Kimi K2 (Thinking)57.4
10DeepSeek V3.256.9
11Gemini 3.1 Pro (Preview)55.6
12Grok 4 Fast55.3
13DeepSeek V4 Pro54.1
14Grok 4.1 Fast52.2
15GPT-5 Nano51.9

Interactive version: theaggregate.ai/benchmark?slug=cais-risk-index · How the rankings work · Data refreshed daily, snapshot 2026-07-22.