SORRY-Bench — leaderboard
Safety refusal benchmark with a balanced taxonomy of 44 safety categories and linguistic mutation tests for measuring when models comply with or refuse unsafe requests.
Source: sorry-bench.github.io.
Interactive version: theaggregate.ai/benchmark?slug=sorry-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.