SORRY-Bench: leaderboard
Safety refusal benchmark with a balanced taxonomy of 44 safety categories and linguistic mutation tests for measuring when models comply with or refuse unsafe requests.
Source: sorry-bench.github.io.
Interactive version: theaggregate.ai/benchmark?slug=sorry-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.