AIRTBench: leaderboard

AI red teaming benchmark evaluating language models' ability to autonomously discover and exploit AI/ML security vulnerabilities across 70 security challenges.

Source: huggingface.co.

Interactive version: theaggregate.ai/benchmark?slug=airtbench · How It Works · Data refreshed daily, snapshot 2026-09-05.