OSGuard: leaderboard

Metric: Macro-F1 (0-100) over the allowed, unrelated and unsafe labels on the OSGuard action-level benchmark: 324 human-labelled candidate actions from OSWorld-derived computer-use trajectories, each judged by a prompted multimodal guardrail from the original instruction, a screenshot plus accessibility tree and the proposed action, and classified as allowed, unrelated or unsafe; printed 0-1 and x100 here; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 3 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)80
2GPT-5.162
3Claude Sonnet 4.560

Interactive version: theaggregate.ai/benchmark?slug=osguard · How It Works · Data refreshed daily, snapshot 2026-09-29.