OSGuard: leaderboard
Metric: Macro-F1 (0-100) over the allowed, unrelated and unsafe labels on the OSGuard action-level benchmark: 324 human-labelled candidate actions from OSWorld-derived computer-use trajectories, each judged by a prompted multimodal guardrail from the original instruction, a screenshot plus accessibility tree and the proposed action, and classified as allowed, unrelated or unsafe; printed 0-1 and x100 here; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 80 |
| 2 | GPT-5.1 | 62 |
| 3 | Claude Sonnet 4.5 | 60 |
Interactive version: theaggregate.ai/benchmark?slug=osguard · How It Works · Data refreshed daily, snapshot 2026-09-29.