ImmersedPrivacy - Inferred Privacy Planning (Audio as Text): leaderboard

Metric: Exact match (%; Tier 3: share of cases where the model selects exactly the two task actions that leave alone the item an observed person wanted kept private; audio replaced by a text description of the soundscape (Tier 2) or the dialogue transcript (Tier 3)). Source: arxiv.org. Saturation forecast: Around March 2027. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)53
2Gemini 3.1 Pro (Preview) (Low)50
3Qwen 3.5 27B (Non-reasoning)30
4Qwen 3.5 27B (Thinking)18
5Gemini 3 Flash (Preview) (High)18
6Gemini 3 Flash (Preview) (Low)11

Interactive version: theaggregate.ai/benchmark?slug=immersedprivacy-inferred-privacy-planning-audio-as-text · How It Works · Data refreshed daily, snapshot 2026-09-26.