IHBench - Task Fulfillment: leaderboard
Metric: Win rate against GPT-4o Audio (%; share of 428 interruption points where a GPT-5.4-mini judge (high reasoning) prefers the model's next turn for advancing the workflow, audio input, mean of three epochs). Source: arxiv.org. Saturation forecast: Around December 2026. 27 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT Audio | 64.4 |
| 2 | Gemini 2.5 Flash (Thinking) | 58.6 |
| 3 | Gemini 3.1 Pro (Preview) | 58.2 |
| 4 | Gemini 2.5 Pro | 52.6 |
| 5 | Gemma 4 12B (Reasoning) | 51.1 |
| 6 | Gemma 4 12B (Non-reasoning) | 50.5 |
| 7 | Gemini 2.5 Flash (Non-reasoning) | 48.8 |
| 8 | GPT Audio Mini | 48.4 |
| 9 | Voxtral-Small-24B-2507 | 30.8 |
| 10 | Qwen3 Omni 30B A3B Instruct | 30.4 |
| 11 | Qwen2.5-Omni-7B | 18.1 |
| 12 | Phi-4 Multimodal Instruct | 10.4 |
| 13 | Qwen2-Audio-7B-Instruct | 4.4 |
Interactive version: theaggregate.ai/benchmark?slug=ihbench-task-fulfillment · How It Works · Data refreshed daily, snapshot 2026-09-26.