KAware - Internal Function Pass Rate: leaderboard
Metric: The end-to-end pass rate (%): share of tasks whose trajectory and final answer a gpt-4.1 judge marks correct, on the internal function setting, on the KAware tool-use tasks, where tools are always available but the task may need them (external function, 310 tasks), mix them with parametric knowledge (hybrid composition, 368 tasks) or be solvable from knowledge alone (internal function, 398 tasks); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 (Thinking) | 96.23 |
| 2 | Qwen 3 Max (Preview) | 95.98 |
| 3 | Qwen 3 235B A22B 2507 (Thinking) | 93.22 |
| 4 | Qwen 3 235B A22B 2507 Instruct | 90.45 |
| 5 | Claude Sonnet 4.5 | 87.94 |
| 6 | GPT-5.5 | 86.68 |
| 7 | DeepSeek V3.2 (Non-reasoning) | 86.18 |
| 8 | Qwen 3.5 397B A17B | 84.92 |
| 9 | Qwen 3.5 397B A17B (Non-reasoning) | 84.67 |
| 10 | DeepSeek V4 Flash (Non-reasoning) | 83.42 |
| 11 | GPT-5 | 60.8 |
| 12 | O4 Mini (2025-04-16) | 60.05 |
| 13 | GPT-4o (2024-11-20) | 59.3 |
| 14 | GPT-4.1 | 59.05 |
| 15 | Gemini 3 Flash (Preview) | 59.05 |
Interactive version: theaggregate.ai/benchmark?slug=kaware-internal-function-pass-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.