KAware - Hybrid Composition: leaderboard
Metric: The knowing-acting score (%, KAS): harmonic mean of knowing accuracy (Jaccard overlap between the tools the agent says the task needs and the gold tool set) and acting accuracy (Jaccard overlap between the tools it actually calls and the gold set), on the hybrid composition setting, on the KAware tool-use tasks, where tools are always available but the task may need them (external function, 310 tasks), mix them with parametric knowledge (hybrid composition, 368 tasks) or be solvable from knowledge alone (internal function, 398 tasks); higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 64.64 |
| 2 | GPT-5.5 | 63.14 |
| 3 | Claude Sonnet 4.5 (Thinking) | 58.61 |
| 4 | Qwen 3.5 397B A17B | 53.42 |
| 5 | Qwen 3 Max (Preview) | 53.22 |
| 6 | DeepSeek V4 Flash (Non-reasoning) | 51.6 |
| 7 | GPT-5 | 51.33 |
| 8 | Qwen 3.5 397B A17B (Non-reasoning) | 51.12 |
| 9 | O4 Mini (2025-04-16) | 50.32 |
| 10 | Qwen 3 235B A22B 2507 (Thinking) | 49.54 |
| 11 | DeepSeek V3.2 (Non-reasoning) | 49.26 |
| 12 | Qwen 3 235B A22B 2507 Instruct | 48.23 |
| 13 | Claude Sonnet 4.5 | 47.62 |
| 14 | DeepSeek R1 0528 | 46.9 |
| 15 | Gemini 3 Flash (Preview) | 45.31 |
Interactive version: theaggregate.ai/benchmark?slug=kaware-hybrid-composition · How It Works · Data refreshed daily, snapshot 2026-09-29.