Finding the Right Fit - TUA-Bench (OpenHands): leaderboard
Metric: Task score (%), OpenHands harness, high reasoning effort, one run per task over 120 tasks. Source: huggingface.co. Saturation forecast: Around 2029. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (High) | 65.23 |
| 2 | GLM-5.3 (High) | 62.86 |
| 3 | GPT-6 Astra (High) | 61.01 |
| 4 | DeepSeek V4 Pro (High) | 56.49 |
| 5 | Kimi K3 (High) | 50.54 |
Interactive version: theaggregate.ai/benchmark?slug=finding-the-right-fit-tua-bench-openhands · How It Works · Data refreshed daily, snapshot 2026-10-05.