GUI-Owl-7B: benchmark results
Provider: Other.
Unified ELO 1549 ± 23, rank #594 of 1605 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LabOSBench - Preparation & Sample Handling | 85.3 | Subtask Success Rate (%; mean over the category's operation- | 68.8 |
| LabOSBench | 64 | Subtask Success Rate (%; subtask-averaged per instrument, ma | 62.5 |
| LabOSBench - Measurement Configuration | 64.8 | Subtask Success Rate (%; mean over the category's operation- | 62.5 |
| DocOS - Easy | 9.88 | Task completion rate (%; share of tasks whose final applicat | 60 |
| LabOSBench - Post-processing & Completion | 68.8 | Subtask Success Rate (%; mean over the category's operation- | 56.2 |
| DocOS - Hard | 3.85 | Task completion rate (%; share of tasks whose final applicat | 50 |
| LabOSBench - Experimental Execution | 50 | Subtask Success Rate (%; mean over the category's operation- | 43.8 |
| AndroidWorld | 66.4 | Success Rate pass@1 (%) | 41.5 |
| DocOS - Medium | 6.15 | Task completion rate (%; share of tasks whose final applicat | 40 |
| TVWorld-N | 5.4 | Success Rate (%; 500 tasks: 5 Google TV graphs x 50 start-go | 25 |
| AndroidIntent - Complete Instructions | 50.4 | Cumulative success rate (%; offline step-wise execution of 7 | 20 |
| GUI-RobustEval - Error Depth 0 | 28.7 | Post-error success rate (%; the agent takes over a desktop t | 20 |
Interactive version: theaggregate.ai/model?slug=gui-owl-7b · How It Works · Data refreshed daily, snapshot 2026-09-26.