GUI-Owl-7B: benchmark results

Provider: Other.

Unified ELO 1549 ± 23, rank #594 of 1605 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LabOSBench - Preparation & Sample Handling85.3Subtask Success Rate (%; mean over the category's operation-68.8
LabOSBench64Subtask Success Rate (%; subtask-averaged per instrument, ma62.5
LabOSBench - Measurement Configuration64.8Subtask Success Rate (%; mean over the category's operation-62.5
DocOS - Easy9.88Task completion rate (%; share of tasks whose final applicat60
LabOSBench - Post-processing & Completion68.8Subtask Success Rate (%; mean over the category's operation-56.2
DocOS - Hard3.85Task completion rate (%; share of tasks whose final applicat50
LabOSBench - Experimental Execution50Subtask Success Rate (%; mean over the category's operation-43.8
AndroidWorld66.4Success Rate pass@1 (%)41.5
DocOS - Medium6.15Task completion rate (%; share of tasks whose final applicat40
TVWorld-N5.4Success Rate (%; 500 tasks: 5 Google TV graphs x 50 start-go25
AndroidIntent - Complete Instructions50.4Cumulative success rate (%; offline step-wise execution of 720
GUI-RobustEval - Error Depth 028.7Post-error success rate (%; the agent takes over a desktop t20

Interactive version: theaggregate.ai/model?slug=gui-owl-7b · How It Works · Data refreshed daily, snapshot 2026-09-26.