Skip to content
Advertisement
Image

OSWorld-Eval-Results

OSWorld-Eval-Results

OSWorld-Eval-Results Evaluation results of the UI-MOPD trained model (Qwen3-VL-8B-Thinking) on the OSWorld benchmark. Contains full execution trajectories including screenshots, action logs, and task outcomes. Evaluation Summary Metric Value Model Qwen3-VL-8B-Thinking Total Tasks 359 Successful 126 Success Rate 35.1% Action Space pyautogui Observation Screenshot (1920×1080) Max Steps 50 Coordinate Relative Per-Application… See the full description on the dataset page:

Source: Hugging Face Hub (UI-MOPD/OSWorld-Eval-Results). Metadata imported from the dataset’s Hub tags.

Advertisement