OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward finds that VLM judges often approve failed computer-use runs. Its benchmark and open reward models offer a more grounded way to evaluate trajectories without paying frontier-model costs.
OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed.
Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring.
OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed. Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring. The authors report commercial-level judging at **30–60% lower cost** than frontier models, but the supplied material gives no per-platform scores or evidence that these reward models generalize to a builder’s own UI and task mix.
OSReward identifies false approval as a specific failure mode in computer-use judging and makes human-calibrated failure detection a prerequisite for trusting automated trajectory scores. It strengthens the prior case for replayable, step-aware evaluation while narrowing it: even a reproducible rollout is misleading if its judge systematically accepts failed runs.