Sign InOpen Brain
arXivPaperNeeds Review

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

OSReward finds that VLM judges often approve failed computer-use runs. Its benchmark and open reward models offer a more grounded way to evaluate trajectories without paying frontier-model costs.

arXiv · Jul 30, 2026
Open Source Open MarkdownOpen JSON
Source Summary

OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed.

Practical Implication

Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring.

Agent-Ready Context
OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed.

Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring.

The authors report commercial-level judging at **30–60% lower cost** than frontier models, but the supplied material gives no per-platform scores or evidence that these reward models generalize to a builder’s own UI and task mix.
Connected Context · Feed7 Judgment

OSReward identifies false approval as a specific failure mode in computer-use judging and makes human-calibrated failure detection a prerequisite for trusting automated trajectory scores. It strengthens the prior case for replayable, step-aware evaluation while narrowing it: even a reproducible rollout is misleading if its judge systematically accepts failed runs.

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?Desktop-Delta Bench exposes missed GUI transitions; OSReward adds the consequence for evaluation pipelines, showing that trajectory judges themselves may approve failures and therefore need calibration.Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude InstituteHarbor’s reproducible rollout loop provides the evaluation setting, while OSReward shows that outcome verification within that loop must be checked against human-labeled failures.The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AIOSReward supplies evidence for retaining calibrated checks around model-based trajectory analysis: more flexible judges do not remove the risk of systematic leniency.From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AIReplayable production traces can test reward models under a team’s actual UI and task mix, addressing OSReward’s unresolved generalization beyond its evaluated platforms.
Context Map
benchmarkcoding#computer-use#agent-evals#agent-reliability
Uncertainty
The authors report commercial-level judging at **30–60% lower cost** than frontier models, but the supplied material gives no per-platform scores or evidence that these reward models generalize to a builder’s own UI and task mix.