Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
An analysis of OpenClaw discussions suggests agent evaluations miss what users value around delegation: cost, access, bounded reach, reviewability, and oversight.
Researchers used LLM assistance to analyze **73,093 first-person Reddit posts** about OpenClaw across **21 values in six groups**. Values clustered more around operating conditions than outputs; fulfillment appeared mostly in delivery descriptions, while unmet values concentrated in supervision descriptions.
When evaluating coding agents, measure the delegation envelope as well as task completion: cost, access, oversight, reviewability, and limits on agent reach can determine whether a run works for the user.
Researchers used LLM assistance to analyze **73,093 first-person Reddit posts** about OpenClaw across **21 values in six groups**. Values clustered more around operating conditions than outputs; fulfillment appeared mostly in delivery descriptions, while unmet values concentrated in supervision descriptions. When evaluating coding agents, measure the delegation envelope as well as task completion: cost, access, oversight, reviewability, and limits on agent reach can determine whether a run works for the user. The evidence comes from interpreted Reddit posts about one agent ecosystem, not controlled observations of agent runs. The abstract does not report annotation accuracy or establish that the patterns generalize to solo software development.
This expands agent evaluation from whether a task finished to whether delegation remained acceptable: access, cost, oversight, reviewability, and reach are part of success. It supports trace-based and production-failure evaluation, while narrowing generalization because the evidence is interpreted self-reports from one ecosystem rather than controlled runs or direct evidence about solo development.