Sign InOpen Brain
arXivPaperNeeds Review

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

An analysis of OpenClaw discussions suggests agent evaluations miss what users value around delegation: cost, access, bounded reach, reviewability, and oversight.

arXiv · Sep 18, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Researchers used LLM assistance to analyze **73,093 first-person Reddit posts** about OpenClaw across **21 values in six groups**. Values clustered more around operating conditions than outputs; fulfillment appeared mostly in delivery descriptions, while unmet values concentrated in supervision descriptions.

Practical Implication

When evaluating coding agents, measure the delegation envelope as well as task completion: cost, access, oversight, reviewability, and limits on agent reach can determine whether a run works for the user.

Agent-Ready Context
Researchers used LLM assistance to analyze **73,093 first-person Reddit posts** about OpenClaw across **21 values in six groups**. Values clustered more around operating conditions than outputs; fulfillment appeared mostly in delivery descriptions, while unmet values concentrated in supervision descriptions.

When evaluating coding agents, measure the delegation envelope as well as task completion: cost, access, oversight, reviewability, and limits on agent reach can determine whether a run works for the user.

The evidence comes from interpreted Reddit posts about one agent ecosystem, not controlled observations of agent runs. The abstract does not report annotation accuracy or establish that the patterns generalize to solo software development.
Connected Context · Feed7 Judgment

This expands agent evaluation from whether a task finished to whether delegation remained acceptable: access, cost, oversight, reviewability, and reach are part of success. It supports trace-based and production-failure evaluation, while narrowing generalization because the evidence is interpreted self-reports from one ecosystem rather than controlled runs or direct evidence about solo development.

From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AIReplayable simulations already measure cost, latency, retries, and outcomes; this study argues that access, oversight, reviewability, and reach should join those release gates.Quantifying Overclaiming Propensity in Frontier LLM AgentsOverclaiming makes reviewability operationally important: users cannot supervise delegation reliably when completion reports conceal unread files or unsupported coverage.Inside 847 Production Clinical AI Notes — Sebastian Fox, ComposoThe clinical evidence reinforces the supervision finding by showing that plausible outputs and generic judges can miss consequential failures requiring expert review.Designing Agents (The Floor Is the Frontier) — Ben Hylak, RaindropProduction-failure monitoring supplies a complementary method for turning value concerns reported by users into measurable regressions with onset and affected-user scope.
Context Map
benchmark#agent-evals#agent-reliability#adoption
Uncertainty
The evidence comes from interpreted Reddit posts about one agent ecosystem, not controlled observations of agent runs. The abstract does not report annotation accuracy or establish that the patterns generalize to solo software development.