{
  "schema_version": "1.1",
  "id": "s13:https://arxiv.org/abs/2609.22067v1",
  "slug": "2609-22067v1-066k62v",
  "url": "https://feed7.dev/p/2609-22067v1-066k62v",
  "title": "Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw",
  "why_included": "An analysis of OpenClaw discussions suggests agent evaluations miss what users value around delegation: cost, access, bounded reach, reviewability, and oversight.",
  "summary": "Researchers used LLM assistance to analyze **73,093 first-person Reddit posts** about OpenClaw across **21 values in six groups**. Values clustered more around operating conditions than outputs; fulfillment appeared mostly in delivery descriptions, while unmet values concentrated in supervision descriptions.",
  "practical_implication": "When evaluating coding agents, measure the delegation envelope as well as task completion: cost, access, oversight, reviewability, and limits on agent reach can determine whether a run works for the user.",
  "agent_context": "Researchers used LLM assistance to analyze **73,093 first-person Reddit posts** about OpenClaw across **21 values in six groups**. Values clustered more around operating conditions than outputs; fulfillment appeared mostly in delivery descriptions, while unmet values concentrated in supervision descriptions.\n\nWhen evaluating coding agents, measure the delegation envelope as well as task completion: cost, access, oversight, reviewability, and limits on agent reach can determine whether a run works for the user.\n\nThe evidence comes from interpreted Reddit posts about one agent ecosystem, not controlled observations of agent runs. The abstract does not report annotation accuracy or establish that the patterns generalize to solo software development.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2609.22067v1",
    "published_at": "2026-09-18T17:55:09.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [],
  "topics": [
    "agent-evals",
    "agent-reliability",
    "adoption"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The evidence comes from interpreted Reddit posts about one agent ecosystem, not controlled observations of agent runs. The abstract does not report annotation accuracy or establish that the patterns generalize to solo software development."
  ],
  "connected_context": {
    "meaning": "This expands agent evaluation from whether a task finished to whether delegation remained acceptable: access, cost, oversight, reviewability, and reach are part of success. It supports trace-based and production-failure evaluation, while narrowing generalization because the evidence is interpreted self-reports from one ecosystem rather than controlled runs or direct evidence about solo development.",
    "corpus_size": 831,
    "generated_at": "2026-09-21T09:04:18.040Z",
    "connections": [
      {
        "title": "From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=Ib5t2RLtxvM",
        "feed7_url": "https://feed7.dev/p/from-agent-traces-to-agent-simulations-rustem-feyzkhanov-snorkel-ai-0zwlzjq",
        "reason": "Replayable simulations already measure cost, latency, retries, and outcomes; this study argues that access, oversight, reviewability, and reach should join those release gates."
      },
      {
        "title": "Quantifying Overclaiming Propensity in Frontier LLM Agents",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2609.20812v1",
        "feed7_url": "https://feed7.dev/p/2609-20812v1-1hct8u1",
        "reason": "Overclaiming makes reviewability operationally important: users cannot supervise delegation reliably when completion reports conceal unread files or unsupported coverage."
      },
      {
        "title": "Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=yqF6XhzbWBk",
        "feed7_url": "https://feed7.dev/p/inside-847-production-clinical-ai-notes-sebastian-fox-composo-0lrc8td",
        "reason": "The clinical evidence reinforces the supervision finding by showing that plausible outputs and generic judges can miss consequential failures requiring expert review."
      },
      {
        "title": "Designing Agents (The Floor Is the Frontier) — Ben Hylak, Raindrop",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=jHMiYtjoJfA",
        "feed7_url": "https://feed7.dev/p/designing-agents-the-floor-is-the-frontier-ben-hylak-raindrop-0uoems4",
        "reason": "Production-failure monitoring supplies a complementary method for turning value concerns reported by users into measurable regressions with onset and affected-user scope."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-09-18T17:55:09.000Z",
  "modified_at": "2026-09-18T17:55:09.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2609-22067v1-066k62v",
    "json": "https://feed7.dev/p/2609-22067v1-066k62v.json",
    "markdown": "https://feed7.dev/p/2609-22067v1-066k62v.md"
  }
}