{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2608.21325v1",
  "slug": "2608-21325v1-0ebpnes",
  "url": "https://feed7.dev/p/2608-21325v1-0ebpnes",
  "title": "Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy",
  "why_included": "A tool-exposed ontology steered models closer to human therapy patterns without fine-tuning, showing how explicit action vocabularies can improve agent behavior.",
  "summary": "The researchers define **10 therapeutic moves**, validated with **5 licensed psychologists**. Frontier models used inquiry at up to **3× the human rate**, while exposing the moves as tools improved turn-level alignment by **7–9 percentage points**.",
  "practical_implication": "For builders, this suggests representing desired behavior as explicit, callable actions rather than relying only on prose instructions. A compact action ontology can make agent behavior measurable and steerable without fine-tuning.",
  "agent_context": "The researchers define **10 therapeutic moves**, validated with **5 licensed psychologists**. Frontier models used inquiry at up to **3× the human rate**, while exposing the moves as tools improved turn-level alignment by **7–9 percentage points**.\n\nFor builders, this suggests representing desired behavior as explicit, callable actions rather than relying only on prose instructions. A compact action ontology can make agent behavior measurable and steerable without fine-tuning.\n\nThe evidence concerns psychotherapy, where behavioral alignment is safety-sensitive and human distributions are not automatically ideal outcomes. The material does not establish whether the method transfers to coding agents.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.21325v1",
    "published_at": "2026-08-21T17:32:38.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "agent",
  "domains": [
    "research"
  ],
  "topics": [
    "tool-use",
    "agent-evals",
    "agent-reliability"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The evidence concerns psychotherapy, where behavioral alignment is safety-sensitive and human distributions are not automatically ideal outcomes. The material does not establish whether the method transfers to coding agents."
  ],
  "connected_context": {
    "meaning": "This converts an expert-defined behavioral taxonomy into both an evaluation surface and a steering interface. It strengthens the case for encoding domain practice as observable actions rather than trusting prose instructions or model introspection, while narrowing the claim to turn-level therapeutic behavior: matching human move distributions does not itself establish clinical quality or transfer to other agents.",
    "corpus_size": 551,
    "generated_at": "2026-08-24T10:04:42.349Z",
    "connections": [
      {
        "title": "Trading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=Yphdry8ttAQ",
        "feed7_url": "https://feed7.dev/p/trading-desks-to-clinical-trials-parallels-in-applied-vertical-ai-ayush-1hwvsg3",
        "reason": "The expert-validated move ontology provides a concrete implementation of the vertical-AI requirement to encode expert workflow, while preserving that experts—not self-grading—must define useful behavior."
      },
      {
        "title": "Metacognition in LLMs: Foundations, Progress, and Opportunities",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.11881v1",
        "feed7_url": "https://feed7.dev/p/2607-11881v1-157hpaj",
        "reason": "Exposing therapeutic moves as callable actions offers an externally measurable control mechanism where the metacognition survey warns that internal self-inspection may be unreliable."
      },
      {
        "title": "MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.02520v1",
        "feed7_url": "https://feed7.dev/p/2608-02520v1-16negmk",
        "reason": "MedPRESS evaluates whether safe behavior persists across user pressure, complementing this paper’s move-level alignment with a multi-turn test of whether steering remains reliable under interaction."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-21T17:32:38.000Z",
  "modified_at": "2026-08-21T17:32:38.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-21325v1-0ebpnes",
    "json": "https://feed7.dev/p/2608-21325v1-0ebpnes.json",
    "markdown": "https://feed7.dev/p/2608-21325v1-0ebpnes.md"
  }
}