{
  "schema_version": "1.1",
  "id": "s13:https://arxiv.org/abs/2609.02859v1",
  "slug": "2609-02859v1-121q3oo",
  "url": "https://feed7.dev/p/2609-02859v1-121q3oo",
  "title": "User Feedback Provides a Unique Signal that LLMs Can not Detect",
  "why_included": "User feedback helps models repair targeted faults, but LLM judges often miss those repairs. Agent evals should retain human feedback and avoid treating model preference as ground truth.",
  "summary": "Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it.",
  "practical_implication": "Preserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers.",
  "agent_context": "Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it.\n\nPreserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers.\n\n**LLM judges** often favored inferior baselines when only feedback revealed the necessary fix. The abstract reports significance but provides no effect sizes or task-level results.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2609.02859v1",
    "published_at": "2026-09-02T17:42:44.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [],
  "topics": [
    "agent-evals",
    "benchmark-integrity",
    "agent-reliability"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "**LLM judges** often favored inferior baselines when only feedback revealed the necessary fix. The abstract reports significance but provides no effect sizes or task-level results."
  ],
  "connected_context": {
    "meaning": "This identifies user feedback as an evaluation signal that cannot safely be collapsed into an LLM’s preference between complete answers. It strengthens failure-driven evaluation by requiring the reported defect to be preserved and tested directly, while further limiting judge-only pipelines. Because the abstract omits effect sizes and task-level results, it establishes the direction and significance of the gap rather than its practical magnitude.",
    "corpus_size": 669,
    "generated_at": "2026-09-03T10:01:24.920Z",
    "connections": [
      {
        "title": "Verifiable Environments for AI in Biology — Kenny Workman, LatchBio",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=3ZMUiFaQ3qg",
        "feed7_url": "https://feed7.dev/p/verifiable-environments-for-ai-in-biology-kenny-workman-latchbio-1vs6y66",
        "reason": "The finding reinforces LatchBio’s evidence that human review catches valid distinctions brittle automated graders miss, while generalizing the concern from scientific workflow plurality to user-identified response defects."
      },
      {
        "title": "Demystifying evals for AI agents",
        "source_name": "Anthropic",
        "source_url": "https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents",
        "feed7_url": "https://feed7.dev/p/demystifying-evals-for-ai-agents-1kh2tdz",
        "reason": "It gives a specific reason to build evaluation tasks from real failures: the original user report may contain information that an LLM grader cannot reconstruct from competing final responses."
      },
      {
        "title": "Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.11871v1",
        "feed7_url": "https://feed7.dev/p/2607-11871v1-17vejh0",
        "reason": "Both undermine unqualified reliance on LLM judges: this target shows judges can miss feedback-revealed fixes, while the candidate locates systematic judge failures in steerable internal representations."
      },
      {
        "title": "Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=2aS7aKoXn64",
        "feed7_url": "https://feed7.dev/p/rethinking-environments-for-long-horizon-work-rayan-garg-theta-software-11r7wbx",
        "reason": "The target sharpens this candidate’s warning that flexible model judges are fallible by showing that even when comparing revisions, they may prefer an inferior answer unless the user’s failure signal is evaluated directly."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-09-02T17:42:44.000Z",
  "modified_at": "2026-09-02T17:42:44.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2609-02859v1-121q3oo",
    "json": "https://feed7.dev/p/2609-02859v1-121q3oo.json",
    "markdown": "https://feed7.dev/p/2609-02859v1-121q3oo.md"
  }
}