Sign InOpen Brain
arXivPaperNeeds Review

User Feedback Provides a Unique Signal that LLMs Can not Detect

User feedback helps models repair targeted faults, but LLM judges often miss those repairs. Agent evals should retain human feedback and avoid treating model preference as ground truth.

arXiv · Sep 2, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it.

Practical Implication

Preserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers.

Agent-Ready Context
Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it.

Preserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers.

**LLM judges** often favored inferior baselines when only feedback revealed the necessary fix. The abstract reports significance but provides no effect sizes or task-level results.
Connected Context · Feed7 Judgment

This identifies user feedback as an evaluation signal that cannot safely be collapsed into an LLM’s preference between complete answers. It strengthens failure-driven evaluation by requiring the reported defect to be preserved and tested directly, while further limiting judge-only pipelines. Because the abstract omits effect sizes and task-level results, it establishes the direction and significance of the gap rather than its practical magnitude.

Verifiable Environments for AI in Biology — Kenny Workman, LatchBioThe finding reinforces LatchBio’s evidence that human review catches valid distinctions brittle automated graders miss, while generalizing the concern from scientific workflow plurality to user-identified response defects.Demystifying evals for AI agentsIt gives a specific reason to build evaluation tasks from real failures: the original user report may contain information that an LLM grader cannot reconstruct from competing final responses.Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiasBoth undermine unqualified reliance on LLM judges: this target shows judges can miss feedback-revealed fixes, while the candidate locates systematic judge failures in steerable internal representations.Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta SoftwareThe target sharpens this candidate’s warning that flexible model judges are fallible by showing that even when comparing revisions, they may prefer an inferior answer unless the user’s failure signal is evaluated directly.
Context Map
benchmark#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
**LLM judges** often favored inferior baselines when only feedback revealed the necessary fix. The abstract reports significance but provides no effect sizes or task-level results.