User Feedback Provides a Unique Signal that LLMs Can not Detect
User feedback helps models repair targeted faults, but LLM judges often miss those repairs. Agent evals should retain human feedback and avoid treating model preference as ground truth.
Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it.
Preserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers.
Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it. Preserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers. **LLM judges** often favored inferior baselines when only feedback revealed the necessary fix. The abstract reports significance but provides no effect sizes or task-level results.
This identifies user feedback as an evaluation signal that cannot safely be collapsed into an LLM’s preference between complete answers. It strengthens failure-driven evaluation by requiring the reported defect to be preserved and tested directly, while further limiting judge-only pipelines. Because the abstract omits effect sizes and task-level results, it establishes the direction and significance of the gap rather than its practical magnitude.