# User Feedback Provides a Unique Signal that LLMs Can not Detect

Source: [arXiv](https://arxiv.org/abs/2609.02859v1)  
Feed7 permalink: https://feed7.dev/p/2609-02859v1-121q3oo  
Published: 2026-09-02T17:42:44.000Z  
Trust: Needs Review (needs_review)

## Why Included

User feedback helps models repair targeted faults, but LLM judges often miss those repairs. Agent evals should retain human feedback and avoid treating model preference as ground truth.

## Source Summary

Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it.

## Practical Implication

Preserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers.

## Agent-Ready Context

Across **synthetic ground-truth and naturalistic data**, revisions given user feedback fixed targeted issues more often than revisions produced without it.

Preserve user reports as a distinct signal when improving an agent. Test the reported failure directly instead of asking another model only which complete response it prefers.

**LLM judges** often favored inferior baselines when only feedback revealed the necessary fix. The abstract reports significance but provides no effect sizes or task-level results.

## Connected Context

Feed7 judgment across 669 accumulated Signals:

This identifies user feedback as an evaluation signal that cannot safely be collapsed into an LLM’s preference between complete answers. It strengthens failure-driven evaluation by requiring the reported defect to be preserved and tested directly, while further limiting judge-only pipelines. Because the abstract omits effect sizes and task-level results, it establishes the direction and significance of the gap rather than its practical magnitude.

- [Verifiable Environments for AI in Biology — Kenny Workman, LatchBio](https://feed7.dev/p/verifiable-environments-for-ai-in-biology-kenny-workman-latchbio-1vs6y66) — The finding reinforces LatchBio’s evidence that human review catches valid distinctions brittle automated graders miss, while generalizing the concern from scientific workflow plurality to user-identified response defects.
- [Demystifying evals for AI agents](https://feed7.dev/p/demystifying-evals-for-ai-agents-1kh2tdz) — It gives a specific reason to build evaluation tasks from real failures: the original user report may contain information that an LLM grader cannot reconstruct from competing final responses.
- [Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias](https://feed7.dev/p/2607-11871v1-17vejh0) — Both undermine unqualified reliance on LLM judges: this target shows judges can miss feedback-revealed fixes, while the candidate locates systematic judge failures in steerable internal representations.
- [Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software](https://feed7.dev/p/rethinking-environments-for-long-horizon-work-rayan-garg-theta-software-11r7wbx) — The target sharpens this candidate’s warning that flexible model judges are fallible by showing that even when comparing revisions, they may prefer an inferior answer unless the user’s failure signal is evaluated directly.

## Context Map

- Layer: benchmark
- Domains: None
- Topics: agent-evals, benchmark-integrity, agent-reliability

## Uncertainty

- **LLM judges** often favored inferior baselines when only feedback revealed the necessary fix. The abstract reports significance but provides no effect sizes or task-level results.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
