User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
This presentation examines how implicit user feedback in real human-LLM conversations can reveal user intent and dissatisfaction, but proves surprisingly unreliable as a training signal. We explore when and why users provide feedback, the challenges of detecting it automatically, and the paradox that feedback-aware response generation helps on simple tasks but often degrades performance on complex ones—suggesting that stronger-model distillation may be more reliable than trying to learn from noisy human corrections.Script
When users chat with language models, their follow-up messages often reveal hidden signals: rephrased requests, corrections, or clarifications that expose what went wrong. The researchers behind this paper ask whether these implicit feedback signals, already embedded in conversation logs, can be detected automatically and used to improve models without interrupting the user experience.
Negative feedback becomes dramatically more common as conversations lengthen. In multi-turn dialogues, users repeatedly correct, rephrase, or ask for clarification, and the researchers found this pattern holds across thousands of real interactions from both Chatbot Arena and a free GPT service.
Here is where things get counterintuitive. Positive feedback turned out to be the least reliable signal of all. The authors discovered that users often praised models for completing jailbreak requests or generating inappropriate content, making positive feedback a dangerous training target rather than a safe reward.
The researchers tested whether a stronger model could use user feedback to generate better corrections. Surprisingly, responses generated with explicit feedback did not consistently beat responses regenerated from scratch using only the original prompt, suggesting that simply asking a better model may be more reliable than trying to incorporate noisy user corrections.
Training on feedback-aware responses improved performance on MT-Bench, a set of short, human-designed questions. But on WildBench, with its longer and more complex real-world tasks, the same approach often degraded performance. Models trained on feedback struggled to follow clear instructions compared to models distilled from stronger systems without feedback.
User feedback in natural conversations is both a lens and a trap. It reveals what users wanted and where models failed, but its noise, context-dependence, and potential to reward harmful behavior make it unreliable as a standalone learning signal. Visit EmergentMind.com to explore this research further and create your own video presentations about the latest papers.