Learning From Free-Text Human Feedback: Collect New Datasets or Extend Existing Ones?
This presentation examines whether existing dialogue datasets can be extended with free-text human feedback annotations instead of collecting entirely new feedback datasets from scratch. The authors analyze six widely used dialogue corpora, develop refined error and user-response taxonomies, and demonstrate that human-bot, open-domain, and knowledge-grounded datasets contain substantial implicit feedback that can be extracted and used to improve conversational systems. The work offers practical guidance for researchers seeking to leverage existing dialogue data for learning from user corrections, clarifications, and alternative suggestions.Script
Should you collect a brand-new feedback dataset from scratch, or can you extend what already exists? The authors of this paper analyzed six major dialogue datasets and found that many already contain rich, implicit user feedback just waiting to be annotated.
Free-text human feedback is any user utterance that explains what went wrong, expresses dissatisfaction, provides a correction, adds knowledge, or suggests an alternative response. The challenge is identifying where these moments occur in existing datasets and whether they happen often enough to be useful.
The authors examined task-oriented, open-domain, and knowledge-grounded dialogues across human-human and human-bot interactions. They discovered that error and feedback patterns vary dramatically depending on dialogue type and whether users are talking to another person or a bot.
To make annotation efficient, the researchers developed an automatic filtering method using sentence similarity. By comparing user sentences to a set of error-indicating phrases, they concentrated errors into about 25 percent of the data, improving recall to 72 percent and dramatically reducing manual effort.
The authors refined an existing error taxonomy down to 10 clear categories and created a new five-category user-response taxonomy. The most useful feedback types are repeat or rephrase, make aware with correction, and ask for clarification, because they contain actionable information the system can learn from.
Human-bot, open-domain, and knowledge-grounded datasets already contain substantial implicit feedback that can be surfaced through targeted annotation. If you want to explore how existing dialogue data can be transformed into rich learning resources, visit EmergentMind.com to dive deeper into this research and create your own video summaries.