Papers
Topics
Authors
Recent
Detailed Answer
Quick Answer
Concise responses based on abstracts
Detailed Answer
Thorough responses based on abstracts and some paper content
Custom Instructions Pro
Preferences or requirements that you'd like Emergent Mind to consider when generating responses
Gemini 2.5 Flash
Gemini 2.5 Flash
116 tokens/sec
GPT-4o
74 tokens/sec
Gemini 2.5 Pro Pro
62 tokens/sec
o3 Pro
18 tokens/sec
GPT-4.1 Pro
74 tokens/sec
DeepSeek R1 via Azure Pro
24 tokens/sec
2000 character limit reached

Look at Me When I Talk to You: A Video Dataset to Enable Voice Assistants to Recognize Errors (2104.07153v1)

Published 14 Apr 2021 in cs.HC

Abstract: People interacting with voice assistants are often frustrated by voice assistants' frequent errors and inability to respond to backchannel cues. We introduce an open-source video dataset of 21 participants' interactions with a voice assistant, and explore the possibility of using this dataset to enable automatic error recognition to inform self-repair. The dataset includes clipped and labeled videos of participants' faces during free-form interactions with the voice assistant from the smart speaker's perspective. To validate our dataset, we emulated a machine learning classifier by asking crowdsourced workers to recognize voice assistant errors from watching soundless video clips of participants' reactions. We found trends suggesting it is possible to determine the voice assistant's performance from a participant's facial reaction alone. This work posits elicited datasets of interactive responses as a key step towards improving error recognition for repair for voice assistants in a wide variety of applications.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Andrea Cuadra (3 papers)
  2. Hansol Lee (11 papers)
  3. Jason Cho (9 papers)
  4. Wendy Ju (35 papers)
Citations (8)