Papers
Topics
Authors
Recent
Search
2000 character limit reached

AQuA: Automated Question-Answering in Software Tutorial Videos with Visual Anchors

Published 8 Mar 2024 in cs.HC | (2403.05213v1)

Abstract: Tutorial videos are a popular help source for learning feature-rich software. However, getting quick answers to questions about tutorial videos is difficult. We present an automated approach for responding to tutorial questions. By analyzing 633 questions found in 5,944 video comments, we identified different question types and observed that users frequently described parts of the video in questions. We then asked participants (N=24) to watch tutorial videos and ask questions while annotating the video with relevant visual anchors. Most visual anchors referred to UI elements and the application workspace. Based on these insights, we built AQuA, a pipeline that generates useful answers to questions with visual anchors. We demonstrate this for Fusion 360, showing that we can recognize UI elements in visual anchors and generate answers using GPT-4 augmented with that visual information and software documentation. An evaluation study (N=16) demonstrates that our approach provides better answers than baseline methods.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (79)
  1. Benjamin Alcott. 2017. Does Teacher Encouragement Influence Students’ Educational Progress? A Propensity-Score Matching Analysis. Research in Higher Education 58, 7 (Jan 2017), 773–804. https://doi.org/10.1007/s11162-017-9446-2
  2. Amazon. 2023. Amazon Transcribe. https://aws.amazon.com/transcribe/. Accessed: 2023-09-09.
  3. Autodesk. 2023a. Autodesk Screencast. https://www.autodesk.com/support/technical/article/caas/tsarticles/ts/71QzAbqSskV6l5Hm0ULMMp.html. Accessed: 2023-09-10.
  4. Autodesk. 2023b. Fusion 360 Product Documentation. https://help.autodesk.com/view/fusion360/ENU. Accessed: 2023-09-10.
  5. Waken: Reverse Engineering Usage Information and Interface Structure from Software Videos. In Proceedings of the 25th Annual ACM Symposium on User Interface Software and Technology (Cambridge, Massachusetts, USA) (UIST ’12). Association for Computing Machinery, New York, NY, USA, 83–92. https://doi.org/10.1145/2380116.2380129
  6. VideoSticker: A Tool for Active Viewing and Visual Note-Taking from Videos. In 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, NY, USA, 672–690. https://doi.org/10.1145/3490099.3511132
  7. RubySlippers: Supporting Content-Based Voice Navigation for How-to Videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 97, 14 pages. https://doi.org/10.1145/3411764.3445131
  8. How to Design Voice Based Navigation for How-To Videos. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland UK) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–11. https://doi.org/10.1145/3290605.3300931
  9. Towards Complete Icon Labeling in Mobile Applications. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 387, 14 pages. https://doi.org/10.1145/3491102.3502073
  10. DemoCut: Generating Concise Instructional Videos for Physical Demonstrations. In Proceedings of the 26th Annual ACM Symposium on User Interface Software and Technology (St. Andrews, Scotland, United Kingdom) (UIST ’13). Association for Computing Machinery, New York, NY, USA, 141–150. https://doi.org/10.1145/2501988.2502052
  11. Korero: Facilitating Complex Referencing of Visual Materials in Asynchronous Discussion Interface. Proc. ACM Hum.-Comput. Interact. 1, CSCW, Article 34 (Dec 2017), 19 pages. https://doi.org/10.1145/3134669
  12. Beyond Show of Hands: Engaging Viewers via Expressive and Scalable Visual Communication in Live Streaming. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 109, 14 pages. https://doi.org/10.1145/3411764.3445419
  13. TutorialVQA: Question Answering Dataset for Tutorial Videos. In Proceedings of the Twelfth Language Resources and Evaluation Conference. European Language Resources Association, Marseille, France, 5450–5455. https://aclanthology.org/2020.lrec-1.670
  14. Rico: A Mobile App Dataset for Building Data-Driven Design Applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology (Québec City, QC, Canada) (UIST ’17). Association for Computing Machinery, New York, NY, USA, 845–854. https://doi.org/10.1145/3126594.3126651
  15. Temporal Segmentation of Creative Live Streams. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3313831.3376437
  16. ReMap: Multimodal Help-Seeking. In Adjunct Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology (New Orleans, LA, USA) (UIST ’19 Adjunct). Association for Computing Machinery, New York, NY, USA, 96–98. https://doi.org/10.1145/3332167.3356884
  17. RePlay: Contextually Presenting Learning Videos Across Software Applications (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3290605.3300527
  18. Mudslide: A Spatially Anchored Census of Student Confusion for Online Lecture Videos. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15). Association for Computing Machinery, New York, NY, USA, 1555–1564. https://doi.org/10.1145/2702123.2702304
  19. Google. 2023a. Cloud Vision API. https://cloud.google.com/vision/docs. Accessed: 2023-09-10.
  20. Google. 2023b. YouTube Data API. https://developers.google.com/youtube/v3. Accessed: 2023-09-10.
  21. MathBot: Transforming Online Resources for Learning Math into Conversational Interactions. https://api.semanticscholar.org/CorpusID:236143850
  22. Chronicle: Capture, Exploration, and Playback of Document Workflow Histories. In Proceedings of the 23nd Annual ACM Symposium on User Interface Software and Technology (New York, New York, USA) (UIST ’10). Association for Computing Machinery, New York, NY, USA, 143–152. https://doi.org/10.1145/1866029.1866054
  23. How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection. arXiv preprint arxiv:2301.07597 (2023).
  24. MicroMentor: Peer-to-Peer Software Help Sessions in Three Minutes or Less. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3313831.3376230
  25. Beyond ”One-Size-Fits-All”: Understanding the Diversity in How Software Newcomers Discover and Make Use of Help Resources. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland UK) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–14. https://doi.org/10.1145/3290605.3300570
  26. Crowdsourcing Step-by-Step Information Extraction to Enhance Existing How-to Videos. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada) (CHI ’14). Association for Computing Machinery, New York, NY, USA, 4017–4026. https://doi.org/10.1145/2556288.2556986
  27. HyperButton: In-Video Question Answering via Interactive Buttons and Hyperlinks. In Asian CHI Symposium 2021 (Yokohama, Japan) (Asian CHI Symposium 2021). Association for Computing Machinery, New York, NY, USA, 48–52. https://doi.org/10.1145/3429360.3468179
  28. Exploring Chart Question Answering for Blind and Low Vision Users. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 828, 15 pages. https://doi.org/10.1145/3544548.3581532
  29. Winder: Linking Speech and Visual Objects to Support Communication in Asynchronous Collaboration. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 453, 17 pages. https://doi.org/10.1145/3411764.3445686
  30. Understanding the Roles and Uses of Web Tutorials. Proceedings of the International AAAI Conference on Web and Social Media 7, 1 (Aug. 2021), 303–310. https://doi.org/10.1609/icwsm.v7i1.14413
  31. DAPIE: Interactive Step-by-Step Explanatory Dialogues to Answer Children’s Why and How Questions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 450, 22 pages. https://doi.org/10.1145/3544548.3581369
  32. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages.
  33. Gang Li and Yang Li. 2023. Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus. In Proceedings of the Eleventh International Conference on Learning Representations (ICLR’23).
  34. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). Article 814, 13 pages.
  35. Stargazer: An Interactive Camera Robot for Capturing How-To Videos Based on Subtle Instructor Cues. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 800, 16 pages. https://doi.org/10.1145/3544548.3580896
  36. Screencast Tutorial Video Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  37. Hero: Hierarchical Encoder for Video+Language Omni-representation Pre-training. In Conference on Empirical Methods in Natural Language Processing. https://api.semanticscholar.org/CorpusID:218470055
  38. Screen2Vec: Semantic Embedding of GUI Screens and GUI Components. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 578, 15 pages. https://doi.org/10.1145/3411764.3445049
  39. Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 761, 17 pages. https://doi.org/10.1145/3544548.3581006
  40. Joshua Lochner. 2023. Chat Downloader. https://chat-downloader.readthedocs.io/en/latest/. Accessed: 2023-09-10.
  41. StreamSketch: Exploring Multi-Modal Interactions in Creative Live Streams. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 58 (Apr 2021), 26 pages. https://doi.org/10.1145/3449132
  42. A classification scheme for content analyses of YouTube video comments. J. Documentation 69 (2013), 693–714. https://api.semanticscholar.org/CorpusID:206402932
  43. Supercharging Trial-and-Error for Learning Complex Software Applications. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 381, 13 pages. https://doi.org/10.1145/3491102.3501895
  44. Ambient Help. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Vancouver, BC, Canada) (CHI ’11). Association for Computing Machinery, New York, NY, USA, 2751–2760. https://doi.org/10.1145/1978942.1979349
  45. IP-QAT: In-Product Questions, Answers, & Tips. In Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology (Santa Barbara, California, USA) (UIST ’11). Association for Computing Machinery, New York, NY, USA, 175–184. https://doi.org/10.1145/2047196.2047218
  46. Teaching language models to support answers with verified quotes. arXiv:2203.11147 [cs.CL]
  47. Cuong Nguyen and Feng Liu. 2015. Making Software Tutorial Video Responsive (CHI ’15). Association for Computing Machinery, New York, NY, USA, 1565–1568. https://doi.org/10.1145/2702123.2702209
  48. OpenAI. 2023a. ChatGPT. https://openai.com/blog/chatgpt. Accessed: 2023-09-09.
  49. OpenAI. 2023b. GPT-4. https://openai.com/gpt-4. Accessed: 2023-09-09.
  50. OpenAI. 2023c. text-embedding-ada-002. https://openai.com/blog/new-and-improved-embedding-model. Accessed: 2023-09-09.
  51. OpenCV. 2023a. Feature Matching. https://docs.opencv.org/4.x/dc/dc3/tutorial_py_matcher.html. Accessed: 2023-09-09.
  52. OpenCV. 2023b. Template Matching. https://docs.opencv.org/4.x/d4/dc6/tutorial_py_template_matching.html. Accessed: 2023-09-09.
  53. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Association for Computing Machinery, New York, NY, USA, Article 2, 22 pages. https://doi.org/10.1145/3586183.3606763
  54. Analyzing User Comments on YouTube Coding Tutorial Videos. In 2017 IEEE/ACM 25th International Conference on Program Comprehension (ICPC). 196–206. https://doi.org/10.1109/ICPC.2017.26
  55. Rhitabrat Pokharel and Dixit Bhatta. 2021. Classifying YouTube Comments Based on Sentiment and Type of Sentence. arXiv:2111.01908 [cs.IR]
  56. Pause-and-Play: Automatically Linking Screencast Video Tutorials with Applications (UIST ’11). Association for Computing Machinery, New York, NY, USA, 135–144. https://doi.org/10.1145/2047196.2047213
  57. Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). Article 1182, 27 pages.
  58. A Survey of Evaluation Metrics Used for NLG Systems. ACM Comput. Surv. 55, 2, Article 26 (Jan 2022), 39 pages. https://doi.org/10.1145/3485766
  59. Leave a Comment! An In-Depth Analysis of User Comments on YouTube. In Wirtschaftsinformatik. https://api.semanticscholar.org/CorpusID:11966671
  60. Understanding the Effect of In-Video Prompting on Learners and Instructors. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3173574.3173893
  61. Visual Transcripts: Lecture Notes from Blackboard-Style Lecture Videos. ACM Trans. Graph. 34, 6, Article 240 (Nov 2015), 10 pages. https://doi.org/10.1145/2816795.2818123
  62. GVQA: Learning to Answer Questions about Graphs with Visualizations via Knowledge Base. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 464, 16 pages. https://doi.org/10.1145/3544548.3581067
  63. Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup Videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 108, 16 pages. https://doi.org/10.1145/3411764.3445721
  64. Enabling Conversational Interaction with Mobile UI Using Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 432, 17 pages. https://doi.org/10.1145/3544548.3580895
  65. Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association for Computing Machinery, New York, NY, USA, 498–510. https://doi.org/10.1145/3472749.3474765
  66. Sara, the Lecturer: Improving Learning in Online Education with a Scaffolding-Based Conversational Agent. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–14. https://doi.org/10.1145/3313831.3376781
  67. WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 286, 14 pages. https://doi.org/10.1145/3544548.3581158
  68. UIED: A Hybrid Tool for GUI Element Detection. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Virtual Event, USA) (ESEC/FSE 2020). Association for Computing Machinery, New York, NY, USA, 1655–1659. https://doi.org/10.1145/3368089.3417940
  69. Just Ask: Learning to Answer Questions from Millions of Narrated Videos. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2020), 1666–1677. https://api.semanticscholar.org/CorpusID:227238996
  70. Saelyne Yang and Juho Kim. 2020. What Makes It Hard for Users to Follow Software Tutorial Videos?. In Proceedings of HCI Korea 2020. The HCI Society of KOREA, South Korea, 531–536.
  71. Beyond Instructions: A Taxonomy of Information Types in How-to Videos. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 797, 21 pages. https://doi.org/10.1145/3544548.3581126
  72. Snapstream: Snapshot-Based Interaction in Live Streaming for Visual Art. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3313831.3376390
  73. SoftVideo: Improving the Learning Experience of Software Tutorial Videos with Collective Interaction Data. In 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, NY, USA, 646–660. https://doi.org/10.1145/3490099.3511106
  74. ”Can You Believe [1:21]?!”: Content and Time-Based Reference Patterns in Video Comments. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland UK) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3290605.3300719
  75. Sikuli: Using GUI Screenshots for Search and Automation. In Proceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology (Victoria, BC, Canada) (UIST ’09). Association for Computing Machinery, New York, NY, USA, 183–192. https://doi.org/10.1145/1622176.1622213
  76. Screen Recognition: Creating Accessibility Metadata for Mobile Applications from Pixels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 275, 15 pages. https://doi.org/10.1145/3411764.3445186
  77. Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv preprint arXiv:2309.01219 (2023).
  78. Video Question Answering on Screencast Tutorials. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (Yokohama, Yokohama, Japan) (IJCAI’20). Article 148, 8 pages.
  79. “Rewind to the Jiggling Meat Part”: Understanding Voice Control of Instructional Videos in Everyday Tasks. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 58, 11 pages. https://doi.org/10.1145/3491102.3502036
Citations (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 3 tweets with 52 likes about this paper.