AQuA: Automated Question-Answering in Software Tutorial Videos with Visual Anchors
Abstract: Tutorial videos are a popular help source for learning feature-rich software. However, getting quick answers to questions about tutorial videos is difficult. We present an automated approach for responding to tutorial questions. By analyzing 633 questions found in 5,944 video comments, we identified different question types and observed that users frequently described parts of the video in questions. We then asked participants (N=24) to watch tutorial videos and ask questions while annotating the video with relevant visual anchors. Most visual anchors referred to UI elements and the application workspace. Based on these insights, we built AQuA, a pipeline that generates useful answers to questions with visual anchors. We demonstrate this for Fusion 360, showing that we can recognize UI elements in visual anchors and generate answers using GPT-4 augmented with that visual information and software documentation. An evaluation study (N=16) demonstrates that our approach provides better answers than baseline methods.
- Benjamin Alcott. 2017. Does Teacher Encouragement Influence Students’ Educational Progress? A Propensity-Score Matching Analysis. Research in Higher Education 58, 7 (Jan 2017), 773–804. https://doi.org/10.1007/s11162-017-9446-2
- Amazon. 2023. Amazon Transcribe. https://aws.amazon.com/transcribe/. Accessed: 2023-09-09.
- Autodesk. 2023a. Autodesk Screencast. https://www.autodesk.com/support/technical/article/caas/tsarticles/ts/71QzAbqSskV6l5Hm0ULMMp.html. Accessed: 2023-09-10.
- Autodesk. 2023b. Fusion 360 Product Documentation. https://help.autodesk.com/view/fusion360/ENU. Accessed: 2023-09-10.
- Waken: Reverse Engineering Usage Information and Interface Structure from Software Videos. In Proceedings of the 25th Annual ACM Symposium on User Interface Software and Technology (Cambridge, Massachusetts, USA) (UIST ’12). Association for Computing Machinery, New York, NY, USA, 83–92. https://doi.org/10.1145/2380116.2380129
- VideoSticker: A Tool for Active Viewing and Visual Note-Taking from Videos. In 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, NY, USA, 672–690. https://doi.org/10.1145/3490099.3511132
- RubySlippers: Supporting Content-Based Voice Navigation for How-to Videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 97, 14 pages. https://doi.org/10.1145/3411764.3445131
- How to Design Voice Based Navigation for How-To Videos. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland UK) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–11. https://doi.org/10.1145/3290605.3300931
- Towards Complete Icon Labeling in Mobile Applications. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 387, 14 pages. https://doi.org/10.1145/3491102.3502073
- DemoCut: Generating Concise Instructional Videos for Physical Demonstrations. In Proceedings of the 26th Annual ACM Symposium on User Interface Software and Technology (St. Andrews, Scotland, United Kingdom) (UIST ’13). Association for Computing Machinery, New York, NY, USA, 141–150. https://doi.org/10.1145/2501988.2502052
- Korero: Facilitating Complex Referencing of Visual Materials in Asynchronous Discussion Interface. Proc. ACM Hum.-Comput. Interact. 1, CSCW, Article 34 (Dec 2017), 19 pages. https://doi.org/10.1145/3134669
- Beyond Show of Hands: Engaging Viewers via Expressive and Scalable Visual Communication in Live Streaming. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 109, 14 pages. https://doi.org/10.1145/3411764.3445419
- TutorialVQA: Question Answering Dataset for Tutorial Videos. In Proceedings of the Twelfth Language Resources and Evaluation Conference. European Language Resources Association, Marseille, France, 5450–5455. https://aclanthology.org/2020.lrec-1.670
- Rico: A Mobile App Dataset for Building Data-Driven Design Applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology (Québec City, QC, Canada) (UIST ’17). Association for Computing Machinery, New York, NY, USA, 845–854. https://doi.org/10.1145/3126594.3126651
- Temporal Segmentation of Creative Live Streams. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3313831.3376437
- ReMap: Multimodal Help-Seeking. In Adjunct Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology (New Orleans, LA, USA) (UIST ’19 Adjunct). Association for Computing Machinery, New York, NY, USA, 96–98. https://doi.org/10.1145/3332167.3356884
- RePlay: Contextually Presenting Learning Videos Across Software Applications (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3290605.3300527
- Mudslide: A Spatially Anchored Census of Student Confusion for Online Lecture Videos. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15). Association for Computing Machinery, New York, NY, USA, 1555–1564. https://doi.org/10.1145/2702123.2702304
- Google. 2023a. Cloud Vision API. https://cloud.google.com/vision/docs. Accessed: 2023-09-10.
- Google. 2023b. YouTube Data API. https://developers.google.com/youtube/v3. Accessed: 2023-09-10.
- MathBot: Transforming Online Resources for Learning Math into Conversational Interactions. https://api.semanticscholar.org/CorpusID:236143850
- Chronicle: Capture, Exploration, and Playback of Document Workflow Histories. In Proceedings of the 23nd Annual ACM Symposium on User Interface Software and Technology (New York, New York, USA) (UIST ’10). Association for Computing Machinery, New York, NY, USA, 143–152. https://doi.org/10.1145/1866029.1866054
- How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection. arXiv preprint arxiv:2301.07597 (2023).
- MicroMentor: Peer-to-Peer Software Help Sessions in Three Minutes or Less. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3313831.3376230
- Beyond ”One-Size-Fits-All”: Understanding the Diversity in How Software Newcomers Discover and Make Use of Help Resources. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland UK) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–14. https://doi.org/10.1145/3290605.3300570
- Crowdsourcing Step-by-Step Information Extraction to Enhance Existing How-to Videos. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada) (CHI ’14). Association for Computing Machinery, New York, NY, USA, 4017–4026. https://doi.org/10.1145/2556288.2556986
- HyperButton: In-Video Question Answering via Interactive Buttons and Hyperlinks. In Asian CHI Symposium 2021 (Yokohama, Japan) (Asian CHI Symposium 2021). Association for Computing Machinery, New York, NY, USA, 48–52. https://doi.org/10.1145/3429360.3468179
- Exploring Chart Question Answering for Blind and Low Vision Users. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 828, 15 pages. https://doi.org/10.1145/3544548.3581532
- Winder: Linking Speech and Visual Objects to Support Communication in Asynchronous Collaboration. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 453, 17 pages. https://doi.org/10.1145/3411764.3445686
- Understanding the Roles and Uses of Web Tutorials. Proceedings of the International AAAI Conference on Web and Social Media 7, 1 (Aug. 2021), 303–310. https://doi.org/10.1609/icwsm.v7i1.14413
- DAPIE: Interactive Step-by-Step Explanatory Dialogues to Answer Children’s Why and How Questions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 450, 22 pages. https://doi.org/10.1145/3544548.3581369
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages.
- Gang Li and Yang Li. 2023. Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus. In Proceedings of the Eleventh International Conference on Learning Representations (ICLR’23).
- BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). Article 814, 13 pages.
- Stargazer: An Interactive Camera Robot for Capturing How-To Videos Based on Subtle Instructor Cues. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 800, 16 pages. https://doi.org/10.1145/3544548.3580896
- Screencast Tutorial Video Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
- Hero: Hierarchical Encoder for Video+Language Omni-representation Pre-training. In Conference on Empirical Methods in Natural Language Processing. https://api.semanticscholar.org/CorpusID:218470055
- Screen2Vec: Semantic Embedding of GUI Screens and GUI Components. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 578, 15 pages. https://doi.org/10.1145/3411764.3445049
- Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 761, 17 pages. https://doi.org/10.1145/3544548.3581006
- Joshua Lochner. 2023. Chat Downloader. https://chat-downloader.readthedocs.io/en/latest/. Accessed: 2023-09-10.
- StreamSketch: Exploring Multi-Modal Interactions in Creative Live Streams. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 58 (Apr 2021), 26 pages. https://doi.org/10.1145/3449132
- A classification scheme for content analyses of YouTube video comments. J. Documentation 69 (2013), 693–714. https://api.semanticscholar.org/CorpusID:206402932
- Supercharging Trial-and-Error for Learning Complex Software Applications. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 381, 13 pages. https://doi.org/10.1145/3491102.3501895
- Ambient Help. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Vancouver, BC, Canada) (CHI ’11). Association for Computing Machinery, New York, NY, USA, 2751–2760. https://doi.org/10.1145/1978942.1979349
- IP-QAT: In-Product Questions, Answers, & Tips. In Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology (Santa Barbara, California, USA) (UIST ’11). Association for Computing Machinery, New York, NY, USA, 175–184. https://doi.org/10.1145/2047196.2047218
- Teaching language models to support answers with verified quotes. arXiv:2203.11147 [cs.CL]
- Cuong Nguyen and Feng Liu. 2015. Making Software Tutorial Video Responsive (CHI ’15). Association for Computing Machinery, New York, NY, USA, 1565–1568. https://doi.org/10.1145/2702123.2702209
- OpenAI. 2023a. ChatGPT. https://openai.com/blog/chatgpt. Accessed: 2023-09-09.
- OpenAI. 2023b. GPT-4. https://openai.com/gpt-4. Accessed: 2023-09-09.
- OpenAI. 2023c. text-embedding-ada-002. https://openai.com/blog/new-and-improved-embedding-model. Accessed: 2023-09-09.
- OpenCV. 2023a. Feature Matching. https://docs.opencv.org/4.x/dc/dc3/tutorial_py_matcher.html. Accessed: 2023-09-09.
- OpenCV. 2023b. Template Matching. https://docs.opencv.org/4.x/d4/dc6/tutorial_py_template_matching.html. Accessed: 2023-09-09.
- Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Association for Computing Machinery, New York, NY, USA, Article 2, 22 pages. https://doi.org/10.1145/3586183.3606763
- Analyzing User Comments on YouTube Coding Tutorial Videos. In 2017 IEEE/ACM 25th International Conference on Program Comprehension (ICPC). 196–206. https://doi.org/10.1109/ICPC.2017.26
- Rhitabrat Pokharel and Dixit Bhatta. 2021. Classifying YouTube Comments Based on Sentiment and Type of Sentence. arXiv:2111.01908 [cs.IR]
- Pause-and-Play: Automatically Linking Screencast Video Tutorials with Applications (UIST ’11). Association for Computing Machinery, New York, NY, USA, 135–144. https://doi.org/10.1145/2047196.2047213
- Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). Article 1182, 27 pages.
- A Survey of Evaluation Metrics Used for NLG Systems. ACM Comput. Surv. 55, 2, Article 26 (Jan 2022), 39 pages. https://doi.org/10.1145/3485766
- Leave a Comment! An In-Depth Analysis of User Comments on YouTube. In Wirtschaftsinformatik. https://api.semanticscholar.org/CorpusID:11966671
- Understanding the Effect of In-Video Prompting on Learners and Instructors. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3173574.3173893
- Visual Transcripts: Lecture Notes from Blackboard-Style Lecture Videos. ACM Trans. Graph. 34, 6, Article 240 (Nov 2015), 10 pages. https://doi.org/10.1145/2816795.2818123
- GVQA: Learning to Answer Questions about Graphs with Visualizations via Knowledge Base. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 464, 16 pages. https://doi.org/10.1145/3544548.3581067
- Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup Videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 108, 16 pages. https://doi.org/10.1145/3411764.3445721
- Enabling Conversational Interaction with Mobile UI Using Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 432, 17 pages. https://doi.org/10.1145/3544548.3580895
- Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association for Computing Machinery, New York, NY, USA, 498–510. https://doi.org/10.1145/3472749.3474765
- Sara, the Lecturer: Improving Learning in Online Education with a Scaffolding-Based Conversational Agent. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–14. https://doi.org/10.1145/3313831.3376781
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 286, 14 pages. https://doi.org/10.1145/3544548.3581158
- UIED: A Hybrid Tool for GUI Element Detection. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Virtual Event, USA) (ESEC/FSE 2020). Association for Computing Machinery, New York, NY, USA, 1655–1659. https://doi.org/10.1145/3368089.3417940
- Just Ask: Learning to Answer Questions from Millions of Narrated Videos. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2020), 1666–1677. https://api.semanticscholar.org/CorpusID:227238996
- Saelyne Yang and Juho Kim. 2020. What Makes It Hard for Users to Follow Software Tutorial Videos?. In Proceedings of HCI Korea 2020. The HCI Society of KOREA, South Korea, 531–536.
- Beyond Instructions: A Taxonomy of Information Types in How-to Videos. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 797, 21 pages. https://doi.org/10.1145/3544548.3581126
- Snapstream: Snapshot-Based Interaction in Live Streaming for Visual Art. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3313831.3376390
- SoftVideo: Improving the Learning Experience of Software Tutorial Videos with Collective Interaction Data. In 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, NY, USA, 646–660. https://doi.org/10.1145/3490099.3511106
- ”Can You Believe [1:21]?!”: Content and Time-Based Reference Patterns in Video Comments. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland UK) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3290605.3300719
- Sikuli: Using GUI Screenshots for Search and Automation. In Proceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology (Victoria, BC, Canada) (UIST ’09). Association for Computing Machinery, New York, NY, USA, 183–192. https://doi.org/10.1145/1622176.1622213
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from Pixels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 275, 15 pages. https://doi.org/10.1145/3411764.3445186
- Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv preprint arXiv:2309.01219 (2023).
- Video Question Answering on Screencast Tutorials. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (Yokohama, Yokohama, Japan) (IJCAI’20). Article 148, 8 pages.
- “Rewind to the Jiggling Meat Part”: Understanding Voice Control of Instructional Videos in Everyday Tasks. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 58, 11 pages. https://doi.org/10.1145/3491102.3502036
Paper Prompts
Sign up for free to create and run prompts on this paper.