Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
38 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

ChatPipe: Orchestrating Data Preparation Program by Optimizing Human-ChatGPT Interactions (2304.03540v1)

Published 7 Apr 2023 in cs.DB, cs.AI, cs.HC, and cs.LG

Abstract: Orchestrating a high-quality data preparation program is essential for successful ML, but it is known to be time and effort consuming. Despite the impressive capabilities of LLMs like ChatGPT in generating programs by interacting with users through natural language prompts, there are still limitations. Specifically, a user must provide specific prompts to iteratively guide ChatGPT in improving data preparation programs, which requires a certain level of expertise in programming, the dataset used and the ML task. Moreover, once a program has been generated, it is non-trivial to revisit a previous version or make changes to the program without starting the process over again. In this paper, we present ChatPipe, a novel system designed to facilitate seamless interaction between users and ChatGPT. ChatPipe provides users with effective recommendation on next data preparation operations, and guides ChatGPT to generate program for the operations. Also, ChatPipe enables users to easily roll back to previous versions of the program, which facilitates more efficient experimentation and testing. We have developed a web application for ChatPipe and prepared several real-world ML tasks from Kaggle. These tasks can showcase the capabilities of ChatPipe and enable VLDB attendees to easily experiment with our novel features to rapidly orchestrate a high-quality data preparation program.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Sibei Chen (4 papers)
  2. Hanbing Liu (20 papers)
  3. Weiting Jin (1 paper)
  4. Xiangyu Sun (16 papers)
  5. Xiaoyao Feng (1 paper)
  6. Ju Fan (26 papers)
  7. Xiaoyong Du (40 papers)
  8. Nan Tang (63 papers)
Citations (3)