CodeS: Towards Building Open-source Language Models for Text-to-SQL (2402.16347v1)

Published 26 Feb 2024 in cs.CL and cs.DB

Abstract: LLMs have shown promising performance on the task of translating natural language questions into SQL queries (Text-to-SQL). However, most of the state-of-the-art (SOTA) approaches rely on powerful yet closed-source LLMs, such as ChatGPT and GPT-4, which may have the limitations of unclear model architectures, data privacy risks, and expensive inference overheads. To address the limitations, we introduce CodeS, a series of pre-trained LLMs with parameters ranging from 1B to 15B, specifically designed for the text-to-SQL task. CodeS is a fully open-source LLM, which achieves superior accuracy with much smaller parameter sizes. This paper studies the research challenges in building CodeS. To enhance the SQL generation abilities of CodeS, we adopt an incremental pre-training approach using a specifically curated SQL-centric corpus. Based on this, we address the challenges of schema linking and rapid domain adaptation through strategic prompt construction and a bi-directional data augmentation technique. We conduct comprehensive evaluations on multiple datasets, including the widely used Spider benchmark, the newly released BIRD benchmark, robustness-diagnostic benchmarks such as Spider-DK, Spider-Syn, Spider-Realistic, and Dr.Spider, as well as two real-world datasets created for financial and academic applications. The experimental results show that our CodeS achieves new SOTA accuracy and robustness on nearly all challenging text-to-SQL benchmarks.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

References (84)

Authors (10)

Haoyang Li (95 papers)
Jing Zhang (730 papers)
Hanbing Liu (20 papers)
Ju Fan (26 papers)
Xiaokang Zhang (42 papers)
Jun Zhu (424 papers)
Renjie Wei (9 papers)
Hongyan Pan (1 paper)
Cuiping Li (42 papers)
Hong Chen (230 papers)

Citations (46)

View on Semantic Scholar

CodeS: Towards Building Open-source Language Models for Text-to-SQL (2402.16347v1)

Related Papers