Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

NL2Formula: Generating Spreadsheet Formulas from Natural Language Queries (2402.14853v1)

Published 20 Feb 2024 in cs.CL and cs.AI

Abstract: Writing formulas on spreadsheets, such as Microsoft Excel and Google Sheets, is a widespread practice among users performing data analysis. However, crafting formulas on spreadsheets remains a tedious and error-prone task for many end-users, particularly when dealing with complex operations. To alleviate the burden associated with writing spreadsheet formulas, this paper introduces a novel benchmark task called NL2Formula, with the aim to generate executable formulas that are grounded on a spreadsheet table, given a Natural Language (NL) query as input. To accomplish this, we construct a comprehensive dataset consisting of 70,799 paired NL queries and corresponding spreadsheet formulas, covering 21,670 tables and 37 types of formula functions. We realize the NL2Formula task by providing a sequence-to-sequence baseline implementation called fCoder. Experimental results validate the effectiveness of fCoder, demonstrating its superior performance compared to the baseline models. Furthermore, we also compare fCoder with an initial GPT-3.5 model (i.e., text-davinci-003). Lastly, through in-depth error analysis, we identify potential challenges in the NL2Formula task and advocate for further investigation.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Wei Zhao (309 papers)
  2. Zhitao Hou (2 papers)
  3. Siyuan Wu (18 papers)
  4. Yan Gao (157 papers)
  5. Haoyu Dong (55 papers)
  6. Yao Wan (70 papers)
  7. Hongyu Zhang (147 papers)
  8. Yulei Sui (29 papers)
  9. Haidong Zhang (29 papers)
Citations (4)