Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Question Answering via Web Extracted Tables and Pipelined Models (1903.07113v2)

Published 17 Mar 2019 in cs.CL

Abstract: In this paper, we describe a dataset and baseline result for a question answering that utilizes web tables. It contains commonly asked questions on the web and their corresponding answers found in tables on websites. Our dataset is novel in that every question is paired with a table of a different signature. In particular, the dataset contains two classes of tables: entity-instance tables and the key-value tables. Each QA instance comprises a table of either kind, a natural language question, and a corresponding structured SQL query. We build our model by dividing question answering into several tasks, including table retrieval and question element classification, and conduct experiments to measure the performance of each task. We extract various features specific to each task and compose a full pipeline which constructs the SQL query from its parts. Our work provides qualitative results and error analysis for each task, and identifies in detail the reasoning required to generate SQL expressions from natural language questions. This analysis of reasoning informs future models based on neural machine learning.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Bhavya Karki (1 paper)
  2. Fan Hu (29 papers)
  3. Nithin Haridas (1 paper)
  4. Suhail Barot (3 papers)
  5. Zihua Liu (14 papers)
  6. Lucile Callebert (1 paper)
  7. Matthias Grabmair (33 papers)
  8. Anthony Tomasic (8 papers)
Citations (2)