Embedding-based Zero-shot Retrieval through Query Generation (2009.10270v1)

Published 22 Sep 2020 in cs.IR

Abstract: Passage retrieval addresses the problem of locating relevant passages, usually from a large corpus, given a query. In practice, lexical term-matching algorithms like BM25 are popular choices for retrieval owing to their efficiency. However, term-based matching algorithms often miss relevant passages that have no lexical overlap with the query and cannot be finetuned to downstream datasets. In this work, we consider the embedding-based two-tower architecture as our neural retrieval model. Since labeled data can be scarce and because neural retrieval models require vast amounts of data to train, we propose a novel method for generating synthetic training data for retrieval. Our system produces remarkable results, significantly outperforming BM25 on 5 out of 6 datasets tested, by an average of 2.45 points for Recall@1. In some cases, our model trained on synthetic data can even outperform the same model trained on real data

PDF Abstract

Summarize Bookmark Chat (Pro)

Authors (7)

Davis Liang (15 papers)
Peng Xu (357 papers)
Siamak Shakeri (29 papers)
Cicero Nogueira dos Santos (31 papers)
Ramesh Nallapati (38 papers)
Zhiheng Huang (33 papers)
Bing Xiang (74 papers)

Citations (40)

View on Semantic Scholar

Embedding-based Zero-shot Retrieval through Query Generation (2009.10270v1)

Related Papers