Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Exploring the Practicality of Generative Retrieval on Dynamic Corpora (2305.18952v5)

Published 27 May 2023 in cs.IR and cs.AI

Abstract: Benchmarking the performance of information retrieval (IR) is mostly conducted with a fixed set of documents (static corpora). However, in realistic scenarios, this is rarely the case and the documents to be retrieved are constantly updated and added. In this paper, we focus on Generative Retrievals (GR), which apply autoregressive LLMs to IR problems, and explore their adaptability and robustness in dynamic scenarios. We also conduct an extensive evaluation of computational and memory efficiency, crucial factors for real-world deployment of IR systems handling vast and ever-changing document collections. Our results on the StreamingQA benchmark demonstrate that GR is more adaptable to evolving knowledge (4-11%), robust in learning knowledge with temporal information, and efficient in terms of inference FLOPs (x2), indexing time (x6), and storage footprint (x4) compared to Dual Encoders (DE), which are commonly used in retrieval systems. Our paper highlights the potential of GR for future use in practical IR systems within dynamic environments.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Soyoung Yoon (8 papers)
  2. Chaeeun Kim (5 papers)
  3. Hyunji Lee (19 papers)
  4. Joel Jang (30 papers)
  5. Sohee Yang (23 papers)
  6. Minjoon Seo (82 papers)
Citations (2)

Summary

We haven't generated a summary for this paper yet.