Papers

Topics

Authors

Recent

View all

Detailed Answer

Quick Answer

Concise responses based on abstracts only

Detailed Answer

Well-researched responses based on abstracts and relevant paper content.

Custom Instructions Pro

Preferences or requirements that you'd like Emergent Mind to consider when generating responses

Gemini 2.5 Flash

Gemini 2.5 Flash 75 tok/s

Gemini 2.5 Pro 42 tok/s Pro

GPT-5 Medium 31 tok/s Pro

GPT-5 High 24 tok/s Pro

GPT-4o 98 tok/s Pro

Kimi K2 226 tok/s Pro

GPT OSS 120B 447 tok/s Pro

Claude Sonnet 4 38 tok/s Pro

2000 character limit reached

Upcycling Candidate Tokens of Large Language Models for Query Expansion (2509.02377v1)

Published 2 Sep 2025 in cs.IR

Abstract: Query Expansion (QE) improves retrieval performance by enriching queries with related terms. Recently, LLMs have been used for QE, but existing methods face a trade-off: generating diverse terms boosts performance but increases computational cost. To address this challenge, we propose Candidate Token Query Expansion (CTQE), which extracts diverse and relevant terms from a single LLM decoding pass by leveraging unselected candidate tokens. These tokens, though not part of the final output, are conditioned on the full query and capture useful information. By aggregating them, CTQE achieves both relevance and diversity without extra inference, reducing overhead and latency. Experiments show that CTQE delivers strong retrieval performance with significantly lower cost, outperforming or comparable to more expensive methods. Code is available at: https://github.com/bluejeans8/CTQE