Single selector across chunk sizes

Develop a single selector that generalizes across different chunk sizes at inference time for Doc-REFRAG, avoiding the need to train and route among separate selectors for each fixed chunk size.

Background

Doc-REFRAG fixes the chunk size k as a training hyperparameter, and each selector is trained for one specific value of k. Supporting multiple chunk sizes therefore requires an additional routing step to dispatch each input to the corresponding selector, which can undermine the framework’s inference-efficiency objective. The paper identifies a single selector capable of operating across chunk sizes as an unresolved direction for future work, particularly because adaptive granularity could improve coverage for documents with dense layouts while avoiding the latency and deployment costs of multiple specialized selectors.

References

We leave a single selector that generalizes across chunk sizes to future work.

— Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation  (2608.30163 - Hu et al., 31 Aug 2026) in Limitations section

Three gaps remain: all three datasets are Wikipedia-derived, so domain diversity (legal, biomedical, code) is untested; the generative validation covers a single provider model; and our negative result concerns the specific selectors we implemented, so a learned or reader-aware selector could still beat first-non-empty.

— ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines  (2609.29828 - Nautiyal et al., 24 Sep 2026) in Section Limitations, item 6, “Evaluation scope”