Controlled evaluation of interleaved multimodal queries

Investigate the isolated contribution of interleaved multimodal input sequences to Omni-Embed-Mini’s retrieval performance through a controlled experimental study.

Background

Omni-Embed-Mini interleaves tokens from multiple modalities into a single causal-transformer sequence, making multimodal queries native to the architecture rather than requiring late fusion. The paper identifies this architectural capability as a distinction from ImageBind and LanguageBind but does not experimentally separate its effect from the other components of the training recipe.

A controlled ablation would determine whether interleaved multimodal queries independently improve retrieval quality, rather than attributing any observed performance to the frozen text anchor, dense-caption distillation, modality projectors, or other design choices.

References

Interleaved multi-modal queries are therefore native to the architecture rather than an added capability; we do not isolate their contribution experimentally, and leave a controlled study to future work.

— Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation  (2610.02148 - Kurpath et al., 1 Oct 2026) in Section 2, paragraph “Anchored vs. contrastive joint alignment”