Unified Multimodal Embedding Space for Direct Cross-Modal Search
Construct a unified embedding space spanning text, images, audio, and video that enables direct multimodal search without intermediary conversion modules (e.g., automatic speech recognition), thereby improving alignment and retrieval in multimodal retrieval-augmented generation.
References
Despite some progress, mapping multimodal knowledge into a unified space remains an open challenge with significant potential.
No method constructs a single grounded semantic embedding space spanning text, image, haptic, and radar modalities that remains valid across three orders of magnitude of device compute capability (Cortex-M33 to Snapdragon-class UE, Table~\ref{tab:hardware}) -- the specific research gap this survey identifies is not cross-modal alignment in general (CLIP-scale models already do this for two modalities), but cross-modal alignment at TinyLM scale, where no foundation model has been trained jointly across more than two modalities under a sub-10~MB budget.