---
title: 'VLM2Rec: Resolving Modality Collapse'
url: https://www.emergentmind.com/topics/vlm2rec
type: topic
---

# VLM2Rec: Resolving Modality Collapse

Searching arXiv for the primary paper and closely related contextual works.
arXiv paper: 2603.17450
Title: VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation
Authors: Chenglin Song, Saeed Hamidpour, Ruihong Qiu, Tiankang Yu, Jiawei Chen, Licheng Jiao, Lirong Wang
Published: 2026-03-18
URL: http://arxiv.org/abs/2603.17450v1

Abstract:
Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity and prevents Collaborative Filtering (CF) signals from being fully integrated into item representations. Inspired by the recent success of Large Language Models (LLMs) as high-capacity embedders, we investigate the use of Vision-Language Models (VLMs) as CF-aware multimodal encoders for SR. However, we find that standard contrastive supervised fine-tuning (SFT), which adapts VLMs for embedding generation and injects CF signals, can amplify its inherent modality collapse. In this state, optimization is dominated by a single modality while the other degrades, ultimately undermining recommendation accuracy. To address this, we propose VLM2Rec, a VLM embedder-based framework for multimodal sequential recommendation designed to ensure balanced modality utilization. Specifically, we introduce Weak-modality Penalized Contrastive Learning to rectify gradient imbalance during optimization and Cross-Modal Relational Topology Regularization to preserve geometric consistency between modalities. Extensive experiments demonstrate that VLM2Rec consistently outperforms state-of-the-art baselines in both accuracy and robustness across diverse scenarios.
Searching for contextual papers that use “VLM2Rec” more loosely or provide adjacent formulations.
arXiv paper: 2501.18271
Title: Vision-Language Model Selection and Reuse for Downstream Adaptation
Authors: Zhixiang Chi, Guangyi Fan, Yunge Zhang, Tao Pu, Xiangyu Zhao, Jianye Hao, Zhenya Huang
Published: 2025-01-30
URL: http://arxiv.org/abs/2501.18271v1

Abstract:
Pre-trained Vision-Language Models (VLMs) are becoming increasingly popular across various visual tasks, and several open-sourced VLM variants have been released. However, selecting the best-performing pre-trained VLM for a specific downstream task is challenging since no single VLM can achieve promising performance on all downstream tasks, and evaluating all available VLMs is impossible due to time and data limitations. To address this problem, this paper proposes a novel paradigm to select and reuse VLM for downstream tasks, called Model Label Learning (MLL). The proposal contains three key modules: model labeling, which assigns labels to each VLM to describe their specialty and utility; model selection, which matches the requirements of the target task with model labels; and model reuse, which applies selected VLMs to the target task in an ensemble manner. The proposal is highly computationally efficient and growable since the model labeling process is completed target task independent and the ability could grow with the number of candidate VLMs. We also introduce a new benchmark for evaluating VLM selection methods, including 49 VLMs and 17 target task datasets. Experimental results clearly demonstrate the effectiveness of the proposed method for selecting and reusing VLMs.

VLM2Rec is a **VLM embedder-based framework for multimodal sequential recommendation** that uses a large Vision-Language Model as a **CF-aware multimodal encoder** and is explicitly designed to prevent **modality collapse** during contrastive supervised fine-tuning. In the formulation introduced in “VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation” [2603.17450], the central claim is that standard InfoNCE-style SFT can make optimization collapse onto a single modality—typically text—so that the weaker modality degrades and fused recommendation quality is undermined. VLM2Rec addresses this failure mode with two objective-level components, **Weak-modality Penalized Contrastive Learning** and **Cross-Modal Relational Topology Regularization**, while retaining a deliberately simple fusion design.

## 1. Problem setting and conceptual scope

VLM2Rec is defined in the setting of **multimodal sequential recommendation (SR)**. There is a user set \(\mathcal{U}\), an item set \(\mathcal{I}\), and for each item \(i \in \mathcal{I}\), both text \(t_i\) and image \(v_i\) are available. Each user \(u\) is associated with a chronological interaction sequence
\[
S_u = [i_1, i_2, \dots, i_{|S_u|}],
\]
and the task is to predict the next item \(i_{|S_u|+1}\) [2603.17450].

The framework is motivated by two limitations of conventional multimodal SR pipelines. First, these pipelines often rely on **small frozen pretrained encoders**, which impose **low semantic capacity**. Second, because those encoders are frozen, **Collaborative Filtering signals cannot reshape the semantic space**, so the sequential backbone must operate over representations that are not co-adapted to user-item interaction structure. VLM2Rec therefore repurposes a large VLM as both a **sequence encoder** and a **CF-aware embedder**.

A common misconception is that simply replacing frozen encoders with a stronger VLM is sufficient. The paper argues the opposite: the crucial problem is not merely capacity, but the optimization pathology induced by standard contrastive fine-tuning. The resulting **Paradox of SFT** is that the very objective used to inject CF signals can amplify an intrinsic modality imbalance, so that the multimodal encoder behaves as a text-dominated model rather than a balanced visual-textual recommender [2603.17450].

## 2. Encoder architecture and representation design

The backbone in VLM2Rec is a pretrained VLM, specifically **Qwen2.5-VL-3B**, used as a **sequence encoder**. For text-only and image-only user sequences, the model produces
\[
\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad
\mathbf{z}_u^V = \Phi(P_V(S_u^V)),
\]
and for each item \(i\),
\[
\mathbf{e}_i^T = \Phi(P_T(i\text{-text})), \qquad
\mathbf{e}_i^V = \Phi(P_V(i\text{-image})).
\]
The representation is taken from the **last-token hidden state** of the final Transformer layer. For efficiency, the framework uses **LoRA** with rank \(16\), \(\alpha = 32\), and dropout \(0.2\), so the main VLM weights remain effectively frozen [2603.17450].

Two prompting strategies are studied. In **internal (interleaved) fusion**, text and image tokens are interleaved in a single prompt. In **external (separate) fusion**, text and image sequences are encoded through separate prompts. The framework adopts **external fusion** as the base design, because it decouples modality-specific encoding paths and keeps modality balancing at the **objective level** rather than embedding it in a more complex fusion architecture.

Fusion itself is intentionally minimal. User and item representations are formed by element-wise summation:
\[
\mathbf{z}_u = \mathbf{z}_u^T + \mathbf{z}_u^V, \qquad
\mathbf{e}_i = \mathbf{e}_i^T + \mathbf{e}_i^V.
\]
No extra fusion parameters are introduced. This design makes the empirical effects of modality balancing attributable to the training objective rather than to a learned fusion block [2603.17450].

The framework is used in two modes. In **Task 1**, it operates as a direct recommender by ranking items through sequence-item similarity in embedder space. In **Task 2**, it acts as a **plug-and-play embedding generator**: learned item embeddings are projected through a one-layer linear adapter to hidden size \(d = 128\) and used to initialize downstream SR backbones such as **GRU4Rec** and **SASRec**.

## 3. Modality collapse and the failure of standard contrastive SFT

The paper identifies **modality collapse** as the key optimization pathology in VLM-based multimodal SR. The starting point is standard fused-embedding InfoNCE:
\[
\mathcal{L}_{\text{SFT}} =
- \sum_{u \in \mathcal{B}}
\log
\frac{\exp\left(s(\mathbf{z}_u, \mathbf{e}_{i^+}) / \tau\right)}
{\exp\left(s(\mathbf{z}_u, \mathbf{e}_{i^+}) / \tau\right)
+ \sum_{i^- \in \mathcal{N}_u}
\exp\left(s(\mathbf{z}_u, \mathbf{e}_{i^-}) / \tau\right)}.
\]
Here, \(\mathbf{z}_u\) and \(\mathbf{e}_i\) are fused sequence and item embeddings, \(s(\cdot,\cdot)\) is cosine similarity, and \(\tau\) is a temperature [2603.17450].

Empirically, the problem is that pretrained VLMs already exhibit an **intrinsic modality gap**: text is often the easier modality and vision the harder one. Under fused InfoNCE, optimization is free to solve the task through whichever modality is easiest. The consequence is that gradients become dominated by text, while the vision branch loses discriminative power. The paper diagnoses this with three forms of evidence.

First, **dropout-style modality tests** compare fused-to-fused, text-to-fused, and vision-to-fused performance. After SFT, **vision-to-fused degrades**, and **fused-to-fused can become worse than text-to-fused**, meaning that the image branch contributes negative transfer rather than complementary signal.

Second, **gradient analysis** shows that
\[
\cos(\mathbf{g}_{\text{total}}, \mathbf{g}_T) \to 1,
\]
while \(\cos(\mathbf{g}_{\text{total}}, \mathbf{g}_V)\) drops quickly. Optimization thus follows the text gradient almost exclusively.

Third, the paper analyzes embedding geometry through alignment, uniformity, and separability:
\[
A_{\text{pos}}^{m}
\triangleq
\mathbb{E}_{(s,i^+) \sim \mathcal{D}}
\big[\|\Phi_m(s)-\Phi_m(i^+)\|_2^2\big],
\]
\[
A_{\text{neg}}^{m}
\triangleq
\mathbb{E}_{(s,i^+) \sim \mathcal{D}}
\mathbb{E}_{i^- \sim \mathcal{P}_n(\cdot \mid s)}
\big[\|\Phi_m(s)-\Phi_m(i^-)\|_2^2\big],
\]
\[
U^{m}
\triangleq
\log
\mathbb{E}_{(s,i)\sim \mathcal{P}_u}
\big[\exp\!\big(-2\|\Phi_m(s)-\Phi_m(i)\|_2^2\big)\big],
\qquad
S^m \triangleq \frac{A^m_{\text{neg}}}{A^m_{\text{pos}}+\epsilon}.
\]
After standard SFT, the vision modality’s separability \(S^V\) collapses to approximately \(1\) or lower, indicating that negatives are no longer separated from positives in the visual space [2603.17450].

## 4. Weak-modality Penalized Contrastive Learning and CRTR

VLM2Rec introduces **Weak-modality Penalized Contrastive Learning (WPCL)** to rectify the gradient imbalance. For each user \(u\) and modality \(m \in \{T,V\}\), the discriminative margin is defined as
\[
\mathcal{M}_{u,m} =
s(\mathbf{z}_u^m, \mathbf{e}_{i^+}^m)
-
\frac{1}{|\mathcal{N}_u|}
\sum_{i^- \in \mathcal{N}_u}
s(\mathbf{z}_u^m, \mathbf{e}_{i^-}^m).
\]
The stronger and weaker modality margins are then identified by \(\max\) and \(\min\), and the **modality gap** is defined as
\[
\Delta_{u,\text{gap}} =
\text{sg}[\mathcal{M}_{u,\text{strong}}] - \mathcal{M}_{u,\text{weak}},
\]
where \(\text{sg}[\cdot]\) is the stop-gradient operator [2603.17450].

This gap controls a dynamic penalty weight,
\[
w_{u,\mathrm{pen}} = 1+\beta \cdot \mathrm{Softplus}(\alpha \cdot \Delta_{u,\mathrm{gap}}),
\]
which rescales the negative term of the fused contrastive loss:
\[
\mathcal{L}_{\text{WPCL}} =
-\sum_{u \in \mathcal{B}}
\log
\frac{\exp\left(s(\mathbf{z}_u, \mathbf{e}_{i^+}) / \tau_{\text{WPCL}}\right)}
{\exp\left(s(\mathbf{z}_u, \mathbf{e}_{i^+}) / \tau_{\text{WPCL}}\right)
+
w_{u,\mathrm{pen}}
\sum_{i^- \in \mathcal{N}_u}
\exp\left(s(\mathbf{z}_u, \mathbf{e}_{i^-}) / \tau_{\text{WPCL}}\right)}.
\]
The stop-gradient is critical: it prevents the model from reducing the gap by degrading the strong modality. Instead, the optimization pressure is placed on improving the **weak modality’s negative separation**.

Because aggressive negative pushing can distort geometry, VLM2Rec adds **Cross-Modal Relational Topology Regularization (CRTR)**. For each modality, in-batch similarities are converted into ranking distributions \(\mathbf{P}_i^T\) and \(\mathbf{P}_i^V\), and consistency is enforced with symmetric KL divergence:
\[
\mathcal{L}_{\text{CRTR}} =
\frac{1}{2B}
\sum_{i=1}^{B}
\big(
\mathrm{KL}(\mathbf{P}_i^T \,\|\, \mathbf{P}_i^V)
+
\mathrm{KL}(\mathbf{P}_i^V \,\|\, \mathbf{P}_i^T)
\big).
\]
This regularizer preserves **relational topology** rather than point-wise alignment, so the modalities retain distinct geometry while sharing neighborhood structure. The final objective is
\[
\mathcal{L} = \mathcal{L}_{\text{WPCL}} + \lambda \cdot \mathcal{L}_{\text{CRTR}}.
\]
A plausible implication is that VLM2Rec is less an architectural innovation than an **objective-level intervention**: the paper deliberately keeps fusion simple so that the method’s contribution lies in how the multimodal space is optimized [2603.17450].

## 5. Evaluation protocol and empirical results

The experiments use four **Amazon 5-core** domains—**Toys**, **Beauty**, **Clothing**, and **Sports**—with titles and images, and with sequence length truncated to \(10\) [2603.17450]. The reported dataset statistics are: Toys with \(15{,}921\) users, \(8{,}383\) items, and \(108{,}336\) interactions; Beauty with \(19{,}757\) users, \(9{,}311\) items, and \(137{,}300\) interactions; Clothing with \(30{,}757\) users, \(17{,}087\) items, and \(196{,}614\) interactions; and Sports with \(32{,}127\) users, \(14{,}820\) items, and \(222{,}591\) interactions. Evaluation uses leave-one-out splitting, \(100\) sampled negatives per test instance, and the metrics **H@10**, **H@20**, **N@10**, and **N@20**.

In **Task 1** direct recommendation, VLM2Rec outperforms all baselines on all four datasets. On **Toys**, the best baseline, \(\text{VLM}_{\text{SFT}(\text{Ext.})}\), reaches \(H@10 = 0.4160\), \(H@20 = 0.5897\), \(N@10 = 0.2209\), and \(N@20 = 0.2647\), while **VLM2Rec** reaches \(H@10 = 0.5225\), \(H@20 = 0.6476\), \(N@10 = 0.3578\), and \(N@20 = 0.3893\). On **Beauty**, VLM2Rec achieves \(N@20 = 0.4121\), compared with a best baseline range of \(0.3498\)–\(0.3531\). On **Clothing**, VLM2Rec reaches \(N@20 = 0.3851\), versus \(0.3531\) for the best baseline. On **Sports**, it reaches \(N@20 = 0.4052\), versus \(0.3265\) [2603.17450].

In **Task 2** downstream SR initialization, VLM2Rec also produces the best embeddings for **GRU4Rec** and **SASRec**. For **GRU4Rec** on Beauty, the best baseline reports \(N@20 = 0.3296\) for \(\text{VLM}_{\text{SFT}(\text{Ext.})}\) and \(0.3253\) for \(\text{LLM}_{\text{SFT}}\), while VLM2Rec reports \(0.3669\). For **SASRec** on Sports, the best baseline gives \(N@20 = 0.3441\), and VLM2Rec gives \(0.3657\). On **Beauty** with SASRec, the best baseline gives \(N@20 = 0.3570\), and VLM2Rec gives \(0.3932\).

The robustness experiments reinforce the same pattern. In **cross-domain recommendation**, training on Clothing and testing on Beauty yields \(N@20 = 0.3372\) for VLM2Rec in Task 1, versus \(0.2676\) for LLMEmb and \(0.2566\) for \(\text{VLM}_{\text{SFT}}\). Training on Sports and testing on Clothing yields \(N@20 = 0.3154\) for VLM2Rec, versus \(0.2362\) for \(\text{VLM}_{\text{SFT}}\). In **few-shot training**, even with \(K = 128\) sampled users, VLM2Rec beats many full-training baselines, and with \(K = 1024\) users it nearly matches full-data baselines while reducing epoch time on Beauty from \(118\) minutes to \(9\) [2603.17450].

Ablation results show that **WPCL is the central component**. Removing \(\mathcal{L}_{\text{WPCL}}\) causes the largest drop; for example, on Beauty Task 1, \(N@20\) falls from \(0.4121\) to \(0.2592\). Removing \(\mathcal{L}_{\text{CRTR}}\) causes a smaller decline, from \(0.4121\) to \(0.4058\), indicating that CRTR acts primarily as a stabilizer rather than the principal source of gain.

## 6. Interpretation, broader usage of the term, and limitations

Within recommendation research, **VLM2Rec** specifically denotes the modality-collapse-aware sequential recommendation framework of [2603.17450]. However, adjacent literature sometimes uses the expression more loosely to mean **recommending or reusing VLMs**, or more generally **VLM-based recommendation or retrieval**. For example, “Vision-Language Model Selection and Reuse for Downstream Adaptation” formulates VLM selection and reuse as a recommendation problem over a model hub [2501.18271]. This suggests that the term has both a **narrow** meaning, tied to multimodal sequential recommendation with collapse mitigation, and a **broader** informal meaning, tied to recommendation-like VLM selection and reuse.

Several limitations are explicit. VLM2Rec remains computationally heavier than text-only LLM methods, because image tokens make VLM training expensive. Its effectiveness depends on reasonably informative titles and images. The method is designed around **two modalities**, text and image, and its sequence length is capped at \(10\). Extending the user-adaptive margin and topology regularization to additional modalities such as audio, video, or structured attributes is identified as future work [2603.17450].

Another possible misunderstanding is that VLM2Rec’s contribution is a new multimodal fusion module. The paper instead argues for the opposite design choice: **element-wise summation** is kept intentionally simple so that the framework’s gains can be attributed to balanced optimization rather than fusion complexity. In that sense, VLM2Rec is best understood as a method for making large VLM embedders *usable* for multimodal SR by preventing text-dominant optimization from collapsing the visual branch.

Taken together, VLM2Rec establishes a specific research agenda within multimodal recommendation: large VLM embedders are not merely stronger encoders, but unstable optimizers unless their modality imbalance is addressed directly. The framework’s central technical claim is therefore not that vision-language models should replace conventional encoders, but that **balanced modality utilization** must be enforced if CF-aware VLM embeddings are to improve recommendation accuracy and robustness at all [2603.17450].

Source: https://www.emergentmind.com/topics/vlm2rec