Papers
Topics
Authors
Recent
Search
2000 character limit reached

VLM2Rec: Resolving Modality Collapse

Updated 16 July 2026
  • The paper introduces VLM2Rec, a CF-aware framework that prevents modality collapse by balancing vision and text signals in sequential recommendations.
  • It employs Weak-modality Penalized Contrastive Learning and Cross-Modal Relational Topology Regularization to optimize and preserve distinct modality features.
  • Extensive experiments demonstrate that VLM2Rec outperforms state-of-the-art baselines in accuracy and robustness across diverse recommendation scenarios.

Searching arXiv for the primary paper and closely related contextual works. arXiv paper: (Kim et al., 18 Mar 2026) Title: VLM2Rec: Resolving Modality Collapse in Vision-LLM Embedders for Multimodal Sequential Recommendation Authors: Chenglin Song, Saeed Hamidpour, Ruihong Qiu, Tiankang Yu, Jiawei Chen, Licheng Jiao, Lirong Wang Published: 2026-03-18 URL: http://arxiv.org/abs/([2603.17450](/papers/2603.17450))v1

Abstract: Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity and prevents Collaborative Filtering (CF) signals from being fully integrated into item representations. Inspired by the recent success of LLMs as high-capacity embedders, we investigate the use of Vision-LLMs (VLMs) as CF-aware multimodal encoders for SR. However, we find that standard contrastive supervised fine-tuning (SFT), which adapts VLMs for embedding generation and injects CF signals, can amplify its inherent modality collapse. In this state, optimization is dominated by a single modality while the other degrades, ultimately undermining recommendation accuracy. To address this, we propose VLM2Rec, a VLM embedder-based framework for multimodal sequential recommendation designed to ensure balanced modality utilization. Specifically, we introduce Weak-modality Penalized Contrastive Learning to rectify gradient imbalance during optimization and Cross-Modal Relational Topology Regularization to preserve geometric consistency between modalities. Extensive experiments demonstrate that VLM2Rec consistently outperforms state-of-the-art baselines in both accuracy and robustness across diverse scenarios. Searching for contextual papers that use “VLM2Rec” more loosely or provide adjacent formulations. arXiv paper: (Tan et al., 30 Jan 2025) Title: Vision-LLM Selection and Reuse for Downstream Adaptation Authors: Zhixiang Chi, Guangyi Fan, Yunge Zhang, Tao Pu, Xiangyu Zhao, Jianye Hao, Zhenya Huang Published: 2025-01-30 URL: http://arxiv.org/abs/([2501.18271](/papers/2501.18271))v1

Abstract: Pre-trained Vision-LLMs (VLMs) are becoming increasingly popular across various visual tasks, and several open-sourced VLM variants have been released. However, selecting the best-performing pre-trained VLM for a specific downstream task is challenging since no single VLM can achieve promising performance on all downstream tasks, and evaluating all available VLMs is impossible due to time and data limitations. To address this problem, this paper proposes a novel paradigm to select and reuse VLM for downstream tasks, called Model Label Learning (MLL). The proposal contains three key modules: model labeling, which assigns labels to each VLM to describe their specialty and utility; model selection, which matches the requirements of the target task with model labels; and model reuse, which applies selected VLMs to the target task in an ensemble manner. The proposal is highly computationally efficient and growable since the model labeling process is completed target task independent and the ability could grow with the number of candidate VLMs. We also introduce a new benchmark for evaluating VLM selection methods, including 49 VLMs and 17 target task datasets. Experimental results clearly demonstrate the effectiveness of the proposed method for selecting and reusing VLMs.

VLM2Rec is a VLM embedder-based framework for multimodal sequential recommendation that uses a large Vision-LLM as a CF-aware multimodal encoder and is explicitly designed to prevent modality collapse during contrastive supervised fine-tuning. In the formulation introduced in “VLM2Rec: Resolving Modality Collapse in Vision-LLM Embedders for Multimodal Sequential Recommendation” (Kim et al., 18 Mar 2026), the central claim is that standard InfoNCE-style SFT can make optimization collapse onto a single modality—typically text—so that the weaker modality degrades and fused recommendation quality is undermined. VLM2Rec addresses this failure mode with two objective-level components, Weak-modality Penalized Contrastive Learning and Cross-Modal Relational Topology Regularization, while retaining a deliberately simple fusion design.

1. Problem setting and conceptual scope

VLM2Rec is defined in the setting of multimodal sequential recommendation (SR). There is a user set U\mathcal{U}, an item set I\mathcal{I}, and for each item i∈Ii \in \mathcal{I}, both text tit_i and image viv_i are available. Each user uu is associated with a chronological interaction sequence

Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],

and the task is to predict the next item i∣Su∣+1i_{|S_u|+1} (Kim et al., 18 Mar 2026).

The framework is motivated by two limitations of conventional multimodal SR pipelines. First, these pipelines often rely on small frozen pretrained encoders, which impose low semantic capacity. Second, because those encoders are frozen, Collaborative Filtering signals cannot reshape the semantic space, so the sequential backbone must operate over representations that are not co-adapted to user-item interaction structure. VLM2Rec therefore repurposes a large VLM as both a sequence encoder and a CF-aware embedder.

A common misconception is that simply replacing frozen encoders with a stronger VLM is sufficient. The paper argues the opposite: the crucial problem is not merely capacity, but the optimization pathology induced by standard contrastive fine-tuning. The resulting Paradox of SFT is that the very objective used to inject CF signals can amplify an intrinsic modality imbalance, so that the multimodal encoder behaves as a text-dominated model rather than a balanced visual-textual recommender (Kim et al., 18 Mar 2026).

2. Encoder architecture and representation design

The backbone in VLM2Rec is a pretrained VLM, specifically Qwen2.5-VL-3B, used as a sequence encoder. For text-only and image-only user sequences, the model produces

zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),

and for each item ii,

I\mathcal{I}0

The representation is taken from the last-token hidden state of the final Transformer layer. For efficiency, the framework uses LoRA with rank I\mathcal{I}1, I\mathcal{I}2, and dropout I\mathcal{I}3, so the main VLM weights remain effectively frozen (Kim et al., 18 Mar 2026).

Two prompting strategies are studied. In internal (interleaved) fusion, text and image tokens are interleaved in a single prompt. In external (separate) fusion, text and image sequences are encoded through separate prompts. The framework adopts external fusion as the base design, because it decouples modality-specific encoding paths and keeps modality balancing at the objective level rather than embedding it in a more complex fusion architecture.

Fusion itself is intentionally minimal. User and item representations are formed by element-wise summation: I\mathcal{I}4 No extra fusion parameters are introduced. This design makes the empirical effects of modality balancing attributable to the training objective rather than to a learned fusion block (Kim et al., 18 Mar 2026).

The framework is used in two modes. In Task 1, it operates as a direct recommender by ranking items through sequence-item similarity in embedder space. In Task 2, it acts as a plug-and-play embedding generator: learned item embeddings are projected through a one-layer linear adapter to hidden size I\mathcal{I}5 and used to initialize downstream SR backbones such as GRU4Rec and SASRec.

3. Modality collapse and the failure of standard contrastive SFT

The paper identifies modality collapse as the key optimization pathology in VLM-based multimodal SR. The starting point is standard fused-embedding InfoNCE: I\mathcal{I}6 Here, I\mathcal{I}7 and I\mathcal{I}8 are fused sequence and item embeddings, I\mathcal{I}9 is cosine similarity, and i∈Ii \in \mathcal{I}0 is a temperature (Kim et al., 18 Mar 2026).

Empirically, the problem is that pretrained VLMs already exhibit an intrinsic modality gap: text is often the easier modality and vision the harder one. Under fused InfoNCE, optimization is free to solve the task through whichever modality is easiest. The consequence is that gradients become dominated by text, while the vision branch loses discriminative power. The paper diagnoses this with three forms of evidence.

First, dropout-style modality tests compare fused-to-fused, text-to-fused, and vision-to-fused performance. After SFT, vision-to-fused degrades, and fused-to-fused can become worse than text-to-fused, meaning that the image branch contributes negative transfer rather than complementary signal.

Second, gradient analysis shows that

i∈Ii \in \mathcal{I}1

while i∈Ii \in \mathcal{I}2 drops quickly. Optimization thus follows the text gradient almost exclusively.

Third, the paper analyzes embedding geometry through alignment, uniformity, and separability: i∈Ii \in \mathcal{I}3

i∈Ii \in \mathcal{I}4

i∈Ii \in \mathcal{I}5

After standard SFT, the vision modality’s separability i∈Ii \in \mathcal{I}6 collapses to approximately i∈Ii \in \mathcal{I}7 or lower, indicating that negatives are no longer separated from positives in the visual space (Kim et al., 18 Mar 2026).

4. Weak-modality Penalized Contrastive Learning and CRTR

VLM2Rec introduces Weak-modality Penalized Contrastive Learning (WPCL) to rectify the gradient imbalance. For each user i∈Ii \in \mathcal{I}8 and modality i∈Ii \in \mathcal{I}9, the discriminative margin is defined as

tit_i0

The stronger and weaker modality margins are then identified by tit_i1 and tit_i2, and the modality gap is defined as

tit_i3

where tit_i4 is the stop-gradient operator (Kim et al., 18 Mar 2026).

This gap controls a dynamic penalty weight,

tit_i5

which rescales the negative term of the fused contrastive loss: tit_i6 The stop-gradient is critical: it prevents the model from reducing the gap by degrading the strong modality. Instead, the optimization pressure is placed on improving the weak modality’s negative separation.

Because aggressive negative pushing can distort geometry, VLM2Rec adds Cross-Modal Relational Topology Regularization (CRTR). For each modality, in-batch similarities are converted into ranking distributions tit_i7 and tit_i8, and consistency is enforced with symmetric KL divergence: tit_i9 This regularizer preserves relational topology rather than point-wise alignment, so the modalities retain distinct geometry while sharing neighborhood structure. The final objective is

viv_i0

A plausible implication is that VLM2Rec is less an architectural innovation than an objective-level intervention: the paper deliberately keeps fusion simple so that the method’s contribution lies in how the multimodal space is optimized (Kim et al., 18 Mar 2026).

5. Evaluation protocol and empirical results

The experiments use four Amazon 5-core domains—Toys, Beauty, Clothing, and Sports—with titles and images, and with sequence length truncated to viv_i1 (Kim et al., 18 Mar 2026). The reported dataset statistics are: Toys with viv_i2 users, viv_i3 items, and viv_i4 interactions; Beauty with viv_i5 users, viv_i6 items, and viv_i7 interactions; Clothing with viv_i8 users, viv_i9 items, and uu0 interactions; and Sports with uu1 users, uu2 items, and uu3 interactions. Evaluation uses leave-one-out splitting, uu4 sampled negatives per test instance, and the metrics H@10, H@20, N@10, and N@20.

In Task 1 direct recommendation, VLM2Rec outperforms all baselines on all four datasets. On Toys, the best baseline, uu5, reaches uu6, uu7, uu8, and uu9, while VLM2Rec reaches Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],0, Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],1, Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],2, and Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],3. On Beauty, VLM2Rec achieves Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],4, compared with a best baseline range of Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],5–Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],6. On Clothing, VLM2Rec reaches Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],7, versus Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],8 for the best baseline. On Sports, it reaches Su=[i1,i2,…,i∣Su∣],S_u = [i_1, i_2, \dots, i_{|S_u|}],9, versus i∣Su∣+1i_{|S_u|+1}0 (Kim et al., 18 Mar 2026).

In Task 2 downstream SR initialization, VLM2Rec also produces the best embeddings for GRU4Rec and SASRec. For GRU4Rec on Beauty, the best baseline reports i∣Su∣+1i_{|S_u|+1}1 for i∣Su∣+1i_{|S_u|+1}2 and i∣Su∣+1i_{|S_u|+1}3 for i∣Su∣+1i_{|S_u|+1}4, while VLM2Rec reports i∣Su∣+1i_{|S_u|+1}5. For SASRec on Sports, the best baseline gives i∣Su∣+1i_{|S_u|+1}6, and VLM2Rec gives i∣Su∣+1i_{|S_u|+1}7. On Beauty with SASRec, the best baseline gives i∣Su∣+1i_{|S_u|+1}8, and VLM2Rec gives i∣Su∣+1i_{|S_u|+1}9.

The robustness experiments reinforce the same pattern. In cross-domain recommendation, training on Clothing and testing on Beauty yields zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),0 for VLM2Rec in Task 1, versus zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),1 for LLMEmb and zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),2 for zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),3. Training on Sports and testing on Clothing yields zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),4 for VLM2Rec, versus zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),5 for zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),6. In few-shot training, even with zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),7 sampled users, VLM2Rec beats many full-training baselines, and with zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),8 users it nearly matches full-data baselines while reducing epoch time on Beauty from zuT=Φ(PT(SuT)),zuV=Φ(PV(SuV)),\mathbf{z}_u^T = \Phi(P_T(S_u^T)), \qquad \mathbf{z}_u^V = \Phi(P_V(S_u^V)),9 minutes to ii0 (Kim et al., 18 Mar 2026).

Ablation results show that WPCL is the central component. Removing ii1 causes the largest drop; for example, on Beauty Task 1, ii2 falls from ii3 to ii4. Removing ii5 causes a smaller decline, from ii6 to ii7, indicating that CRTR acts primarily as a stabilizer rather than the principal source of gain.

6. Interpretation, broader usage of the term, and limitations

Within recommendation research, VLM2Rec specifically denotes the modality-collapse-aware sequential recommendation framework of (Kim et al., 18 Mar 2026). However, adjacent literature sometimes uses the expression more loosely to mean recommending or reusing VLMs, or more generally VLM-based recommendation or retrieval. For example, “Vision-LLM Selection and Reuse for Downstream Adaptation” formulates VLM selection and reuse as a recommendation problem over a model hub (Tan et al., 30 Jan 2025). This suggests that the term has both a narrow meaning, tied to multimodal sequential recommendation with collapse mitigation, and a broader informal meaning, tied to recommendation-like VLM selection and reuse.

Several limitations are explicit. VLM2Rec remains computationally heavier than text-only LLM methods, because image tokens make VLM training expensive. Its effectiveness depends on reasonably informative titles and images. The method is designed around two modalities, text and image, and its sequence length is capped at ii8. Extending the user-adaptive margin and topology regularization to additional modalities such as audio, video, or structured attributes is identified as future work (Kim et al., 18 Mar 2026).

Another possible misunderstanding is that VLM2Rec’s contribution is a new multimodal fusion module. The paper instead argues for the opposite design choice: element-wise summation is kept intentionally simple so that the framework’s gains can be attributed to balanced optimization rather than fusion complexity. In that sense, VLM2Rec is best understood as a method for making large VLM embedders usable for multimodal SR by preventing text-dominant optimization from collapsing the visual branch.

Taken together, VLM2Rec establishes a specific research agenda within multimodal recommendation: large VLM embedders are not merely stronger encoders, but unstable optimizers unless their modality imbalance is addressed directly. The framework’s central technical claim is therefore not that vision-LLMs should replace conventional encoders, but that balanced modality utilization must be enforced if CF-aware VLM embeddings are to improve recommendation accuracy and robustness at all (Kim et al., 18 Mar 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to VLM2Rec.