VLM2Rec: Resolving Modality Collapse
- The paper introduces VLM2Rec, a CF-aware framework that prevents modality collapse by balancing vision and text signals in sequential recommendations.
- It employs Weak-modality Penalized Contrastive Learning and Cross-Modal Relational Topology Regularization to optimize and preserve distinct modality features.
- Extensive experiments demonstrate that VLM2Rec outperforms state-of-the-art baselines in accuracy and robustness across diverse recommendation scenarios.
Searching arXiv for the primary paper and closely related contextual works. arXiv paper: (Kim et al., 18 Mar 2026) Title: VLM2Rec: Resolving Modality Collapse in Vision-LLM Embedders for Multimodal Sequential Recommendation Authors: Chenglin Song, Saeed Hamidpour, Ruihong Qiu, Tiankang Yu, Jiawei Chen, Licheng Jiao, Lirong Wang Published: 2026-03-18 URL: http://arxiv.org/abs/([2603.17450](/papers/2603.17450))v1
Abstract: Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity and prevents Collaborative Filtering (CF) signals from being fully integrated into item representations. Inspired by the recent success of LLMs as high-capacity embedders, we investigate the use of Vision-LLMs (VLMs) as CF-aware multimodal encoders for SR. However, we find that standard contrastive supervised fine-tuning (SFT), which adapts VLMs for embedding generation and injects CF signals, can amplify its inherent modality collapse. In this state, optimization is dominated by a single modality while the other degrades, ultimately undermining recommendation accuracy. To address this, we propose VLM2Rec, a VLM embedder-based framework for multimodal sequential recommendation designed to ensure balanced modality utilization. Specifically, we introduce Weak-modality Penalized Contrastive Learning to rectify gradient imbalance during optimization and Cross-Modal Relational Topology Regularization to preserve geometric consistency between modalities. Extensive experiments demonstrate that VLM2Rec consistently outperforms state-of-the-art baselines in both accuracy and robustness across diverse scenarios. Searching for contextual papers that use “VLM2Rec” more loosely or provide adjacent formulations. arXiv paper: (Tan et al., 30 Jan 2025) Title: Vision-LLM Selection and Reuse for Downstream Adaptation Authors: Zhixiang Chi, Guangyi Fan, Yunge Zhang, Tao Pu, Xiangyu Zhao, Jianye Hao, Zhenya Huang Published: 2025-01-30 URL: http://arxiv.org/abs/([2501.18271](/papers/2501.18271))v1
Abstract: Pre-trained Vision-LLMs (VLMs) are becoming increasingly popular across various visual tasks, and several open-sourced VLM variants have been released. However, selecting the best-performing pre-trained VLM for a specific downstream task is challenging since no single VLM can achieve promising performance on all downstream tasks, and evaluating all available VLMs is impossible due to time and data limitations. To address this problem, this paper proposes a novel paradigm to select and reuse VLM for downstream tasks, called Model Label Learning (MLL). The proposal contains three key modules: model labeling, which assigns labels to each VLM to describe their specialty and utility; model selection, which matches the requirements of the target task with model labels; and model reuse, which applies selected VLMs to the target task in an ensemble manner. The proposal is highly computationally efficient and growable since the model labeling process is completed target task independent and the ability could grow with the number of candidate VLMs. We also introduce a new benchmark for evaluating VLM selection methods, including 49 VLMs and 17 target task datasets. Experimental results clearly demonstrate the effectiveness of the proposed method for selecting and reusing VLMs.
VLM2Rec is a VLM embedder-based framework for multimodal sequential recommendation that uses a large Vision-LLM as a CF-aware multimodal encoder and is explicitly designed to prevent modality collapse during contrastive supervised fine-tuning. In the formulation introduced in “VLM2Rec: Resolving Modality Collapse in Vision-LLM Embedders for Multimodal Sequential Recommendation” (Kim et al., 18 Mar 2026), the central claim is that standard InfoNCE-style SFT can make optimization collapse onto a single modality—typically text—so that the weaker modality degrades and fused recommendation quality is undermined. VLM2Rec addresses this failure mode with two objective-level components, Weak-modality Penalized Contrastive Learning and Cross-Modal Relational Topology Regularization, while retaining a deliberately simple fusion design.
1. Problem setting and conceptual scope
VLM2Rec is defined in the setting of multimodal sequential recommendation (SR). There is a user set , an item set , and for each item , both text and image are available. Each user is associated with a chronological interaction sequence
and the task is to predict the next item (Kim et al., 18 Mar 2026).
The framework is motivated by two limitations of conventional multimodal SR pipelines. First, these pipelines often rely on small frozen pretrained encoders, which impose low semantic capacity. Second, because those encoders are frozen, Collaborative Filtering signals cannot reshape the semantic space, so the sequential backbone must operate over representations that are not co-adapted to user-item interaction structure. VLM2Rec therefore repurposes a large VLM as both a sequence encoder and a CF-aware embedder.
A common misconception is that simply replacing frozen encoders with a stronger VLM is sufficient. The paper argues the opposite: the crucial problem is not merely capacity, but the optimization pathology induced by standard contrastive fine-tuning. The resulting Paradox of SFT is that the very objective used to inject CF signals can amplify an intrinsic modality imbalance, so that the multimodal encoder behaves as a text-dominated model rather than a balanced visual-textual recommender (Kim et al., 18 Mar 2026).
2. Encoder architecture and representation design
The backbone in VLM2Rec is a pretrained VLM, specifically Qwen2.5-VL-3B, used as a sequence encoder. For text-only and image-only user sequences, the model produces
and for each item ,
0
The representation is taken from the last-token hidden state of the final Transformer layer. For efficiency, the framework uses LoRA with rank 1, 2, and dropout 3, so the main VLM weights remain effectively frozen (Kim et al., 18 Mar 2026).
Two prompting strategies are studied. In internal (interleaved) fusion, text and image tokens are interleaved in a single prompt. In external (separate) fusion, text and image sequences are encoded through separate prompts. The framework adopts external fusion as the base design, because it decouples modality-specific encoding paths and keeps modality balancing at the objective level rather than embedding it in a more complex fusion architecture.
Fusion itself is intentionally minimal. User and item representations are formed by element-wise summation: 4 No extra fusion parameters are introduced. This design makes the empirical effects of modality balancing attributable to the training objective rather than to a learned fusion block (Kim et al., 18 Mar 2026).
The framework is used in two modes. In Task 1, it operates as a direct recommender by ranking items through sequence-item similarity in embedder space. In Task 2, it acts as a plug-and-play embedding generator: learned item embeddings are projected through a one-layer linear adapter to hidden size 5 and used to initialize downstream SR backbones such as GRU4Rec and SASRec.
3. Modality collapse and the failure of standard contrastive SFT
The paper identifies modality collapse as the key optimization pathology in VLM-based multimodal SR. The starting point is standard fused-embedding InfoNCE: 6 Here, 7 and 8 are fused sequence and item embeddings, 9 is cosine similarity, and 0 is a temperature (Kim et al., 18 Mar 2026).
Empirically, the problem is that pretrained VLMs already exhibit an intrinsic modality gap: text is often the easier modality and vision the harder one. Under fused InfoNCE, optimization is free to solve the task through whichever modality is easiest. The consequence is that gradients become dominated by text, while the vision branch loses discriminative power. The paper diagnoses this with three forms of evidence.
First, dropout-style modality tests compare fused-to-fused, text-to-fused, and vision-to-fused performance. After SFT, vision-to-fused degrades, and fused-to-fused can become worse than text-to-fused, meaning that the image branch contributes negative transfer rather than complementary signal.
Second, gradient analysis shows that
1
while 2 drops quickly. Optimization thus follows the text gradient almost exclusively.
Third, the paper analyzes embedding geometry through alignment, uniformity, and separability: 3
4
5
After standard SFT, the vision modality’s separability 6 collapses to approximately 7 or lower, indicating that negatives are no longer separated from positives in the visual space (Kim et al., 18 Mar 2026).
4. Weak-modality Penalized Contrastive Learning and CRTR
VLM2Rec introduces Weak-modality Penalized Contrastive Learning (WPCL) to rectify the gradient imbalance. For each user 8 and modality 9, the discriminative margin is defined as
0
The stronger and weaker modality margins are then identified by 1 and 2, and the modality gap is defined as
3
where 4 is the stop-gradient operator (Kim et al., 18 Mar 2026).
This gap controls a dynamic penalty weight,
5
which rescales the negative term of the fused contrastive loss: 6 The stop-gradient is critical: it prevents the model from reducing the gap by degrading the strong modality. Instead, the optimization pressure is placed on improving the weak modality’s negative separation.
Because aggressive negative pushing can distort geometry, VLM2Rec adds Cross-Modal Relational Topology Regularization (CRTR). For each modality, in-batch similarities are converted into ranking distributions 7 and 8, and consistency is enforced with symmetric KL divergence: 9 This regularizer preserves relational topology rather than point-wise alignment, so the modalities retain distinct geometry while sharing neighborhood structure. The final objective is
0
A plausible implication is that VLM2Rec is less an architectural innovation than an objective-level intervention: the paper deliberately keeps fusion simple so that the method’s contribution lies in how the multimodal space is optimized (Kim et al., 18 Mar 2026).
5. Evaluation protocol and empirical results
The experiments use four Amazon 5-core domains—Toys, Beauty, Clothing, and Sports—with titles and images, and with sequence length truncated to 1 (Kim et al., 18 Mar 2026). The reported dataset statistics are: Toys with 2 users, 3 items, and 4 interactions; Beauty with 5 users, 6 items, and 7 interactions; Clothing with 8 users, 9 items, and 0 interactions; and Sports with 1 users, 2 items, and 3 interactions. Evaluation uses leave-one-out splitting, 4 sampled negatives per test instance, and the metrics H@10, H@20, N@10, and N@20.
In Task 1 direct recommendation, VLM2Rec outperforms all baselines on all four datasets. On Toys, the best baseline, 5, reaches 6, 7, 8, and 9, while VLM2Rec reaches 0, 1, 2, and 3. On Beauty, VLM2Rec achieves 4, compared with a best baseline range of 5–6. On Clothing, VLM2Rec reaches 7, versus 8 for the best baseline. On Sports, it reaches 9, versus 0 (Kim et al., 18 Mar 2026).
In Task 2 downstream SR initialization, VLM2Rec also produces the best embeddings for GRU4Rec and SASRec. For GRU4Rec on Beauty, the best baseline reports 1 for 2 and 3 for 4, while VLM2Rec reports 5. For SASRec on Sports, the best baseline gives 6, and VLM2Rec gives 7. On Beauty with SASRec, the best baseline gives 8, and VLM2Rec gives 9.
The robustness experiments reinforce the same pattern. In cross-domain recommendation, training on Clothing and testing on Beauty yields 0 for VLM2Rec in Task 1, versus 1 for LLMEmb and 2 for 3. Training on Sports and testing on Clothing yields 4 for VLM2Rec, versus 5 for 6. In few-shot training, even with 7 sampled users, VLM2Rec beats many full-training baselines, and with 8 users it nearly matches full-data baselines while reducing epoch time on Beauty from 9 minutes to 0 (Kim et al., 18 Mar 2026).
Ablation results show that WPCL is the central component. Removing 1 causes the largest drop; for example, on Beauty Task 1, 2 falls from 3 to 4. Removing 5 causes a smaller decline, from 6 to 7, indicating that CRTR acts primarily as a stabilizer rather than the principal source of gain.
6. Interpretation, broader usage of the term, and limitations
Within recommendation research, VLM2Rec specifically denotes the modality-collapse-aware sequential recommendation framework of (Kim et al., 18 Mar 2026). However, adjacent literature sometimes uses the expression more loosely to mean recommending or reusing VLMs, or more generally VLM-based recommendation or retrieval. For example, “Vision-LLM Selection and Reuse for Downstream Adaptation” formulates VLM selection and reuse as a recommendation problem over a model hub (Tan et al., 30 Jan 2025). This suggests that the term has both a narrow meaning, tied to multimodal sequential recommendation with collapse mitigation, and a broader informal meaning, tied to recommendation-like VLM selection and reuse.
Several limitations are explicit. VLM2Rec remains computationally heavier than text-only LLM methods, because image tokens make VLM training expensive. Its effectiveness depends on reasonably informative titles and images. The method is designed around two modalities, text and image, and its sequence length is capped at 8. Extending the user-adaptive margin and topology regularization to additional modalities such as audio, video, or structured attributes is identified as future work (Kim et al., 18 Mar 2026).
Another possible misunderstanding is that VLM2Rec’s contribution is a new multimodal fusion module. The paper instead argues for the opposite design choice: element-wise summation is kept intentionally simple so that the framework’s gains can be attributed to balanced optimization rather than fusion complexity. In that sense, VLM2Rec is best understood as a method for making large VLM embedders usable for multimodal SR by preventing text-dominant optimization from collapsing the visual branch.
Taken together, VLM2Rec establishes a specific research agenda within multimodal recommendation: large VLM embedders are not merely stronger encoders, but unstable optimizers unless their modality imbalance is addressed directly. The framework’s central technical claim is therefore not that vision-LLMs should replace conventional encoders, but that balanced modality utilization must be enforced if CF-aware VLM embeddings are to improve recommendation accuracy and robustness at all (Kim et al., 18 Mar 2026).