Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

Published 13 Jun 2026 in cs.AI | (2606.15038v1)

Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift. We introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data, designed to generalize across tasks and institutions. CT and EHR modalities are encoded independently using domain-specific foundation models and aligned in a shared latent space through four principled fusion strategies: late fusion, contrastive alignment, cross-attention, and co-attention. We evaluate two clinically distinct TTE tasks: pulmonary embolism (PE) mortality and cardiovascular disease (CVD) outcomes, on large-scale multi-institutional cohorts (PE: N=3,099 train; 1,098 internal; 435 external; CVD: N=2,951 train; 837 internal; 682 external). Fusion consistently improves concordance index by 1.5-5.4% over unimodal baselines when modalities contribute comparably. Overall, contrastive multimodal fusion, particularly with CLMBR representations, provided the most consistent and statistically robust improvements, especially for PE mortality prediction. For MACE, cross-attention (one-hot) achieved the highest internal performance and image-guided co-attention achieved the best external performance. We therefore introduce a generalizable foundation model-based cross-modal alignment framework and provide the first systematic analysis of fusion behavior under modality imbalance in TTE prediction. Our results establish task-aware multimodal alignment as a necessary design principle for robust generalization and scalable clinical deployment.

Summary

  • The paper demonstrates that task-aware cross-modal fusion, including contrastive and attention-based strategies, significantly enhances TTE prediction in clinical settings.
  • It employs frozen foundation models for CT (MedImageInsight) and EHR (CLMBR-T-base) with four fusion paradigms to align multimodal representations in a shared latent space.
  • Experimental results across PE mortality and MACE tasks show that adaptive fusion outperforms unimodal approaches, improving generalization and interpretability.

Cross-Modal Representation Alignment for Time-to-Event Modeling: Task-Specific Fusion Strategies

Introduction

The paper "Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling" (2606.15038) introduces a principled multimodal framework that aligns CT imaging and longitudinal EHR data in a shared latent space to predict time-to-event (TTE) outcomes in clinical settings. By leveraging domain-adapted foundation models and exploring the effect of cross-modal fusion mechanisms, the work systematically investigates the dependency of optimal fusion strategies on task and modality characteristics. The experimental setup spans large-scale multi-institutional cohorts targeting pulmonary embolism (PE) mortality and major adverse cardiac event (MACE) prediction, elucidating the interplay between modality contribution, fusion strategies, and generalization under distributional shift. Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1: Overview of the multimodal survival framework and fusion strategies. Cross-modal alignment of chest CT and longitudinal EHR using domain-specific foundation encoders and four latent-space fusion approaches for TTE prediction.

Methodological Framework

Foundation Model Encoders

The approach employs frozen foundation models for modality-specific embedding extraction: MedImageInsight for CT imaging, producing averaged 1×10241 \times 1024 slice-based representations, and the CLMBR-T-base for EHR, generating 768-dimensional patient-level embeddings via autoregressive pretraining on longitudinal clinical codes.

The paper notes the inclusion of hand-curated one-hot EHR encodings for benchmark comparisons, reflecting real-world clinical risk scoring practices and offering a controlled assessment of fusion efficacy.

Fusion Mechanisms

The study contrasts four fusion paradigms:

  • Traditional Concatenation: Direct linear joining of modality embeddings without further alignment.
  • Contrastive Alignment: Two-step training where symmetric contrastive loss aligns paired image and EHR embeddings, enforcing separation from negative pairs and subsequently optimizing for survival prediction.
  • Cross-Attention: Bidirectional supervised attention mechanisms allowing each modality to attend selectively to relevant features from the other, improving adaptivity to task relevance.
  • Co-Attention: Guiding modality-based embedding refinement, allowing one modality's representation to serve as the query for attention computation, thus focusing on contextually significant features.

All fusion variants are integrated upstream of an MLP prediction head with negative log-likelihood optimization, enabling end-to-end alignment for the TTE objective with right-censoring.

Experimental Evaluation

Cohorts and TTE Tasks

Two benchmarks were constructed:

  • PE Mortality Task: Multi-institutional datasets with acute PE confirmed on CTPA studies and linked standardized EHR, externally validated on the INSPECT cohort.
  • MACE Task: Retrospective and external collections of thoracic CTs paired with cardiovascular outcomes (myocardial infarction, stroke, heart failure, cardiac mortality), with temporal EHR features per AHA recommendations.

Quantitative Results

Multimodal fusion models robustly outperformed unimodal baselines in both tasks. Notably, CLMBR-based contrastive fusion achieved the highest internal concordance for PE mortality (AUC 0.862), with consistently strong external performance (AUC 0.743), demonstrating statistically significant improvements over both image-only and concatenation. For MACE, cross-attention (one-hot) and image-guided co-attention emerged as superior performers for internal and external validation, respectively. The results substantiate the claim that fusion strategies must be adapted to the task, modality balance, and target cohort, rather than employing a uniform approach.

EHR-only predictions were significantly inferior to image-only models across settings, highlighting the necessity of capturing modality complementarity in fusion.

Saliency Analysis

Gradient-based saliency mapping revealed that multimodal fusion altered spatial attention in imaging: Figure 2

Figure 2: Comparative saliency maps for MACE and PE outcome prediction using image-only and multimodal fusion models. Highlighted regions correspond to predicted risk for adverse events.

Fusion shifted emphasis toward clinically relevant anatomical structures—embolic burden for PE and pulmonary artery dilation for MACE—indicating that EHR integration dynamically guides imaging features toward pathophysiologically pertinent zones.

Discussion and Implications

Supervised cross-modal fusion mechanisms, particularly contrastive alignment and attention-based strategies, provide superior adaptivity and discrimination compared to baseline concatenation and pre-aligned spaces. Task and modality dependence is pronounced: CT delivers discriminatory value in MACE, while longitudinal EHR excels for PE mortality. The paper underscores a limitation of generic foundation models, as representations optimized for unsupervised objectives do not reliably encode survival-predictive signals, requiring explicit supervised realignment.

Attention-based fusion mechanisms outperform simple concatenation, particularly under modality imbalance and distribution shift, by dynamically weighting source contributions. The saliency findings demonstrate that incorporating EHR context enables imaging models to focus on anatomically relevant features, directly enhancing interpretability and clinical applicability.

Persistent generalization gaps, despite external validation, suggest that robust domain adaptation and harmonization of foundational encoders are prerequisites for scalable deployment. Limitations include event-rate imbalance, cohort size, and residual confounding, necessitating prospective validation and calibration before clinical translation.

Conclusion

This paper rigorously demonstrates that multimodal fusion for TTE prediction is inherently task-dependent; the optimal alignment strategy varies across outcome, modality, and cohort. Fusion consistently enhances predictive accuracy compared to unimodal baselines, with contrastive and attention-based approaches providing the most significant gains, particularly in external cohorts. The findings advocate for supervised, task-aware cross-modal alignment as a foundational principle in clinical survival modeling, challenging the prevailing use of static fusion heuristics and underscoring the need for further research in adaptive fusion, resilience to domain shifts, and harmonization of multimodal encoders in medical AI.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.