Papers
Topics
Authors
Recent
Search
2000 character limit reached

impuTMAE: Multimodal Imputation & Survival

Updated 8 July 2026
  • impuTMAE is a transformer-based end-to-end model designed for imputing missing modalities in heterogeneous cancer data.
  • It leverages modality-specific encoders, patch-wise masking, and a shared multimodal decoder to reconstruct genetic, imaging, and clinical inputs.
  • The model employs masked pre-training followed by survival fine-tuning, achieving state-of-the-art glioma survival predictions on TCGA-GBM/LGG and BraTS datasets.

Searching arXiv for the impuTMAE paper and closely related multimodal masked autoencoder work. impuTMAE is a transformer-based end-to-end approach for missing modalities imputation in cancer survival prediction. It combines multimodal masked pre-training with downstream survival modeling, and is designed for heterogeneous, incomplete medical data comprising genetic, imaging, and clinical inputs. The reported implementation is pre-trained on incomplete multimodal data and fine-tuned for glioma survival prediction using TCGA-GBM/LGG and BraTS, integrating five modalities: DNA methylation (DNAm), RNA-seq, magnetic resonance imaging (MRI), whole-slide imaging (WSI), and clinical variables (Boyko et al., 8 Aug 2025).

1. Problem setting and model objective

The method addresses a central difficulty in multimodal medical learning: prognostic datasets are often incomplete at the modality level. The paper frames this as a joint representation-learning and imputation problem, rather than treating imputation as a separate preprocessing stage. In this formulation, missing modalities are handled during pre-training through masked reconstruction, and the same machinery is retained for downstream survival prediction. The stated goal is to learn both inter-modal and intra-modal interactions while simultaneously imputing missing modalities by reconstructing masked patches (Boyko et al., 8 Aug 2025).

This design places impuTMAE in the masked-autoencoding family, but with a specifically multimodal and clinical orientation. The model is described as using a single multimodal Transformer decoder to jointly reconstruct all medical modalities and impute missing data in a scalable manner. A plausible implication is that the decoder is not merely a denoising component; it functions as a shared cross-modal reconstruction operator whose outputs regularize the latent space toward modality-complete patient representations.

2. Architectural organization

The architecture is composed of modality-specific encoders, patch-wise masking, embedding fusion, and a shared multimodal decoder. Each modality is processed with a tailored encoder rather than a single homogeneous frontend, reflecting the heterogeneity of omics, volumetric imaging, histopathology, and clinical variables (Boyko et al., 8 Aug 2025).

Modality Encoder Patch specification
RNA-seq 1D convolutional patchifier + 6-layer transformer Patch size 512
DNAm 1D convolutional patchifier + 6-layer transformer Patch size 1024
MRI 2 x 3D convolutions + 4-layer transformer Kernel/stride 4; patch size 16×16×1616\times16\times16
WSI 2D convolutional patchifier + 4-layer, 4-head transformer Patch size 16×1616\times16
Clinical data Age, radiation, pharmaceutical treatment Combined downstream

Inputs are split into non-overlapping patches. For RNA-seq and DNAm, the patchification is one-dimensional; for MRI and WSI, it is spatial. Patch embeddings are augmented with positional encodings and a [CLS][CLS] token per modality, then concatenated. The decoder is a shared 3-layer transformer that reconstructs modality-specific outputs in parallel. The reconstruction heads are linear layers for RNA and DNAm, two 3D transposed convolutions for MRI, and a 2D transposed convolution for WSI (Boyko et al., 8 Aug 2025).

The paper distinguishes intra-modal from inter-modal interaction learning. Intra-modal dependencies are modeled within each encoder through attention over the modality’s own patch sequence. Inter-modal dependencies are learned after concatenation of modality embeddings, through the multimodal decoder and a later fusion attention block. This division of labor is central to the method’s imputation behavior: modality-specific encoders extract structured local representations, while the decoder leverages cross-modal context to reconstruct masked or absent content.

3. Masked pre-training and imputation mechanism

Masked pre-training is the defining training strategy. During pre-training, 50% of patches are randomly masked, and entire missing modalities are regarded as fully masked. Reconstruction is supervised with a multimodal mean-squared-error objective computed only on masked patches of non-missing modalities (Boyko et al., 8 Aug 2025):

MSEmultimodal=∑m=1MMSE(m)\text{MSE}_\text{multimodal} = \sum_{m=1}^M \text{MSE}^{(m)}

MSE(m)=1Nm∑i=1Nm(yi(m)−y^i(m))2\text{MSE}^{(m)} = \frac{1}{N_m} \sum_{i=1}^{N_m} \left(y_i^{(m)} - \hat{y}_i^{(m)}\right)^2

For truly missing modalities, the data is omitted from loss computation. At the same time, the decoder is described as capable of reconstructing missing or masked modalities, enabling the use of incomplete datasets in both pretraining and downstream inference. Read together, these statements indicate that the model learns reconstruction from available multimodal context even though supervision is restricted to cases where a ground-truth target is present.

This masking scheme differs materially from pipelines that require complete multimodal tuples during pre-training. The paper explicitly contrasts impuTMAE with contrastive pretraining, stating that it can use any available subset of modalities by treating missing modalities as masked and reconstructing them for downstream or clinical use. This suggests that data efficiency is not only a computational property but also a sampling property: incomplete patients remain useful during representation learning rather than being discarded.

4. Survival modeling and downstream fine-tuning

After pre-training, the encoders are reused for glioma survival prediction. The reported fine-tuning protocol uses the pretrained encoders as fixed or partially frozen feature extractors: MRI and WSI encoders are fully frozen, and for DNAm and RNA, 5 out of 6 transformer layers are frozen (Boyko et al., 8 Aug 2025).

The downstream pipeline preserves the imputation stage. The multimodal decoder imputes missing modalities so that a unified patient representation can be formed even when input data remain incomplete. Concatenated modality features, with WSI patch embeddings averaged, are then processed by a single self-attention fusion block with latent size 256. A linear projection outputs discrete-time hazard predictions over T=20T=20 intervals, following the reported formulation (Boyko et al., 8 Aug 2025):

L=−1n∑i=1n∑t=1κ(ti)(yitlog⁡h(τt∣xi)+(1−yit)log⁡(1−h(τt∣xi)))\mathcal{L} = - \frac{1}{n} \sum_{i=1}^{n} \sum_{t=1}^{\kappa(t_i)} \left( y_{it} \log h(\tau_t \mid x_i) + (1 - y_{it}) \log(1 - h(\tau_t \mid x_i)) \right)

Here h(τt∣xi)h(\tau_t \mid x_i) denotes the hazard at interval tt, and yity_{it} is the event indicator. The use of discrete-time hazards places impuTMAE within neural survival analysis rather than pure risk scoring. This suggests that the model’s multimodal representation is optimized not simply for ranking but for interval-wise hazard estimation.

5. Data, optimization protocol, and reported performance

The experimental setting combines TCGA-GBM/LGG with BraTS. The integrated modalities are specified as follows: DNAm with 25,978 Beta values, RNA-seq with 16,304 genes using FPKM-UQ normalization, MRI using T1-w 16×1616\times160 voxel crops from BraTS, WSI using 10 top 16×1616\times161-pixel histopathology patches per subject, and clinical variables comprising age, radiation therapy, and pharmaceutical treatment (Boyko et al., 8 Aug 2025).

The reported pre-training optimizer is AdamW with learning rate 16×1616\times162, batch size 24, 800 epochs, and cosine annealing. Fine-tuning is conducted for 20 epochs with dropout 0.1. Different learning rates are used depending on the modality combination, with 16×1616\times163 given as an example for RNA+DNA. Evaluation uses 5-fold cross-validation on 80/20 train/test splits (Boyko et al., 8 Aug 2025).

Performance is reported using Concordance Index (C-index) and CS-score, where

16×1616\times164

The paper states that impuTMAE surpasses prior multimodal approaches and reports state-of-the-art performance for glioma survival prediction (Boyko et al., 8 Aug 2025).

Model Modalities Reported scores
DRIM RNA, DNA, MRI, WSI C-index 0.825; CS-score 0.856; CS-score* 0.863
MultiSurv RNA, DNA, WSI, CLN C-index 0.790; CS-score 0.711; CS-score* 0.880
impuTMAE (all) RNA, DNA, MRI, WSI, CLN C-index 0.831; CS-score 0.864; CS-score* 0.863
impuTMAE (RNA+DNA+MRI+CLN) RNA, DNA, MRI, CLN C-index 0
Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to impuTMAE.