Papers
Topics
Authors
Recent
Search
2000 character limit reached

Masked and Predictive Self-Supervised Foundation Models for 3D Brain MRI

Published 11 Jun 2026 in cs.CV and eess.IV | (2606.13315v1)

Abstract: Self-supervised foundation models have shown strong promise in medical imaging. However, existing MRI foundation-model studies have primarily emphasized segmentation and dense prediction tasks, while systematic investigation of self-supervised foundation models for MRI-based disease detection remains limited. In this work, we investigate two major self-supervised pretraining paradigms for MRI-based disease detection: reconstruction-based learning via Masked Autoencoders (MAE) and predictive representation learning via Joint Embedding Predictive Architectures (JEPA). We study the role of auxiliary objectives by introducing a novel spectral-domain reconstruction loss for MAE to enhance sensitivity to fine-grained anatomical structure, and by integrating variance--covariance regularization (VCR) within our JEPA framework to encourage decorrelated latent representations. Our models are pretrained on heterogeneous single-contrast MRI volumes in a contrast-agnostic setting, without modality concatenation. Across five downstream disease detection tasks, our results highlight the importance of self-supervised objective design for medical foundation model pretraining, demonstrating that the downstream benefit of each objective is determined by its relevance to the task's structure. Specifically, spectral regularization yields the largest improvements when the downstream discriminative signal is characterized by strong high-frequency anatomical structures, while covariance regularization is most beneficial when discriminative information spans multiple decorrelated feature dimensions. MAE with spectral-domain supervision consistently achieves superior downstream performance for MRI-based disease detection. These findings suggest that self-supervised objectives in medical imaging encode specific biases, and their downstream benefit is fundamentally conditioned on the task's structure.

Summary

  • The paper systematically compares masked autoencoders and JEPA for 3D brain MRI, finding that spectral-loss MAE achieved a 0.945 AUC for AD versus controls on ADNI and generally outperformed JEPA.
  • The study shows auxiliary objectives should match disease-signal structure: spectral regularization benefits high-frequency anatomical tasks, while covariance regularization helps when information is distributed across independent features.
  • The models support practical low-label learning, with spectral MAE outperforming supervised 3D ResNet-50 baselines in few-shot tests and frozen features sometimes exceeding fine-tuning under distribution shift.

Overview

This paper presents a systematic comparison of two self-supervised learning (SSL) paradigms—Masked Autoencoders (MAE) and Joint Embedding Predictive Architectures (JEPA)—as pretraining frameworks for 3D structural brain MRI, with a specific focus on disease detection rather than the segmentation tasks that dominate prior MRI foundation-model literature. The authors pretrain ViT-Base 3D backbones on 58,781 heterogeneous MRI volumes from approximately 8,769 subjects spanning seven datasets (ADNI, NACC/SCAN, PPMI, OASIS-3, IXI, MOOD, BraTS-2024), using single-contrast inputs without modality concatenation. Two domain-motivated auxiliary objectives are introduced: a spectral-domain reconstruction loss for MAE (SL-MAE) computed on high-pass filtered log-magnitude spectra via 3D FFT, and variance–covariance regularization (VCR) integrated into JEPA to encourage decorrelated latent representations.

The central empirical claim is that the downstream benefit of any auxiliary SSL objective is conditioned on the statistical structure of the downstream task: spectral regularization helps most when discriminative signal is high-frequency anatomical structure, while covariance regularization helps most when discriminative information is distributed across multiple decorrelated feature dimensions. Notably, despite JEPA's recent momentum in natural image modeling, MAE with spectral supervision consistently outperformed JEPA variants across the evaluated disease detection tasks.

Pretraining framework and design choices

A distinctive aspect of the framework is its contrast-agnostic, minimally normalized preprocessing pipeline. Volumes retain their native in-plane acquisition orientation (randomly sampled among axial, coronal, or sagittal for isotropic data), are cropped to 160×160160 \times 160 in-plane, depth-interpolated to a fixed slice count, and divided into 838^3 tubelets. The authors deliberately avoid spatial normalization, bias-field correction, and intensity harmonization beyond percentile clipping and per-volume standardization. This design accommodates incomplete contrast series and heterogeneous acquisition protocols directly, at the cost of forgoing the benefits that harmonization might confer—a trade-off the paper does not quantify against a harmonized baseline.

Both MAE and JEPA adopt spatiotemporal block masking adapted from V-JEPA, with a 75% masking ratio. The SL-MAE objective augments pixel-space MSE with a loss on high-pass filtered log-magnitude 3D FFT spectra; because full-volume FFT computation is expensive, the spectral term is amortized by computing it every nn iterations and distributing gradients over a cached buffer. The JEPA variant follows the student–teacher EMA formulation with an L1 prediction loss and stop-gradient, augmented with VCR applied jointly to concatenated context and target representations; for efficiency, a single covariance matrix is computed treating all token positions as independent samples rather than per-token covariances as in the original V-JEPA work.

Downstream evaluation uses an Attentive Classifier head under both fine-tuned and frozen settings across five tasks: NC vs. AD and NC vs. MCI on ADNI (5-fold cross-validation) and NACC/SCAN, binary and multiclass glioma grading on UCSF-PDGM, and autism classification on ABIDE. UCSF and ABIDE were excluded from pretraining, providing cross-domain evaluation. Strict subject-level splitting is enforced throughout, and all splits are publicly released—an important contribution given the field's lack of standardized evaluation protocols.

Results: spectral regularization in MAE

The strongest result in the paper is the effect of spectral supervision on ADNI NC vs. AD classification under fine-tuning: mean AUC improves from 0.866 ± 0.130 to 0.945 ± 0.024, with fold-wise variance dropping more than fivefold. This gain propagates to NACC/SCAN NC vs. AD (0.761 → 0.841) and UCSF tumor grading (binary AUC 0.770 → 0.833), while gains on NC vs. MCI and ABIDE autism detection are marginal or absent. The paper supports its mechanistic interpretation quantitatively: UCSF tumor volumes contain approximately 1.93× higher mean high-frequency power spectral density than ABIDE volumes, aligning the magnitude of SL-MAE's benefit with the high-frequency content of each task's discriminative signal. The advantage also shrinks on UCSF multiclass grading (0.786 → 0.797), consistent with the observation that multiclass grading requires cues beyond texture alone.

Under frozen evaluation, SL-MAE retains its advantage on ADNI NC vs. AD (0.922 vs. 0.914) and both UCSF tasks, but plain MAE leads on NC vs. MCI and ABIDE—evidence that spectral pretraining imposes a representational bias toward high-frequency structure that can reduce frozen-feature discriminability when such structure is uninformative.

An additional finding concerns frozen versus fine-tuned transfer. On NACC/SCAN NC vs. MCI (0.702 vs. 0.677) and ABIDE (0.616 vs. 0.580), frozen MAE representations exceed their fine-tuned counterparts, which the authors attribute to overfitting and distribution shift during adaptation—ABIDE's predominantly adolescent cohort differs substantially from the older ADNI-based pretraining population. This echoes prior reports that fine-tuning can distort pretrained features under distribution shift, and suggests frozen probing remains a meaningful evaluation protocol for structural brain MRI.

Results: covariance regularization in JEPA

VCR-JEPA improves over plain JEPA most substantially on ADNI NC vs. AD (fine-tuned AUC 0.854 vs. 0.759; frozen 0.709 vs. 0.680) and on UCSF tumor grading under frozen evaluation (binary 0.673 vs. 0.590; multiclass 0.640 vs. 0.599). The pattern reverses for NC vs. MCI, where plain JEPA achieves higher frozen AUC (0.716 vs. 0.694) with notably lower fold variance (std 0.057 vs. 0.106), indicating that aggressive decorrelation introduces instability when the discriminative signal is spatially concentrated in medial temporal structures. The authors interpret this through the dimensionality of each task's pathological signal: AD and glioma grading involve multiple partially independent anatomical systems that benefit from distributed latent representations, whereas MCI-related changes are low-dimensional and localized.

Across nearly all comparisons, MAE-based models outperform JEPA-based models. The paper explicitly challenges the common assumption that reconstruction-based methods suit only dense prediction while predictive/contrastive methods suit semantic classification, arguing that with a deliberately lightweight decoder, MAE's encoder must absorb global semantic context to reconstruct missing anatomy successfully.

Comparisons, clustering analysis, and few-shot performance

Against BrainFound, a DINOv2-based brain MRI foundation model using multimodal input and a ViT-Large backbone, SL-MAE achieves higher AUC on ADNI NC vs. AD (0.945 vs. 0.810) and UCSF tumor grading (0.833 vs. 0.780) despite training from scratch with single-contrast inputs and smaller capacity; BrainFound leads on NACC/SCAN NC vs. AD (0.883 vs. 0.841). The comparison is approximate given differing splits and protocols, a caveat the authors state plainly.

Clustering metrics on held-out ADNI embeddings corroborate the AUC findings: after fine-tuning, SL-MAE attains the best Silhouette Score (0.9218) and NMI (0.8977) in the raw 768-dimensional space, with VCR-JEPA improving over plain JEPA analogously. Pretrained encoders show near-zero separability, as expected from label-agnostic SSL. The authors note that t-SNE projections can misrepresent VCR's benefits due to distortion of global geometry, and therefore emphasize metrics computed in the original embedding space.

In few-shot experiments against a fully supervised ResNet50-3D trained from scratch, SL-MAE consistently wins across k=16,32,64k = 16, 32, 64 labeled samples per class on both NC vs. AD (e.g., 0.618 vs. 0.598 at k=16k{=}16; 0.775 vs. 0.617 at k=64k{=}64) and NC vs. MCI (0.791 vs. 0.765 at k=16k{=}16). VCR-JEPA surpasses the supervised baseline except at k=32k{=}32. The gap is largest in the lowest-data regime, supporting the clinical relevance of SSL pretraining where annotations are scarce.

Limitations and open questions

The paper concedes several constraints. Evaluations are restricted to ViT-Base scale and to structural brain MRI disease detection; comparisons with modern contrastive frameworks such as DINOv3 remain unperformed. Several downstream results rest on single train/test splits (NACC/SCAN, UCSF, ABIDE) rather than cross-validation, so fold-level variance is unavailable for those tasks. The claimed mechanism linking auxiliary objectives to task structure is supported by correlational evidence—the PSD analysis and per-task gain patterns—rather than by direct intervention on representation frequency content, and the interpretation that MCI signals are "low-dimensional" is a hypothesis consistent with, but not proven by, the data. JEPA-based downstream adaptation was observed to be hyperparameter-sensitive, raising the possibility that JEPA's weaker showing partly reflects tuning difficulty rather than fundamental representational inferiority. Whether the task-dependence of auxiliary-objective benefit generalizes to other imaging domains, larger multimodal corpora, and newer JEPA/DINO variants is left open, as is the question of whether additional spectral, probabilistic, or information-theoretic objectives would yield further gains.

Conclusion

This work provides the first controlled side-by-side comparison of MAE and JEPA pretraining for 3D structural MRI disease detection, and demonstrates that auxiliary objectives encode specific inductive biases whose downstream value depends on the pathological signal's structure. The headline quantitative results—SL-MAE reaching 0.945 AUC on ADNI NC vs. AD, consistent few-shot superiority over supervised training, and the demonstrated alignment between high-frequency image content and spectral-loss benefit—are substantive. Equally important is the negative result that neither objective helps uniformly, and that frozen representations can outperform fine-tuning under distribution shift. Combined with strict subject-level splitting and public release of all data partitions, the study offers a reproducible benchmark and a clear message: self-supervised objectives in medical imaging should be selected based on the statistical properties of the target clinical task, not treated as interchangeable defaults.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.