- The paper introduces MUST-PET, a multimodal, multi-tracer framework for PET/CT lesion segmentation that achieves superior performance by leveraging paired FDG and PSMA tracers.
- Key results showed a reduction in MAE for PSMA test sets, which supports the model’s ability to improve generalization and cross-domain transfer, especially in scenarios with limited labeled data.
- Segmentation accuracy, assessed using Dice and lesion sensitivity metrics, was best with MUST-PET pretraining for PSMA scans from unseen data distributions, which outperformed baselines such as SwinUNETR scratch and nnU-Net scratch.
MUST-PET addresses a concrete gap in whole-body PET/CT lesion segmentation: existing self-supervised pretraining approaches for PET/CT have been confined either to single tracers (predominantly FDG) or to small datasets. The authors propose a multimodal, multi-tracer masked-reconstruction pretraining framework trained on 5,910 pan-cancer PET/CT scans spanning FDG and PSMA radiotracers, and demonstrate that the resulting representations transfer across tracers, institutions, and labeled-data budgets (2608.19666).
Pretraining objective and architecture
The framework uses a SwinUNETR backbone for both pretraining and fine-tuning. During pretraining, two-channel PET/CT patches of size 96×128×128 are extracted, with each modality partitioned into non-overlapping masking blocks of 12×16×16. For each training sample, one modality (PET or CT) is chosen with equal probability; 50% of its blocks are randomly masked via zero-mean imputation while the spatially aligned complementary modality remains fully visible, and the decoder reconstructs the masked voxels. This design forces the model to exploit cross-modal context — anatomical structure from CT and functional uptake from PET — rather than relying on within-modality redundancy alone. Training runs for 200 epochs with AdamW, cosine annealing (1×10−4 initial learning rate), weight decay 1×10−2, and a batch size of 3, using a weighted global reconstruction loss.
Data
Pretraining draws on four sources: AutoPET-III training/validation splits (FDG and PSMA), an internal Dartmouth Hitchcock Medical Center (DHMC) cohort of 1,599 unlabeled PSMA scans, SPADE (Stanford FDG), and VI-MED (Vietnamese FDG with non-standard CT intensity ranges and raw-count PET storage). Fine-tuning uses AutoPET-III labels only. All volumes are resampled to (2,2,3) mm voxel spacing, body-cropped, and instance-normalized, following the preprocessing of the FDG-only foundation model of Liu et al. (Liu et al., 20 May 2026). Evaluation uses three held-out test sets: AutoPET-III-test (N=321), Deep-PSMA (N=200, metastatic prostate cancer pre-LuPSMA therapy), and DHMC (N=100).
Cross-tracer reconstruction quality
The central evidence for multi-tracer generalization comes from reconstruction MAE on unseen data:
| Model |
FDG (AutoPET) |
PSMA (AutoPET) |
PSMA (DHMC) |
| FDG-only baseline |
0.2431 |
0.2982 |
0.2970 |
| MUST-PET |
0.2709 |
0.2908 |
0.2860 |
The authors state plainly that this comparison is asymmetric: the split used by Liu et al. is unknown, so some AutoPET FDG test scans may overlap that model's training data, biasing it toward FDG reconstruction. The PSMA and DHMC cases were not seen by either model, making those comparisons fair. On those fair comparisons, MUST-PET reduces MAE on both PSMA test sets. Notably, MUST-PET's FDG MAE (0.2709) is worse than the FDG-only baseline even under favorable conditions — a trade-off consistent with capacity being shared across tracers rather than specialized to one. This concession strengthens, rather than undermines, the paper's claim: the benefit of multi-tracer pretraining appears specifically in out-of-distribution tracer settings.
Label-efficient segmentation
Fine-tuning compares three strategies — full fine-tuning (Full FT), encoder+decoder fine-tuning, and decoder-only fine-tuning — against scratch-trained SwinUNETR and nnU-Net baselines, optimized with equally weighted Dice plus cross-entropy loss per the nnU-Net protocol of Rokuss et al. (Rokuss et al., 2024). Across 1%, 5%, 10%, 20%, 50%, and 100% labeled-data budgets, pretrained models improve lesion detection on both test sets, with the largest relative gains at low-label budgets. This is the practically salient result: annotation of whole-body PET/CT lesions is expensive, and the curves indicate that pretraining substitutes meaningfully for labeled data.
With 100% of labeled data, results are as follows (FPVol/FNVol in mL):
| Dataset |
Model / Strategy |
Dice |
Lesion Sensitivity |
| AutoPET III |
SwinUNETR scratch |
0.526 |
0.773 |
| AutoPET III |
nnUNet scratch |
0.566 |
0.773 |
| AutoPET III |
MUST-PET Full FT |
0.576 |
0.778 |
| AutoPET III |
MUST-PET Decoder FT |
0.565 |
0.794 |
| Deep-PSMA |
SwinUNETR scratch |
0.547 |
0.711 |
| Deep-PSMA |
nnUNet scratch |
0.561 |
0.757 |
| Deep-PSMA |
MUST-PET Decoder FT |
0.601 |
0.801 |
| Deep-PSMA |
MUST-PET Enc.+Dec. FT |
0.600 |
0.779 |
Two observations bear emphasis. First, the out-of-distribution gains are larger than in-distribution ones: Decoder FT improves Dice by +0.054 over scratch SwinUNETR on Deep-PSMA but only +0.039 on AutoPET III, supporting the claim that multi-tracer pretraining aids domain transfer. Second, no single fine-tuning strategy dominates: Full FT gives the best AutoPET Dice and lowest FP volume (14.112 mL) but the highest FN volume among MUST-PET variants there, while Decoder FT is best on Deep-PSMA across most metrics including FP volume (47.959 mL vs. ~73–74 mL from scratch). Qualitatively, MUST-PET reduces both false positives and false negatives relative to scratch models.
Limitations and open questions
Several caveats are acknowledged by the authors themselves. The comparison against the FDG-only foundation model is limited to reconstruction quality because of unknown training-split overlap; no direct lesion-segmentation head-to-head was possible. The FDG-only model retains an advantage on in-distribution FDG reconstruction, so the multi-tracer advantage should be read as a generalization property, not uniform superiority. The evaluation of downstream segmentation covers two test sets and primarily prostate-centric external data (Deep-PSMA); performance on other cancer types outside AutoPET-III's distribution is untested. Finally, whether the learned representations support broader downstream tasks beyond lesion segmentation — the criterion typically applied to "foundation model" claims — remains open, as the authors note explicitly.
Conclusion
MUST-PET demonstrates that multimodal, context-aware masked reconstruction over paired FDG/PSMA PET/CT yields representations that transfer across tracers, institutions, and label budgets. Its strongest quantitative result is on the unseen Deep-PSMA set (Dice 0.601 vs. 0.547 scratch), and its strongest qualitative claim — better cross-domain generalization than an FDG-only foundation model — is supported by reconstruction MAE on fair held-out PSMA comparisons while conceding worse FDG reconstruction. The paper leaves open direct benchmarking against prior foundation models under controlled splits and evaluation on additional downstream tasks.