---
title: 'PUUMA: Dual-Branch MRI Predictor'
url: https://www.emergentmind.com/topics/puuma
type: topic
---

# PUUMA: Dual-Branch MRI Predictor

Searching arXiv for the PUUMA paper and closely related work on U-Mamba, fetal MRI, and preterm birth prediction.
Searching arXiv for: "PUUMA U-Mamba placental patch whole-Uterus fetal MRI gestational age birth preterm risk".
PUUMA, the **Placental patch and whole-Uterus dual-branch U-Mamba-based Architecture**, is a deep-learning framework for predicting gestational age at birth and preterm risk from T2* functional MRI of the pregnant uterus. It was introduced as a fully automated MRI-based pipeline operating on 295 pregnancies and designed to combine global whole-uterus context with local placental detail in a single end-to-end trainable model [2509.07042]. In the reported evaluation, PUUMA and linear regression on cervical length achieved comparable mean absolute errors of about 3 weeks, and both obtained sensitivity of 0.67 for detecting preterm birth, despite pronounced class imbalance in the dataset [2509.07042].

## 1. Clinical task and prediction targets

PUUMA addresses two linked prediction problems. The first is **gestational age at birth** regression from T2* fetal MRI. The second is **preterm risk** classification, reported as binary term versus preterm discrimination, with accuracy, sensitivity, and specificity used for assessment [2509.07042].

The framework is positioned in a context where preterm birth is described as a major cause of mortality and lifelong morbidity in childhood, while current clinical predictors are limited by the complex and multifactorial origins of preterm delivery [2509.07042]. Within that setting, PUUMA uses functional MRI rather than only anatomical measurements. The reported study also benchmarked the model against linear regression using cervical length measurements obtained by experienced clinicians from anatomical MRI, as well as against other deep-learning architectures [2509.07042].

A central feature of the problem formulation is that prediction is not based solely on placental appearance or solely on whole-organ context. Instead, the model integrates both global whole-uterus information and local placental features. This design reflects the stated objective of leveraging organ-level context and high-resolution sub-volumes of the placenta simultaneously [2509.07042].

## 2. Cohort, inputs, and preprocessing pipeline

The study excluded scans acquired at or beyond 37 weeks at scan and scans missing gestational age at birth, yielding **295 cases** [2509.07042]. The training protocol describes the dataset as **295 subjects (15–40 weeks at scan; all <37 weeks at birth for preterm cases)**, split into **243 train / 26 validation / 26 test**, with stratification intended to maintain equal proportions of **extremely preterm (EPT)**, **very preterm (VPT)**, **late preterm (LPT)**, and **term (T)** cases in each set [2509.07042].

Preprocessing begins from **T2* relaxometry volumes**. These are **mono-exponentially fitted**, **clipped to 300 ms**, and **resampled to 128×128×64** [2509.07042]. A **placental mask via nnU-Net segmentation** is then used both as an auxiliary target and as a constraint for local sampling [2509.07042].

The two input streams are constructed differently. The **global branch** receives a down-sampled T2* volume of the entire uterus at **128×128×64 voxels** together with the automatically generated placenta mask. The **local branch** receives multiple randomly sampled placental patches of size **16×16×16 voxels**, each required to overlap the placenta mask by **at least 33% tissue** [2509.07042].

Data augmentation is applied to the whole-uterus volume **before patch extraction**. The reported augmentations are **affine/elastic transforms, zoom, contrast, simulated bias fields, and Gaussian noise** [2509.07042]. Class imbalance is addressed by oversampling under-represented preterm categories with sampling frequency inversely proportional to class prevalence [2509.07042].

## 3. Dual-branch U-Mamba architecture

PUUMA is a **dual-branch** architecture with a global whole-uterus pathway, a local placental-patch pathway, and a late fusion stage [2509.07042].

The **global branch** uses a **U-Mamba encoder-decoder** that mirrors a standard U-Net scaffold but replaces convolutional blocks with **Mamba blocks (state-space sequence modules)**. Its input is the whole-uterus T2* volume and placenta mask. It produces two outputs: a **placental segmentation map** through a sigmoid head on the decoder’s final features, and a **bottleneck feature vector** $h_G \in \mathbb{R}^{D_G}$ with $D_G = 5{,}120$, which is passed through a fully connected layer for initial gestational-age regression $\hat y_G$ and preterm classification logits $z_G \in \mathbb{R}^4$ for **EPT/VPT/LPT/T** categories [2509.07042].

The **local branch** uses only the **U-Mamba encoder portion**, with **depth = 3 levels**, on the high-resolution placental patches. It produces patch features $h_L \in \mathbb{R}^{D_L}$ with $D_L = 4{,}096$, together with direct gestational-age regression $\hat y_L$ and classification logits $z_L$ [2509.07042].

Fusion is performed by concatenating the branch-level outputs together with the known **gestational age at scan**:
$$
[\hat y_G, \hat y_L, z_G, z_L, y_{\text{scan}}].
$$
This joint feature vector is passed through a final fully connected layer to yield the ultimate gestational-age prediction $\hat y$ and **preterm risk probabilities** $p \in [0,1]^2$ for **term vs. preterm** [2509.07042].

At the block level, each encoder stage applies a **Mamba block** that combines a convolution, a state-space sequence layer, and a residual connection, followed by downsampling via strided convolution. In the decoder, upsampling is followed by a symmetric Mamba block and a skip connection that concatenates encoder features to preserve fine spatial detail. The paper states that the state-space layers inside Mamba blocks capture **long-range 3D context more efficiently than self-attention** [2509.07042].

## 4. Mathematical formulation and optimization

The model is formalized as a mapping
$$
f_\theta:\bigl(X,\{P_j\},y_{\text{scan}}\bigr)\mapsto\bigl(\hat y,\,p\bigr),
$$
where $X$ is the down-sampled uterus volume, $\{P_j\}_{j=1}^M$ is the set of placenta patches, $y$ is the true gestational age at birth, and $\tilde y \in \{0,1\}$ is the binary term/preterm label [2509.07042].

The reported regression objective is **mean squared error**:
$$
\mathcal{L}_{\mathrm{MSE}}=\frac{1}{N}\sum_{i=1}^N(\hat y_i-y_i)^2.
$$
For reporting, the study also uses **mean absolute error**:
$$
\mathcal{L}_{\mathrm{MAE}}=\frac{1}{N}\sum_{i=1}^N|\hat y_i-y_i|.
$$
Placental segmentation uses a combined **Dice + binary cross-entropy** objective,
$$
\mathcal{L}_{\mathrm{seg}}
=
\lambda_{\mathrm{Dice}}\left(1-\frac{2\sum \hat m\,m}{\sum \hat m+\sum m}\right)
+
\lambda_{\mathrm{BCE}}
\left[
-\frac{1}{V}\sum_v
\bigl(m_v\log \hat m_v+(1-m_v)\log(1-\hat m_v)\bigr)
\right],
$$
and preterm classification uses **categorical cross-entropy**:
$$
\mathcal{L}_{\mathrm{CE}}
=
-\frac{1}{N}\sum_{i=1}^N\sum_{c\in\{\mathrm{T},\mathrm{PT}\}} y_{i,c}\log p_{i,c}.
$$
The total loss is
$$
\mathcal{L}_{\mathrm{total}}
=
\alpha\,\mathcal{L}_{\mathrm{MSE}}
+
\beta\,\mathcal{L}_{\mathrm{seg}}
+
\gamma\,\mathcal{L}_{\mathrm{CE}}
+
\delta\,\|\theta\|_2^2,
$$
with $\alpha,\beta,\gamma,\delta$ balancing regression, segmentation, classification, and weight decay [2509.07042].

Training used **Adam** with **initial learning rate $1\times10^{-4}$**, decayed on plateau. The **batch size** was **1 subject**, defined as one full-uterus volume together with its patches. Training continued until validation loss plateaued, at approximately **50 epochs**, with checkpoints saved every **50 batches** and then every **5** during fine-tuning [2509.07042].

## 5. Empirical results and benchmark comparisons

Evaluation on the test set was reported as **mean ± SD over 26 subjects** [2509.07042]. The main benchmark comparisons are summarized below.

| Model | MAE (weeks) | Accuracy / Sensitivity / Specificity |
|---|---:|---:|
| PUUMA | 3.05 ± 3.10 | 0.65 / 0.67 / 0.65 |
| Whole-uterus U-Mamba | 2.95 ± 3.65 | 0.73 / 0.33 / 0.85 |
| Linear regression on cervical length | 2.94 ± 2.59 | 0.77 / 0.67 / 0.80 |
| U-Net (global only) | 3.98 ± 3.60 | 0.65 / 0.50 / 0.70 |

Several comparative observations are explicit in the reported results [2509.07042]. **PUUMA achieves comparable gestational-age MAE to cervical-length regression, at about 3 weeks**, and its **sensitivity of 0.67** is substantially higher than that of the **whole-uterus U-Mamba alone**, which reached **0.33**. The **U-Net baseline** underperforms both PUUMA and the cervical-length model. The paper also notes that **no formal statistical tests were reported** [2509.07042].

The results therefore do not indicate uniform dominance of PUUMA across all metrics. The **linear cervical-length model** achieved the best reported **MAE** and **accuracy**, while the **whole-uterus U-Mamba** achieved the highest **specificity** but much lower sensitivity. PUUMA’s empirical profile is instead characterized by a more balanced trade-off between sensitivity and specificity within the reported comparison set [2509.07042].

## 6. Clinical interpretation, misconceptions, and future directions

The study presents PUUMA as a **proof of concept** for automated prediction of gestational age at birth directly from functional MRI, and as evidence for the value of **whole-uterus functional imaging** in identifying pregnancies at risk of preterm birth [2509.07042]. It also emphasizes that **manual, high-definition cervical length measurements derived from MRI**, although **not currently routine in clinical practice**, provide valuable predictive information [2509.07042].

A common misconception would be to treat PUUMA as having clearly surpassed simpler clinical baselines. The reported numbers do not support that interpretation. **Linear regression on cervical length** achieved **2.94 ± 2.59 weeks MAE**, **0.77 accuracy**, **0.67 sensitivity**, and **0.80 specificity**, which is competitive with or better than PUUMA on several metrics [2509.07042]. Another misconception would be to regard the whole-uterus branch alone as sufficient; in the reported evaluation it achieved higher accuracy and specificity than PUUMA, but its sensitivity fell to **0.33**, whereas PUUMA reached **0.67** [2509.07042]. This suggests that the dual-branch design is particularly relevant when sensitivity to preterm cases is a priority.

The principal limitations stated in the work concern **cohort size**, **class imbalance**, and **generalisability** [2509.07042]. Future work is described as focusing on **expanding the cohort size**, **incorporating additional organ-specific imaging**, **refining regularisation and loss weighting**, and **exploring uncertainty quantification for clinical deployment** [2509.07042]. The paper gives examples of possible multimodal extensions, including **fetal lung or brain volumes** and **placental diffusion/relaxometry** [2509.07042].

In that form, PUUMA occupies a specific methodological position: it is not presented as a routine clinical tool, but as a technically defined dual-branch U-Mamba system showing that automated MRI-based prenatal risk stratification can be performed with performance comparable to a strong MRI-derived cervical-length baseline, while exploiting both placental microstructure and whole-uterus context in a unified model [2509.07042].

Source: https://www.emergentmind.com/topics/puuma