Med-CTX: Multimodal Ultrasound Segmentation
- Med-CTX is a fully transformer-based multimodal framework that fuses ultrasound images and structured BI-RADS text for breast cancer segmentation.
- It employs a dual-branch visual encoder with global and local feature fusion, enhanced by uncertainty-aware cross-modal attention for robust calibration.
- Empirical results on BUS-BRA show state-of-the-art Dice, IoU, and explanation metrics, demonstrating its clinical and research significance.
Med-CTX is a fully transformer-based multimodal framework designed for explainable breast cancer ultrasound segmentation, integrating both medical images and clinical radiology report text for optimal accuracy, calibration, and interpretability. Med-CTX leverages clinical language structured via BI-RADS semantics, a dual-branch visual encoder (Vision Transformer and Swin Transformer), and a cross-modal attention fusion mechanism with explicit uncertainty gating. The system jointly generates precise segmentation masks, calibrated uncertainty maps, and model-generated natural language rationales, facilitating trustworthy and transparent computer-assisted diagnosis in breast ultrasound imaging (Adahada et al., 19 Aug 2025).
1. System Overview and Architectural Components
Med-CTX comprises three primary processing stages, each engineered to extract, align, and decode both visual and textual medical information:
- Dual-Branch Visual Encoder:
- The ViT branch divides the input ultrasound image into patches, projects them to embedding tokens, and passes them through layers of global self-attention.
- The Swin branch processes the same image using a hierarchical transformer with shifted windows, enabling high-resolution local context modeling.
- At each level , global () and local () features are adaptively combined:
Text Encoder:
- BI-RADS metadata (category, laterality, histology) and free-text notes are concatenated, tokenized, and encoded by BioClinicalBERT, yielding a text embedding tensor where , .
- Uncertainty-Aware Cross-Modal Fusion:
- Visual features and projected text features are integrated by cross-modal attention modulated by an uncertainty gate 0, effectively weighting contributions according to their estimated confidence.
- Fused representations drive multi-task predictions.
2. Mathematical Formulation
Several core algorithmic components underpin the system's multimodal integration and explainability:
- Cross-Modal Attention: Given visual 1 and textual 2 embeddings:
3
4
5
- Uncertainty Estimation: Pixelwise uncertainty is estimated via:
6
The estimated uncertainty informs cross-modal fusion through 7.
- Loss Functions: The optimization objective combines multiple terms:
8
Main loss components include binary cross-entropy, Dice similarity, uncertainty regularization, CLIP-style contrastive loss, clinical classification, and confidence alignment.
- Calibration and Alignment: Model calibration is quantitated via Expected Calibration Error (ECE):
9
Multimodal (image-text) representation alignment uses CLIP score:
0
3. Output Modalities and Interpretability
Med-CTX generates multiple output modalities per forward pass:
- Segmentation Logits: Per-pixel lesion prediction.
- Uncertainty Maps: Calibrated probability maps indicating model confidence for each spatial location.
- Structured Clinical Prediction: BI-RADS class and pathology via global pooling and classification.
- Textual Rationale: Free-text explanation generated by a GRU-based language decoder, conditioned on fused multimodal features, delivering model-internal diagnostic justification.
This design enables simultaneous interpretability (via rationale generation and uncertainty quantification) and predictive accuracy in the clinical workflow.
4. Training Protocol and Evaluation Methodology
The training scheme comprises three successive phases:
- Contrastive Pretraining: NT-Xent loss on unlabeled images—10 epochs, learning rate 1, AdamW optimizer, batch size 32.
- Modality Alignment: CLIP-style alignment with vision encoder frozen and text encoder unfrozen—10 epochs, learning rate 2.
- Supervised Fine-tuning: Full end-to-end training—150 epochs, vision LR 3, text LR 4, cosine annealing, gradient accumulation (effective batch size 16), and early stopping (patience 15).
Preprocessing includes resizing images and masks to 5, normalizing intensities, tokenizing and padding/truncating text to 128 tokens, and extensive data augmentation. All experiments are conducted on the BUS-BRA breast ultrasound dataset using four A100 GPUs over 18 hours.
5. Empirical Results and Ablation Analysis
Med-CTX attains leading metrics on BUS-BRA validation:
- Dice: 0.9879
- IoU: 0.9518
- Pixel Accuracy: 0.9842
- CIDEr (Text Explanation): 0.58
- BLEU-4: 0.42
- METEOR: 0.39
- BI-RADS Accuracy: 0.84
- CLIP Score: 0.854
- ECE: 3.2% (post temperature scaling, 6)
Ablation studies highlight the significance of each system component:
| Configuration | Dice | CIDEr | ECE | CLIP |
|---|---|---|---|---|
| Full Med-CTX | 0.9879 | 0.58 | 3.2% | 0.854 |
| w/o Swin branch | 0.9649 | 0.57 | 4.1% | 0.847 |
| w/o BI-RADS structuring | 0.9712 | 0.37 | 5.8% | 0.832 |
| w/o uncertainty fusion | 0.9699 | 0.56 | 15.8% | 0.841 |
| w/o CLIP contrastive | 0.9699 | 0.55 | 6.9% | 0.784 |
| w/o any clinical text | 0.9339 | 0.27 | 18.3% | 0.721 |
Removal of clinical text representation yields a -5.4% decline in Dice and -31% in CIDEr, substantiating the critical role of textual input for both segmentation and explainability. Uncertainty-aware fusion and multimodal alignment (CLIP loss) are pivotal for calibration (ECE) and alignment.
6. Clinical and Research Significance
Med-CTX advances the paradigm of multimodal, explainable medical image segmentation by establishing:
- State-of-the-art lesion delineation, outperforming U-Net, ViT, and Swin baselines.
- Clinically grounded explanations by aligning natural language output with both image evidence and structured text.
- Model confidence calibration suitable for clinical deployment (ECE as low as 3.2%).
- Demonstrated critical impact of clinical language incorporation for both machine interpretability and performance.
Implementation reveals that the dual-branch vision backbone, BI-RADS-aware language encoding, and uncertainty gating are architecturally and statistically essential, setting a reference performance benchmark for reliable, multimodal computer-assisted diagnosis in breast cancer imaging (Adahada et al., 19 Aug 2025).