XBusNet: Dual-Prompt BUS Segmentation
- XBusNet is a text-guided breast ultrasound segmentation model that integrates clinical text cues with imaging features for improved lesion localization.
- It employs a dual-prompt, dual-branch design, combining a global CLIP-based transformer and a local U-Net–style pathway to enhance boundary detection.
- Evaluation on the BLU dataset shows high performance with a mean Dice of 0.8765 and IoU of 0.8149, notably improving segmentation of small, low-contrast lesions.
XBusNet is a text-guided breast ultrasound segmentation model that addresses the difficulty of segmenting small or low-contrast lesions with fuzzy margins and speckle noise by combining image features with clinically grounded text in a dual-prompt, dual-branch multimodal design. It merges a global, language-conditioned view of the entire ultrasound frame with a local, boundary-sensitive U-Net–style path, and fuses the two through lightweight, prompt-driven feature modulation. In reported evaluation on the Breast Lesions USG (BLU) dataset, XBusNet achieves a mean Dice of 0.8765 and IoU of 0.8149 under five-fold cross-validation, with the largest gains on the smallest lesions (Mallina et al., 8 Sep 2025).
1. Clinical problem setting and design rationale
Precise breast ultrasound (BUS) segmentation supports reliable measurement, quantitative analysis, and downstream classification, yet remains difficult for small or low-contrast lesions with fuzzy margins and speckle noise. The central motivation for XBusNet is that text prompts can add clinical context, but directly applying weakly localized text-image cues, including CAM- or CLIP-derived signals, tends to produce coarse, blob-like responses that smear boundaries unless additional mechanisms recover fine edges.
XBusNet is positioned against that limitation through a dual-prompt, dual-branch multimodal model. One branch captures whole-image semantics conditioned on lesion size and location, while the other emphasizes precise boundaries and is modulated by prompts that describe shape, margin, and Breast Imaging Reporting and Data System (BI-RADS) terms. Prompts are assembled automatically from structured metadata, requiring no manual clicks.
This architecture brings together two complementary streams of information: a global, language-conditioned view of the entire ultrasound frame and a local, boundary-sensitive U-Net–style path. A plausible implication is that the model is designed to separate scene-level semantic localization from contour-level delineation rather than forcing a single representation to serve both roles.
2. Global pathway and scene-level semantic conditioning
The global branch is a Global Feature Extractor (GFE) built on a frozen CLIP Vision Transformer. It processes the input breast ultrasound image in patch-token form and uses a small prompt-conditioning head to inject lesion size and coarse location cues into the transformer's token stream. The global prompt , exemplified by text such as “small lesion in upper outer quadrant,” is encoded by CLIP’s text tower into .
At a chosen transformer depth, patch tokens are affinely adjusted as
after which the modulated representation is reshaped and up-projected to form the global feature map .
The backbone is frozen so that global semantics remain stable, while lightweight projection layers learn how size and quadrant cues steer high-level attention toward plausible lesion regions. In the formulation provided for XBusNet, the global pathway therefore contributes “where” and “how large” information at the image level rather than attempting boundary recovery directly.
3. Local pathway, transformer augmentation, and boundary sensitivity
The local branch is a Local Feature Extractor (LFE) based on a ResNet50-backboned U-Net with deep transformer blocks. It begins with a ResNet50 encoder augmented by multi-head self-attention blocks at the deepest levels and is followed by a U-Net decoder that upsamples through skip connections. This branch yields a local feature map .
Its prompt is designed to encode lesion attributes rather than scene layout. The local prompt can take forms such as “irregular shape, microlobulated margin, BI-RADS 4,” and it informs shape and margin priors. At four key stages—mid-encoder and three decoder up-blocks—the Semantic Feature Adjustment (SFA) module applies feature-wise linear modulation in the same form as the global branch:
The stated role of this mechanism is to preserve and sharpen boundary details while biasing the network toward edges and shapes consistent with the BI-RADS descriptors. This suggests that the local path is the principal source of contour continuity and fine-grained lesion morphology, especially where purely semantic activation maps would otherwise remain spatially coarse.
4. Prompt construction, FiLM-based modulation, and fusion
All prompts are assembled automatically from the BLU dataset’s structured metadata, requiring no manual clicks or report parsing. The global feature context prompt encodes lesion size discretized into small, medium, and large via training-set quantiles, together with breast quadrant, computed as the centroid of the true mask during training or from a coarse first-pass mask at test time. The local feature prompt encodes shape, margin, and BI-RADS category: shape is round, oval, or irregular; margin is circumscribed, microlobulated, spiculated, or indistinct; BI-RADS category is 3, 4, or 5.
Examples given for assembled prompts are “medium lesion in lower inner quadrant” for the global branch and “irregular shape, indistinct margin, BI-RADS 5” for the local branch. These prompts are tokenized and embedded by the frozen CLIP text encoder to 0 and 1, which then drive the SFA modules in each branch.
Text-image fusion is implemented through channel-wise Feature-Wise Linear Modulation (FiLM). For a feature tensor 2 and a prompt embedding 3, SFA predicts modulation parameters through a small MLP 4:
5
After modulating both branches, the model concatenates them and refines the joint representation:
6
Here 7 is processed by two residual blocks, denoted 8, and then projected with a 9 convolution to per-pixel logits 0. The resulting formulation makes the multimodal interaction explicitly channel-wise and branch-specific rather than relying on a single shared conditioning pathway.
The same source also specifies standard overlap metrics for evaluation:
1
For training objectives, the paper notes that segmentation networks of this style typically optimize a weighted sum of binary cross-entropy, Dice loss, and IoU loss,
2
so that both region overlap and boundary fidelity are learned. Because this statement is framed as a typical formulation, it should be read as a stylistic training description rather than as a uniquely identifying property of XBusNet.
5. Dataset, protocol, and empirical performance
XBusNet was validated on the Breast Lesions USG (BLU) dataset, which comprises 252 single-lesion ultrasound images, including 154 benign and 98 malignant cases, with expert-verified pixel masks and structured BI-RADS metadata (Mallina et al., 8 Sep 2025). Images were resized to 3, and models were trained under five-fold cross-validation with 80/20 splits, 1,000 iterations per fold, batch size 4, AdamW at 4, and a cosine learning-rate schedule, using mixed precision on NVIDIA A100 GPUs. Segmentation probabilities were thresholded at 5 to produce binary masks.
Across the five folds, the reported performance is mean Dice 6, mean IoU 7, FPR 8, and FNR 9. The model is described as outperforming six strong baselines on BLU.
The size-stratified analysis is particularly central to the model’s claims. For the smallest lesions, defined as 0–1 px, Dice rose from approximately 2 for the best baseline to 3, and IoU rose from 4 to 5. The accompanying interpretation is that text-steered global context plus local attribute cues helps recover tiny, low-contrast masses that typically vanish under purely pixel-driven pipelines.
The reported qualitative significance is correspondingly specific: small lesions show the largest gains, with fewer missed regions and fewer spurious activations. In this framing, the empirical contribution is not merely higher overlap metrics, but improved robustness under the lesion regimes where BUS segmentation is most error-prone.
6. Component ablations, limitations, and future directions
Three ablations were reported to isolate the roles of the global branch, the local branch, and prompt-driven modulation.
| Variant | Dice | IoU |
|---|---|---|
| Full model | 0.8765 | 0.8149 |
| –LFE (no local branch) | 0.8572 | 0.7865 |
| –GFE (no global branch) | 0.8453 | 0.7772 |
| –SFA (no prompt-driven modulation) | 0.8600 | 0.8068 |
The full model outperformed each partial variant, supporting the claim that global semantics, fine-grain boundary modeling, and explicit prompt modulation contribute in complementary fashion. The relative reductions also indicate that removing either branch degrades performance more than removing prompt-driven modulation alone, although all three components remain beneficial in the reported configuration.
The limitations are explicitly stated. XBusNet relies on a single curated dataset, BLU, and lacks cross-center validation, leaving open questions about generalization to different scanners or operators (Mallina et al., 8 Sep 2025). The relatively elevated FPR compared to some under-segmenting baselines also suggests that operating thresholds may need careful clinical calibration.
Future work is described along several directions: incorporating radiology-report-derived prompts or knowledge graphs to enrich the prompt vocabulary; adding detection or instance-segmentation heads to handle multiple lesions per image; exploring boundary-aware losses such as Hausdorff or boundary losses to further sharpen edges; conducting a multi-center study to confirm robustness across institutions; and tightly integrating with structured reporting workflows, where the same prompts that guide segmentation can feed downstream classification or risk assessment.
Taken together, these points define XBusNet less as a generic vision-language segmentation model than as a specific BUS framework in which automatically assembled clinical text cues are used to combine vision-language pretraining with U-Net–style precision in a unified, end-to-end trainable framework.