SAM-DCE: Enhancing Medical Segmentation
- SAM-DCE is a prompt-free medical segmentation framework that augments SAM’s encoder–decoder pipeline with ML-DCE to mitigate semantic over-smoothing and token uniformity.
- It integrates mask-conditioned (MCC) and image-conditioned (ICC) modules to restore class-specific details and enhance inter-class separability.
- Empirical results on the Synapse CT benchmark demonstrate improved Mean Dice and reduced HD95, underlining its potential in clinical segmentation tasks.
Searching arXiv for the cited SAM-DCE papers and related records. SAM-DCE denotes a prompt-free medical segmentation framework that adapts the Segment Anything Model (SAM) by targeting feature-level failures inside SAM’s decoding process, especially semantic over-smoothing and token uniformity (Hu et al., 21 Sep 2025). In its primary and formal usage, introduced in "SAM-DCE: Addressing Token Uniformity and Semantic Over-Smoothing in Medical Segmentation," the method augments SAM’s encoder–decoder pipeline with a Multi-Level Decoupled Class Embedding (ML-DCE) module that balances local discrimination and global semantics, enhances inter-class separability, and enriches mask decoding with fine-grained, consistent representations (Hu et al., 21 Sep 2025). In adjacent contemporaneous usage, the label also appears as shorthand for SAM-based decoder-enhanced segmentation in 3D medical imaging and as an encoder-centric robustness perspective for downstream SAM systems, situating the term within a broader family of decoder enhancement, prompt-free adaptation, and foundation-model auditing efforts (Chen et al., 24 Nov 2025, Qin et al., 8 Aug 2025).
1. Problem formulation and motivating failure modes
SAM’s transfer from natural images to medical imaging is constrained by three coupled issues: domain shift, anatomical variability, and reliance on prompts. Medical images differ strongly from natural images in texture statistics, contrast, noise characteristics, and object appearance; organs and lesions are often low-contrast, elongated or irregular, and embedded in complex backgrounds. Under these conditions, SAM’s image encoder and mask decoder, tuned on natural scenes, do not explicitly model organ-level semantics or spatial priors, and prompt-dependent operation requires expert intervention that is often impractical in clinical settings (Hu et al., 21 Sep 2025).
SAM-DCE isolates a more specific internal failure mode than prompt dependence alone. Its central claim is that SAM’s transformer-based decoding stack exhibits semantic over-smoothing: through repeated self- and cross-attention, deep mask tokens become progressively homogeneous, token variance is reduced layer by layer, and the decoder emphasizes highly discriminative regions while underrepresenting subtle boundaries and low-frequency anatomical classes. The associated token uniformity means that class-level tokens intended to represent different organs acquire high pairwise similarity, weakening inter-class separability and degrading multi-organ segmentation, especially for small structures and ambiguous borders (Hu et al., 21 Sep 2025).
This framing distinguishes SAM-DCE from prompt-free adaptations that remove user interaction but adapt SAM only superficially. A common misconception is that eliminating prompts is sufficient for robust medical transfer. SAM-DCE instead argues that the internal semantics of decoder tokens remain a separate and limiting bottleneck even in prompt-free settings, and that the primary intervention should occur at the level of class-aware token formation inside the decoder (Hu et al., 21 Sep 2025).
2. Core architecture: SAM plus Multi-Level Decoupled Class Embedding
SAM-DCE is built on SAM’s canonical three-part structure: a ViT-based image encoder producing patch embeddings
a prompt encoder, and a mask decoder. In SAM-DCE, the system operates prompt-free in practice through hyper-prompting and LoRA adapters, while the principal architectural change is the insertion of ML-DCE into the decoding pipeline (Hu et al., 21 Sep 2025).
ML-DCE is a dual-path module operating on two representational sources. The first path consumes deep mask tokens from the decoder; the second path consumes intermediate encoder embeddings from SAM’s image encoder. These paths are designed to address different deficits. The decoder-side path targets class-specific semantics and boundary information that have been homogenized by deep attention. The encoder-side path injects global spatial and structural priors, including organ location, shape, and contextual relationships, that remain more diverse in encoder embeddings than in late decoder tokens (Hu et al., 21 Sep 2025).
The decoder still outputs 9 class logits on the Synapse benchmark, corresponding to 8 organs + background, but the internal representation is no longer based solely on vanilla deep tokens. Instead, ML-DCE produces a new enriched class-token sequence , formed by residual fusion of original foreground mask tokens with decoder-conditioned and encoder-conditioned class embeddings. This design is notable because SAM-DCE does not rely on an explicit decorrelation loss or a contrastive objective; token diversification is enforced architecturally rather than by an auxiliary regularizer (Hu et al., 21 Sep 2025).
3. MCC, ICC, and residual class-semantic enhancement
The decoder-side branch of ML-DCE is the Mask-Conditioned Class Token (MCC) module. Its input is the decoder’s deep mask-token sequence
where is the background token and are foreground class tokens. Foreground class queries are initialized as
These class queries attend over mask-token keys and values through cross-attention,
followed by a two-layer MLP:
The role of MCC is to extract class-aligned semantics from homogenized deep mask tokens and to restore local discrimination and boundary-aware semantics (Hu et al., 21 Sep 2025).
The encoder-side branch is the Image-Conditioned Class Token (ICC) module. It aggregates class-aware information from encoder embeddings
again using class queries, multi-head cross-attention, and a two-layer MLP:
ICC contributes global context, spatial priors, and structurally grounded semantics that complement the purely decoder-derived representation (Hu et al., 21 Sep 2025).
Foreground enhancement is then performed by separating the background token from the foreground subsequence 0 and applying residual fusion with learnable scalars 1 and 2:
3
This preserves explicit background handling while enriching foreground tokens with both deep class semantics and encoder-grounded context. Empirically, the full MCC+ICC configuration reduces mean pairwise cosine similarity between class tokens to 0.476, which the reported analysis interprets as improved class discriminability and reduced token uniformity (Hu et al., 21 Sep 2025).
4. Optimization, benchmark setting, and quantitative performance
SAM-DCE is trained with a multi-resolution segmentation objective,
4
where 5 is pixel-wise cross-entropy and 6 is Dice loss. The decoder, including ML-DCE and the mask prediction head, is trained from scratch to output 9-class logits, while SAM’s encoder is adapted via LoRA with randomly initialized adapter weights. A Hyper-Prompting Adapter is used for prompt conditioning, although no user input is required in practice. Training is reported for up to 300 epochs on an NVIDIA RTX 4090 with AdamW, batch size 12, and initial learning rate 0.0005 (Hu et al., 21 Sep 2025).
The primary benchmark is MICCAI 2015 Synapse multi-organ CT, comprising 3,779 contrast-enhanced abdominal CT slices, with 18 training and 12 testing cases, 2,212 slices used for training, inputs resized to 7, and 8 target organs: Aorta, Gallbladder, Spleen, Left Kidney, Right Kidney, Liver, Pancreas, Stomach (Hu et al., 21 Sep 2025).
On Synapse, SAM-DCE is compared with fully supervised medical segmentation models and earlier SAM adaptations. The main reported metrics are Mean Dice [%] and HD95 [mm].
| Method | Mean Dice [%] | HD95 [mm] |
|---|---|---|
| AutoSAM | 62.08 | 27.56 |
| SAM Adapter | 72.80 | 33.08 |
| SAMed | 78.37 | 27.81 |
| TransDeepLab | 80.16 | 21.25 |
| SAM-DCE | 80.54 | 13.66 |
These results place SAM-DCE above SAMed by +2.17% in Mean Dice and markedly below it in HD95, indicating stronger boundary delineation. Reported per-organ Dice values for SAM-DCE are Spleen 83.45%, Kidney (R) 87.10%, Kidney (L) 91.82%, Gallbladder 58.31%, Liver 94.57%, Stomach 78.56%, Aorta 85.21%, and Pancreas 65.30% (Hu et al., 21 Sep 2025).
Ablation further separates the effects of MCC and ICC:
| Config | MCC | ICC | Mean Dice [%] |
|---|---|---|---|
| I | 63.65 | ||
| II | ✓ | 68.61 | |
| III | ✓ | 70.24 | |
| IV | ✓ | ✓ | 80.54 |
The baseline decoder without MCC or ICC yields 63.65% Mean Dice. Adding MCC alone raises performance to 68.61%, adding ICC alone to 70.24%, and combining both reaches 80.54%. The reported gain of the full model exceeds the sum of the individual gains, which is presented as evidence of strong synergy between local class discrimination and global structural context (Hu et al., 21 Sep 2025).
5. Relation to adjacent uses of the term
Within the supplied literature, “SAM-DCE” is not used exclusively for the ML-DCE framework. In DEAP-3DSAM, the label is explicitly associated with SAM-based decoder-enhanced segmentation for 3D medical imaging, where the emphasis shifts from class-token decoupling to a Feature Enhanced Decoder (FED) and a Dual Attention Prompter (DAP) (Chen et al., 24 Nov 2025). In that setting, the decoder is a 3D CNN-based mask decoder that fuses multi-scale SAM encoder features with original 3D image features, while the prompter replaces manual points or boxes with spatial and channel attention. The reported four-dataset evaluation spans KiTS21, Pancreas, LiTS17, and Colon, and the method is described as matching or surpassing manual-prompt SAM adaptations on several of those datasets without manual prompts (Chen et al., 24 Nov 2025).
This neighboring usage is conceptually related but technically distinct. The formal SAM-DCE of (Hu et al., 21 Sep 2025) addresses token uniformity and semantic over-smoothing inside the decoder through class-aware token construction. DEAP-3DSAM, by contrast, addresses spatial feature loss induced by pseudo-3D adaptation and the impracticality of manual prompting in volumetric segmentation. The shared denominator is the claim that vanilla SAM decoding is inadequate for medical transfer and that decoder-side enhancement is necessary, but the mechanisms differ substantially: MCC/ICC residual fusion in one case, 3D feature fusion and dual-attention auto-prompting in the other (Hu et al., 21 Sep 2025, Chen et al., 24 Nov 2025).
A further adjacent usage appears in the adversarial-robustness work VeSCA, where “SAM-DCE” is used as a broader encoder-centered evaluation and defense lens rather than as a specific medical segmentation architecture (Qin et al., 8 Aug 2025). There, the emphasis is on the SAM image encoder as a shared attack surface for downstream systems, and the proposed Vertex-Refining Simplicial Complex Attack (VeSCA) reports up to 12.71% improvement over state-of-the-art attack baselines across downstream tasks and datasets. This usage broadens the scope of SAM-DCE from decoder enhancement to ecosystem-level analysis of encoder vulnerabilities and downstream consequences (Qin et al., 8 Aug 2025).
6. Limitations, misconceptions, and future directions
SAM-DCE adds architectural complexity relative to vanilla SAM. ML-DCE introduces extra cross-attention blocks and MLPs, increasing computational cost during decoding and memory usage for class tokens and attention maps. The method also depends on a known set of foreground categories because class queries are tied to predefined semantic classes; scaling to many more classes or open-set segmentation would likely require dynamic query generation or a more flexible class-embedding mechanism. In addition, the detailed empirical setup is centered on Synapse CT multi-organ segmentation, so broader claims across MRI, histopathology, or other modalities remain more suggestive than fully established within the reported evidence (Hu et al., 21 Sep 2025).
Several misconceptions are explicitly contradicted by the design. SAM-DCE does not completely redesign SAM’s decoder; it injects ML-DCE into the token pipeline and leaves the overall encoder–decoder paradigm intact. It is also not a method based on explicit token decorrelation or contrastive regularization. The reported class separability improvements arise from MCC + ICC + residual fusion under standard segmentation supervision, not from an auxiliary loss such as 8 (Hu et al., 21 Sep 2025).
The future directions named around SAM-DCE are structurally coherent. One line is to extend decoupled class embeddings beyond medical imaging, including natural-image multi-object segmentation, remote sensing, and industrial inspection. Another is to introduce explicit token decorrelation or contrastive objectives on top of the present architecture. Additional possibilities include dynamic or hierarchical class queries, multi-modal extension such as CT+MRI or CT+PET fusion, and integration with more volumetric backbones, including 3D ViTs, as related SAM-DCE-style work in 3D medical imaging has suggested (Hu et al., 21 Sep 2025, Chen et al., 24 Nov 2025).
Taken in its primary sense, SAM-DCE is best understood as a class-aware, prompt-free extension of SAM that attacks decoder token homogenization directly. Its central contribution is not merely better adaptation to medical imagery, but a more specific claim: that medical segmentation performance can be improved by reconstituting class semantics inside the decoder through separate mask-conditioned and image-conditioned pathways, thereby reducing token uniformity and improving boundary fidelity and inter-class separation (Hu et al., 21 Sep 2025).