Papers
Topics
Authors
Recent
Search
2000 character limit reached

GO-MLVTON: Diffusion-Based Multi-Layer VTON

Updated 27 January 2026
  • GO-MLVTON is a multi-layer virtual try-on approach that explicitly learns pixel-level occlusion relationships between inner and outer garments.
  • It integrates a Garment Occlusion Learning module with a StableDiffusion-based morphing module to achieve realistic deformation and artifact-free layering.
  • Evaluated on the MLG dataset using the LACD metric, GO-MLVTON outperforms previous methods in FID, SSIM, and overall visual coherence.

GO-MLVTON (Garment Occlusion-Aware Multi-Layer Virtual Try-On with Diffusion Models) is a methodology for multi-layer garment virtual try-on (ML-VTON) that jointly addresses the challenges of garment occlusion modeling, spatial deformation, and layer-to-layer visual coherence. By explicitly learning pixel-level occlusion relationships between inner and outer garments, GO-MLVTON produces artifact-free, realistic multi-layer try-on results, addressing limitations in prior single-layer or multi-garment VTON approaches. The system integrates a dedicated Garment Occlusion Learning module, a StableDiffusion-based morphing and fitting module, and is benchmarked on the newly introduced MLG dataset with a layer- and edge-sensitive evaluation metric, LACD (Yu et al., 20 Jan 2026).

1. Multi-Layer Virtual Try-On and the Challenge of Occlusion

Image-based virtual try-on (VTON) research has traditionally focused on either single-garment try-on (SG-VTON), synthesizing one target garment onto a person, or multi-garment (MG-VTON), compositing non-overlapping garments without enforcing physical overlap or occlusion. Real-world usage, however, typically involves dressing multiple, sometimes overlapping, layers (e.g., T-shirt under a jacket), where physical realism demands:

  • Accurate spatial occlusion—outer garments must mask overlapping inner garment regions.
  • Realistic deformation—proper draping and fitting of both inner and outer layers to the target body and pose.
  • Artifact-free compositing—elimination of “bleed-through” or ghosting from occluded inner garment pixels.

Previous methods lack explicit pixel-level occlusion reasoning or the ability to deform and layer garments in a physically plausible manner. GO-MLVTON directly addresses these by introducing two interlinked components: the Garment Occlusion Learning (GOL) module and the Garment Morphing & Fitting (GMF) module built atop a diffusion model.

2. Garment Occlusion Learning (GOL) Module

The GOL module is designed to infer a spatially-varying attention mask that determines which pixels of an inner garment remain visible after layering with an outer garment. The process involves:

  • Dual-encoder feature extraction: Separate encoders EiE_i and EoE_o, each with five layers (composed of downsample convolutions followed by residual blocks), process the inner (gig_i) and outer (gog_o) garment images to yield feature maps fif_i and fof_o.
  • Mapping to occlusion attention: Channel-concatenated features [fo∥fi][f_o\Vert f_i] are passed through a mapping network UU and a final 1×11\times 1 convolution, followed by a sigmoid nonlinearity, producing the occlusion mask A∈[0,1]1×H′×W′A \in [0,1]^{1 \times H' \times W'}.
  • Occlusion-weighted inner garment latent: This attention mask is applied element-wise to the latent representation of the inner garment (EoE_o0 from the shared VAE encoder), yielding EoE_o1.

Loss function: A reconstruction loss enforces that EoE_o2 matches the visible inner garment as cropped from the real person image EoE_o3 via

EoE_o4

with EoE_o5 in the overall objective. This direct supervision ensures the GOL module learns to suppress occluded regions effectively. Backpropagating EoE_o6 leads to optimal occlusion-aware masking, especially at garment boundaries.

3. StableDiffusion-Based Garment Morphing & Fitting (GMF) Module

Following occlusion reasoning, the GMF module synthesizes the final try-on output, adapting diffusion modeling to multi-layer conditioning:

  • Latent conditioning: The core conditional input is the concatenation EoE_o7 where EoE_o8 is the clothing-agnostic person latent (body pose/shape), EoE_o9 is the outer garment latent, and gig_i0 is the occlusion-weighted inner garment latent.
  • Diffusion UNet: A StableDiffusion v1.5 UNet, with cross-attention layers removed for architectural alignment with ML-VTON tasks, jointly models the diffusion denoising process conditioned on these latents as well as a binary inpainting mask for masked person areas.
  • Training objective: The denoising loss per diffusion step gig_i1 is

gig_i2

and the combined objective is

gig_i3

  • Optimization and inference: Model initialization starts from InstructPix2Pix-pretrained StableDiffusion weights; most weights remain frozen except for GOL and self-attention layers. AdamW is used with learning rate gig_i4, batch size gig_i5, and gig_i6 conditional dropout. Inference incorporates classifier-free guidance (CFG) with scale gig_i7 for optimal fidelity/diversity balance.

4. Multi-Layer Garment (MLG) Dataset

MLG is a dataset specifically constructed for the ML-VTON task, supporting robust training and evaluation:

  • Composition: 3,538 samples, split into 2,783 training and 755 test quadruplets gig_i8.
    • gig_i9: inner garment product shot
    • gog_o0: outer garment product shot
    • gog_o1: person image wearing both garments
    • gog_o2: clothing-agnostic image (gog_o3 with upper-body region masked using SCHP parsing)
  • Garment and pose diversity: Categories include T-shirts, vests, cardigans, jackets, dresses, and coats; poses cover standing, walking, and dynamic actions in varied environments.
  • Annotation schema: Product crops obtained or segmented via SAM; person parsing via SCHP provides upper-body mask gog_o4; for each layer gog_o5, pixelwise masks gog_o6, boundary bands gog_o7, and interiors gog_o8 are defined for metric evaluation.

5. Evaluation Metric: Layered Appearance Coherence Difference (LACD)

Standard metrics such as FID, SSIM, and LPIPS provide limited sensitivity to inter-layer artifacts and occlusion errors. LACD is introduced to address these limitations by directly measuring per-layer and seam consistency:

  • Definitions:
    • For each garment layer gog_o9, fif_i0 is its pixel region, fif_i1 is the seam band adjoining layer fif_i2, and fif_i3 is the interior.
  • Layer discrepancy:

fif_i4

where fif_i5 upweights seam artifacts.

  • Overall metric:

fif_i6

  • Interpretation: The LACD emphasizes seam fidelity and inter-layer coherence, penalizing unrealistic edge transitions and “leakage” of occluded layers, making it highly discriminative for ML-VTON scenarios.

6. Experimental Results and Ablations

Extensive experiments on MLG benchmark the quantitative and qualitative advantages of GO-MLVTON:

Quantitative Comparison:

Method FID ↓ KID ↓ SSIM ↑ LPIPS ↓ LACD ↓
CAT-DM (CVPR’24) 32.36 5.97 0.845 0.138 0.719
MV-VTON (AAAI’25) 36.31 7.60 0.830 0.175 0.973
CATVTON (ICLR’25) 30.54 4.61 0.841 0.127 0.626
GO-MLVTON 22.82 0.35 0.858 0.108 0.623

GO-MLVTON outperforms all baselines, achieving significant improvements (e.g., a 7.7-point drop in FID, 4.26 lower KID versus the best baseline, and consistent gains in SSIM, LPIPS, and LACD).

GOL Module Ablation:

Ablation studies reveal that the GOL module alone (without fif_i7) may leave occluded inner-layer details visible. Full supervision with fif_i8 is necessary for optimal suppression and minimal seam and LACD errors.

Classifier-Free Guidance (CFG) Hyperparameter:

The optimal CFG scale for the StableDiffusion UNet is fif_i9, producing the best accuracy/diversity trade-off.

Qualitative Assessment:

Visual results confirm the elimination of inner garment bleed-through, preservation of garment boundary sharpness, and physically coherent layer interaction and deformation.

7. GO-MLVTON Pipeline Summary

A high-level pseudocode outline of the GO-MLVTON process is as follows: fof_o0

GO-MLVTON thus establishes an effective paradigm for multi-layer virtual try-on, innovating across occlusion modeling, diffusion-based compositing, and dedicated evaluation for fashion-oriented generative vision tasks (Yu et al., 20 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to GO-MLVTON.