---
title: 'InstructVTON: Instruction-Guided Virtual Try-On'
url: https://www.emergentmind.com/topics/instructvton
type: topic
---

# InstructVTON: Instruction-Guided Virtual Try-On

InstructVTON denotes an instruction-following interactive virtual try-on system for inpainting-based virtual try-on that replaces manual mask drawing with natural-language-guided, automatically generated masks and agentic execution logic. Introduced by Julien Han, Shuwen Qiu, Qi Li, Xingzi Xu, Mehmet Saygin Seyfioglu, Kavosh Asadi, and Karim Bouyarmane, with affiliations listed as Amazon, UCLA, and Duke University, the method takes a human model image \(I_\text{src}\), a set of target garment images \(S_\text{ref}\), and optionally a free-text instruction \(T_\text{instruction}\), and produces a try-on image \(I_\text{try-on}\). The manuscript states “Work submitted in November 2024 to CVPR 2025,” and describes a training-free framework that is interoperable with existing inpainting-based virtual try-on backbones “without the need for retraining or fine-tuning” [2509.20524].

## 1. Problem formulation and scope

InstructVTON is situated in the now-common image-guided or image-conditioned inpainting formulation of virtual try-on. In that formulation, a model takes a person image, a target garment image, and a binary mask, then generates a realistic image of the person wearing the target garment. The paper isolates the mask as the principal practical bottleneck: in inpainting-based virtual try-on, the binary mask determines the spatial editing region, how much of the original person image is preserved, and the layout bias supplied to the inpainting model [2509.20524].

Within this setting, “instruction-following interactive virtual try-on” means that free-form natural language replaces manual mask design as the user-facing control mechanism. The instructions can describe not only garment replacement but also styling preferences and layering configurations, including sleeves rolled up, jacket open, shirt tucked in, and multi-garment outfit composition. The system interprets these instructions, determines the try-on order for multiple garments if needed, automatically constructs the necessary inpainting masks, and may execute multiple rounds of generation if one pass cannot realize the requested style.

A central misconception addressed by the formulation is that improved masking alone is sufficient. The paper argues that some desired edits are not expressible with a single binary mask. Its canonical example is a long-sleeve shirt with “sleeves rolled up” styling on a person already wearing a long-sleeve top: if the whole sleeve region is masked, a standard one-shot VTON model tends to generate the new long-sleeve shirt with sleeves down. This frames InstructVTON not merely as an auto-masking method, but as an orchestration system for instruction-conditioned multi-step inference.

## 2. System architecture and execution logic

The architecture has two main components: a **Top level Agent** and a **VTO Agent**. It also relies on **AutoMasker** and an underlying inpainting VTO model. The top-level agent plans the order to try on each garment and summarizes corresponding style instructions, while the VTO agent performs masking and executes the VTO model [2509.20524].

The end-to-end pipeline proceeds in stages. Given \(I_\text{src}\), \(S_\text{ref}\), and \(T_\text{instruction}\), the top-level agent reorganizes the problem into a sequence of single-garment VTO tasks. Its duties are to reorder garments in the correct try-on sequence, summarize or paraphrase the instruction for each garment, and produce an ordered list of garment-specific try-on actions. The paper does not provide a formal grammar or exact prompt template for this plan.

Execution is autoregressive across garments. For the first garment, the source image is the original human image \(I_\text{src}\); for each subsequent garment, the source image is the output of the previous step. For each action, the VTO agent sends the current person image, the current garment image, and the garment-specific style instruction to AutoMasker. AutoMasker segments body parts and existing clothing, reasons about which regions must be masked according to garment type and style instruction, and outputs a binary inpainting mask. The VTO agent then calls the existing inpainting-based VTON backbone with the current person image, current target garment image, and generated binary mask.

This organization is significant because InstructVTON does not alter the core generator architecture. It acts as an external controller that changes conditioning inputs and execution schedule rather than the learned generator weights. The paper explicitly characterizes the framework as training-free at inference time and notes no architectural additions such as adapter tuning, LoRA layers, prompt injection into the diffusion denoiser, or architecture modification inside the VTON network.

## 3. Automatic mask generation and the meaning of “optimal”

AutoMasker uses two segmentation sources: a **Body Parts Segmentation Map (BPSM)** and a **Clothing Segmentation Map (CSM)**. The provided text does not specify exact model names. Functionally, BPSM outputs body-part regions including face, upper torso, lower torso, upper arms, lower arms, hands, upper legs, lower legs, and feet, while CSM provides segmentation of all clothing pieces [2509.20524].

The paper introduces a formal view of mask generation. Let \(\mathcal{B}=\{b_1,b_2,\ldots\}\) denote body-part segments and \(\mathcal{C}=\{c_0,c_1,c_2,\ldots\}\) denote clothing segments, where \(c_0\) is the unclothed area. Both partition the person region \(B\):
$$
\bigcup b_i = B,\quad \forall i,j\; b_i\cap b_j=\emptyset,
$$
$$
\bigcup c_i = B,\quad \forall i,j\; c_i\cap c_j=\emptyset.
$$

Let \(\bar{v}\) denote the ideal position of the target garment in the output image. The paper defines the B-trace as the set of body-part segments intersecting \(\bar v\), and the C-trace as the set of existing clothing segments intersecting \(\bar v\). It then defines the ideal mask \(\bar m\) as covering all existing clothing that overlaps the target garment’s eventual placement together with the target garment’s intended support region. Because \(\bar v\) is unknown a priori, the practical system estimates the needed traces from garment type and style instruction, and the estimated minimally invasive mask \(\hat m\) becomes the union of selected clothing segments and selected body-part segments inferred from garment type and instruction.

The term “optimal” does not denote a globally solved optimization problem. It denotes minimal invasiveness and mask efficiency: the mask should cover the smallest possible area necessary to realize the desired try-on effect while preserving the rest of the person image. The paper operationalizes this with the scalar metric
$$
\text{Mask efficiency} = 1-\frac{\text{masked area}}{\text{total image area}}.
$$
Higher values indicate that more of the source image is preserved.

Instruction-to-region translation is implemented through parts inclusion rules driven by target garment classification—upper, lower, or overall—together with target garment structured attributes such as sleeve length, leg length, and closure type, plus structured styling instructions. The paper gives several examples. For overcoat or dress cases, the leg area may be made convex so the garment occupies a natural silhouette between the legs. For open-chest style, a vertical stripe in the torso center is removed from the mask to encourage the model to leave the underlying shirt visible and generate an open jacket or coat. For rolled sleeves, upper-arm areas are included while lower-arm areas are avoided in the final desired mask.

## 4. Natural-language style control and multi-round generation

The supported instruction space includes garment replacement, open or closed outerwear styling, shirt tucked in, sleeve rolled up, multi-garment outfit assembly, and single-garment style variation [2509.20524]. From the examples and discussion, style attributes can include sleeve length or rolled-up sleeves, closure state, layering relationships, and garment order.

A defining feature is the distinction between cases feasible with a single binary mask and cases impossible or ineffective with a single binary mask. When direct inpainting is insufficient, the VTO agent performs a multi-round strategy. The paper’s primary example is rolled-up sleeves. In that case, the system first uses a dummy garment, such as a tank top, to create an intermediate source image in which the lower arms are uncovered. It then uses the original target garment in a second VTO pass with a smaller, style-appropriate mask that preserves lower-arm visibility, allowing the model to synthesize a rolled-up sleeve appearance.

The same planning logic is extended to multiple garments. For a request involving shirt, pants, and jacket, with the instruction “try on the shirt tucked in, jacket open,” the top-level agent linearizes the problem into an ordered sequence of single-garment actions and the VTO agent executes them in turn. This converts multi-garment styling into ordered inference rather than monolithic generation.

The paper does not specify explicit mechanisms for resolving conflicting instructions, confidence estimation, or disambiguation strategies. It does state that the system uses VLM reasoning to infer the “most intuitive layout” when instruction is empty. A plausible implication is that style control is strongest when the instruction can be translated into garment-specific masking and ordering decisions rather than semantically ambiguous global edits.

## 5. Empirical evaluation, interoperability, and practical behavior

The experiments use DressCode—evaluated by dresses, upper body, and lower body—and VITON-HD. The stated sampling protocol is 50 pairs of human and garment images from each DressCode category and 100 pairs from VITON-HD. For the mask efficiency comparison, InstructVTON is run with empty instruction. Quantitative comparisons are reported against CatVTON and IDM-VTON, two open-source VTO systems with auto-masking [2509.20524].

The strongest quantitative pattern is mask efficiency. On DressCode dresses, upper body, lower body, and VITON-HD, InstructVTON reports 0.8269, 0.8924, 0.8653, and 0.7808, compared with CatVTON’s 0.6876, 0.8379, 0.8179, and 0.6877, and IDM-VTON’s 0.7334, 0.8196, 0.8238, and 0.6889. These values support the paper’s claim that the masks are more minimally invasive.

The quality metrics are more nuanced. For SSIM, InstructVTON achieves 0.9078 on DressCode dresses and 0.9370 on DressCode upper body, exceeding both baselines in those settings; on DressCode lower body it reports 0.9213, below CatVTON’s 0.9294 but above IDM-VTON’s 0.9174; on VITON-HD it reports 0.8887, lower than both baselines. For LPIPS, InstructVTON obtains the best values on all three DressCode categories—0.0689, 0.0478, and 0.0678—and 0.0874 on VITON-HD, improving over CatVTON’s 0.0918 but not IDM-VTON’s 0.0706. The paper therefore does not establish uniform dominance across all metrics and datasets; rather, it shows that smaller masks can coexist with competitive or improved try-on quality.

Qualitatively, the authors emphasize instruction-aware masks, natural open-jacket generation through center-stripe mask removal, improved overcoat and dress layout by appropriate masking between the legs, multi-garment composition through planned garment order, and recovery on difficult prior failure cases. The paper also reports a practical latency figure: for three target garments, total runtime is around 1 minute using state-of-the-art tools, due to multiple segmentation, VLM, and VTO calls. It does not specify exact VLM model name, exact segmentation model names, prompt templates, number of diffusion steps, image resolution, GPU type, or average number of rounds per example.

## 6. Relation to adjacent research and recognized limitations

Within the broader literature, InstructVTON occupies a distinct position: it is an instruction-following and agentic inference framework rather than a new garment-transfer backbone. This is clearest when contrasted with adjacent systems. DH-VTON is described as a “deep text-driven virtual try-on” model, but the underlying task remains garment-image-based try-on with implicit rather than explicit text conditioning at inference. Its relevance lies in semantically enriched garment features, hybrid attention, and a diffusion editing backbone built on Paint-by-Example, yet the paper does not specify a user-facing natural-language instruction interface in the usual sense of free-form edit commands [2410.12501]. In that comparison, InstructVTON’s defining novelty is not deeper garment representation alone but the replacement of mask-following interaction with instruction-following orchestration.

A second adjacent line is training-free spatial correction during diffusion sampling. “Training-free Clothing Region of Interest Self-correction for Virtual Try-On” introduces CSC, an inference-time attention correction mechanism and the VTID metric. It is not an instruction-following editor, but it addresses the spatial localization gap by constraining attention to a clothing region of interest during denoising. This suggests a complementary relationship: InstructVTON supplies instruction understanding, ordering, and automatic masking, while CSC supplies ROI-aware attention correction and a VTON-specific evaluation philosophy centered on person preservation and clothing consistency [2512.07126].

The limitations stated for InstructVTON are clear. First, latency is approximately 1 minute for a three-garment scenario and is not suitable for real-time interaction. Second, current body-part segmentation is too coarse for nuanced instructions such as “roll sleeves to three-quarter length,” because the mask can only include or exclude coarse regions like lower arms. Third, the planning process is open-loop: errors in early planning or execution propagate through later steps, and no closed-loop correction mechanism is described. The paper also does not provide a human study table or an explicit user evaluation protocol.

Taken together, these properties define InstructVTON as a training-free, interoperable, natural-language-guided control layer for inpainting-based virtual try-on. Its central technical proposition is that much of the usability gap in mask-based VTON can be addressed at inference time through instruction understanding, minimally invasive automatic masking, garment-order planning, and multi-round execution, rather than through retraining a new generator from scratch [2509.20524].

Source: https://www.emergentmind.com/topics/instructvton