Papers
Topics
Authors
Recent
Search
2000 character limit reached

EditGarment: Garment Editing Dataset

Updated 18 July 2026
  • EditGarment is defined as the first instruction-based dataset for standalone garment editing, combining image-instruction pairs to generate precise garment modifications.
  • It employs a two-stage pipeline that integrates multimodal synthesis with a semantic-aware Fashion Edit Score to ensure high-quality, reliable edits.
  • The dataset introduces a taxonomy of six editing categories, providing clear control over garment attributes like material, color, and structure in fashion applications.

Searching arXiv for the target paper and closely related garment-editing works to ground the article. Search query: arXiv (Yin et al., 5 Aug 2025) EditGarment garment editing dataset semantic-aware evaluation EditGarment denotes an instruction-based garment-editing dataset constructed through an automated pipeline that combines multimodal LLM synthesis with a semantic-aware evaluation mechanism tailored to apparel editing (Yin et al., 5 Aug 2025). It is defined around standalone garment editing rather than general scene manipulation or virtual try-on, and is motivated by the observation that garment editing requires understanding garment-specific semantics and attribute dependencies, while progress has been limited by the scarcity of high-quality instruction-image pairs and by the limited suitability of generic supervision and evaluation for fashion-specific edits (Yin et al., 5 Aug 2025).

1. Definition and scope

EditGarment is described as the first instruction-based dataset tailored to standalone garment editing, and as the first large-scale, instruction-based garment-editing dataset built through a fully automated two-stage pipeline (Yin et al., 5 Aug 2025). Its central unit is the instruction-image triplet: an original garment image, an edit instruction expressed in natural language, and an edited result whose semantic fidelity is assessed through a fashion-aware scoring procedure (Yin et al., 5 Aug 2025).

The dataset is explicitly positioned against two neighboring categories of resources. First, existing image-editing datasets such as InstructPix2Pix, HQ-EDIT, and MAGICBRUSH focus on general-purpose scenes and employ generic prompts, whereas garment-specific tasks require fine-grained control over components, materials, and silhouettes (Yin et al., 5 Aug 2025). Second, virtual-try-on and garment-generation corpora such as DressCode, VITON-HD, CM-Fashion, and MMDGarment are not structured for iterative, instruction-driven editing (Yin et al., 5 Aug 2025).

A plausible implication is that EditGarment occupies a distinct role in the fashion-editing literature: it treats the garment itself as the editing target, rather than treating the task as person-centric try-on synthesis, latent full-body manipulation, or generic image editing. This interpretation is consistent with adjacent work that addresses text-guided garment manipulation in full-body human images (Yoshikawa et al., 2023), multimodal fashion image editing with region masks and garment texture images (Wang et al., 2024), and retrieval-augmented fashion editing from textual specifications (Sanguigni et al., 18 Apr 2025).

2. Automated construction pipeline

EditGarment is built with a two-stage pipeline consisting of an “Auto-Synthesis” stage and a “Fashion Edit Score” evaluation stage (Yin et al., 5 Aug 2025). In the first stage, real garment photographs from the MMDGarment corpus serve as the visual foundation (Yin et al., 5 Aug 2025). For each image, a vision-LLM, Qwen-VL, is prompted with one of six carefully designed edit-type templates, each reflecting a common operation in professional fashion workflows, and generates a textual triplet comprising an Original Description, an Edit Instruction, and an Edited Description (Yin et al., 5 Aug 2025).

These textual triplets are then combined with the original garment images and passed to an image-editing generator, Gemini-2.0-Flash, to produce candidate edited images (Yin et al., 5 Aug 2025). By cycling each source image through all six edit templates, the pipeline synthesizes 52,257 instruction-image triplets (Yin et al., 5 Aug 2025). The abstract states that these candidate triplets are then filtered, and 20,596 high-quality triplets are retained to build the final dataset (Yin et al., 5 Aug 2025).

This automated formulation is designed to address two bottlenecks identified in the work: manual annotation is costly and hard to scale, and MLLM-based synthesis for garment editing is otherwise constrained by imprecise instruction modeling and by a lack of fashion-specific supervisory signals (Yin et al., 5 Aug 2025). This suggests that the dataset is as much a data-construction method as it is a benchmark, with the semantic filtering stage functioning as the critical quality-control mechanism.

3. Editing taxonomy

The dataset defines six editing instruction categories aligned with real-world fashion workflows (Yin et al., 5 Aug 2025). These categories are intended to produce balanced and diverse instruction-image triplets (Yin et al., 5 Aug 2025).

Category Operation Representative example
Object Removal eliminate a component “Remove the front pocket from the denim jacket.”
Object Replacement swap one element for another “Replace the zipper closure with a row of buttons.”
Object Addition introduce a new detail “Add patch pockets to the sides of the cargo shorts.”
Material Replacement change the textile “Convert the cotton blouse to a silk fabric.”
Color Alteration adjust hue without altering texture “Turn the olive green coat into a beige coat.”
Structural Alteration modify the silhouette or cut “Transform the straight-leg trousers into a flared-leg style.”

These categories formalize a garment-specific edit space centered on garment components, materials, color, and silhouette (Yin et al., 5 Aug 2025). In contrast to generic edit taxonomies, the categories encode dependencies that are particularly important in fashion, such as the relation between a component and its attributes, or between silhouette changes and the persistence of unrelated details (Yin et al., 5 Aug 2025).

A plausible implication is that this taxonomy offers a more operational notion of edit controllability than systems that optimize only for text-image agreement. Related methods such as StyleHumanCLIP target text-guided manipulation in full-body human images through an attention-based latent code mapper and feature-space masking (Yoshikawa et al., 2023), while DPDEdit combines text prompts, region masks, human pose images, and garment texture images in a latent diffusion architecture (Wang et al., 2024). EditGarment, by contrast, encodes the edit space directly at the dataset level through instruction categories and semantically structured supervision (Yin et al., 5 Aug 2025).

4. Fashion Edit Score

The second stage of the pipeline introduces Fashion Edit Score, a semantic-aware evaluation metric intended to capture semantic dependencies between garment attributes and to provide reliable supervision during dataset construction (Yin et al., 5 Aug 2025). The motivation is explicit: MLLM-generated outputs can drift from the intended edits or introduce unwanted changes, and generic evaluation metrics such as SSIM, LPIPS, or CLIP similarity cannot reliably detect whether domain-specific edits such as fabric swaps or pocket removals have been faithfully executed (Yin et al., 5 Aug 2025).

The scoring procedure begins by parsing each Edited Description using a scene-graph-style procedure that identifies garment components and their attributes and builds a hierarchical dependency graph (Yin et al., 5 Aug 2025). From this graph, three classes of binary visual-question-answering checks are derived: Instruction-Critical Questions, Instruction-Dependent Questions, and Context-Preserving Questions (Yin et al., 5 Aug 2025). The questions are answered on the edited image using Qwen-VL, and a question is counted as correct only if it and all its parents in the dependency graph are affirmed (Yin et al., 5 Aug 2025).

The final score is the normalized weighted sum

FEditScore=∑i∈ICQwICQ⋅δi+∑j∈IDQwIDQ(lj)⋅δj+∑k∈CPQwCPQ⋅δk∑i∈ICQwICQ+∑j∈IDQwIDQ(lj)+∑k∈CPQwCPQ,\mathrm{FEditScore} = \frac{ \sum_{i\in \mathrm{ICQ}} w_{\mathrm{ICQ}}\cdot \delta_i + \sum_{j\in \mathrm{IDQ}} w_{\mathrm{IDQ}}(l_j)\cdot \delta_j + \sum_{k\in \mathrm{CPQ}} w_{\mathrm{CPQ}}\cdot \delta_k }{ \sum_{i\in \mathrm{ICQ}} w_{\mathrm{ICQ}} + \sum_{j\in \mathrm{IDQ}} w_{\mathrm{IDQ}}(l_j) + \sum_{k\in \mathrm{CPQ}} w_{\mathrm{CPQ}} },

with

wIDQ(l)=1+wICQ⋅tdecayl.w_{\mathrm{IDQ}}(l)=1+w_{\mathrm{ICQ}}\cdot t_{\mathrm{decay}}^l.

The specified hyperparameters are wICQ=3w_{\mathrm{ICQ}}=3, wCPQ=1w_{\mathrm{CPQ}}=1, tdecay=0.3t_{\mathrm{decay}}=0.3, and threshold α=0.8\alpha=0.8 (Yin et al., 5 Aug 2025). Only triplets whose edited image scores at least α\alpha are retained (Yin et al., 5 Aug 2025).

This metric is interpretable at the attribute level and is designed to enforce semantic dependencies tailored to apparel (Yin et al., 5 Aug 2025). That design sharply contrasts with evaluation regimes in adjacent work. For example, Fashion-RAG reports LPIPS, SSIM, FID, KID, CLIP-T, and CLIP-I on Dress Code (Sanguigni et al., 18 Apr 2025), while DPDEdit reports FID, LPIPS, CLIP-I, and CLIP-S on an extended VITON-HD benchmark (Wang et al., 2024). EditGarment’s contribution is not another image-level similarity score, but a domain-specific semantic filter that is embedded into the dataset construction process itself (Yin et al., 5 Aug 2025).

5. Dataset composition and relation to neighboring benchmarks

The construction pipeline synthesizes 52,257 candidate triplets and retains 20,596 high-quality triplets, corresponding to approximately 39.4% of the candidates (Yin et al., 5 Aug 2025). The retained subset is defined by the Fashion Edit Score threshold and is intended to provide high-quality instruction-image pairs for garment editing (Yin et al., 5 Aug 2025).

EditGarment’s scope is narrower than later benchmarks that unify editing with other fashion-generation tasks. Dress-ED, for example, is presented as the first large-scale benchmark that unifies virtual try-on, virtual try-off, and text-guided garment editing within a single framework, and contains over 146k verified quadruplets spanning three garment categories and seven edit types (Sanguigni et al., 23 Mar 2026). EditGarment instead remains centered on standalone garment editing and on the semantic structure of garment-specific instructions (Yin et al., 5 Aug 2025).

The distinction from try-on resources is especially important. M&M VTO is a mix-and-match virtual try-on method that takes multiple garment images, a text description for garment layout, and an image of a person, and produces a visualization of those garments on that person (Zhu et al., 2024). “Wearing the Same Outfit in Different Ways -- A Controllable Virtual Try-on Method” focuses on controlling drape while preserving garment identity and allows instance independent editing of garment drape such as tucking a shirt or wearing a jacket open or closed (Li et al., 2022). These are garment-related editing problems, but they operate in a person-conditioned visualization setting rather than in the standalone garment-editing regime that defines EditGarment (Yin et al., 5 Aug 2025).

A plausible implication is that EditGarment provides a complementary benchmark layer beneath such systems: it isolates the semantics of garment modification before those modifications are composed with body pose, person identity, or full-scene rendering.

6. Research uses, implications, and limitations

EditGarment is described as suitable for bootstrapping the training of specialized garment-editing models, enabling designers to iterate virtually on prototypes without reshooting photographs, underpinning interactive customization tools on e-commerce platforms, and serving as a benchmark for future research in fashion-aware generative models that preserve garment structure while accommodating complex, domain-specific instructions (Yin et al., 5 Aug 2025).

This research role connects naturally to several methodological directions already visible in adjacent work. StyleHumanCLIP studies text-guided garment manipulation for StyleGAN-Human through an attention-based latent code mapper and feature-space masking (Yoshikawa et al., 2023). Fashion-RAG retrieves multiple garments from a database and incorporates their attributes into Stable Diffusion inpainting through textual inversion (Sanguigni et al., 18 Apr 2025). DPDEdit integrates text prompts, region masks, human pose images, and garment texture images, using Grounded-SAM for region prediction and a texture injection and refinement mechanism (Wang et al., 2024). These systems target controllable image editing, but EditGarment contributes the missing instruction-image substrate specialized to garment semantics (Yin et al., 5 Aug 2025).

The paper’s stated limitation is structural rather than algorithmic: progress in instruction-based garment editing had been limited by the scarcity of high-quality instruction-image pairs, and automated MLLM synthesis alone was insufficient because of imprecise instruction modeling and the absence of fashion-specific supervisory signals (Yin et al., 5 Aug 2025). EditGarment addresses those limits with automated synthesis plus semantic-aware filtering, but the need for a thresholded selection step also indicates that candidate generation remains noisy (Yin et al., 5 Aug 2025).

A common misconception would be to treat garment editing as merely a special case of generic instruction-based image editing. The dataset’s design argues against that view: garment editing requires understanding garment-specific semantics and attribute dependencies, and generic metrics cannot reliably detect whether edits such as material replacement or component removal were faithfully executed (Yin et al., 5 Aug 2025). In that sense, EditGarment formalizes garment editing as a domain with its own instruction taxonomy, dependency structure, and evaluation logic.

7. Position within the broader evolution of garment editing

Within the broader literature, EditGarment can be situated between earlier controllable editing systems and later unified fashion-editing benchmarks. Earlier work emphasized controllability within generative models or virtual try-on pipelines: latent-space editing for full-body garment manipulation (Yoshikawa et al., 2023), detail-preserved diffusion for multimodal fashion image editing (Wang et al., 2024), and controllable drape editing in virtual try-on (Li et al., 2022). Later work extends instruction-driven editing into larger multimodal ecosystems, such as Dress-ED, which couples edited garments with corresponding edited person images and uses a unified multimodal diffusion framework for instruction-guided virtual try-on and virtual try-off (Sanguigni et al., 23 Mar 2026).

EditGarment’s distinctive contribution is therefore not a new editing backbone, but a dataset-construction methodology in which semantic structure governs both synthesis and filtering (Yin et al., 5 Aug 2025). This suggests a broader methodological shift in fashion AI: from relying on generic prompts and generic similarity measures toward domain-aware instruction schemas and semantic verification. For research on fashion-aware generative models, that shift is likely to be consequential because it changes not only how models are evaluated, but also what kinds of instruction following can be learned in the first place (Yin et al., 5 Aug 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to EditGarment.