---
title: 'EditGarment: Garment Editing Dataset'
url: https://www.emergentmind.com/topics/editgarment
type: topic
---

# EditGarment: Garment Editing Dataset

Searching arXiv for the target paper and closely related garment-editing works to ground the article.
Search query: arXiv 2508.03497 EditGarment garment editing dataset semantic-aware evaluation
EditGarment denotes an instruction-based garment-editing dataset constructed through an automated pipeline that combines multimodal large language model synthesis with a semantic-aware evaluation mechanism tailored to apparel editing [2508.03497]. It is defined around standalone garment editing rather than general scene manipulation or virtual try-on, and is motivated by the observation that garment editing requires understanding garment-specific semantics and attribute dependencies, while progress has been limited by the scarcity of high-quality instruction-image pairs and by the limited suitability of generic supervision and evaluation for fashion-specific edits [2508.03497].

## 1. Definition and scope

EditGarment is described as the first instruction-based dataset tailored to standalone garment editing, and as the first large-scale, instruction-based garment-editing dataset built through a fully automated two-stage pipeline [2508.03497]. Its central unit is the instruction-image triplet: an original garment image, an edit instruction expressed in natural language, and an edited result whose semantic fidelity is assessed through a fashion-aware scoring procedure [2508.03497].

The dataset is explicitly positioned against two neighboring categories of resources. First, existing image-editing datasets such as InstructPix2Pix, HQ-EDIT, and MAGICBRUSH focus on general-purpose scenes and employ generic prompts, whereas garment-specific tasks require fine-grained control over components, materials, and silhouettes [2508.03497]. Second, virtual-try-on and garment-generation corpora such as DressCode, VITON-HD, CM-Fashion, and MMDGarment are not structured for iterative, instruction-driven editing [2508.03497].

A plausible implication is that EditGarment occupies a distinct role in the fashion-editing literature: it treats the garment itself as the editing target, rather than treating the task as person-centric try-on synthesis, latent full-body manipulation, or generic image editing. This interpretation is consistent with adjacent work that addresses text-guided garment manipulation in full-body human images [2305.16759], multimodal fashion image editing with region masks and garment texture images [2409.01086], and retrieval-augmented fashion editing from textual specifications [2504.14011].

## 2. Automated construction pipeline

EditGarment is built with a two-stage pipeline consisting of an “Auto-Synthesis” stage and a “Fashion Edit Score” evaluation stage [2508.03497]. In the first stage, real garment photographs from the MMDGarment corpus serve as the visual foundation [2508.03497]. For each image, a vision-language model, Qwen-VL, is prompted with one of six carefully designed edit-type templates, each reflecting a common operation in professional fashion workflows, and generates a textual triplet comprising an Original Description, an Edit Instruction, and an Edited Description [2508.03497].

These textual triplets are then combined with the original garment images and passed to an image-editing generator, Gemini-2.0-Flash, to produce candidate edited images [2508.03497]. By cycling each source image through all six edit templates, the pipeline synthesizes 52,257 instruction-image triplets [2508.03497]. The abstract states that these candidate triplets are then filtered, and 20,596 high-quality triplets are retained to build the final dataset [2508.03497].

This automated formulation is designed to address two bottlenecks identified in the work: manual annotation is costly and hard to scale, and MLLM-based synthesis for garment editing is otherwise constrained by imprecise instruction modeling and by a lack of fashion-specific supervisory signals [2508.03497]. This suggests that the dataset is as much a data-construction method as it is a benchmark, with the semantic filtering stage functioning as the critical quality-control mechanism.

## 3. Editing taxonomy

The dataset defines six editing instruction categories aligned with real-world fashion workflows [2508.03497]. These categories are intended to produce balanced and diverse instruction-image triplets [2508.03497].

| Category | Operation | Representative example |
|---|---|---|
| Object Removal | eliminate a component | “Remove the front pocket from the denim jacket.” |
| Object Replacement | swap one element for another | “Replace the zipper closure with a row of buttons.” |
| Object Addition | introduce a new detail | “Add patch pockets to the sides of the cargo shorts.” |
| Material Replacement | change the textile | “Convert the cotton blouse to a silk fabric.” |
| Color Alteration | adjust hue without altering texture | “Turn the olive green coat into a beige coat.” |
| Structural Alteration | modify the silhouette or cut | “Transform the straight-leg trousers into a flared-leg style.” |

These categories formalize a garment-specific edit space centered on garment components, materials, color, and silhouette [2508.03497]. In contrast to generic edit taxonomies, the categories encode dependencies that are particularly important in fashion, such as the relation between a component and its attributes, or between silhouette changes and the persistence of unrelated details [2508.03497].

A plausible implication is that this taxonomy offers a more operational notion of edit controllability than systems that optimize only for text-image agreement. Related methods such as StyleHumanCLIP target text-guided manipulation in full-body human images through an attention-based latent code mapper and feature-space masking [2305.16759], while DPDEdit combines text prompts, region masks, human pose images, and garment texture images in a latent diffusion architecture [2409.01086]. EditGarment, by contrast, encodes the edit space directly at the dataset level through instruction categories and semantically structured supervision [2508.03497].

## 4. Fashion Edit Score

The second stage of the pipeline introduces Fashion Edit Score, a semantic-aware evaluation metric intended to capture semantic dependencies between garment attributes and to provide reliable supervision during dataset construction [2508.03497]. The motivation is explicit: MLLM-generated outputs can drift from the intended edits or introduce unwanted changes, and generic evaluation metrics such as SSIM, LPIPS, or CLIP similarity cannot reliably detect whether domain-specific edits such as fabric swaps or pocket removals have been faithfully executed [2508.03497].

The scoring procedure begins by parsing each Edited Description using a scene-graph-style procedure that identifies garment components and their attributes and builds a hierarchical dependency graph [2508.03497]. From this graph, three classes of binary visual-question-answering checks are derived: Instruction-Critical Questions, Instruction-Dependent Questions, and Context-Preserving Questions [2508.03497]. The questions are answered on the edited image using Qwen-VL, and a question is counted as correct only if it and all its parents in the dependency graph are affirmed [2508.03497].

The final score is the normalized weighted sum

$$
\mathrm{FEditScore}
=
\frac{
\sum_{i\in \mathrm{ICQ}} w_{\mathrm{ICQ}}\cdot \delta_i
+
\sum_{j\in \mathrm{IDQ}} w_{\mathrm{IDQ}}(l_j)\cdot \delta_j
+
\sum_{k\in \mathrm{CPQ}} w_{\mathrm{CPQ}}\cdot \delta_k
}{
\sum_{i\in \mathrm{ICQ}} w_{\mathrm{ICQ}}
+
\sum_{j\in \mathrm{IDQ}} w_{\mathrm{IDQ}}(l_j)
+
\sum_{k\in \mathrm{CPQ}} w_{\mathrm{CPQ}}
},
$$

with

$$
w_{\mathrm{IDQ}}(l)=1+w_{\mathrm{ICQ}}\cdot t_{\mathrm{decay}}^l.
$$

The specified hyperparameters are $w_{\mathrm{ICQ}}=3$, $w_{\mathrm{CPQ}}=1$, $t_{\mathrm{decay}}=0.3$, and threshold $\alpha=0.8$ [2508.03497]. Only triplets whose edited image scores at least $\alpha$ are retained [2508.03497].

This metric is interpretable at the attribute level and is designed to enforce semantic dependencies tailored to apparel [2508.03497]. That design sharply contrasts with evaluation regimes in adjacent work. For example, Fashion-RAG reports LPIPS, SSIM, FID, KID, CLIP-T, and CLIP-I on Dress Code [2504.14011], while DPDEdit reports FID, LPIPS, CLIP-I, and CLIP-S on an extended VITON-HD benchmark [2409.01086]. EditGarment’s contribution is not another image-level similarity score, but a domain-specific semantic filter that is embedded into the dataset construction process itself [2508.03497].

## 5. Dataset composition and relation to neighboring benchmarks

The construction pipeline synthesizes 52,257 candidate triplets and retains 20,596 high-quality triplets, corresponding to approximately 39.4% of the candidates [2508.03497]. The retained subset is defined by the Fashion Edit Score threshold and is intended to provide high-quality instruction-image pairs for garment editing [2508.03497].

EditGarment’s scope is narrower than later benchmarks that unify editing with other fashion-generation tasks. Dress-ED, for example, is presented as the first large-scale benchmark that unifies virtual try-on, virtual try-off, and text-guided garment editing within a single framework, and contains over 146k verified quadruplets spanning three garment categories and seven edit types [2603.22607]. EditGarment instead remains centered on standalone garment editing and on the semantic structure of garment-specific instructions [2508.03497].

The distinction from try-on resources is especially important. M&M VTO is a mix-and-match virtual try-on method that takes multiple garment images, a text description for garment layout, and an image of a person, and produces a visualization of those garments on that person [2406.04542]. “Wearing the Same Outfit in Different Ways -- A Controllable Virtual Try-on Method” focuses on controlling drape while preserving garment identity and allows instance independent editing of garment drape such as tucking a shirt or wearing a jacket open or closed [2211.16989]. These are garment-related editing problems, but they operate in a person-conditioned visualization setting rather than in the standalone garment-editing regime that defines EditGarment [2508.03497].

A plausible implication is that EditGarment provides a complementary benchmark layer beneath such systems: it isolates the semantics of garment modification before those modifications are composed with body pose, person identity, or full-scene rendering.

## 6. Research uses, implications, and limitations

EditGarment is described as suitable for bootstrapping the training of specialized garment-editing models, enabling designers to iterate virtually on prototypes without reshooting photographs, underpinning interactive customization tools on e-commerce platforms, and serving as a benchmark for future research in fashion-aware generative models that preserve garment structure while accommodating complex, domain-specific instructions [2508.03497].

This research role connects naturally to several methodological directions already visible in adjacent work. StyleHumanCLIP studies text-guided garment manipulation for StyleGAN-Human through an attention-based latent code mapper and feature-space masking [2305.16759]. Fashion-RAG retrieves multiple garments from a database and incorporates their attributes into Stable Diffusion inpainting through textual inversion [2504.14011]. DPDEdit integrates text prompts, region masks, human pose images, and garment texture images, using Grounded-SAM for region prediction and a texture injection and refinement mechanism [2409.01086]. These systems target controllable image editing, but EditGarment contributes the missing instruction-image substrate specialized to garment semantics [2508.03497].

The paper’s stated limitation is structural rather than algorithmic: progress in instruction-based garment editing had been limited by the scarcity of high-quality instruction-image pairs, and automated MLLM synthesis alone was insufficient because of imprecise instruction modeling and the absence of fashion-specific supervisory signals [2508.03497]. EditGarment addresses those limits with automated synthesis plus semantic-aware filtering, but the need for a thresholded selection step also indicates that candidate generation remains noisy [2508.03497].

A common misconception would be to treat garment editing as merely a special case of generic instruction-based image editing. The dataset’s design argues against that view: garment editing requires understanding garment-specific semantics and attribute dependencies, and generic metrics cannot reliably detect whether edits such as material replacement or component removal were faithfully executed [2508.03497]. In that sense, EditGarment formalizes garment editing as a domain with its own instruction taxonomy, dependency structure, and evaluation logic.

## 7. Position within the broader evolution of garment editing

Within the broader literature, EditGarment can be situated between earlier controllable editing systems and later unified fashion-editing benchmarks. Earlier work emphasized controllability within generative models or virtual try-on pipelines: latent-space editing for full-body garment manipulation [2305.16759], detail-preserved diffusion for multimodal fashion image editing [2409.01086], and controllable drape editing in virtual try-on [2211.16989]. Later work extends instruction-driven editing into larger multimodal ecosystems, such as Dress-ED, which couples edited garments with corresponding edited person images and uses a unified multimodal diffusion framework for instruction-guided virtual try-on and virtual try-off [2603.22607].

EditGarment’s distinctive contribution is therefore not a new editing backbone, but a dataset-construction methodology in which semantic structure governs both synthesis and filtering [2508.03497]. This suggests a broader methodological shift in fashion AI: from relying on generic prompts and generic similarity measures toward domain-aware instruction schemas and semantic verification. For research on fashion-aware generative models, that shift is likely to be consequential because it changes not only how models are evaluated, but also what kinds of instruction following can be learned in the first place [2508.03497].

Source: https://www.emergentmind.com/topics/editgarment