LUIVITON: Automated 3D Virtual Try-On
- LUIVITON is an end-to-end system for 3D virtual try-on that transfers arbitrarily complex and multilayer garments onto humanoid bodies without 2D sewing patterns.
- It decomposes garment fitting into clothing-to-SMPL and body-to-SMPL correspondence problems using diffusion-based features and energy-driven noise filtering.
- The system employs a two-stage registration with SMPL and SMPL+D interpolation followed by neural cloth simulation to achieve low interpenetration and high-quality draping.
Searching arXiv for the cited LUIVITON paper and related methods mentioned in the provided source material. First, I’ll verify the main LUIVITON paper. Searching for: (Cao et al., 5 Sep 2025) LUIVITON, short for Learned Universal Interoperable VIrtual Try-ON, is an end-to-end system for fully automated 3D virtual try-on that targets arbitrary rest-posed 3D garments, including complex, multilayer, and potentially non-manifold meshes, and transfers them onto diverse humanoid characters with arbitrary body shapes and poses, without requiring 2D sewing patterns or pre-skinned garment assets (Cao et al., 5 Sep 2025). Its central design choice is to use SMPL as a universal proxy representation and to decompose garment-to-body draping into two correspondence problems—clothing-to-SMPL and body-to-SMPL—followed by registration of SMPL or SMPL+D and simulation-based garment transfer. The system is intended for automated, practical workflows in which high-quality 3D fittings, low interpenetration, and post-drape customization are all required.
1. Problem formulation and system scope
LUIVITON addresses a specific formulation of virtual try-on: the input is a rest-posed 3D garment mesh of arbitrary topology and a target 3D humanoid body mesh with arbitrary pose and shape, together with an SMPL proxy model and camera parameters for multi-view rendering during body correspondence (Cao et al., 5 Sep 2025). The output is a draped garment transferred to the target body through a sequence of correspondence prediction, registration, interpolation, and neural cloth simulation.
The system’s stated goal is fully automated 3D virtual try-on for arbitrary rest-posed garments on diverse humanoid characters. The target domain explicitly includes humans, robots, cartoon subjects, creatures, and aliens. The garment domain explicitly includes complex geometries, multilayer constructions, and non-manifold meshes. The method is also designed to function without 2D clothing sewing patterns and without pre-skinned garment assets.
A central modeling assumption is that SMPL provides a common parametric body space. SMPL is parameterized by pose and shape , while SMPL+D augments SMPL with per-vertex displacement and axis-aligned scale . The role of SMPL and SMPL+D is not merely representational; they mediate interoperability between garment geometry and arbitrary target bodies by making both correspondence problems relative to the same proxy space.
This decomposition is significant because the geometric and semantic difficulties of garment fitting and body matching are different. Clothing-to-SMPL is treated as a partial-to-complete surface correspondence problem in UV space, whereas body-to-SMPL is treated as a multi-view semantic correspondence problem regularized by geometry. This suggests that LUIVITON is less a single predictor than a modular transfer stack whose robustness depends on isolating heterogeneous sources of error.
2. Pipeline architecture and proxy-based decomposition
The pipeline consists of four stages: (1) clothing-to-SMPL correspondence prediction, (2) body-to-SMPL correspondence prediction, (3) registration, and (4) clothing fitting through SMPL-to-SMPL+D interpolation and ContourCraft simulation (Cao et al., 5 Sep 2025).
In the first stage, the system predicts garment-to-SMPL correspondences using a geometric learning approach. In the second, it predicts body-to-SMPL correspondences using multi-view consistent diffusion features, DINOv2 semantic features, and an energy-based noise filter. In the third, it solves two distinct registration problems: SMPL to clothing and SMPL+D to body. In the fourth, it transfers the garment by interpolating between the garment-aligned SMPL state and the body-aligned SMPL+D state, then draping through ContourCraft.
The decomposition is operationally important. The clothing side benefits from a surface-based correspondence model that can tolerate non-manifold and multilayer topology. The body side benefits from view-conditioned semantic features that can generalize across stylized and nonstandard humanoids. Registration then turns these correspondences into physically usable geometry, while simulation resolves collisions and preserves cloth topology.
LUIVITON also includes post-drape customization. It supports a Default Mode, an Auto-resized Mode, and a Customized Mode. In Auto-resized Mode, the optimized body scale from SMPL+D registration is used to scale the garment’s rest geometry along axes , , and for better fit. In Customized Mode, the garment rest geometry is scaled by user-defined axis-aligned factors, and this does not require recomputing correspondences or registration. Reported efficiency for such size adjustments is approximately 15 seconds per clothing size adjustment after correspondences and registrations have been precomputed.
3. Clothing-to-SMPL correspondence
The clothing-to-SMPL module is trained on a curated dataset of 300 garments manually draped by artists onto a canonical SMPL using Marvelous Designer (Cao et al., 5 Sep 2025). For each garment vertex, the nearest point on SMPL is computed and then mapped to SMPL UV space, yielding dense 3D-to-2D correspondences. The learned representation is a mapping
from garment vertex coordinates to SMPL UV coordinates.
The architecture is an adapted DiffusionNet operating on surfaces with feature diffusion, spatial gradients, and a per-vertex MLP. It predicts UV coordinates for each garment vertex. The UV representation is used as the correspondence space because it defines garment-to-SMPL relations on the canonical proxy surface rather than directly on an incomplete garment mesh.
The training objective is defined as
where 0 is the 1-th garment vertex in 3D and 2 is its ground-truth 2D SMPL UV coordinate.
The main reported advantage of this construction is robustness to partial coverage, front/back ambiguities, multilayer stitches, and non-manifold meshes. The data attributes this to functional learning on surfaces combined with UV-based supervision. A plausible implication is that UV supervision gives the network a topologically stable target even when the source garment surface is incomplete or structurally irregular.
Quantitatively, on clothing-SMPL correspondence and registration over 100 GarmentCode pairs, the reported comparison against CorrPredNet is as follows (Cao et al., 5 Sep 2025):
| Method | MGE / MEE | IR / No-Penetration |
|---|---|---|
| CorrPredNet | 0.2290 / 0.0993 | 1.02% / 5% |
| LUIVITON | 0.0499 / 0.0274 | 0.17% / 59% |
These figures indicate that the correspondence stage is not only more accurate in geodesic and Euclidean terms but also materially improves downstream penetration behavior after registration.
4. Body-to-SMPL correspondence and noise filtering
The body-to-SMPL module starts from multi-view depth maps rendered from the target body mesh (Cao et al., 5 Sep 2025). These depth images are passed to SyncMVD to synthesize consistent view textures and to extract multi-scale UNet diffusion features. Features are extracted from the last three UNet layers at denoising steps 3 and fused over time according to
4
with
5
and 6.
The consistent multi-view images are then processed by DINOv2, and the diffusion features are aggregated with DINOv2 features through a Feature Aggregation Module. The fused per-pixel features are unprojected to mesh vertices using camera intrinsics and extrinsics, producing per-vertex 3D embeddings for both the body and SMPL. Initial correspondences are computed using cosine similarity of these 3D features.
Because feature matching alone is vulnerable to outliers and left-right ambiguity, LUIVITON applies an energy-based iterative noise filter. For a vertex 7 and neighborhood 8, with neighbor offsets
9
where 0 and 1 are the corresponding SMPL vertices, an optimal rotation 2 is computed by SVD of the covariance of 3 and 4. The local deformation energy is
5
Vertices with energy above
6
are filtered out, with the neighborhood size increased as 7 for up to four iterations or until convergence.
This filter is presented as essential for removing outliers and resolving left-right ambiguities before registration. In body-SMPL+D registration, the reported Chamfer Distance 8 improves from 1.46 to 1.19 when the filter is enabled in the full method, and from 4.99 to 4.49 in the Diff3f baseline. The paper also reports that body-SMPL correspondence on a benchmark of 8 stylized characters with 10 poses each achieves MEE 9 and MGE 0, outperforming GeomFmaps, ULRSSM, Hybrid methods, DiffusionNet, and Diff3f (Cao et al., 5 Sep 2025).
The body correspondence subsystem is therefore both semantic and geometric: semantics come from diffusion and DINOv2 features, while geometric consistency is imposed by the local energy filter. This suggests that LUIVITON treats semantic features as proposal generators rather than final correspondences.
5. Registration, interpolation, and cloth simulation
After both correspondence sets are predicted, LUIVITON performs two-stage registration (Cao et al., 5 Sep 2025). The first stage aligns SMPL to the garment by optimizing garment-side SMPL parameters 1. The objective is
2
Here, 3 enforces garment-SMPL alignment, 4 and 5 regularize SMPL parameters, 6 promotes smoothness, and 7 penalizes body-garment interpenetrations.
The second stage aligns SMPL+D to the target body by optimizing 8, where 9 is per-vertex displacement and 0 is axis-aligned scale. The objective is
1
In this expression, 2 and 3 are bidirectional point-to-mesh distances, 4 enforces the filtered correspondences, 5 regularizes per-vertex displacements, 6 regularizes scale, and 7 again enforces smoothness.
Garment transfer is then realized by interpolating from the garment-aligned SMPL parameters to the body-aligned SMPL+D parameters. The paper explicitly notes that SMPL is a special case of SMPL+D:
8
A motion sequence is generated via parametric interpolation between these states, and ContourCraft simulates the garment draping over that sequence. The simulation stage is responsible for resolving intersections and preserving fine topology, yielding a physically plausible fit with low interpenetration.
Material customization is exposed through the simulator. Garment stiffness is controlled via Lamé parameters 9 and a bending coefficient. The paper states that increasing 0 and bending yields stiffer fabrics with better shape retention. No explicit cloth energy formulation is provided in the main text.
6. Evaluation, ablations, and limitations
The reported evaluations cover Cloth3D, GarmentCode, and a new benchmark of 8 stylized characters with 10 poses each (Cao et al., 5 Sep 2025). On Cloth3D, using 100 top-trouser pairs and 4 poses each, the metrics are Chamfer Distance (CD), Point-to-Mesh (P2M), and Interpenetration Ratio (IR). For tops, LUIVITON reports CD 1, P2M 2, and IR 3, compared with DrapeNet CD 4 and ISP CD 5. For trousers, LUIVITON reports CD 6, P2M 7, and IR 8, compared with DrapeNet CD 9 and ISP CD 0.
On body-SMPL+D registration, Chamfer Distance 1 is reported as 13.05 for NICP, 4.99 for Diff3f without filtering, 4.49 for Diff3f with filtering, 1.46 for LUIVITON without filtering, and 1.19 for LUIVITON with filtering. The accompanying analysis states that NICP fails on extreme stylization, Diff3f degrades under large pose changes, and that noise filtering and per-vertex displacements are critical.
The ablations identify three sensitivity factors. First, viewpoint choice for body correspondence matters: the best azimuth density uses 36 views, with 18 per elevation, and elevations of 2, reducing MGE by 0.94%–14.98% for elevation choice and 2.33%–7.89% for azimuth coverage. Second, multi-scale diffusion layer features combined with DINOv2 features provide 2.33%–62.03% lower MGE than alternatives. Third, registration quality deteriorates when the noise filter is removed, when per-vertex displacements 3 are removed, or when both are removed; learnable scale 4 improves alignment for extreme proportions.
The system’s limitations are explicit. It does not currently support non-rest garment inputs. It fails on non-humanoid shapes such as airplanes because semantic features do not yield meaningful correspondences. Hard, segmented armor remains difficult because of simulator and physical constraints. Training schedules, hardware details, memory usage, throughput, code availability, exact versions, and some module specifics are not provided in the main paper, and certain definitions are deferred to supplementary material.
A common misconception would be to treat LUIVITON as a pure correspondence model. The reported design instead combines learned correspondence, optimization-based registration, and neural cloth simulation. Another misconception would be to assume that sewing patterns are required; the system is explicitly designed to operate without 2D sewing patterns. A plausible implication is that the method’s practical value derives from this hybridization: correspondence alone establishes transfer structure, but registration and simulation are what make the output geometrically clean and physically plausible.