---
title: 'SeamCrafter: Autoregressive Mesh Seam Generation'
url: https://www.emergentmind.com/topics/seamcrafter
type: topic
---

# SeamCrafter: Autoregressive Mesh Seam Generation

SeamCrafter is an autoregressive system for generating artist-quality mesh seams for UV unwrapping and texture mapping. It was introduced as a response to a persistent trade-off in seam placement: too few or poorly aligned seams induce UV distortion, whereas too many or poorly organized seams fragment the UV atlas into numerous small islands [2509.20725]. In adjacent literature, the same name also functions as a conceptual label for seam-driven garment state estimation and configuration alignment in robotics, and as a broader seam-centric framework derived from seam-informed garment unfolding methods [2606.15171] [2409.06990]. Across these usages, the common premise is that seams are not merely geometric boundaries: they are compact carriers of topology, structure, and control-relevant information.

## 1. Definition and scope

In its primary graphics sense, SeamCrafter targets seam generation on 3D meshes. Seams are represented as cuts along which a surface is flattened into UV islands, and the central optimization target is a balance between low distortion and low fragmentation [2509.20725]. The method is explicitly conditioned on point cloud inputs and trained to produce seam sets that preserve topological consistency and visual fidelity while improving UV parameterization quality.

The same seam-centric logic appears in garment perception and manipulation. In robotic configuration alignment, seams are treated as structural cues that encode how garment parts connect; partial seam observations are lifted into a topology-encoded structural skeleton graph for real-time state estimation [2606.15171]. In dual-arm T-shirt unfolding, seams and seam crossings constrain the grasp search space and guide the selection of dual-arm grasps that reliably flatten and orient a shirt with few episode steps [2409.06990]. This broader literature suggests that the term *SeamCrafter* denotes not a single invariant architecture, but a family of methods that elevate seams to first-class computational primitives.

A related but distinct research direction is seam-aware garment reconstruction. ReWeaver reconstructs seams, panels, and their connectivities in both 2D UV space and 3D space from sparse multi-view RGB, yielding structured 2D–3D garment representations suitable for physical simulation and robotic manipulation [2601.16672]. AutoSew, by contrast, predicts stitch correspondences directly from 2D pattern contours using geometry-only graph matching and differentiable optimal transport [2602.22052]. Together, these systems situate SeamCrafter within a larger seam-centric ecosystem spanning UV unwrapping, garment structure recovery, and robotic control.

## 2. Core formulation for mesh seam generation

SeamCrafter represents seams as ordered edge segments in 3D. Each segment $s_i = [p_{i,1}, p_{i,2}] \in \mathbb{R}^{2 \times 3}$ contains the 3D coordinates of its two endpoints, and the full seam set is serialized as a hierarchical autoregressive token sequence over segments, endpoints, and coordinates [2509.20725]. Each 3D endpoint coordinate is quantized into 1024 discrete bins, ordered in a yzx scheme; endpoints within a segment are sorted lexicographically, and segments are ordered by the first endpoint. This quantized sequence becomes the action space of the autoregressive policy.

Conditioning is provided by a dual-branch point-cloud encoder. One branch samples points on the vertex–edge skeleton of the mesh to capture connectivity structure, producing $P_t^\mathcal{M} \in \mathbb{R}^{N_t \times 3}$; the other samples points uniformly over the surface to capture geometric coverage, producing $P_g^\mathcal{M} \in \mathbb{R}^{N_g \times 3}$ [2509.20725]. These two point sets are encoded by separate VecSet encoders with identical architecture but distinct parameters, and their latent tokens are concatenated into a single conditioning sequence. The geometry branch is initialized with pretrained weights and kept fixed, whereas the topology branch is trained jointly with the decoder.

The decoder is a GPT-style hourglass transformer. It predicts quantized coordinate tokens under causal masking, using a hierarchical sequence abstraction in which coordinate codes are downsampled by factor 3 at the coordinate level and factor 2 at the segment–endpoint level, then upsampled by symmetric factors [2509.20725]. Shape tokens are injected by cross-attention in the first layer of each transformer level. The model has 24 layers, with a pattern of three self-attention layers followed by one cross-attention layer conditioned on the fused point-cloud tokens.

Validity is not enforced purely by token-level modeling. After prediction, endpoints are snapped to nearest mesh vertices, and a shortest geodesic path on the mesh graph connects endpoints, marking all edges along the path as seam edges [2509.20725]. This projection-to-graph stage ensures that predicted seams respect mesh connectivity, avoid non-manifold edge selections, and remain actionable for downstream UV unwrapping.

## 3. Preference learning, evaluation, and reported performance

SeamCrafter’s post-training stage is organized around a seam-evaluation framework that scores candidate seam sets primarily by UV distortion and fragmentation density [2509.20725]. Distortion is computed in UV space using an area-normalized measure with per-triangle singular values of the deformation gradient, whereas fragmentation density is the number of UV islands induced by the seams. For each input point cloud, the pretrained model generates five candidate seam sets, and pairwise preferences are assigned only under strict Pareto dominance:

$$
S_\mathcal{M}^i \succ S_\mathcal{M}^j
\iff
\mathrm{Distortion}(S_\mathcal{M}^i) < \mathrm{Distortion}(S_\mathcal{M}^j)
\text{ and }
\mathrm{Density}(S_\mathcal{M}^i) < \mathrm{Density}(S_\mathcal{M}^j).
$$

This procedure yields 4,000 automatically curated preference pairs. Direct Preference Optimization (DPO) is then used to align the autoregressive policy to these pairwise preferences while regularizing against a frozen reference policy. The paper describes DPO as a preference-based RL alignment method without an explicit reward model or policy gradient rollouts [2509.20725].

Pretraining uses a corpus of approximately 700K meshes curated primarily from Objaverse, Objaverse-XL, and 3D-FUTURE, with y-axis rotations in 10° increments, uniform scaling in $[0.9, 1.1]$, and translations in $[-0.1, 0.1]$ [2509.20725]. Each branch samples $N_g = N_t = 30{,}720$ points and outputs $\ell = 3{,}072$ tokens of dimension $d = 1{,}024$. Pretraining runs for 200K steps on 64 NVIDIA H20 GPUs with batch size 128; DPO post-training uses 16 H20 GPUs for 2,500 steps at learning rate $1 \times 10^{-6}$.

The reported benchmark results show consistent gains from DPO. In the table below, each entry is reported as Distortion / Fragments / Runtime.

| Benchmark | w/o DPO | w DPO |
|---|---:|---:|
| Toys4k | 1.41 / 11.57 / 51.43 s | 1.39 / 7.28 / 50.67 s |
| ShapeNet | 34.19 / 100.78 / 17.46 s | 32.60 / 91.28 / 22.07 s |
| FAM-benchmark | 11.85 / 27.33 / 57.13 s | 10.07 / 10.05 / 59.66 s |
| AIGC-100 | 13.65 / 45.12 / 8.73 s | 10.63 / 33.72 / 13.55 s |

The ablation results clarify the role of the preference construction. Using only Distortion for preferences reduces distortion but causes excessive fragmentation, exemplified by 127.20 islands; using only Fragmentation reduces island count but causes high distortion, exemplified by 72.80 distortion [2509.20725]. Joint DPO over both metrics yields the intended trade-off.

## 4. Relation to other seam-generation paradigms

A central comparison point is MeshTailor, described as the first mesh-native generative framework for synthesizing edge-aligned seams on 3D surfaces [2603.27309]. In MeshTailor’s related work, SeamCrafter is mentioned as a reinforcement learning approach that refines seam quality; MeshTailor differs by traversing the existing mesh graph directly with an autoregressive pointer layer and therefore avoiding projection from Euclidean coordinates back to edges [2603.27309]. Its ChainingSeams hierarchy orders loop cuts before local detail seams, and its neighborhood-masked pointer mechanism enforces edge-aligned moves by construction.

This contrast is methodological rather than merely nominal. SeamCrafter predicts quantized endpoint coordinates and then restores mesh validity through nearest-vertex snapping and shortest geodesic routing [2509.20725]. MeshTailor instead treats the seam as a vertex walk on the mesh graph, with a masked local pointer over one-ring neighborhoods and a dual-stream encoder that fuses graph connectivity with pretrained point-cloud semantics [2603.27309]. MeshTailor reports that predicting coordinates expands the candidate space and that snapping introduces jagged, misaligned boundaries; this is precisely the failure mode its mesh-native design is intended to avoid.

Another adjacent line of work is CraftMesh, which addresses seamless fusion after generative mesh editing rather than seam placement for UV charting. CraftMesh decomposes editing into 2D reference image manipulation, region-specific 3D mesh generation, and Poisson Geometric Fusion plus Poisson Texture Harmonization to eliminate visible seams in geometry and appearance [2509.13688]. Although this is not a seam-generation method in the UV-unwrapping sense, it shares the broader concern of turning seams from artifacts into controlled, mathematically regularized interfaces.

## 5. SeamCrafter in robotic perception and manipulation

In robotic garment manipulation, *SeamCrafter* is used conceptually for a seam-driven system that converts partial seam observations into control-oriented garment state estimates [2606.15171]. The sensing stage uses $K$ RGB-D cameras that provide RGB images $I_k$ and depth maps $D_k$, with SAM2 for garment segmentation and a CNN-based seam segment detector that yields 2D line segments with visual categories. Masked depth pixels are fused into a 3D garment point cloud $P$, and segment endpoints are back-projected into 3D seam segments $S = \{(p^1, p^2, \text{category})\}$.

The Seam-to-Graph network then maps $P$, $S$, and a garment-type template skeleton graph $G_{\text{temp}} = (V, E, X_{\text{temp}})$ to a topology-encoded structural skeleton graph in the world frame [2606.15171]. Its point cloud branch is PointNet-style, producing global and local garment features; its seam branch encodes each 3D seam segment using endpoints, midpoint, normalized direction, length, and a learned embedding of the visual category, followed by a transformer over segment features. Node features are grounded first in point-cloud neighborhoods by cross-attention, then in seam segment features by a second cross-attention, and finally regularized by topology-aware message passing with a graph attention network. A residual regression head predicts vertex displacements relative to the template skeleton.

A topology-constrained post-refinement step attracts predicted skeleton edges toward seam endpoints and nearby surface points while preserving predicted edge lengths [2606.15171]. The refined graph drives a deformation-aware hierarchical visual servoing controller that unfolds skeleton regions onto the platen plane, estimates per-region planar pose errors in $SE(2)$ via Kabsch, and applies an inner PBVS loop plus an outer region coordinator to generate bimanual end-effector commands. The system runs at 9.9 Hz on an NVIDIA RTX 5090 and is implemented on a bimanual robot system with four RGB-D cameras, SAM2 segmentation, and ROS2 integration [2606.15171].

The experimental target is screen-printing preparation: loading a garment onto a platen and aligning it precisely. Real-robot experiments show that the robot achieves human-level alignment accuracy with reduced variance in alignment error and remains robust across short-sleeve and long-sleeve T-shirts not seen during training [2606.15171]. Synthetic ablations over 11k unseen samples report $L_{\text{vtx}} = 0.3357 \times 10^{-2}$ and MV2E distance $= 2.6854 \times 10^{-2}$ for the full model, compared with $0.6168 \times 10^{-2}$ and $3.9506 \times 10^{-2}$ without seams, which isolates the structural value of seam information [2606.15171].

A precursor to this formulation is SIS, the Seam-Informed Strategy for T-shirt unfolding [2409.06990]. SIS decomposes the policy into SFEM, which detects oriented seam line segments and seam crossing segments from RGB images, and DMIM, which scores candidate dual-arm grasp pairs by seam-type combination using a decision matrix initialized from human demonstrations and updated by robot executions. Under the joint success threshold $\mathrm{ncov} \ge 0.85$ and $\mathrm{IoU} \ge 0.85$, success rates over 20 trials progress from 20% at step 1 to 85%, 90%, 95%, and 95% at steps 2–5 [2409.06990]. The paper explicitly proposes adapting SIS into a broader framework called SeamCrafter.

## 6. Structured garment reconstruction, stitching prediction, and limitations

SeamCrafter also sits near methods that reconstruct or operationalize seams in structured garment pipelines. ReWeaver predicts seams as 3D curves, panels as 3D patches, and their connectivities in both 2D UV space and 3D from as few as four RGB views [2601.16672]. It produces flattened panels with edge loops, explicit seam connectivities, and 3D geometry aligned to images, and it is trained on GCD-TS, a synthetic dataset with over 100,000 samples comprising four-view RGB, textured human bodies, 3D garment geometries, and annotated sewing patterns [2601.16672]. On topology and geometry metrics, ReWeaver reports $\mathrm{Acc}_p = 0.9210$, $\mathrm{Acc}_e = 0.7175$, $\mathrm{Acc}_o = 0.6608$, $\mathrm{CD}_e = 0.0391$, and $\mathrm{IoU} = 0.8221$, outperforming an AIpparel-MV baseline [2601.16672].

AutoSew addresses a different stage of the pipeline: stitch prediction from 2D pattern geometry. It formulates sewing as a graph matching problem over contour edges, uses a GNN to derive contextual edge embeddings, and employs an entropic optimal transport solver with length-capacity marginals so that one long edge can receive mass from multiple short edges [2602.22052]. On the extended M-E.GarmentCodeData with 18,003 patterns and realistic multi-edge annotations, AutoSew reports an edge-level matching F1 of approximately 96% and full-garment assembly success of approximately 73.3% [2602.22052]. This geometry-only formulation is explicitly designed to avoid dependence on panel labels or handcrafted heuristics.

In physical manufacturing, robotic seam creation has also been framed as a three-stage pipeline of pose estimation, temporary joining, and closed-loop visual servoing [2503.00249]. The system uses Sewbo’s water-soluble posing agent to temporarily stiffen garments, ultrasonic spot welding to join panels, and a sewing controller that keeps the stitch path within seam tolerance from the fabric edge. On denim, the reported seam error metric is 1.37 mm for open-loop without disturbance, 9.775 mm for open-loop with disturbance, 1.585 mm for closed-loop without disturbance, and 3.11 mm for closed-loop with disturbance, relative to an industry standard of seam error below 3 mm [2503.00249]. This is not SeamCrafter in the naming sense, but it extends seam-centric computation into industrial sewing.

Across these domains, the principal limitations are consistent. The mesh-seam SeamCrafter assumes reasonably manifold meshes and can degrade on very thin structures or extremely high-genus shapes; projection and geodesic routing may also suffer on poor triangulations [2509.20725]. MeshTailor is tuned for meshes up to about 2,000 triangles and remains challenged by hair-like geometry, thin sheets, and autoregressive decoding errors [2603.27309]. Robotic seam-to-graph systems are sensitive to highly wrinkled or occluded seams, garments without visible seam lines, and calibration or depth artifacts [2606.15171]. ReWeaver notes remaining sim-to-real gaps and limited availability of complex, photorealistic 3D garments, even with GCD-TS [2601.16672].

The cumulative picture is that SeamCrafter names a seam-first computational strategy whose clearest formalization is the DPO-aligned autoregressive seam generator for UV unwrapping [2509.20725], but whose conceptual reach extends into robotic manipulation, structured garment reconstruction, and geometric garment assembly [2606.15171] [2409.06990] [2601.16672] [2602.22052]. In each case, seam information acts as a compressed structural prior: for meshes, it defines how surfaces should be cut; for garments, it encodes topology, correspondence, and control-relevant geometry.

Source: https://www.emergentmind.com/topics/seamcrafter