SeamCrafter: Autoregressive Mesh Seam Generation
- SeamCrafter is an autoregressive system that generates artist-quality mesh seams for UV unwrapping, balancing distortion and fragmentation.
- It employs a dual-branch point cloud encoder and a GPT-style transformer to predict quantized endpoints, ensuring geometric fidelity and topological consistency.
- The approach extends into robotic garment manipulation and structured reconstruction by treating seams as compact carriers of critical structural and control information.
SeamCrafter is an autoregressive system for generating artist-quality mesh seams for UV unwrapping and texture mapping. It was introduced as a response to a persistent trade-off in seam placement: too few or poorly aligned seams induce UV distortion, whereas too many or poorly organized seams fragment the UV atlas into numerous small islands (Xu et al., 25 Sep 2025). In adjacent literature, the same name also functions as a conceptual label for seam-driven garment state estimation and configuration alignment in robotics, and as a broader seam-centric framework derived from seam-informed garment unfolding methods (Huang et al., 13 Jun 2026, Huang et al., 2024). Across these usages, the common premise is that seams are not merely geometric boundaries: they are compact carriers of topology, structure, and control-relevant information.
1. Definition and scope
In its primary graphics sense, SeamCrafter targets seam generation on 3D meshes. Seams are represented as cuts along which a surface is flattened into UV islands, and the central optimization target is a balance between low distortion and low fragmentation (Xu et al., 25 Sep 2025). The method is explicitly conditioned on point cloud inputs and trained to produce seam sets that preserve topological consistency and visual fidelity while improving UV parameterization quality.
The same seam-centric logic appears in garment perception and manipulation. In robotic configuration alignment, seams are treated as structural cues that encode how garment parts connect; partial seam observations are lifted into a topology-encoded structural skeleton graph for real-time state estimation (Huang et al., 13 Jun 2026). In dual-arm T-shirt unfolding, seams and seam crossings constrain the grasp search space and guide the selection of dual-arm grasps that reliably flatten and orient a shirt with few episode steps (Huang et al., 2024). This broader literature suggests that the term SeamCrafter denotes not a single invariant architecture, but a family of methods that elevate seams to first-class computational primitives.
A related but distinct research direction is seam-aware garment reconstruction. ReWeaver reconstructs seams, panels, and their connectivities in both 2D UV space and 3D space from sparse multi-view RGB, yielding structured 2D–3D garment representations suitable for physical simulation and robotic manipulation (Li et al., 23 Jan 2026). AutoSew, by contrast, predicts stitch correspondences directly from 2D pattern contours using geometry-only graph matching and differentiable optimal transport (RÃos-Navarro et al., 25 Feb 2026). Together, these systems situate SeamCrafter within a larger seam-centric ecosystem spanning UV unwrapping, garment structure recovery, and robotic control.
2. Core formulation for mesh seam generation
SeamCrafter represents seams as ordered edge segments in 3D. Each segment contains the 3D coordinates of its two endpoints, and the full seam set is serialized as a hierarchical autoregressive token sequence over segments, endpoints, and coordinates (Xu et al., 25 Sep 2025). Each 3D endpoint coordinate is quantized into 1024 discrete bins, ordered in a yzx scheme; endpoints within a segment are sorted lexicographically, and segments are ordered by the first endpoint. This quantized sequence becomes the action space of the autoregressive policy.
Conditioning is provided by a dual-branch point-cloud encoder. One branch samples points on the vertex–edge skeleton of the mesh to capture connectivity structure, producing ; the other samples points uniformly over the surface to capture geometric coverage, producing (Xu et al., 25 Sep 2025). These two point sets are encoded by separate VecSet encoders with identical architecture but distinct parameters, and their latent tokens are concatenated into a single conditioning sequence. The geometry branch is initialized with pretrained weights and kept fixed, whereas the topology branch is trained jointly with the decoder.
The decoder is a GPT-style hourglass transformer. It predicts quantized coordinate tokens under causal masking, using a hierarchical sequence abstraction in which coordinate codes are downsampled by factor 3 at the coordinate level and factor 2 at the segment–endpoint level, then upsampled by symmetric factors (Xu et al., 25 Sep 2025). Shape tokens are injected by cross-attention in the first layer of each transformer level. The model has 24 layers, with a pattern of three self-attention layers followed by one cross-attention layer conditioned on the fused point-cloud tokens.
Validity is not enforced purely by token-level modeling. After prediction, endpoints are snapped to nearest mesh vertices, and a shortest geodesic path on the mesh graph connects endpoints, marking all edges along the path as seam edges (Xu et al., 25 Sep 2025). This projection-to-graph stage ensures that predicted seams respect mesh connectivity, avoid non-manifold edge selections, and remain actionable for downstream UV unwrapping.
3. Preference learning, evaluation, and reported performance
SeamCrafter’s post-training stage is organized around a seam-evaluation framework that scores candidate seam sets primarily by UV distortion and fragmentation density (Xu et al., 25 Sep 2025). Distortion is computed in UV space using an area-normalized measure with per-triangle singular values of the deformation gradient, whereas fragmentation density is the number of UV islands induced by the seams. For each input point cloud, the pretrained model generates five candidate seam sets, and pairwise preferences are assigned only under strict Pareto dominance:
This procedure yields 4,000 automatically curated preference pairs. Direct Preference Optimization (DPO) is then used to align the autoregressive policy to these pairwise preferences while regularizing against a frozen reference policy. The paper describes DPO as a preference-based RL alignment method without an explicit reward model or policy gradient rollouts (Xu et al., 25 Sep 2025).
Pretraining uses a corpus of approximately 700K meshes curated primarily from Objaverse, Objaverse-XL, and 3D-FUTURE, with y-axis rotations in 10° increments, uniform scaling in , and translations in (Xu et al., 25 Sep 2025). Each branch samples points and outputs tokens of dimension . Pretraining runs for 200K steps on 64 NVIDIA H20 GPUs with batch size 128; DPO post-training uses 16 H20 GPUs for 2,500 steps at learning rate .
The reported benchmark results show consistent gains from DPO. In the table below, each entry is reported as Distortion / Fragments / Runtime.
| Benchmark | w/o DPO | w DPO |
|---|---|---|
| Toys4k | 1.41 / 11.57 / 51.43 s | 1.39 / 7.28 / 50.67 s |
| ShapeNet | 34.19 / 100.78 / 17.46 s | 32.60 / 91.28 / 22.07 s |
| FAM-benchmark | 11.85 / 27.33 / 57.13 s | 10.07 / 10.05 / 59.66 s |
| AIGC-100 | 13.65 / 45.12 / 8.73 s | 10.63 / 33.72 / 13.55 s |
The ablation results clarify the role of the preference construction. Using only Distortion for preferences reduces distortion but causes excessive fragmentation, exemplified by 127.20 islands; using only Fragmentation reduces island count but causes high distortion, exemplified by 72.80 distortion (Xu et al., 25 Sep 2025). Joint DPO over both metrics yields the intended trade-off.
4. Relation to other seam-generation paradigms
A central comparison point is MeshTailor, described as the first mesh-native generative framework for synthesizing edge-aligned seams on 3D surfaces (Ma et al., 28 Mar 2026). In MeshTailor’s related work, SeamCrafter is mentioned as a reinforcement learning approach that refines seam quality; MeshTailor differs by traversing the existing mesh graph directly with an autoregressive pointer layer and therefore avoiding projection from Euclidean coordinates back to edges (Ma et al., 28 Mar 2026). Its ChainingSeams hierarchy orders loop cuts before local detail seams, and its neighborhood-masked pointer mechanism enforces edge-aligned moves by construction.
This contrast is methodological rather than merely nominal. SeamCrafter predicts quantized endpoint coordinates and then restores mesh validity through nearest-vertex snapping and shortest geodesic routing (Xu et al., 25 Sep 2025). MeshTailor instead treats the seam as a vertex walk on the mesh graph, with a masked local pointer over one-ring neighborhoods and a dual-stream encoder that fuses graph connectivity with pretrained point-cloud semantics (Ma et al., 28 Mar 2026). MeshTailor reports that predicting coordinates expands the candidate space and that snapping introduces jagged, misaligned boundaries; this is precisely the failure mode its mesh-native design is intended to avoid.
Another adjacent line of work is CraftMesh, which addresses seamless fusion after generative mesh editing rather than seam placement for UV charting. CraftMesh decomposes editing into 2D reference image manipulation, region-specific 3D mesh generation, and Poisson Geometric Fusion plus Poisson Texture Harmonization to eliminate visible seams in geometry and appearance (Jincheng et al., 17 Sep 2025). Although this is not a seam-generation method in the UV-unwrapping sense, it shares the broader concern of turning seams from artifacts into controlled, mathematically regularized interfaces.
5. SeamCrafter in robotic perception and manipulation
In robotic garment manipulation, SeamCrafter is used conceptually for a seam-driven system that converts partial seam observations into control-oriented garment state estimates (Huang et al., 13 Jun 2026). The sensing stage uses 0 RGB-D cameras that provide RGB images 1 and depth maps 2, with SAM2 for garment segmentation and a CNN-based seam segment detector that yields 2D line segments with visual categories. Masked depth pixels are fused into a 3D garment point cloud 3, and segment endpoints are back-projected into 3D seam segments 4.
The Seam-to-Graph network then maps 5, 6, and a garment-type template skeleton graph 7 to a topology-encoded structural skeleton graph in the world frame (Huang et al., 13 Jun 2026). Its point cloud branch is PointNet-style, producing global and local garment features; its seam branch encodes each 3D seam segment using endpoints, midpoint, normalized direction, length, and a learned embedding of the visual category, followed by a transformer over segment features. Node features are grounded first in point-cloud neighborhoods by cross-attention, then in seam segment features by a second cross-attention, and finally regularized by topology-aware message passing with a graph attention network. A residual regression head predicts vertex displacements relative to the template skeleton.
A topology-constrained post-refinement step attracts predicted skeleton edges toward seam endpoints and nearby surface points while preserving predicted edge lengths (Huang et al., 13 Jun 2026). The refined graph drives a deformation-aware hierarchical visual servoing controller that unfolds skeleton regions onto the platen plane, estimates per-region planar pose errors in 8 via Kabsch, and applies an inner PBVS loop plus an outer region coordinator to generate bimanual end-effector commands. The system runs at 9.9 Hz on an NVIDIA RTX 5090 and is implemented on a bimanual robot system with four RGB-D cameras, SAM2 segmentation, and ROS2 integration (Huang et al., 13 Jun 2026).
The experimental target is screen-printing preparation: loading a garment onto a platen and aligning it precisely. Real-robot experiments show that the robot achieves human-level alignment accuracy with reduced variance in alignment error and remains robust across short-sleeve and long-sleeve T-shirts not seen during training (Huang et al., 13 Jun 2026). Synthetic ablations over 11k unseen samples report 9 and MV2E distance 0 for the full model, compared with 1 and 2 without seams, which isolates the structural value of seam information (Huang et al., 13 Jun 2026).
A precursor to this formulation is SIS, the Seam-Informed Strategy for T-shirt unfolding (Huang et al., 2024). SIS decomposes the policy into SFEM, which detects oriented seam line segments and seam crossing segments from RGB images, and DMIM, which scores candidate dual-arm grasp pairs by seam-type combination using a decision matrix initialized from human demonstrations and updated by robot executions. Under the joint success threshold 3 and 4, success rates over 20 trials progress from 20% at step 1 to 85%, 90%, 95%, and 95% at steps 2–5 (Huang et al., 2024). The paper explicitly proposes adapting SIS into a broader framework called SeamCrafter.
6. Structured garment reconstruction, stitching prediction, and limitations
SeamCrafter also sits near methods that reconstruct or operationalize seams in structured garment pipelines. ReWeaver predicts seams as 3D curves, panels as 3D patches, and their connectivities in both 2D UV space and 3D from as few as four RGB views (Li et al., 23 Jan 2026). It produces flattened panels with edge loops, explicit seam connectivities, and 3D geometry aligned to images, and it is trained on GCD-TS, a synthetic dataset with over 100,000 samples comprising four-view RGB, textured human bodies, 3D garment geometries, and annotated sewing patterns (Li et al., 23 Jan 2026). On topology and geometry metrics, ReWeaver reports 5, 6, 7, 8, and 9, outperforming an AIpparel-MV baseline (Li et al., 23 Jan 2026).
AutoSew addresses a different stage of the pipeline: stitch prediction from 2D pattern geometry. It formulates sewing as a graph matching problem over contour edges, uses a GNN to derive contextual edge embeddings, and employs an entropic optimal transport solver with length-capacity marginals so that one long edge can receive mass from multiple short edges (RÃos-Navarro et al., 25 Feb 2026). On the extended M-E.GarmentCodeData with 18,003 patterns and realistic multi-edge annotations, AutoSew reports an edge-level matching F1 of approximately 96% and full-garment assembly success of approximately 73.3% (RÃos-Navarro et al., 25 Feb 2026). This geometry-only formulation is explicitly designed to avoid dependence on panel labels or handcrafted heuristics.
In physical manufacturing, robotic seam creation has also been framed as a three-stage pipeline of pose estimation, temporary joining, and closed-loop visual servoing (Ajith et al., 28 Feb 2025). The system uses Sewbo’s water-soluble posing agent to temporarily stiffen garments, ultrasonic spot welding to join panels, and a sewing controller that keeps the stitch path within seam tolerance from the fabric edge. On denim, the reported seam error metric is 1.37 mm for open-loop without disturbance, 9.775 mm for open-loop with disturbance, 1.585 mm for closed-loop without disturbance, and 3.11 mm for closed-loop with disturbance, relative to an industry standard of seam error below 3 mm (Ajith et al., 28 Feb 2025). This is not SeamCrafter in the naming sense, but it extends seam-centric computation into industrial sewing.
Across these domains, the principal limitations are consistent. The mesh-seam SeamCrafter assumes reasonably manifold meshes and can degrade on very thin structures or extremely high-genus shapes; projection and geodesic routing may also suffer on poor triangulations (Xu et al., 25 Sep 2025). MeshTailor is tuned for meshes up to about 2,000 triangles and remains challenged by hair-like geometry, thin sheets, and autoregressive decoding errors (Ma et al., 28 Mar 2026). Robotic seam-to-graph systems are sensitive to highly wrinkled or occluded seams, garments without visible seam lines, and calibration or depth artifacts (Huang et al., 13 Jun 2026). ReWeaver notes remaining sim-to-real gaps and limited availability of complex, photorealistic 3D garments, even with GCD-TS (Li et al., 23 Jan 2026).
The cumulative picture is that SeamCrafter names a seam-first computational strategy whose clearest formalization is the DPO-aligned autoregressive seam generator for UV unwrapping (Xu et al., 25 Sep 2025), but whose conceptual reach extends into robotic manipulation, structured garment reconstruction, and geometric garment assembly (Huang et al., 13 Jun 2026, Huang et al., 2024, Li et al., 23 Jan 2026, RÃos-Navarro et al., 25 Feb 2026). In each case, seam information acts as a compressed structural prior: for meshes, it defines how surfaces should be cut; for garments, it encodes topology, correspondence, and control-relevant geometry.