TexMorph: Texture-rich 3D Morph Benchmark
- TexMorph Benchmark is a texture-rich, multi-view 3D morphing evaluation protocol that assesses geometric and appearance interpolation without pre-aligned meshes.
- It employs metrics like MSE-SSIM, Color Consistency (ΔE), and Edge Integrity to quantitatively measure structural stability, color transition, and silhouette continuity.
- The benchmark challenges methods with non-isometric deformations, diverse capture modalities, and complex textures to foster advancements in realistic 3D morphing.
TexMorph Benchmark is a rigorous, texture-rich, multi-view 3D morphing benchmark introduced in connection with GaussianMorphing for evaluating geometric and appearance interpolation between arbitrary object pairs without requiring pre-aligned meshes or manual correspondences (Li et al., 2 Oct 2025). It is designed around morphing scenarios in which only multi-view RGB images of a source object and a target object are provided, and the evaluated system must reconstruct 3D representations and synthesize intermediate 3D shapes across time. Within the underlying paper, TexMorph functions both as a dataset-and-protocol definition and as a quantitative test suite for semantic-aware object morphing under realistic textures, topology variation, and heterogeneous capture conditions.
1. Scope and defining objectives
The stated purpose of TexMorph is to provide a benchmark that tests both geometric and appearance interpolation between arbitrary object pairs without requiring pre-aligned meshes or manual correspondences (Li et al., 2 Oct 2025). This directly distinguishes it from settings in which homeomorphic mappings, aligned templates, or correspondence supervision are assumed in advance.
A defining emphasis is placed on cross-category and non-isometric deformations, with examples including “dog→lion” and “giraffe→elephant.” The benchmark is therefore not restricted to near-isometric shape interpolation or to morphing within narrowly defined object classes. It is explicitly intended to expose failure modes arising from substantial topology variation, texture complexity, and semantic disparity.
Within the paper’s broader argument, TexMorph serves as the evaluation substrate for a method that combines mesh-guided 3D Gaussian Splatting, topology-aware constraints, and unsupervised semantic correspondence. A plausible implication is that the benchmark was constructed to stress-test methods that must preserve both local texture detail and global semantic coherence rather than optimize for geometry alone.
2. Dataset composition and source material
TexMorph is built from multi-view RGB images only, and there are no ground-truth meshes at test time (Li et al., 2 Oct 2025). The training views may be rendered or captured via BlenderNeRF (synthetic), 3D scans (real-world), or hand-held phone captures (in-the-wild). This composition places synthetic, scanned, and casually captured imagery within a single benchmark definition.
The benchmark spans over ten high-level categories, including animals, fruits, furniture, vehicles, etc. The instance pool includes high-fidelity synthetic models with artist-crafted complex textures, real-world scanned objects (Google Scanned Objects) with natural surface details, and in-the-wild consumer phone photos of everyday items. These source types introduce substantial variability in texture statistics, lighting, image quality, and capture artifacts.
TexMorph also defines a curated set of source–target pairs spanning both intra- and cross-category transformations. These pairs are selected to expose challenges in topology variation, texture complexity, and lighting changes. In consequence, the benchmark couples semantic difficulty with appearance difficulty: the same evaluation setting may require substantial structural adaptation while maintaining perceptually smooth texture transitions.
3. Morphing task formulation
The benchmark defines the morphing task as follows: given only multi-view images of a source object and a target object , a method must reconstruct 3D representations and synthesize a sequence of intermediate 3D shapes for (Li et al., 2 Oct 2025). The formulation is therefore explicitly temporal, with intermediate geometry and appearance evaluated along a morph trajectory rather than only at endpoints.
Evaluation is conducted both qualitatively and quantitatively. The qualitative side concerns visual fidelity and plausibility. The quantitative side is defined as a spatio-temporal evaluation over a discrete set of timestamps .
This protocol encodes a particular view of 3D morphing quality: a successful method must not only reconstruct plausible endpoints, but also produce intermediate states whose structure, color, and silhouettes evolve coherently through time. This suggests that TexMorph is intended to measure temporal regularity and perceptual continuity rather than static reconstruction fidelity alone.
4. Metric design and evaluation criteria
For each intermediate shape , TexMorph computes three metrics and averages them across (Li et al., 2 Oct 2025).
The first metric is Structural Stability (MSE-SSIM). It measures how closely the structural similarity between and the endpoints follows an ideal linear schedule. The ideal curves are defined as
The error is
By construction, lower values indicate closer adherence to the prescribed structural interpolation schedule.
The second metric is Color Consistency (0), based on CIELAB 1 for perceptual color difference:
2
In practice, TexMorph measures source consistency, target consistency, temporal consistency per frame and averages them as
3
The benchmark explicitly states that lower 4 indicates smoother, more faithful color transitions.
The third metric is Edge Integrity (EI), which captures edge fragmentation in the rendered mask sequence via Canny edge detection:
5
The subtraction by one is used to exclude the background. The benchmark states that lower EI implies fewer broken edges and better silhouette continuity.
Taken together, these metrics unify structure, appearance, and silhouette continuity. The benchmark summary explicitly identifies this as a strength: it uses perceptually meaningful color and edge measures in addition to structure, rather than collapsing evaluation to a single geometric criterion.
5. Baselines and reported quantitative outcomes
TexMorph reports per-metric averages over the full TexMorph test suite for four prior baselines and GaussianMorphing (Li et al., 2 Oct 2025). The reported values are:
- DiffMorpher (2D diffusion): MSE-SSIM 6, 7 8, EI 9.
- MorphFlow (2.5D NeRF OT): MSE-SSIM 0, 1 2, EI 3.
- NeuroMorph (3D, no texture): MSE-SSIM 4, 5 not reported, EI 6.
- FreeMorph (2D tuning-free): MSE-SSIM 7, 8 9, EI 0.
- GaussianMorphing (Ours): MSE-SSIM 1, 2 3, EI 4.
On these numbers, GaussianMorphing is the best reported method on all three available axes, and the benchmark summary gives two explicit relative improvements over the best prior baseline, MorphFlow: 5 reduced by approximately 6 and EI reduced by approximately 7. It also states that EI is reduced by 8 relative to the next-best prior, NeuroMorph.
A notable textual inconsistency is present in the source material. The abstract states that GaussianMorphing reduces color consistency error (9) by 0 and EI by 1, whereas the detailed TexMorph benchmark summary reports EI reduced by approximately 2 relative to MorphFlow (Li et al., 2 Oct 2025). Both figures appear in the provided material, and the discrepancy is therefore part of the documented record rather than a matter of external interpretation.
The baseline set also illustrates the benchmark’s modality coverage: 2D diffusion, 2.5D NeRF OT, 3D, no texture, and 2D tuning-free methods are all evaluated under the same protocol. A plausible implication is that TexMorph is intended as a cross-paradigm comparison bed rather than a benchmark tailored only to one representation family.
6. Exposed challenges, strengths, and prospective extensions
The benchmark summary identifies three major challenge clusters (Li et al., 2 Oct 2025). The first is large non-isometric deformations, exemplified by shape changes between disparate species. The second is high-frequency texture patterns, including fur, wood grain, painted surfaces. The third is the presence of diverse capture modalities—rendered, scanned, and casual photos—with different noise, lighting, partial occlusions. These conditions collectively make TexMorph a benchmark for realistic morphing rather than for sanitized interpolation settings.
Its stated strengths are correspondingly broad. TexMorph unifies geometry and appearance evaluation in a single protocol; it uses perceptually meaningful color (3) and edge (EI) measures in addition to structure (SSIM); and it covers a wide spectrum of topology, texture, and capture realism. These properties explain why the benchmark is positioned as a standardizing instrument for future work on texture-rich 3D morphing.
The same summary also enumerates specific areas for improvement and extension. Under scalability, it proposes increasing the number of object pairs and adding articulated or dynamic objects. Under multimodal inputs, it suggests incorporating depth, point clouds, or semantic labels. Under enhanced metrics, it suggests temporal perceptual metrics and 3D point-to-point correspondence accuracy. Under interactive editing tasks, it suggests user-guided control points or region-based morph aims.
Several common misconceptions are implicitly addressed by the benchmark definition itself. TexMorph is not a benchmark that assumes test-time access to meshes, because it uses multi-view RGB images only and has no ground-truth meshes at test time. It is not restricted to appearance transfer, because it explicitly evaluates geometry through structural stability and silhouette continuity. It is also not confined to within-category interpolation, since the curated pairs span both intra- and cross-category transformations.
By defining this benchmark and evaluation protocol, TexMorph is presented as a mechanism that enables the community to quantitatively compare future 3D morphing methods under realistic, texture-rich scenarios and paves the way for standardized improvements in both geometry and appearance interpolation (Li et al., 2 Oct 2025).