Papers
Topics
Authors
Recent
Search
2000 character limit reached

MonoRelief V2: 2.5D Relief Recovery Model

Updated 9 July 2026
  • The paper presents MonoRelief V2 as an end-to-end model for predicting a detailed 2.5D relief, directly mapping RGB images to absolute depth maps.
  • It leverages a dual-stage training process using both pseudo-real and real relief datasets, reducing mean depth error from 21.77% to 14.03% and improving normal angular accuracy.
  • The model integrates depth and implicit normal supervision within a composite loss, ensuring high-fidelity relief reconstruction suitable for digital visualization and fabrication.

MonoRelief V2 is an end-to-end model for directly recovering 2.5D reliefs from single images under complex material and illumination variations. In the paper’s formulation, “2.5D relief recovery” denotes prediction of a height field, or depth map, for a relief sculpture from a single RGB image, so that the relief can be treated as a function z(x,y)z(x,y) over the image plane. Given an input image I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}, the model outputs an absolute depth map D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}, scaled in physical units relative to the relief thickness, while a consistent normal field is obtained implicitly by differentiating D^\hat{D}. The intended reconstruction quality is sufficient for CNC machining, 3D printing, and digital visualization (Zhang et al., 27 Aug 2025).

1. Problem definition and representational scope

MonoRelief V2 addresses a specialized variant of monocular geometry recovery: relief reconstruction rather than general monocular depth estimation. The distinction is central. Relief recovery focuses on shallow depth ranges, but it places unusually strong emphasis on fine local geometry, including curvature, ridges, and engravings. By contrast, general monocular depth estimators usually target global scene geometry over diverse scenes and often produce coarse geometry with good occlusion boundaries but limited high-frequency detail. Surface normal estimation occupies a complementary role, because normals emphasize local shape and are more sensitive to fine structure than depth alone (Zhang et al., 27 Aug 2025).

The model is trained as a depth estimator, but normals are structurally embedded throughout the pipeline. Depth labels, both pseudo and real, are generated by fusing depth and normal predictions; supervision uses a composite loss that combines depth regression with a normal-consistency term; and evaluation is reported for both depth and derived normals. Formally, the network predicts depth as D^=fθ(I)\hat{D} = f_\theta(I), and the normal field is derived as

n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).

A common misconception is that MonoRelief V2 is primarily a normal estimator. The paper instead defines depth as the primary output and normals as an implicit representation used for supervision and evaluation (Zhang et al., 27 Aug 2025).

The target inputs span real photographs of architectural reliefs, heritage artifacts, and plaques; AI-generated “relief-style” images via Flux.1; and rendered reliefs with varied BRDFs. Material diversity includes clay, copper, gold, jade, marble, obsidian, plaster, silver, steel, as well as gypsum, masonry, metal, wood, and synthetic renderings with leather, concrete, lacquer, bamboo, and porcelain. Illumination diversity includes indoor and outdoor HDRI environments, varying directions and intensities, strong shadows, soft lighting, and specular highlights. The paper identifies four domain-specific challenges: material-plus-lighting coupling, geometry–texture ambiguity, depth ambiguity from single-view shading, and a domain gap between relief imagery and conventional indoor or outdoor depth-estimation corpora (Zhang et al., 27 Aug 2025).

2. Lineage, prior systems, and methodological position

MonoRelief V2 is presented explicitly as a successor to MonoRelief V1. V1 uses a three-stage pipeline: a synthetic relief dataset with ground-truth normal labels, relief-domain fine-tuning of Omnidata for monocular normal prediction, and depth reconstruction by depth-constrained normal integration. That design established a relief-specific reconstruction workflow, but it remained synthetic-only, multi-stage, and relatively fragile on real photographs and complex materials. On the new benchmark, V1 reports mean depth error 21.77%21.77\% and mean normal angular error about 16.21∘16.21^\circ (Zhang et al., 27 Aug 2025).

V2 changes this formulation at three levels. First, it becomes end-to-end: a single model maps image to depth, with normals derived from the predicted depth. Second, it adds real supervision through two new data sources: approximately 15,000 pseudo-real images generated by Flux.1 with depth pseudo-labels constructed by fusing DepthAnything V2 and MonoRelief V1 predictions, and a small real-world dataset of 800 reliefs captured through multi-view imaging, reconstruction, and detail refinement. Third, it adopts progressive training: full fine-tuning on pseudo-real data followed by LoRA adaptation on real data. These changes reduce mean depth error from 21.77%21.77\% to 14.03%14.03\% and mean normal angular error from I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}0 to I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}1, while also increasing PSNR and SSIM for both depth and normals (Zhang et al., 27 Aug 2025).

Architecturally, V2 is based on DepthAnything V2’s ViT+DPT design: a DINOv2-pretrained Vision Transformer encoder with a DPT-style dense prediction decoder (Yang et al., 2024). MonoRelief V2 uses the ViT-base encoder and decoder configuration from that family, yielding a model of roughly 100M parameters and avoiding explicit normal integration during inference. Relative to V1, the backbone shifts from Omnidata ViT plus separate integration to a single DepthAnything V2-style encoder–decoder; the loss shifts from purely normal prediction plus post-processing to a joint depth-and-normal objective; and domain adaptation shifts from standard fine-tuning on synthetic data only to LoRA fine-tuning on real data (Zhang et al., 27 Aug 2025).

The paper also situates MonoRelief V2 against general monocular depth and normal estimators. DepthAnything V2, DepthPro, and MoGe V2 are described as strong on global scene geometry but inadequate for relief-specific fine detail and material-dependent shading. General normal methods including Omnidata, DSINE, StableNormal, Lotus, Marigold-E2E-FT, and GenPercept are reported to preserve sharp edges but often lose shallow relief detail or misread material-driven appearance as geometry. MonoRelief V2 is characterized as the first end-to-end, relief-focused monocular 2.5D model trained on real relief data and as a state-of-the-art system specifically for relief depth and normal recovery (Zhang et al., 27 Aug 2025).

3. Pseudo-real data generation and real-data acquisition

A central feature of MonoRelief V2 is its supervision pipeline. Because large-scale real-world paired relief image–depth datasets are difficult to acquire, the paper constructs two complementary datasets: a pseudo-real corpus at scale and a smaller high-fidelity real-world corpus (Zhang et al., 27 Aug 2025).

The pseudo-real dataset is generated with Flux.1-schnell, an open-source text-to-image diffusion model. Prompting combines object categories—animals, humans, flora, architecture, landscapes, and miscellaneous cases—with nine target materials: clay, copper, gold, jade, marble, obsidian, plaster, silver, and steel. The prompt template is “A/an <material> relief of a/an <object>”. Images are generated at I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}2 resolution with four denoising steps and random seeds for diversity. From an initial output of about 30,000 images, the authors perform dual-blind expert review by two domain experts and filter samples based on shape fidelity and detail richness, yielding a curated pseudo-real dataset of about 15,000 images (Zhang et al., 27 Aug 2025).

Pseudo-label generation is not a direct depth-regression procedure. The pipeline first predicts relative depth I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}3 using DepthAnything V2 and detailed normals I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}4 using MonoRelief V1. The relative depth is converted to absolute depth by normal-guided global scaling, after which normals are computed from the scaled depth to obtain I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}5. A nonlinear transformation from V1 is then applied to suppress large gradients at occlusion boundaries while preserving subtle variations on smooth surfaces. The refined normal field is produced by a per-component fusion rule,

I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}6

designed to preserve the global structure of I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}7 while transferring high-frequency detail from I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}8. After normalization to unit vectors, the normals are integrated into depth via iterative normal integration, conceptually enforcing consistency between I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}9 and the gradient implied by the normals. The paper states that multiple integration iterations sharpen occlusion edges and reduce artifacts relative to single-pass integration (Zhang et al., 27 Aug 2025).

The real-world dataset is smaller but more controlled. It contains 800 physical reliefs from parks, museums, and heritage sites, with materials including gypsum, masonry, metal, and wood. For each relief, 20–30 multi-view high-resolution images are captured. Polycam is used to reconstruct textured 3D meshes from the multi-view images, and KeyShot is used to render orthographic RGB images with corresponding rough depth maps. Because Polycam reconstructions often miss fine details, the labels are further refined through depth-constrained normal integration. Using the initial MonoRelief V2 model, the rendered image is first converted into a detail-rich normal estimate and an associated gradient field D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}0; the refined depth D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}1 is then obtained by minimizing

D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}2

where D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}3, D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}4 is the rough depth from Polycam and KeyShot, and the two terms respectively inject detail and preserve global structure. The final refined labels are rendered at D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}5, with depth ranges from 12.81 to 445.73 pixels (Zhang et al., 27 Aug 2025).

The paper is explicit that pseudo labels remain imperfect. Structural distortions can enter from DepthAnything V2, and orientation biases can enter from MonoRelief V1. This matters because it frames the rationale for later training decisions: the large pseudo-real corpus supplies statistical coverage and relief-domain priors, while the smaller real dataset corrects systematic biases (Zhang et al., 27 Aug 2025).

4. Network architecture, progressive optimization, and losses

MonoRelief V2 follows the encoder–decoder pattern inherited from DepthAnything V2. The encoder is a ViT-base model from DepthAnything V2, itself a DINOv2-pretrained Vision Transformer with multi-scale token extraction; the decoder is DPT-style and aggregates multi-scale features into high-resolution dense depth predictions (Yang et al., 2024). The paper summarizes the overall mapping as

D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}6

Although no layer-by-layer architectural inventory is given, the paper states that DepthAnything V2 typically uses skip connections between encoder scales and multi-scale fusion and upsampling at feature-map scales such as D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}7, D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}8, D^∈RH×W\hat{D} \in \mathbb{R}^{H \times W}9, and D^\hat{D}0 (Zhang et al., 27 Aug 2025).

Training proceeds in two stages. Stage 1 fully fine-tunes all parameters of the DepthAnything V2 ViT-base backbone and decoder on the approximately 15,000 pseudo-real images and their pseudo-depth labels. Stage 2 freezes most of the network, inserts LoRA modules into the projection matrices of self-attention modules and the fully connected layers of MLP blocks, and fine-tunes only these low-rank adapters on the 800 real samples. The LoRA rank is 8, and only about 1.196% of the total parameters are trainable during the second stage. This design is motivated by the need to incorporate real-data corrections without overfitting or discarding the priors learned from the larger pseudo-real corpus (Zhang et al., 27 Aug 2025).

The composite objective couples absolute depth regression with normal consistency:

D^\hat{D}1

where

D^\hat{D}2

The normal term is equivalent, up to a constant, to maximizing cosine similarity between predicted and ground-truth normals. The schedule for D^\hat{D}3 is fixed: D^\hat{D}4 for the first 10 epochs and D^\hat{D}5 for the remaining 10 epochs, with the stated purpose of prioritizing depth early and reducing sensitivity to depth-label noise later. No additional explicit geometric or shading losses are used; geometry consistency is enforced entirely through normals derived from depth (Zhang et al., 27 Aug 2025).

The implementation details that are stated explicitly are narrow but sufficient to characterize the training setup: inputs and labels are at D^\hat{D}6 resolution, batch size is 4, the learning rate is D^\hat{D}7, each stage runs for 20 epochs, and training is conducted on a single NVIDIA A100 GPU. The optimizer is not explicitly stated in the paper and therefore remains unspecified here (Zhang et al., 27 Aug 2025).

5. Benchmark construction and empirical performance

MonoRelief V2 is evaluated on a custom relief benchmark of 498 samples at D^\hat{D}8 resolution. The synthetic subset contains 252 images derived from 36 hand-crafted relief models with materials including leather, concrete, high-gloss lacquer, bamboo, silver, marble, and porcelain under 10 HDRI environments, five indoor and five outdoor, sampled randomly. The real subset contains 246 images from 41 textured physical relief models from CGTrader, Zeel, and Sketchfab, rendered in KeyShot under six lighting setups; when original geometry lacks details, depth labels are refined via the same depth-constrained integration objective used in real-label construction. Across the benchmark, depth ranges lie between 27.78 and 296.36 pixels (Zhang et al., 27 Aug 2025).

Depth is evaluated by mean depth error percentage D^\hat{D}9, depth PSNR, and depth SSIM. Normals are evaluated by mean angular error D^=fθ(I)\hat{D} = f_\theta(I)0, normal PSNR, normal SSIM, and the percentages of pixels with angular error below D^=fθ(I)\hat{D} = f_\theta(I)1 and D^=fθ(I)\hat{D} = f_\theta(I)2. On the joint depth-and-normal benchmark, the final V2 model achieves D^=fθ(I)\hat{D} = f_\theta(I)3, depth PSNR D^=fθ(I)\hat{D} = f_\theta(I)4, depth SSIM D^=fθ(I)\hat{D} = f_\theta(I)5, D^=fθ(I)\hat{D} = f_\theta(I)6, normal PSNR D^=fθ(I)\hat{D} = f_\theta(I)7, and normal SSIM D^=fθ(I)\hat{D} = f_\theta(I)8. The initial model trained only on pseudo-real data achieves D^=fθ(I)\hat{D} = f_\theta(I)9 and n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).0. The corresponding V1 values are n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).1 and n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).2. DepthAnything V2 reports n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).3 and n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).4; DepthPro and MoGe V2 are substantially worse on this benchmark (Zhang et al., 27 Aug 2025).

On normal-specific evaluation, the final V2 model achieves mean error n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).5, n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).6 of pixels below n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).7, n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).8 below n^=Norm(−∇xD^,−∇yD^,1).\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).9, normal PSNR 21.77%21.77\%0, and normal SSIM 21.77%21.77\%1. The initial model is nearly identical in normal angular error at 21.77%21.77\%2 but slightly lower on the threshold and image-quality metrics. All listed baselines—GenPercept, DSINE, Lotus, StableNormal, Marigold-E2E-FT, and MonoRelief V1-normal—perform worse in mean angular error, PSNR, or SSIM (Zhang et al., 27 Aug 2025).

The ablations are informative because they separate gains attributable to pseudo-real supervision from gains attributable to real-data adaptation. Pseudo-real-only training already surpasses V1 and the general monocular depth methods, especially in normal metrics. Adding LoRA adaptation on the real dataset substantially improves depth error, from 21.77%21.77\%3 to 21.77%21.77\%4, with a smaller improvement in normal angular error, from 21.77%21.77\%5 to 21.77%21.77\%6. The paper also compares LoRA with decoder-only fine-tuning, where only the last five decoder layers are updated; that alternative reduces depth error but degrades generalization and blurs details, whereas LoRA preserves both generalization and fine detail (Zhang et al., 27 Aug 2025).

Qualitative evaluation extends beyond the benchmark. Figures described in the paper show robustness on real photographs, AI-generated images, and software-rendered graphics, as well as cleaner background planes, fewer artifacts, more physically plausible depth distributions, and improved efficiency relative to V1. This suggests that the main empirical advance is not only lower average error but also a better match between relief-specific geometry and deployment conditions (Zhang et al., 27 Aug 2025).

6. Applications, failure modes, and future directions

The downstream uses emphasized in the paper are practical and fabrication-oriented. MonoRelief V2 supports digital heritage and preservation workflows for monuments, temples, and museums; coin, sculpture, and bas-relief digitization; design and engraving pipelines in which relief-style images generated by text-to-image systems are converted into 2.5D reliefs; 3D printing and fabrication workflows of the form image 21.77%21.77\%7 MonoRelief V2 21.77%21.77\%8 relief mesh 21.77%21.77\%9 print; and AR/VR content creation when shallow geometry is sufficient. The paper specifically demonstrates operation on real photos, AI-generated content, and rendered graphics, with fabrication by 3D printing or CNC as a terminal use case (Zhang et al., 27 Aug 2025).

The paper also delineates several limitations. Output quality depends strongly on input image resolution and clarity because no super-resolution module is included. Failure cases arise under low-contrast shading, strong texture patterns, hard shadows, and strong specular highlights, where appearance can be misread as geometry or true geometry can be missed. Limited real data leads to incomplete understanding of perfectly planar regions, so planar areas adjacent to steep gradients may contain geometric noise. Training data depth ranges top out around 400–445 pixels at 16.21∘16.21^\circ0, and for reliefs thicker than 400 pixels the depth error increases markedly. Certain object categories, including natural landscapes, portraits, and animals with low shading contrast or high-frequency texture, can still exhibit substantial geometric deviation even when fine details are captured (Zhang et al., 27 Aug 2025).

The explicit future-work directions in the paper are larger-scale real relief datasets with more diverse materials, lighting conditions, thicknesses, and object categories, together with additional scene priors, including background priors and possibly explicit modeling of planar backgrounds or depth regularities. Potential directions identified as implied in the source include improved generative models for pseudo-real data with better geometry consistency, explicit reflectance modeling or shading-based losses, multi-view inference when multiple images are available, and super-resolution or multi-scale enhancement for finer relief detail. Because these latter items are presented as implied rather than confirmed roadmap elements, they are best understood as extensions suggested by the current design rather than commitments of the published method (Zhang et al., 27 Aug 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MonoRelief V2.