---
title: 'MonoRelief V2: 2.5D Relief Recovery Model'
url: https://www.emergentmind.com/topics/monorelief-v2
type: topic
---

# MonoRelief V2: 2.5D Relief Recovery Model

MonoRelief V2 is an end-to-end model for directly recovering 2.5D reliefs from single images under complex material and illumination variations. In the paper’s formulation, “2.5D relief recovery” denotes prediction of a height field, or depth map, for a relief sculpture from a single RGB image, so that the relief can be treated as a function $z(x,y)$ over the image plane. Given an input image $I \in \mathbb{R}^{H \times W \times 3}$, the model outputs an absolute depth map $\hat{D} \in \mathbb{R}^{H \times W}$, scaled in physical units relative to the relief thickness, while a consistent normal field is obtained implicitly by differentiating $\hat{D}$. The intended reconstruction quality is sufficient for CNC machining, 3D printing, and digital visualization [2508.19555].

## 1. Problem definition and representational scope

MonoRelief V2 addresses a specialized variant of monocular geometry recovery: relief reconstruction rather than general monocular depth estimation. The distinction is central. Relief recovery focuses on shallow depth ranges, but it places unusually strong emphasis on fine local geometry, including curvature, ridges, and engravings. By contrast, general monocular depth estimators usually target global scene geometry over diverse scenes and often produce coarse geometry with good occlusion boundaries but limited high-frequency detail. Surface normal estimation occupies a complementary role, because normals emphasize local shape and are more sensitive to fine structure than depth alone [2508.19555].

The model is trained as a depth estimator, but normals are structurally embedded throughout the pipeline. Depth labels, both pseudo and real, are generated by fusing depth and normal predictions; supervision uses a composite loss that combines depth regression with a normal-consistency term; and evaluation is reported for both depth and derived normals. Formally, the network predicts depth as $\hat{D} = f_\theta(I)$, and the normal field is derived as
$$
\hat{\mathbf{n}} = \mathrm{Norm}(-\nabla_x \hat{D}, -\nabla_y \hat{D}, 1).
$$
A common misconception is that MonoRelief V2 is primarily a normal estimator. The paper instead defines depth as the primary output and normals as an implicit representation used for supervision and evaluation [2508.19555].

The target inputs span real photographs of architectural reliefs, heritage artifacts, and plaques; AI-generated “relief-style” images via Flux.1; and rendered reliefs with varied BRDFs. Material diversity includes clay, copper, gold, jade, marble, obsidian, plaster, silver, steel, as well as gypsum, masonry, metal, wood, and synthetic renderings with leather, concrete, lacquer, bamboo, and porcelain. Illumination diversity includes indoor and outdoor HDRI environments, varying directions and intensities, strong shadows, soft lighting, and specular highlights. The paper identifies four domain-specific challenges: material-plus-lighting coupling, geometry–texture ambiguity, depth ambiguity from single-view shading, and a domain gap between relief imagery and conventional indoor or outdoor depth-estimation corpora [2508.19555].

## 2. Lineage, prior systems, and methodological position

MonoRelief V2 is presented explicitly as a successor to MonoRelief V1. V1 uses a three-stage pipeline: a synthetic relief dataset with ground-truth normal labels, relief-domain fine-tuning of Omnidata for monocular normal prediction, and depth reconstruction by depth-constrained normal integration. That design established a relief-specific reconstruction workflow, but it remained synthetic-only, multi-stage, and relatively fragile on real photographs and complex materials. On the new benchmark, V1 reports mean depth error $21.77\%$ and mean normal angular error about $16.21^\circ$ [2508.19555].

V2 changes this formulation at three levels. First, it becomes end-to-end: a single model maps image to depth, with normals derived from the predicted depth. Second, it adds real supervision through two new data sources: approximately 15,000 pseudo-real images generated by Flux.1 with depth pseudo-labels constructed by fusing DepthAnything V2 and MonoRelief V1 predictions, and a small real-world dataset of 800 reliefs captured through multi-view imaging, reconstruction, and detail refinement. Third, it adopts progressive training: full fine-tuning on pseudo-real data followed by LoRA adaptation on real data. These changes reduce mean depth error from $21.77\%$ to $14.03\%$ and mean normal angular error from $16.21^\circ$ to $14.33^\circ$, while also increasing PSNR and SSIM for both depth and normals [2508.19555].

Architecturally, V2 is based on DepthAnything V2’s ViT+DPT design: a DINOv2-pretrained Vision Transformer encoder with a DPT-style dense prediction decoder [2406.09414]. MonoRelief V2 uses the ViT-base encoder and decoder configuration from that family, yielding a model of roughly 100M parameters and avoiding explicit normal integration during inference. Relative to V1, the backbone shifts from Omnidata ViT plus separate integration to a single DepthAnything V2-style encoder–decoder; the loss shifts from purely normal prediction plus post-processing to a joint depth-and-normal objective; and domain adaptation shifts from standard fine-tuning on synthetic data only to LoRA fine-tuning on real data [2508.19555].

The paper also situates MonoRelief V2 against general monocular depth and normal estimators. DepthAnything V2, DepthPro, and MoGe V2 are described as strong on global scene geometry but inadequate for relief-specific fine detail and material-dependent shading. General normal methods including Omnidata, DSINE, StableNormal, Lotus, Marigold-E2E-FT, and GenPercept are reported to preserve sharp edges but often lose shallow relief detail or misread material-driven appearance as geometry. MonoRelief V2 is characterized as the first end-to-end, relief-focused monocular 2.5D model trained on real relief data and as a state-of-the-art system specifically for relief depth and normal recovery [2508.19555].

## 3. Pseudo-real data generation and real-data acquisition

A central feature of MonoRelief V2 is its supervision pipeline. Because large-scale real-world paired relief image–depth datasets are difficult to acquire, the paper constructs two complementary datasets: a pseudo-real corpus at scale and a smaller high-fidelity real-world corpus [2508.19555].

The pseudo-real dataset is generated with Flux.1-schnell, an open-source text-to-image diffusion model. Prompting combines object categories—animals, humans, flora, architecture, landscapes, and miscellaneous cases—with nine target materials: clay, copper, gold, jade, marble, obsidian, plaster, silver, and steel. The prompt template is “A/an \<material\> relief of a/an \<object\>”. Images are generated at $1024 \times 1024$ resolution with four denoising steps and random seeds for diversity. From an initial output of about 30,000 images, the authors perform dual-blind expert review by two domain experts and filter samples based on shape fidelity and detail richness, yielding a curated pseudo-real dataset of about 15,000 images [2508.19555].

Pseudo-label generation is not a direct depth-regression procedure. The pipeline first predicts relative depth $D_{\text{rel}}(x)$ using DepthAnything V2 and detailed normals $N_2(x)$ using MonoRelief V1. The relative depth is converted to absolute depth by normal-guided global scaling, after which normals are computed from the scaled depth to obtain $N_1(x)$. A nonlinear transformation from V1 is then applied to suppress large gradients at occlusion boundaries while preserving subtle variations on smooth surfaces. The refined normal field is produced by a per-component fusion rule,
$$
N^* =
\begin{cases}
N_1 - (1 - 2N_2)\cdot N_1 \cdot (1 - N_1), & \text{if } N_2 \leq 0.5, \\
N_1 + (2N_2 - 1)\cdot (\sqrt{N_1} - N_1), & \text{if } N_2 > 0.5,
\end{cases}
$$
designed to preserve the global structure of $N_1$ while transferring high-frequency detail from $N_2$. After normalization to unit vectors, the normals are integrated into depth via iterative normal integration, conceptually enforcing consistency between $\nabla z$ and the gradient implied by the normals. The paper states that multiple integration iterations sharpen occlusion edges and reduce artifacts relative to single-pass integration [2508.19555].

The real-world dataset is smaller but more controlled. It contains 800 physical reliefs from parks, museums, and heritage sites, with materials including gypsum, masonry, metal, and wood. For each relief, 20–30 multi-view high-resolution images are captured. Polycam is used to reconstruct textured 3D meshes from the multi-view images, and KeyShot is used to render orthographic RGB images with corresponding rough depth maps. Because Polycam reconstructions often miss fine details, the labels are further refined through depth-constrained normal integration. Using the initial MonoRelief V2 model, the rendered image is first converted into a detail-rich normal estimate and an associated gradient field $g_i$; the refined depth $z$ is then obtained by minimizing
$$
\min_{z} \iint_{\Omega} \left( \nabla z_i - g_i \right)^2 + \mu \cdot \left( z_i - d_i \right)^2,
$$
where $\mu = 0.02$, $d_i$ is the rough depth from Polycam and KeyShot, and the two terms respectively inject detail and preserve global structure. The final refined labels are rendered at $1024 \times 1024$, with depth ranges from 12.81 to 445.73 pixels [2508.19555].

The paper is explicit that pseudo labels remain imperfect. Structural distortions can enter from DepthAnything V2, and orientation biases can enter from MonoRelief V1. This matters because it frames the rationale for later training decisions: the large pseudo-real corpus supplies statistical coverage and relief-domain priors, while the smaller real dataset corrects systematic biases [2508.19555].

## 4. Network architecture, progressive optimization, and losses

MonoRelief V2 follows the encoder–decoder pattern inherited from DepthAnything V2. The encoder is a ViT-base model from DepthAnything V2, itself a DINOv2-pretrained Vision Transformer with multi-scale token extraction; the decoder is DPT-style and aggregates multi-scale features into high-resolution dense depth predictions [2406.09414]. The paper summarizes the overall mapping as
$$
I \xrightarrow{\text{ViT encoder}} \text{multi-scale features} \xrightarrow{\text{DPT decoder}} \hat{D} \rightarrow \hat{\mathbf{n}}.
$$
Although no layer-by-layer architectural inventory is given, the paper states that DepthAnything V2 typically uses skip connections between encoder scales and multi-scale fusion and upsampling at feature-map scales such as $1/32$, $1/16$, $1/8$, and $1/4$ [2508.19555].

Training proceeds in two stages. Stage 1 fully fine-tunes all parameters of the DepthAnything V2 ViT-base backbone and decoder on the approximately 15,000 pseudo-real images and their pseudo-depth labels. Stage 2 freezes most of the network, inserts LoRA modules into the projection matrices of self-attention modules and the fully connected layers of MLP blocks, and fine-tunes only these low-rank adapters on the 800 real samples. The LoRA rank is 8, and only about 1.196% of the total parameters are trainable during the second stage. This design is motivated by the need to incorporate real-data corrections without overfitting or discarding the priors learned from the larger pseudo-real corpus [2508.19555].

The composite objective couples absolute depth regression with normal consistency:
$$
L = \alpha L_{\text{depth}} + L_{\text{normal}},
$$
where
$$
L_{\text{depth}} = \frac{1}{M}\sum_i \left\| d_i^{\text{pred}} - d_i^{\text{gt}} \right\|,
\qquad
L_{\text{normal}} = -\frac{1}{M}\sum_i \left\langle \mathbf{n}_i^{\text{pred}}, \mathbf{n}_i^{\text{gt}} \right\rangle .
$$
The normal term is equivalent, up to a constant, to maximizing cosine similarity between predicted and ground-truth normals. The schedule for $\alpha$ is fixed: $\alpha = 0.1$ for the first 10 epochs and $\alpha = 0.01$ for the remaining 10 epochs, with the stated purpose of prioritizing depth early and reducing sensitivity to depth-label noise later. No additional explicit geometric or shading losses are used; geometry consistency is enforced entirely through normals derived from depth [2508.19555].

The implementation details that are stated explicitly are narrow but sufficient to characterize the training setup: inputs and labels are at $1024 \times 1024$ resolution, batch size is 4, the learning rate is $5 \times 10^{-5}$, each stage runs for 20 epochs, and training is conducted on a single NVIDIA A100 GPU. The optimizer is not explicitly stated in the paper and therefore remains unspecified here [2508.19555].

## 5. Benchmark construction and empirical performance

MonoRelief V2 is evaluated on a custom relief benchmark of 498 samples at $1024 \times 1024$ resolution. The synthetic subset contains 252 images derived from 36 hand-crafted relief models with materials including leather, concrete, high-gloss lacquer, bamboo, silver, marble, and porcelain under 10 HDRI environments, five indoor and five outdoor, sampled randomly. The real subset contains 246 images from 41 textured physical relief models from CGTrader, Zeel, and Sketchfab, rendered in KeyShot under six lighting setups; when original geometry lacks details, depth labels are refined via the same depth-constrained integration objective used in real-label construction. Across the benchmark, depth ranges lie between 27.78 and 296.36 pixels [2508.19555].

Depth is evaluated by mean depth error percentage $\varepsilon_d$, depth PSNR, and depth SSIM. Normals are evaluated by mean angular error $\varepsilon_n$, normal PSNR, normal SSIM, and the percentages of pixels with angular error below $11.25^\circ$ and $22.5^\circ$. On the joint depth-and-normal benchmark, the final V2 model achieves $\varepsilon_d = 14.025$, depth PSNR $17.739$, depth SSIM $0.789$, $\varepsilon_n = 14.327$, normal PSNR $22.602$, and normal SSIM $0.821$. The initial model trained only on pseudo-real data achieves $\varepsilon_d = 19.766$ and $\varepsilon_n = 14.441$. The corresponding V1 values are $\varepsilon_d = 21.766$ and $\varepsilon_n = 16.213$. DepthAnything V2 reports $\varepsilon_d = 16.415$ and $\varepsilon_n = 15.654$; DepthPro and MoGe V2 are substantially worse on this benchmark [2508.19555].

On normal-specific evaluation, the final V2 model achieves mean error $14.327^\circ$, $52.207\%$ of pixels below $11.25^\circ$, $85.714\%$ below $22.5^\circ$, normal PSNR $22.602$, and normal SSIM $0.821$. The initial model is nearly identical in normal angular error at $14.441^\circ$ but slightly lower on the threshold and image-quality metrics. All listed baselines—GenPercept, DSINE, Lotus, StableNormal, Marigold-E2E-FT, and MonoRelief V1-normal—perform worse in mean angular error, PSNR, or SSIM [2508.19555].

The ablations are informative because they separate gains attributable to pseudo-real supervision from gains attributable to real-data adaptation. Pseudo-real-only training already surpasses V1 and the general monocular depth methods, especially in normal metrics. Adding LoRA adaptation on the real dataset substantially improves depth error, from $19.77\%$ to $14.03\%$, with a smaller improvement in normal angular error, from $14.44^\circ$ to $14.33^\circ$. The paper also compares LoRA with decoder-only fine-tuning, where only the last five decoder layers are updated; that alternative reduces depth error but degrades generalization and blurs details, whereas LoRA preserves both generalization and fine detail [2508.19555].

Qualitative evaluation extends beyond the benchmark. Figures described in the paper show robustness on real photographs, AI-generated images, and software-rendered graphics, as well as cleaner background planes, fewer artifacts, more physically plausible depth distributions, and improved efficiency relative to V1. This suggests that the main empirical advance is not only lower average error but also a better match between relief-specific geometry and deployment conditions [2508.19555].

## 6. Applications, failure modes, and future directions

The downstream uses emphasized in the paper are practical and fabrication-oriented. MonoRelief V2 supports digital heritage and preservation workflows for monuments, temples, and museums; coin, sculpture, and bas-relief digitization; design and engraving pipelines in which relief-style images generated by text-to-image systems are converted into 2.5D reliefs; 3D printing and fabrication workflows of the form image $\rightarrow$ MonoRelief V2 $\rightarrow$ relief mesh $\rightarrow$ print; and AR/VR content creation when shallow geometry is sufficient. The paper specifically demonstrates operation on real photos, AI-generated content, and rendered graphics, with fabrication by 3D printing or CNC as a terminal use case [2508.19555].

The paper also delineates several limitations. Output quality depends strongly on input image resolution and clarity because no super-resolution module is included. Failure cases arise under low-contrast shading, strong texture patterns, hard shadows, and strong specular highlights, where appearance can be misread as geometry or true geometry can be missed. Limited real data leads to incomplete understanding of perfectly planar regions, so planar areas adjacent to steep gradients may contain geometric noise. Training data depth ranges top out around 400–445 pixels at $1024 \times 1024$, and for reliefs thicker than 400 pixels the depth error increases markedly. Certain object categories, including natural landscapes, portraits, and animals with low shading contrast or high-frequency texture, can still exhibit substantial geometric deviation even when fine details are captured [2508.19555].

The explicit future-work directions in the paper are larger-scale real relief datasets with more diverse materials, lighting conditions, thicknesses, and object categories, together with additional scene priors, including background priors and possibly explicit modeling of planar backgrounds or depth regularities. Potential directions identified as implied in the source include improved generative models for pseudo-real data with better geometry consistency, explicit reflectance modeling or shading-based losses, multi-view inference when multiple images are available, and super-resolution or multi-scale enhancement for finer relief detail. Because these latter items are presented as implied rather than confirmed roadmap elements, they are best understood as extensions suggested by the current design rather than commitments of the published method [2508.19555].

Source: https://www.emergentmind.com/topics/monorelief-v2