---
title: 'MonoGlass3D: Monocular 3D Glass Detection'
url: https://www.emergentmind.com/topics/monoglass3d
type: topic
---

# MonoGlass3D: Monocular 3D Glass Detection

Searching arXiv for the cited paper and closely related baselines to ground the article in current literature.
MonoGlass3D is a monocular 3D glass detection framework introduced for transparent-surface understanding in real-world environments. It is defined by three coupled contributions: a new real-world glass dataset with precise 3D annotations, an adaptive feature fusion module for context-sensitive representation learning, and a plane regression pipeline that encodes the planar geometry of glass surfaces within a monocular perception system. The method is presented in "MonoGlass3D: Monocular 3D Glass Detection with Plane Regression and Adaptive Feature Fusion" [2509.05599], which reports state-of-the-art results in both glass segmentation and monocular glass depth estimation on the introduced benchmark, and also evaluates segmentation on GDD and GSD.

## 1. Problem formulation and scope

Detecting and localizing glass in 3D environments is difficult because the optical properties of glass hinder conventional sensors from accurately distinguishing glass surfaces. MonoGlass3D addresses this by treating glass perception as a joint problem of segmentation and planar 3D estimation from a single RGB image, rather than as generic monocular depth prediction alone [2509.05599].

The framework is specialized to glass surfaces that can be modeled as planes. In the reported formulation, each glass panel is represented by a plane equation and estimated densely at the pixel level. This design directly exploits the observation that many architectural glass structures—doors, walls, ceilings, and related installations—are piecewise planar. The paper explicitly states that the method assumes glass surfaces are planar and fails on curved or highly irregular glass [2509.05599].

A plausible implication is that MonoGlass3D occupies an intermediate position between semantic glass segmentation and unconstrained monocular depth estimation. Its key distinction is the injection of an explicit geometric prior, namely piecewise planarity, into a jointly optimized segmentation-and-regression architecture.

## 2. Dataset and annotation protocol

The paper introduces a new real-world glass dataset collected from 50 distinct real-world scenes, including corridors, cafés, libraries, offices, doors, walls, ceilings of glass, and escalators [2509.05599]. The scenes were scanned with a LiDAR-visual SLAM setup. In total, 1,437 RGB frames were annotated, with 1,070 for training and 367 for validation.

The annotation pipeline is defined in geometric terms. Let the 3D bounding-box vertices in the world frame be denoted by $P_w$. These are transformed into the camera frame by
$$
P_c = T_{cw} P_w.
$$
Each glass panel is then annotated as a plane
$$
n \cdot v + d = 0,
$$
where $n$ is the unit normal and $d$ the signed distance to the origin. Plane fitting is performed by least squares on the 3D box corners:
$$
\hat n = (P_c^T P_c)^{-1} P_c^T [-1],
$$
$$
n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.
$$
For every pixel $(u_x,u_y)$ inside a glass mask, the depth is assigned by intersecting its camera ray with the fitted plane:
$$
\text{depth} = d \,/\, \left(n^T K^{-1}[u_x\,u_y\,1]^T\right).
$$

The reported diversity statistics are substantial. The number of glass planes per frame ranges from 1 up to 10, and the depth range per frame spans 0.07 m to 17.76 m. The dataset also includes 21 of 50 scenes collected at night, with varied illumination, backgrounds, and glass shapes [2509.05599].

| Dataset property | Reported value |
|---|---|
| Distinct real-world scenes | 50 |
| Annotated RGB frames | 1,437 |
| Train / validation split | 1,070 / 367 |
| Glass planes per frame | 1 to 10 |
| Depth range per frame | 0.07 m to 17.76 m |
| Night scenes | 21 of 50 |

The dataset is described as the first large-scale real-world glass dataset with 3D plane parameters and dense depth [2509.05599]. Since this statement appears in the paper summary, it is most appropriately understood as the paper’s characterization of its own contribution.

## 3. Network architecture

MonoGlass3D uses a DINOv2 ViT-S encoder as backbone, and the last three feature maps of size $C \times H/14 \times W/14$ are processed in a 4-stage cascade [2509.05599]. At each stage $i$, the Adaptive Feature Fusion (AFF) module receives current features $F^i$ with channels projected from $C$ to 256, together with the previous stage’s predictions: plane parameters $p^{i-1}$, segmentation $s^{i-1}$, and centerness $c^{i-1}$.

The AFF module outputs a new centerness prediction $c^i$, fused features $F_s^i$ for segmentation, and fused features $F_p^i$ for plane regression. The final segmentation branch concatenates $F_s^3$ with positional encoding, applies transformer encoder self-attention, and then uses up-convolution to produce the pixel mask. The final plane regression branch takes $F_p^3$ with positional encoding, applies self-attention, then cross-attention with queries from segmentation features, and finally up-convolution to produce per-pixel plane parameters [2509.05599].

This architecture couples region understanding and geometry estimation at multiple stages. The ablation study reported in the paper indicates that adding cross-attention improves Abs Rel from 0.082 to 0.072, adding self-attention further improves it to 0.071, and adding AFF improves it to 0.064 [2509.05599]. This suggests that iterative fusion of contextual and geometric cues is central rather than peripheral to the design.

## 4. Plane regression and centerness-guided fusion

The plane regression module is based on the plane equation
$$
a x + b y + c z + d = 0.
$$
Instead of regressing $(a,b,c,d)$ directly, MonoGlass3D re-parameterizes the unit normal $n = (n_x,n_y,n_z)^T$ by two polar angles $(\theta_1,\theta_2)$ plus intercept $d$:
$$
\theta_1 = \pm \arccos(r_{xz}/r), \quad \theta_2 = \pm \arccos(z/r_{xz}),
$$
with
$$
r_{xz} = \sqrt{x^2+z^2}, \quad r=1.
$$
The resulting representation has exactly three degrees of freedom, $[\theta_1,\theta_2,d]$ [2509.05599].

The network’s plane head uses $\tanh$ activations to predict
$$
\hat\theta_i = \tanh(\cdot)\cdot(\pi/2), \quad \hat d = \tanh(\cdot)\cdot 5.
$$
The final output per pixel is $(\hat\theta_1,\hat\theta_2,\hat d)$, from which the normal is recovered via the inverse mapping while enforcing $n \cdot v + d = 0$ [2509.05599].

The centerness formulation is equally specific. For each pixel inside a ground-truth mask, centerness $C$ measures normalized distance to its boundary:
$$
C = \sqrt{d_{\min}/d_{\max}}.
$$
In each AFF stage, predicted centerness $c^i$ reweights features according to
$$
F_{fused} = F * (1 - c^i).
$$
The stated effect is to suppress the interior surface, which is often low-texture, and emphasize context near boundaries. Multi-scale fusion is achieved through four cascading stages that iteratively refine $c$, $s$, and $p$ [2509.05599].

The paper’s interpretation is that centerness-based AFF supplies soft, instance-level shape context richer than thin boundaries or ad-hoc optical cues. A plausible implication is that the module is designed to compensate for the weak texture and ambiguous appearance characteristic of transparent objects by shifting representational emphasis toward contextual structures that co-occur with glass boundaries.

## 5. Losses and optimization

MonoGlass3D optimizes a composite objective
$$
L = L_c + L_s + L_p.
$$
The centerness loss is binary cross-entropy:
$$
L_c = \mathrm{BCE}(\hat c, c^{gt}).
$$
The segmentation loss combines binary cross-entropy and IoU loss:
$$
L_s = 0.5\,L_{BCE}(\hat s, s^{gt}) + L_{IoU}(\hat s, s^{gt}).
$$
The plane regression loss is
$$
L_p = L_{param} + L_{dist},
$$
where
$$
L_{param} = \|[\hat\theta_1,\hat\theta_2,\hat d] - [\theta_1^{gt},\theta_2^{gt},d^{gt}]\|_1.
$$
The plane-distance term $L_{dist}$ is defined by selecting, for each pixel, four 3D points on the estimated plane around its projected point and summing their distances to the ground-truth plane. The paper states that this measures true geometric discrepancy independent of view direction [2509.05599].

Instance normalization is used for the plane loss. If an image has $N$ glass instances and instance $i$ has $M_i$ pixels, then
$$
L_p = \frac{1}{N}\sum_{i=1}^N \frac{1}{M_i}\sum_{j=1}^{M_i} L_{p,(i,j)}.
$$
The stage-weighted total losses are
$$
L_p^{total} = 0.1L_p^1 + 0.1L_p^2 + 0.2L_p^3 + 0.6L_p^4,
$$
$$
L_s^{total} = 0.1L_s^1 + 0.1L_s^2 + 0.2L_s^3 + 0.6L_s^4,
$$
$$
L_c^{total} = 0.2L_c^1 + 0.3L_c^2 + 0.5L_c^3.
$$

The reported training setup uses input resolution 504×630, AdamW optimization, backbone learning rate $5\times10^{-6}$, other layers learning rate $5\times10^{-5}$, and a scheduler that reduces the learning rate by 0.95 if there is no loss drop for 8 epochs. Training is performed for 230 epochs on 4× NVIDIA RTX 4090, with no special data augmentations beyond standard resizing and cropping [2509.05599].

The ablation study explicitly states that the proposed plane-distance loss significantly outperforms a naive $L_1$ depth loss, with the best configuration reaching Abs Rel 0.063 when AFF and the proposed plane-distance loss are both used [2509.05599]. This suggests that supervising planar geometry directly, rather than supervising only per-pixel depth residuals, is materially important for transparent-surface reconstruction.

## 6. Empirical performance

The evaluation protocol includes segmentation metrics—IoU, MAE (pixel-wise), F1, BER—and depth-estimation metrics—Abs Rel, MAE (m), RMSE (m), and accuracy thresholds $\delta < 1.25$, $1.25^2$, and $1.25^3$ [2509.05599].

On the new real-world dataset over all scenes, MonoGlass3D with 52.4 M parameters is compared with ViT-B Depth Anything V2 with 97.5 M parameters. The reported results are:

| Method / setting | Key reported results |
|---|---|
| MonoGlass3D, all scenes | Abs Rel 0.063; MAE 0.274; RMSE 0.298; $\delta<1.25$ 0.982; $\delta<1.25^2$ 0.994; $\delta<1.25^3$ 0.997 |
| ViT-B Depth Anything V2, all scenes | Abs Rel 0.067; MAE 0.288; RMSE 0.318; $\delta<1.25$ 0.970; $\delta<1.25^2$ 0.993; $\delta<1.25^3$ 0.997 |
| MonoGlass3D, six completely unseen real scenes (180 frames) | Abs Rel 0.083; MAE 0.383; RMSE 0.405; $\delta<1.25$ 0.976 |

On glass segmentation benchmarks, the reported numbers are also strong. On GDD, MonoGlass3D with 34 M parameters achieves IoU = 0.920, F1 = 0.951, MAE = 0.038, and BER = 3.76, which the paper describes as state of the art. On GSD, it achieves IoU = 0.872, F1 = 0.917, MAE = 0.040, and BER = 5.16 [2509.05599].

The qualitative results described in the paper include depth reconstructions and 3D projections of multiple coplanar, multi-angle, and occluded glass surfaces that show flatter, more consistent planes than direct depth baselines, as well as tighter segmentation masks in weak-boundary or low-contrast cases [2509.05599].

The ablation results identify AFF as the largest single contributor among tested architectural additions:

| Variant on the new dataset | Abs Rel |
|---|---|
| Base backbone only | 0.082 |
| + cross-attention | 0.072 |
| + self-attention | 0.071 |
| + Adaptive Fusion (AFF) | 0.064 |
| AFF + proposed plane-distance loss | 0.063 |

These results support the paper’s central claim that combining geometric and contextual cues improves transparent surface understanding [2509.05599].

## 7. Position within glass perception research, limitations, and future directions

MonoGlass3D is situated at the intersection of monocular depth estimation, transparent-object segmentation, and planar scene reconstruction. Its main conceptual move is to reformulate glass depth estimation as plane regression in polar coordinates while simultaneously predicting segmentation masks and centerness-weighted contextual features [2509.05599]. Relative to generic monocular depth predictors such as the cited Depth Anything V2 baseline, the reported gains are associated with stronger structural priors and geometry-aware supervision.

The method’s limitations are explicitly acknowledged. It assumes planar glass surfaces and therefore fails on curved or highly irregular glass. The dataset annotation process is also described as labor-intensive, requiring 3D box annotation, 2D masks, and cross-modal fusion [2509.05599]. These limitations constrain both model scope and data scalability.

The future directions proposed in the paper include extending the method to multi-modal inputs—such as stereo, depth sensors, and polarization—for increased robustness, and exploring non-planar shape modules or mixtures of primitives to cover curved glass [2509.05599]. This suggests two orthogonal research trajectories: one centered on richer sensing, the other on richer geometric parameterizations.

Within the paper’s own summary, MonoGlass3D advances monocular 3D glass detection by introducing a large-scale real-world glass dataset with 3D plane parameters and dense depth, reformulating depth estimation as plane regression in polar coordinates, adaptively fusing multi-scale features via centerness maps, and designing a geometry-aware plane-distance loss [2509.05599]. In aggregate, these elements define a glass-specific perception framework whose reported performance derives from the explicit coupling of segmentation, context modeling, and planar geometry.

Source: https://www.emergentmind.com/topics/monoglass3d