Papers
Topics
Authors
Recent
Search
2000 character limit reached

MonoGlass3D: Monocular 3D Glass Detection

Updated 10 July 2026
  • MonoGlass3D is a monocular 3D glass detection framework that fuses segmentation and plane regression to accurately model transparent, planar glass surfaces.
  • It introduces a new real-world glass dataset with precise 3D plane annotations, enabling robust segmentation and depth estimation evaluations.
  • The method employs an adaptive feature fusion module with centerness guidance to refine geometric and context-aware cues for improved detection performance.

Searching arXiv for the cited paper and closely related baselines to ground the article in current literature. MonoGlass3D is a monocular 3D glass detection framework introduced for transparent-surface understanding in real-world environments. It is defined by three coupled contributions: a new real-world glass dataset with precise 3D annotations, an adaptive feature fusion module for context-sensitive representation learning, and a plane regression pipeline that encodes the planar geometry of glass surfaces within a monocular perception system. The method is presented in "MonoGlass3D: Monocular 3D Glass Detection with Plane Regression and Adaptive Feature Fusion" (Zhang et al., 6 Sep 2025), which reports state-of-the-art results in both glass segmentation and monocular glass depth estimation on the introduced benchmark, and also evaluates segmentation on GDD and GSD.

1. Problem formulation and scope

Detecting and localizing glass in 3D environments is difficult because the optical properties of glass hinder conventional sensors from accurately distinguishing glass surfaces. MonoGlass3D addresses this by treating glass perception as a joint problem of segmentation and planar 3D estimation from a single RGB image, rather than as generic monocular depth prediction alone (Zhang et al., 6 Sep 2025).

The framework is specialized to glass surfaces that can be modeled as planes. In the reported formulation, each glass panel is represented by a plane equation and estimated densely at the pixel level. This design directly exploits the observation that many architectural glass structures—doors, walls, ceilings, and related installations—are piecewise planar. The paper explicitly states that the method assumes glass surfaces are planar and fails on curved or highly irregular glass (Zhang et al., 6 Sep 2025).

A plausible implication is that MonoGlass3D occupies an intermediate position between semantic glass segmentation and unconstrained monocular depth estimation. Its key distinction is the injection of an explicit geometric prior, namely piecewise planarity, into a jointly optimized segmentation-and-regression architecture.

2. Dataset and annotation protocol

The paper introduces a new real-world glass dataset collected from 50 distinct real-world scenes, including corridors, cafés, libraries, offices, doors, walls, ceilings of glass, and escalators (Zhang et al., 6 Sep 2025). The scenes were scanned with a LiDAR-visual SLAM setup. In total, 1,437 RGB frames were annotated, with 1,070 for training and 367 for validation.

The annotation pipeline is defined in geometric terms. Let the 3D bounding-box vertices in the world frame be denoted by PwP_w. These are transformed into the camera frame by

Pc=TcwPw.P_c = T_{cw} P_w.

Each glass panel is then annotated as a plane

nâ‹…v+d=0,n \cdot v + d = 0,

where nn is the unit normal and dd the signed distance to the origin. Plane fitting is performed by least squares on the 3D box corners:

n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],

n=n^/∥n^∥,d=1/∥n^∥.n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.

For every pixel (ux,uy)(u_x,u_y) inside a glass mask, the depth is assigned by intersecting its camera ray with the fitted plane:

depth=d / (nTK−1[ux uy 1]T).\text{depth} = d \,/\, \left(n^T K^{-1}[u_x\,u_y\,1]^T\right).

The reported diversity statistics are substantial. The number of glass planes per frame ranges from 1 up to 10, and the depth range per frame spans 0.07 m to 17.76 m. The dataset also includes 21 of 50 scenes collected at night, with varied illumination, backgrounds, and glass shapes (Zhang et al., 6 Sep 2025).

Dataset property Reported value
Distinct real-world scenes 50
Annotated RGB frames 1,437
Train / validation split 1,070 / 367
Glass planes per frame 1 to 10
Depth range per frame 0.07 m to 17.76 m
Night scenes 21 of 50

The dataset is described as the first large-scale real-world glass dataset with 3D plane parameters and dense depth (Zhang et al., 6 Sep 2025). Since this statement appears in the paper summary, it is most appropriately understood as the paper’s characterization of its own contribution.

3. Network architecture

MonoGlass3D uses a DINOv2 ViT-S encoder as backbone, and the last three feature maps of size C×H/14×W/14C \times H/14 \times W/14 are processed in a 4-stage cascade (Zhang et al., 6 Sep 2025). At each stage Pc=TcwPw.P_c = T_{cw} P_w.0, the Adaptive Feature Fusion (AFF) module receives current features Pc=TcwPw.P_c = T_{cw} P_w.1 with channels projected from Pc=TcwPw.P_c = T_{cw} P_w.2 to 256, together with the previous stage’s predictions: plane parameters Pc=TcwPw.P_c = T_{cw} P_w.3, segmentation Pc=TcwPw.P_c = T_{cw} P_w.4, and centerness Pc=TcwPw.P_c = T_{cw} P_w.5.

The AFF module outputs a new centerness prediction Pc=TcwPw.P_c = T_{cw} P_w.6, fused features Pc=TcwPw.P_c = T_{cw} P_w.7 for segmentation, and fused features Pc=TcwPw.P_c = T_{cw} P_w.8 for plane regression. The final segmentation branch concatenates Pc=TcwPw.P_c = T_{cw} P_w.9 with positional encoding, applies transformer encoder self-attention, and then uses up-convolution to produce the pixel mask. The final plane regression branch takes nâ‹…v+d=0,n \cdot v + d = 0,0 with positional encoding, applies self-attention, then cross-attention with queries from segmentation features, and finally up-convolution to produce per-pixel plane parameters (Zhang et al., 6 Sep 2025).

This architecture couples region understanding and geometry estimation at multiple stages. The ablation study reported in the paper indicates that adding cross-attention improves Abs Rel from 0.082 to 0.072, adding self-attention further improves it to 0.071, and adding AFF improves it to 0.064 (Zhang et al., 6 Sep 2025). This suggests that iterative fusion of contextual and geometric cues is central rather than peripheral to the design.

4. Plane regression and centerness-guided fusion

The plane regression module is based on the plane equation

nâ‹…v+d=0,n \cdot v + d = 0,1

Instead of regressing nâ‹…v+d=0,n \cdot v + d = 0,2 directly, MonoGlass3D re-parameterizes the unit normal nâ‹…v+d=0,n \cdot v + d = 0,3 by two polar angles nâ‹…v+d=0,n \cdot v + d = 0,4 plus intercept nâ‹…v+d=0,n \cdot v + d = 0,5:

nâ‹…v+d=0,n \cdot v + d = 0,6

with

nâ‹…v+d=0,n \cdot v + d = 0,7

The resulting representation has exactly three degrees of freedom, nâ‹…v+d=0,n \cdot v + d = 0,8 (Zhang et al., 6 Sep 2025).

The network’s plane head uses n⋅v+d=0,n \cdot v + d = 0,9 activations to predict

nn0

The final output per pixel is nn1, from which the normal is recovered via the inverse mapping while enforcing nn2 (Zhang et al., 6 Sep 2025).

The centerness formulation is equally specific. For each pixel inside a ground-truth mask, centerness nn3 measures normalized distance to its boundary:

nn4

In each AFF stage, predicted centerness nn5 reweights features according to

nn6

The stated effect is to suppress the interior surface, which is often low-texture, and emphasize context near boundaries. Multi-scale fusion is achieved through four cascading stages that iteratively refine nn7, nn8, and nn9 (Zhang et al., 6 Sep 2025).

The paper’s interpretation is that centerness-based AFF supplies soft, instance-level shape context richer than thin boundaries or ad-hoc optical cues. A plausible implication is that the module is designed to compensate for the weak texture and ambiguous appearance characteristic of transparent objects by shifting representational emphasis toward contextual structures that co-occur with glass boundaries.

5. Losses and optimization

MonoGlass3D optimizes a composite objective

dd0

The centerness loss is binary cross-entropy:

dd1

The segmentation loss combines binary cross-entropy and IoU loss:

dd2

The plane regression loss is

dd3

where

dd4

The plane-distance term dd5 is defined by selecting, for each pixel, four 3D points on the estimated plane around its projected point and summing their distances to the ground-truth plane. The paper states that this measures true geometric discrepancy independent of view direction (Zhang et al., 6 Sep 2025).

Instance normalization is used for the plane loss. If an image has dd6 glass instances and instance dd7 has dd8 pixels, then

dd9

The stage-weighted total losses are

n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],0

n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],1

n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],2

The reported training setup uses input resolution 504×630, AdamW optimization, backbone learning rate n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],3, other layers learning rate n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],4, and a scheduler that reduces the learning rate by 0.95 if there is no loss drop for 8 epochs. Training is performed for 230 epochs on 4× NVIDIA RTX 4090, with no special data augmentations beyond standard resizing and cropping (Zhang et al., 6 Sep 2025).

The ablation study explicitly states that the proposed plane-distance loss significantly outperforms a naive n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],5 depth loss, with the best configuration reaching Abs Rel 0.063 when AFF and the proposed plane-distance loss are both used (Zhang et al., 6 Sep 2025). This suggests that supervising planar geometry directly, rather than supervising only per-pixel depth residuals, is materially important for transparent-surface reconstruction.

6. Empirical performance

The evaluation protocol includes segmentation metrics—IoU, MAE (pixel-wise), F1, BER—and depth-estimation metrics—Abs Rel, MAE (m), RMSE (m), and accuracy thresholds n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],6, n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],7, and n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],8 (Zhang et al., 6 Sep 2025).

On the new real-world dataset over all scenes, MonoGlass3D with 52.4 M parameters is compared with ViT-B Depth Anything V2 with 97.5 M parameters. The reported results are:

Method / setting Key reported results
MonoGlass3D, all scenes Abs Rel 0.063; MAE 0.274; RMSE 0.298; n^=(PcTPc)−1PcT[−1],\hat n = (P_c^T P_c)^{-1} P_c^T [-1],9 0.982; n=n^/∥n^∥,d=1/∥n^∥.n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.0 0.994; n=n^/∥n^∥,d=1/∥n^∥.n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.1 0.997
ViT-B Depth Anything V2, all scenes Abs Rel 0.067; MAE 0.288; RMSE 0.318; n=n^/∥n^∥,d=1/∥n^∥.n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.2 0.970; n=n^/∥n^∥,d=1/∥n^∥.n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.3 0.993; n=n^/∥n^∥,d=1/∥n^∥.n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.4 0.997
MonoGlass3D, six completely unseen real scenes (180 frames) Abs Rel 0.083; MAE 0.383; RMSE 0.405; n=n^/∥n^∥,d=1/∥n^∥.n = \hat n/\|\hat n\|,\quad d = 1/\|\hat n\|.5 0.976

On glass segmentation benchmarks, the reported numbers are also strong. On GDD, MonoGlass3D with 34 M parameters achieves IoU = 0.920, F1 = 0.951, MAE = 0.038, and BER = 3.76, which the paper describes as state of the art. On GSD, it achieves IoU = 0.872, F1 = 0.917, MAE = 0.040, and BER = 5.16 (Zhang et al., 6 Sep 2025).

The qualitative results described in the paper include depth reconstructions and 3D projections of multiple coplanar, multi-angle, and occluded glass surfaces that show flatter, more consistent planes than direct depth baselines, as well as tighter segmentation masks in weak-boundary or low-contrast cases (Zhang et al., 6 Sep 2025).

The ablation results identify AFF as the largest single contributor among tested architectural additions:

Variant on the new dataset Abs Rel
Base backbone only 0.082
+ cross-attention 0.072
+ self-attention 0.071
+ Adaptive Fusion (AFF) 0.064
AFF + proposed plane-distance loss 0.063

These results support the paper’s central claim that combining geometric and contextual cues improves transparent surface understanding (Zhang et al., 6 Sep 2025).

7. Position within glass perception research, limitations, and future directions

MonoGlass3D is situated at the intersection of monocular depth estimation, transparent-object segmentation, and planar scene reconstruction. Its main conceptual move is to reformulate glass depth estimation as plane regression in polar coordinates while simultaneously predicting segmentation masks and centerness-weighted contextual features (Zhang et al., 6 Sep 2025). Relative to generic monocular depth predictors such as the cited Depth Anything V2 baseline, the reported gains are associated with stronger structural priors and geometry-aware supervision.

The method’s limitations are explicitly acknowledged. It assumes planar glass surfaces and therefore fails on curved or highly irregular glass. The dataset annotation process is also described as labor-intensive, requiring 3D box annotation, 2D masks, and cross-modal fusion (Zhang et al., 6 Sep 2025). These limitations constrain both model scope and data scalability.

The future directions proposed in the paper include extending the method to multi-modal inputs—such as stereo, depth sensors, and polarization—for increased robustness, and exploring non-planar shape modules or mixtures of primitives to cover curved glass (Zhang et al., 6 Sep 2025). This suggests two orthogonal research trajectories: one centered on richer sensing, the other on richer geometric parameterizations.

Within the paper’s own summary, MonoGlass3D advances monocular 3D glass detection by introducing a large-scale real-world glass dataset with 3D plane parameters and dense depth, reformulating depth estimation as plane regression in polar coordinates, adaptively fusing multi-scale features via centerness maps, and designing a geometry-aware plane-distance loss (Zhang et al., 6 Sep 2025). In aggregate, these elements define a glass-specific perception framework whose reported performance derives from the explicit coupling of segmentation, context modeling, and planar geometry.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MonoGlass3D.