Papers
Topics
Authors
Recent
Search
2000 character limit reached

CGL-Decoder: Prompt-Guided Wireframe Parsing

Updated 29 January 2026
  • The paper introduces a robust neural module that refines wireframe extraction by integrating prompt-conditioned sparse attention with point-line consistency, cutting endpoint mismatches from 12.4% to 7.8%.
  • Local feature fusion and windowed multi-head cross-attention exchange spatial cues between line and junction prompts, ensuring coherent and context-aware geometry refinement.
  • Empirical evaluations on benchmarks like Wireframe and YorkUrban demonstrate the decoder’s competitive performance, achieving up to 76.8 FPS with improved prediction accuracy.

The Cross-Guidance Line Decoder (CGL-Decoder) is a neural module introduced within the Co-PLNet framework for prompt-guided wireframe parsing. Its design enables collaborative refinement of structured geometry by exchanging spatial cues between lines and junctions using prompt-conditioned, windowed sparse attention. The CGL-Decoder enforces point-line consistency and computational efficiency, resulting in improved accuracy and real-time performance for tasks such as wireframe extraction in images (Wang et al., 26 Jan 2026).

1. Architectural Overview

The CGL-Decoder operates on feature representations and prompt maps derived from preceding feature extraction and Point-Line Prompt Encoder (PLP-Encoder) stages. Its critical architectural elements are:

  • Inputs:
    • Z∈RH×W×CZ \in \mathbb{R}^{H \times W \times C}: Refined feature map from the U-Net backbone.
    • yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}: Line prompt map (Cp=16C_p=16).
    • yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}: Junction prompt map (Cp=16C_p=16).
  • Outputs:
    • yLy_L: Dense line parameter proposals (d,θ,θ1,θ2,r)(d,\theta, \theta_1, \theta_2, r) at subpixel accuracy.
    • yJy_J: Refined junction heatmap and position offset maps.
    • yy: Final set of sparse line segments (endpoints) following non-maximum suppression (NMS) and line-of-interest (LOI) verification.
  • Internal modules:
    • Local Feature Fusion: Concatenates ZZ, yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}0, and yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}1 along the channel axis; processed through two small convolutional branches to obtain yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}2.
    • 1×1 Projection: yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}3 reduces channel count from yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}4 to yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}5 (yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}6) as pre-attention embedding.
    • Window Partitioning: Feature tensors partitioned into non-overlapping spatial windows (yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}7).
    • Sparse Multi-Head Cross-Attention: Each window attends from the line (or junction) branch to backbone features.
    • Gated Residual Fusion: Attended corrections are fused into yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}8 using learnable gating masks yL(0)∈RH×W×Cpy_L^{(0)} \in \mathbb{R}^{H \times W \times C_p}9.
    • Prediction Heads: HAWP/PLNet line and junction heads refine geometry predictions using Cp=16C_p=160.
    • Endpoint Grouping & LOI Scoring: Endpoints associated to nearest junctions, deduplicated, and scored for final selection.

2. Prompt-Conditioned Sparse Attention Mechanism

Within the CGL-Decoder, attention is conditioned on point-line prompts and applied sparsely within local windows. For the line branch, the formalization within spatial window Cp=16C_p=161 is:

Cp=16C_p=162

with Cp=16C_p=163 as learned projections (Cp=16C_p=164). For each head Cp=16C_p=165 and window Cp=16C_p=166: Cp=16C_p=167 The multi-head output is: Cp=16C_p=168 Windows are reassembled to yield the attended feature map Cp=16C_p=169 (likewise for junctions). Gated residual fusion restores the original channel dimension with

yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}0

where yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}1 are learned gating masks, and yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}2 is element-wise multiplication.

3. Integration with the PLP-Encoder and Local-Global Context

The CGL-Decoder leverages coarse spatial prompts from the PLP-Encoder, which generates yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}3 and yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}4 through lightweight convolutional heads. These prompt maps, spatially and channel aligned with yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}5, are concatenated along with yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}6 and passed through two convolutional branches to yield yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}7 and yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}8. This direct injection of geometry prompts enables the decoder to modulate feature fusion prior to global context aggregation by sparse attention, enhancing context-aware refinement.

4. Stepwise Decoding and Refinement Workflow

The complete refinement algorithm is outlined below:

  1. Backbone & PLP-Encoder: Extract multi-scale features (SuperPoint + U-Net), generate coarse proposals yJ(0)∈RH×W×Cpy_J^{(0)} \in \mathbb{R}^{H \times W \times C_p}9, and produce spatial prompts Cp=16C_p=160.
  2. Local Fusion: Input concatenation and convolution to obtain Cp=16C_p=161.
  3. Sparse Attention: Compute Cp=16C_p=162 for each window; perform windowed multi-head attention to derive Cp=16C_p=163; fuse via gated residuals.
  4. Geometry Prediction: Line head on Cp=16C_p=164 predicts line parameters Cp=16C_p=165; junction head on Cp=16C_p=166 predicts heatmaps and offsets.
  5. Post-Processing: Endpoints are snapped to their nearest junction within a threshold and deduplicated via NMS.
  6. Line-of-Interest Verification: Features are sampled along each candidate line and scored by a small MLP; top-Cp=16C_p=167 lines are retained.
  7. Output: The final set of line segments and junctions is produced.

5. Loss Formulations and Optimization Criteria

CGL-Decoder training is end-to-end and employs a composite loss: Cp=16C_p=168 where:

  • Cp=16C_p=169 supervises junction heatmaps and offsets.
  • yLy_L0 penalizes line regression error.
  • yLy_L1 encourages consistency between dense line proposals and ground-truth geometry (e.g., via point-to-line distance losses as in HAWP).
  • yLy_L2 supervises the LOI MLP's output line confidence scores.

6. Implementation Considerations and Computational Analysis

  • All CGL operations are implemented in PyTorch and benchmarked on an RTX 4080 GPU.
  • Window size default is yLy_L3, providing non-overlapping partitions.
  • Attention cost per image is yLy_L4, linear in image area due to sparse windowing.
  • The 1×1 convolution yLy_L5 reduces feature channels (256 to 32) before attention; prompt maps require only 16 channels.
  • Multi-head attention and window partitioning are parallelized for speed, yielding 76.8 FPS on images sized yLy_L6.
  • Dense (full) attention only marginally increased sAP by yLy_L7, but halved the runtime to yLy_L8 FPS, motivating the sparse approach.

7. Empirical Performance and Ablation Insights

On standard wireframe benchmarks:

Dataset sAP⁵ sAP¹⁰ sAP¹⁵ FPS Endpoint Mismatch Rate
Wireframe 68.4 72.3 73.8 76.8 7.8%
YorkUrban 32.7 35.6 36.6 — —

CGL-Decoder reduces the endpoint mismatch rate (the proportion of line endpoints not snapping to any detected junction within 15 px) from 12.4% (baseline PLNet) to 7.8% with full prompts and sparse attention. Ablation studies demonstrate that point-to-line prompts alone drop mismatches to 11.2% and raise sAP¹⁵ to 72.3; inclusion of line-to-point prompts improves sAP¹⁵ to 72.6 and mismatch to 9.6%. Sparse attention brings further gains (sAP¹⁵ to 73.3; mismatch to 7.8%). A window size of yLy_L9 achieves the optimal accuracy-speed trade-off (Wang et al., 26 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Cross-Guidance Line Decoder (CGL-Decoder).