---
title: 'HDMapLaneNet: Collaborative Lane Graph Extraction'
url: https://www.emergentmind.com/topics/hdmaplanenet
type: topic
---

# HDMapLaneNet: Collaborative Lane Graph Extraction

Searching arXiv for the cited papers to ground the article.
HDMapLaneNet is a lane-centric High-Definition (HD) map construction framework whose central object is the lane graph: lane centerline geometry together with connectivity relations required for map-level reasoning. In current arXiv usage, the name most directly denotes a 2025 system that extracts localized lane centerlines from front-facing camera images, models inter-segment relations through Scene Graph Generation, and serializes the resulting directed graph for collaborative V2X transmission and cloud aggregation [2502.10127]. The same name also appears in later literature as a reference to an earlier segmentation-based aerial-imagery pipeline by He et al. (2022), there also called “LaneExtraction,” which converts lane masks into lane graphs for HD-map extraction [2405.00620]. Across these usages, HDMapLaneNet is associated with the construction of the geometric lane layer of HD maps rather than with full semantic map authoring.

## 1. Terminological scope and problem setting

The contemporary formulation of HDMapLaneNet targets collaborative construction of the localized geometric lane layer of HD maps, specifically the extraction of localized lane centerlines and their connectivity, by combining front-facing camera imagery, scene graph generation for association prediction, and V2X communication for cloud aggregation [2502.10127]. The stated motivation is twofold. First, traditional dependence on dedicated mapping vehicles is expensive and refresh cycles lag real-world changes such as temporary lane closures and construction. Second, single-vehicle perception is constrained by occlusion and field of view, which limits local completeness even when onboard perception is accurate [2502.10127].

A distinct but historically relevant usage of the term appears in aerial-imagery lane graph extraction. In that line of work, HDMapLaneNet is described as a segmentation-first system that predicts lane masks and direction maps from aerial RGB imagery and then converts the masks into a graph through thinning, vectorization, pruning, and simplification; intersection handling is treated as a separate subtask [2405.00620]. The two usages therefore differ in modality, graph semantics, and deployment assumptions: the 2025 framework is camera-only, vehicle-borne, directed, and collaborative through V2X, whereas the earlier aerial formulation is segmentation-based, largely undirected in its non-intersection evaluation, and oriented toward large-area extraction from overhead imagery [2405.00620].

This suggests that “HDMapLaneNet” is best understood not as a single immutable architecture, but as a family of lane-layer HD-map construction approaches organized around graph extraction.

## 2. Architectural pipeline in the V2X collaborative formulation

The 2025 HDMapLaneNet pipeline is explicitly modular and contains four stages: image-view feature extraction, lane centerline segment detection and curve encoding, scene graph generation for connectivity, and V2X-based serialization and transmission to a cloud aggregator [2502.10127].

**Image-view feature extraction**: raw monocular images from the frontal camera are processed by DeepLabv3. The role of DeepLabv3 is to capture multi-scale contextual information via an Atrous Spatial Pyramid Pooling encoder and an encoder–decoder segmentation architecture, producing high-level semantic features that contextualize lane geometry [2502.10127].

**Lane centerline segment extraction and Bézier encoding**: Detection Transformer is adapted as the detection head for lane segments and framed as a set prediction problem. A CNN backbone produces low-resolution feature maps, positional encodings following STSU inject Bird’s-eye-view spatial priors, and a transformer encoder–decoder outputs $K$ fixed-length embeddings. A 3-layer feedforward network with ReLU and linear projection then regresses $N$ control points per detected centerline as a Bézier curve. Each detected centerline is represented by a fixed-length vector of 2D control-point coordinates [2502.10127].

**Scene graph generation for association prediction**: connectivity is modeled as directed edge prediction in a graph. Lane centerline segments are nodes, and directed edges encode connectivity patterns such as succession, merging, and splitting. Relation types act as semantic predicates over segment pairs, and a Relational Graph Convolutional Network is used to capture these multi-relational dependencies [2502.10127].

**V2X serialization and cloud aggregation**: the output graph is transformed into a geo-referenced format, specifically GeoJSON, and transmitted via V2X such as C-V2X or ITS-5G/DSRC. The serializer extracts nodes and edges, projects centerlines to real-world coordinates using vehicle localization, stores centerlines as GeoJSON LineStrings with attributes such as lane type and curvature, and transmits per-frame map data for multi-vehicle aggregation into a global HD map [2502.10127].

The design emphasis is therefore not only on lane detection quality but also on making lane-graph outputs compact, transmissible, and aggregable across vehicles.

## 3. Geometric representation and graph-theoretic formulation

In the collaborative formulation, lane centerlines are represented as directed graphs $G = (V, E, R)$, where vertices $V$ correspond to lane centerline segments modeled by Bézier curves, edges $E$ represent connectivity, and $R$ denotes relation types [2502.10127]. The use of parametric curves gives a fixed-size encoding for centerline geometry while preserving support for arbitrary segment length. The Bézier parameterization is

$$
B(t)= \sum_{k=0}^{n} \binom{n}{k} (1-t)^{n-k} t^k P_k
$$

where $t \in [0, 1]$ is the curve parameter, $n$ is the number of control points, and $P_k \in \mathbb{R}^2$ are the control points of a lane centerline segment [2502.10127].

Connectivity is also given a matrix form. Graph edges satisfy

$$
E \subseteq \{(x, y) \mid (x, y) \in V^2 \land x \neq y\},
$$

and the paper represents centerline connectivity through an incidence matrix $I \in \{0,1\}^{m \times m}$ for $m$ lane segments, with $I(i,j)=1$ if vertices $i$ and $j$ are connected and $0$ otherwise [2502.10127].

Association prediction is learned with an RGCN encoder and a DistMult decoder. Given $G=(V,E,R)$, node updates at layer $l$ follow

$$
h^{(l+1)}_i= \sigma \left(\sum_{r_{ij} \in R} \sum_{j \in N_i^r} \frac{1}{c_{i,r}} W_r^{(l)} h_j^{(l)} + W_0^{(l)} h_i^{(l)} \right)
$$

where $h_i^{(l)} \in \mathbb{R}^{d_l}$ is the hidden state of node $v_i$, $N_i^r$ is the set of neighbors of $i$ under relation $r$, $c_{i,r}$ is a normalization constant, $W_r^{(l)}$ and $W_0^{(l)}$ are learnable weight matrices, and $\sigma$ is an activation such as ReLU [2502.10127]. After $L$ stacked RGCN layers, initial features from the DETR/DeepLabv3 stages are converted into relation-aware embeddings $e_i \in \mathbb{R}^d$.

The DistMult edge score is

$$
f (v_i, r_{ij} ,v_j) = e_i^T R_r e_j
$$

where $e_i, e_j \in \mathbb{R}^d$ are node embeddings and $R_r \in \mathbb{R}^{d \times d}$ is a diagonal relation-specific matrix learned for relation $r_{ij}$ [2502.10127]. During inference, thresholding $f$ or ranking candidate edges yields the predicted incidence structure.

Coordinate handling is also a core part of the representation. Ground-truth centerlines from nuScenes are converted from the real-world reference frame into the frontal camera reference frame using camera intrinsics and extrinsics, then resampled at a BEV grid resolution of 25 cm and normalized before control-point extraction. The paper assumes calibrated camera intrinsics/extrinsics and ego pose, but does not provide explicit projection equations beyond stating the use of intrinsics and extrinsics for conversion [2502.10127].

## 4. Training protocol, data processing, and objectives

The collaborative HDMapLaneNet is trained and evaluated on nuScenes, which provides ground-truth lane centerline coordinates for 1000 scenes in Boston and Singapore [2502.10127]. Training data are derived from frontal camera images, with centerline coordinates converted into the camera frame through camera intrinsics and extrinsics, resampled at 25 cm in BEV, normalized, and then converted into Bézier control-point annotations [2502.10127].

For curve detection and matching, the method follows Hungarian matching as in STSU, with matching loss

$$
\mathcal{L}_{matching} = \mathcal{L}_{CE} + \lambda \mathcal{L}_1
$$

where $\mathcal{L}_{CE}$ is the detection cross-entropy loss over set predictions and $\mathcal{L}_1$ is the $L_1$ loss on Bézier control-point locations [2502.10127]. For association training, the paper adopts the RGCN link-prediction formulation with DistMult decoding; however, the explicit link-prediction loss is not stated, and the paper instead specifies the encoder–decoder scoring function used for link prediction [2502.10127].

Implementation details are partial rather than exhaustive. DeepLabv3 and DETR are implemented in PyTorch, the graph module uses PyTorch Geometric, DeepLabv3 is pretrained on Cityscapes, and STSU positional encodings are used to inject BEV spatial context. Training was performed on an HPC cluster with 4 NVIDIA Tesla V100 SXM2 GPUs. The paper does not provide explicit optimizer hyperparameters, training schedules, batch sizes, or the value of $\lambda$, and code availability is not specified [2502.10127].

This partial reproducibility profile is significant for interpretation. The architecture, data preprocessing, and matching strategy are described clearly enough to specify the modeling stack, but exact optimization replication is not fully documented.

## 5. Evaluation methodology and reported performance

Three metric families are used to evaluate the 2025 system: Detection Precision–Recall, Detection Ratio, and Connectivity Precision–Recall [2502.10127]. Detection metrics are computed after Hungarian matching that minimizes $L_1$ error over Bézier control points, followed by dense interpolation of Bézier curves. Detection Ratio measures the fraction of ground-truth centerlines matched to at least one estimated centerline. Connectivity metrics compare the estimated incidence matrix with the ground-truth incidence under the matching between predicted and ground-truth nodes [2502.10127].

On nuScenes, HDMapLaneNet is reported to achieve superior association prediction relative to STSU while maintaining comparable detection metrics [2502.10127].

| Method | Detection metrics | Connectivity metrics |
|---|---|---|
| STSU | D-Precision 60.7, D-Recall 54.4, D-Ratio 60.6 | C-Precision 60.5, C-Recall 52.2 |
| HDMapLaneNet | D-Precision 60.5, D-Recall 54.6, D-Ratio 59.2 | C-Precision 75.9, C-Recall 67.1 |

The paper summarizes these differences as a gain of +15.4 percentage points in connectivity precision and +14.9 points in connectivity recall relative to STSU, with essentially unchanged detection precision and recall and a small drop in detection ratio of −1.4 points [2502.10127]. Qualitative results are reported to show accurate lane graphs in adverse weather and occlusions, high precision on straight centerlines, and robust association prediction [2502.10127].

The empirical pattern is important. Detection quality remains roughly constant, but graph connectivity improves substantially. This indicates that the principal contribution is not merely improved centerline localization, but stronger relation inference between extracted lane segments.

## 6. Geo-referencing, V2X transfer, and cloud aggregation

The serialization layer transforms each per-frame lane graph into GeoJSON for transmission [2502.10127]. The serialized content includes extracted lane waypoints or control points and lane connections; centerline curves are projected into real-world coordinates such as WGS84 using vehicle localization data; centerlines are encoded as GeoJSON LineStrings with attributes such as lane type and curvature; and the geo-referenced graph is transmitted over a V2X interface such as C-V2X or ITS-5G/DSRC [2502.10127].

The cloud-side objective is to merge local GeoJSON lane graphs from multiple vehicles into the global geometric lane layer. However, the paper explicitly states that its current contribution focuses on local construction and association. Efficient multi-vehicle graph merging and validation of the global layer are identified as future work, and no specific alignment or matching algorithms, optimization cost functions, or explicit global-frame transforms beyond standard ego-to-world transformations are detailed [2502.10127].

Several deployment assumptions are therefore left open. The framework assumes communication infrastructure sufficient for delivery to a cloud aggregator, but does not specify message formats, compression, timing synchronization, or bandwidth and latency assumptions [2502.10127]. Runtime per frame, model size, computational complexity, memory footprint, and V2X throughput are also not reported [2502.10127].

These omissions delimit the current scope of the method. HDMapLaneNet demonstrates a local perception-to-graph-to-serialization pipeline, but not a complete, validated, fleet-scale global mapping system.

## 7. Historical lineage, related systems, and limitations

Later work on lane graph extraction from aerial imagery identifies an earlier HDMapLaneNet by He et al. (2022), also called “LaneExtraction,” as a segmentation-based baseline [2405.00620]. In that formulation, aerial RGB imagery is processed by a two-head D-LinkNet that predicts both a binary lane segmentation map and a lane direction map. Post-processing then applies thresholding with $\alpha = 0.5$, Guo–Hall thinning to obtain a 1-pixel-wide skeleton, raster-to-graph conversion using 8-connected neighbors, pruning of small components and spurs, and Douglas–Peucker simplification [2405.00620]. The paper evaluating that baseline focuses on non-intersection undirected graphs and reports a reproduced baseline of GEO F1 = 0.813 and TOPO F1 = 0.713, while a diffusion-refined variant reaches GEO F1 = 0.841 and TOPO F1 = 0.774 [2405.00620]. This older usage is important because it shows that the HDMapLaneNet label has already been associated with lane-graph extraction before the 2025 V2X formulation.

The broader lineage of HD-map lane modeling includes systems that are adjacent rather than nominally identical. "LineNet: a Zoomable CNN for Crowdsourced High Definition Maps Modeling in Urban Environments" proposed a pipeline combining a zoomable CNN and TTLane for HD-map modeling from unordered crowdsourced imagery, and reported an average lane error of 31.3 cm despite approximately 5 m GPS noise [1807.05696]. "Flexible 3D Lane Detection by Hierarchical Shape Matching" addresses precise 3D lane detection from point clouds using adaptive-axis global curves and local segment refinement, and is evaluated under strict 10 cm and 30 cm tolerances for HD-map construction [2408.07163]. These systems differ in modality and representation, but all treat lane geometry as a primary HD-map primitive.

The principal limitations of the 2025 HDMapLaneNet are explicitly stated. Reliance on a single front-facing camera limits coverage under field-of-view constraints and occlusion; multi-view images and multi-sensor fusion such as LiDAR are identified as potential improvements. Projection to WGS84 depends on accurate ego pose and calibrated intrinsics and extrinsics, so localization and calibration errors can degrade curve geometry and link prediction. V2X latency and dropouts can affect timeliness of cloud aggregation, yet message synchronization and time alignment are not discussed. Map freshness depends on fleet density and communication, without specified freshness guarantees or validation protocols. Although GeoJSON lane graphs are less privacy sensitive than raw images, secure channels are still required, and privacy and security are not addressed explicitly [2502.10127].

The future-work trajectory is correspondingly clear: multi-view and multi-sensor fusion, improved association learning beyond RGCN and DistMult, efficient multi-vehicle graph merging and validation, roadside-unit integration with edge/cloud orchestration, and uncertainty propagation from local detection and association into the global HD-map layer [2502.10127]. In that sense, HDMapLaneNet occupies a specific position in the lane-mapping literature: it is a camera-only, graph-centric, V2X-enabled framework that advances association prediction and collaborative lane-layer construction, while leaving large-scale global fusion and operational deployment as open research problems.

Source: https://www.emergentmind.com/topics/hdmaplanenet