HDMapLaneNet: Collaborative Lane Graph Extraction
- The paper introduces a collaborative lane graph extraction method that uses front-facing camera images and Bézier curve encoding to capture lane geometries.
- The methodology involves modular stages including DeepLabv3-based feature extraction, transformer-based lane detection, and RGCN with DistMult for connectivity prediction.
- Empirical results on nuScenes show improved connectivity precision and recall, highlighting robust V2X transmission and efficient local lane graph construction.
Searching arXiv for the cited papers to ground the article. HDMapLaneNet is a lane-centric High-Definition (HD) map construction framework whose central object is the lane graph: lane centerline geometry together with connectivity relations required for map-level reasoning. In current arXiv usage, the name most directly denotes a 2025 system that extracts localized lane centerlines from front-facing camera images, models inter-segment relations through Scene Graph Generation, and serializes the resulting directed graph for collaborative V2X transmission and cloud aggregation (Elghazaly et al., 14 Feb 2025). The same name also appears in later literature as a reference to an earlier segmentation-based aerial-imagery pipeline by He et al. (2022), there also called “LaneExtraction,” which converts lane masks into lane graphs for HD-map extraction (Ruiz et al., 2024). Across these usages, HDMapLaneNet is associated with the construction of the geometric lane layer of HD maps rather than with full semantic map authoring.
1. Terminological scope and problem setting
The contemporary formulation of HDMapLaneNet targets collaborative construction of the localized geometric lane layer of HD maps, specifically the extraction of localized lane centerlines and their connectivity, by combining front-facing camera imagery, scene graph generation for association prediction, and V2X communication for cloud aggregation (Elghazaly et al., 14 Feb 2025). The stated motivation is twofold. First, traditional dependence on dedicated mapping vehicles is expensive and refresh cycles lag real-world changes such as temporary lane closures and construction. Second, single-vehicle perception is constrained by occlusion and field of view, which limits local completeness even when onboard perception is accurate (Elghazaly et al., 14 Feb 2025).
A distinct but historically relevant usage of the term appears in aerial-imagery lane graph extraction. In that line of work, HDMapLaneNet is described as a segmentation-first system that predicts lane masks and direction maps from aerial RGB imagery and then converts the masks into a graph through thinning, vectorization, pruning, and simplification; intersection handling is treated as a separate subtask (Ruiz et al., 2024). The two usages therefore differ in modality, graph semantics, and deployment assumptions: the 2025 framework is camera-only, vehicle-borne, directed, and collaborative through V2X, whereas the earlier aerial formulation is segmentation-based, largely undirected in its non-intersection evaluation, and oriented toward large-area extraction from overhead imagery (Ruiz et al., 2024).
This suggests that “HDMapLaneNet” is best understood not as a single immutable architecture, but as a family of lane-layer HD-map construction approaches organized around graph extraction.
2. Architectural pipeline in the V2X collaborative formulation
The 2025 HDMapLaneNet pipeline is explicitly modular and contains four stages: image-view feature extraction, lane centerline segment detection and curve encoding, scene graph generation for connectivity, and V2X-based serialization and transmission to a cloud aggregator (Elghazaly et al., 14 Feb 2025).
Image-view feature extraction: raw monocular images from the frontal camera are processed by DeepLabv3. The role of DeepLabv3 is to capture multi-scale contextual information via an Atrous Spatial Pyramid Pooling encoder and an encoder–decoder segmentation architecture, producing high-level semantic features that contextualize lane geometry (Elghazaly et al., 14 Feb 2025).
Lane centerline segment extraction and Bézier encoding: Detection Transformer is adapted as the detection head for lane segments and framed as a set prediction problem. A CNN backbone produces low-resolution feature maps, positional encodings following STSU inject Bird’s-eye-view spatial priors, and a transformer encoder–decoder outputs fixed-length embeddings. A 3-layer feedforward network with ReLU and linear projection then regresses control points per detected centerline as a Bézier curve. Each detected centerline is represented by a fixed-length vector of 2D control-point coordinates (Elghazaly et al., 14 Feb 2025).
Scene graph generation for association prediction: connectivity is modeled as directed edge prediction in a graph. Lane centerline segments are nodes, and directed edges encode connectivity patterns such as succession, merging, and splitting. Relation types act as semantic predicates over segment pairs, and a Relational Graph Convolutional Network is used to capture these multi-relational dependencies (Elghazaly et al., 14 Feb 2025).
V2X serialization and cloud aggregation: the output graph is transformed into a geo-referenced format, specifically GeoJSON, and transmitted via V2X such as C-V2X or ITS-5G/DSRC. The serializer extracts nodes and edges, projects centerlines to real-world coordinates using vehicle localization, stores centerlines as GeoJSON LineStrings with attributes such as lane type and curvature, and transmits per-frame map data for multi-vehicle aggregation into a global HD map (Elghazaly et al., 14 Feb 2025).
The design emphasis is therefore not only on lane detection quality but also on making lane-graph outputs compact, transmissible, and aggregable across vehicles.
3. Geometric representation and graph-theoretic formulation
In the collaborative formulation, lane centerlines are represented as directed graphs , where vertices correspond to lane centerline segments modeled by Bézier curves, edges represent connectivity, and denotes relation types (Elghazaly et al., 14 Feb 2025). The use of parametric curves gives a fixed-size encoding for centerline geometry while preserving support for arbitrary segment length. The Bézier parameterization is
where is the curve parameter, is the number of control points, and are the control points of a lane centerline segment (Elghazaly et al., 14 Feb 2025).
Connectivity is also given a matrix form. Graph edges satisfy
0
and the paper represents centerline connectivity through an incidence matrix 1 for 2 lane segments, with 3 if vertices 4 and 5 are connected and 6 otherwise (Elghazaly et al., 14 Feb 2025).
Association prediction is learned with an RGCN encoder and a DistMult decoder. Given 7, node updates at layer 8 follow
9
where 0 is the hidden state of node 1, 2 is the set of neighbors of 3 under relation 4, 5 is a normalization constant, 6 and 7 are learnable weight matrices, and 8 is an activation such as ReLU (Elghazaly et al., 14 Feb 2025). After 9 stacked RGCN layers, initial features from the DETR/DeepLabv3 stages are converted into relation-aware embeddings 0.
The DistMult edge score is
1
where 2 are node embeddings and 3 is a diagonal relation-specific matrix learned for relation 4 (Elghazaly et al., 14 Feb 2025). During inference, thresholding 5 or ranking candidate edges yields the predicted incidence structure.
Coordinate handling is also a core part of the representation. Ground-truth centerlines from nuScenes are converted from the real-world reference frame into the frontal camera reference frame using camera intrinsics and extrinsics, then resampled at a BEV grid resolution of 25 cm and normalized before control-point extraction. The paper assumes calibrated camera intrinsics/extrinsics and ego pose, but does not provide explicit projection equations beyond stating the use of intrinsics and extrinsics for conversion (Elghazaly et al., 14 Feb 2025).
4. Training protocol, data processing, and objectives
The collaborative HDMapLaneNet is trained and evaluated on nuScenes, which provides ground-truth lane centerline coordinates for 1000 scenes in Boston and Singapore (Elghazaly et al., 14 Feb 2025). Training data are derived from frontal camera images, with centerline coordinates converted into the camera frame through camera intrinsics and extrinsics, resampled at 25 cm in BEV, normalized, and then converted into Bézier control-point annotations (Elghazaly et al., 14 Feb 2025).
For curve detection and matching, the method follows Hungarian matching as in STSU, with matching loss
6
where 7 is the detection cross-entropy loss over set predictions and 8 is the 9 loss on Bézier control-point locations (Elghazaly et al., 14 Feb 2025). For association training, the paper adopts the RGCN link-prediction formulation with DistMult decoding; however, the explicit link-prediction loss is not stated, and the paper instead specifies the encoder–decoder scoring function used for link prediction (Elghazaly et al., 14 Feb 2025).
Implementation details are partial rather than exhaustive. DeepLabv3 and DETR are implemented in PyTorch, the graph module uses PyTorch Geometric, DeepLabv3 is pretrained on Cityscapes, and STSU positional encodings are used to inject BEV spatial context. Training was performed on an HPC cluster with 4 NVIDIA Tesla V100 SXM2 GPUs. The paper does not provide explicit optimizer hyperparameters, training schedules, batch sizes, or the value of 0, and code availability is not specified (Elghazaly et al., 14 Feb 2025).
This partial reproducibility profile is significant for interpretation. The architecture, data preprocessing, and matching strategy are described clearly enough to specify the modeling stack, but exact optimization replication is not fully documented.
5. Evaluation methodology and reported performance
Three metric families are used to evaluate the 2025 system: Detection Precision–Recall, Detection Ratio, and Connectivity Precision–Recall (Elghazaly et al., 14 Feb 2025). Detection metrics are computed after Hungarian matching that minimizes 1 error over Bézier control points, followed by dense interpolation of Bézier curves. Detection Ratio measures the fraction of ground-truth centerlines matched to at least one estimated centerline. Connectivity metrics compare the estimated incidence matrix with the ground-truth incidence under the matching between predicted and ground-truth nodes (Elghazaly et al., 14 Feb 2025).
On nuScenes, HDMapLaneNet is reported to achieve superior association prediction relative to STSU while maintaining comparable detection metrics (Elghazaly et al., 14 Feb 2025).
| Method | Detection metrics | Connectivity metrics |
|---|---|---|
| STSU | D-Precision 60.7, D-Recall 54.4, D-Ratio 60.6 | C-Precision 60.5, C-Recall 52.2 |
| HDMapLaneNet | D-Precision 60.5, D-Recall 54.6, D-Ratio 59.2 | C-Precision 75.9, C-Recall 67.1 |
The paper summarizes these differences as a gain of +15.4 percentage points in connectivity precision and +14.9 points in connectivity recall relative to STSU, with essentially unchanged detection precision and recall and a small drop in detection ratio of −1.4 points (Elghazaly et al., 14 Feb 2025). Qualitative results are reported to show accurate lane graphs in adverse weather and occlusions, high precision on straight centerlines, and robust association prediction (Elghazaly et al., 14 Feb 2025).
The empirical pattern is important. Detection quality remains roughly constant, but graph connectivity improves substantially. This indicates that the principal contribution is not merely improved centerline localization, but stronger relation inference between extracted lane segments.
6. Geo-referencing, V2X transfer, and cloud aggregation
The serialization layer transforms each per-frame lane graph into GeoJSON for transmission (Elghazaly et al., 14 Feb 2025). The serialized content includes extracted lane waypoints or control points and lane connections; centerline curves are projected into real-world coordinates such as WGS84 using vehicle localization data; centerlines are encoded as GeoJSON LineStrings with attributes such as lane type and curvature; and the geo-referenced graph is transmitted over a V2X interface such as C-V2X or ITS-5G/DSRC (Elghazaly et al., 14 Feb 2025).
The cloud-side objective is to merge local GeoJSON lane graphs from multiple vehicles into the global geometric lane layer. However, the paper explicitly states that its current contribution focuses on local construction and association. Efficient multi-vehicle graph merging and validation of the global layer are identified as future work, and no specific alignment or matching algorithms, optimization cost functions, or explicit global-frame transforms beyond standard ego-to-world transformations are detailed (Elghazaly et al., 14 Feb 2025).
Several deployment assumptions are therefore left open. The framework assumes communication infrastructure sufficient for delivery to a cloud aggregator, but does not specify message formats, compression, timing synchronization, or bandwidth and latency assumptions (Elghazaly et al., 14 Feb 2025). Runtime per frame, model size, computational complexity, memory footprint, and V2X throughput are also not reported (Elghazaly et al., 14 Feb 2025).
These omissions delimit the current scope of the method. HDMapLaneNet demonstrates a local perception-to-graph-to-serialization pipeline, but not a complete, validated, fleet-scale global mapping system.
7. Historical lineage, related systems, and limitations
Later work on lane graph extraction from aerial imagery identifies an earlier HDMapLaneNet by He et al. (2022), also called “LaneExtraction,” as a segmentation-based baseline (Ruiz et al., 2024). In that formulation, aerial RGB imagery is processed by a two-head D-LinkNet that predicts both a binary lane segmentation map and a lane direction map. Post-processing then applies thresholding with 2, Guo–Hall thinning to obtain a 1-pixel-wide skeleton, raster-to-graph conversion using 8-connected neighbors, pruning of small components and spurs, and Douglas–Peucker simplification (Ruiz et al., 2024). The paper evaluating that baseline focuses on non-intersection undirected graphs and reports a reproduced baseline of GEO F1 = 0.813 and TOPO F1 = 0.713, while a diffusion-refined variant reaches GEO F1 = 0.841 and TOPO F1 = 0.774 (Ruiz et al., 2024). This older usage is important because it shows that the HDMapLaneNet label has already been associated with lane-graph extraction before the 2025 V2X formulation.
The broader lineage of HD-map lane modeling includes systems that are adjacent rather than nominally identical. "LineNet: a Zoomable CNN for Crowdsourced High Definition Maps Modeling in Urban Environments" proposed a pipeline combining a zoomable CNN and TTLane for HD-map modeling from unordered crowdsourced imagery, and reported an average lane error of 31.3 cm despite approximately 5 m GPS noise (Liang et al., 2018). "Flexible 3D Lane Detection by Hierarchical Shape Matching" addresses precise 3D lane detection from point clouds using adaptive-axis global curves and local segment refinement, and is evaluated under strict 10 cm and 30 cm tolerances for HD-map construction (Guan et al., 2024). These systems differ in modality and representation, but all treat lane geometry as a primary HD-map primitive.
The principal limitations of the 2025 HDMapLaneNet are explicitly stated. Reliance on a single front-facing camera limits coverage under field-of-view constraints and occlusion; multi-view images and multi-sensor fusion such as LiDAR are identified as potential improvements. Projection to WGS84 depends on accurate ego pose and calibrated intrinsics and extrinsics, so localization and calibration errors can degrade curve geometry and link prediction. V2X latency and dropouts can affect timeliness of cloud aggregation, yet message synchronization and time alignment are not discussed. Map freshness depends on fleet density and communication, without specified freshness guarantees or validation protocols. Although GeoJSON lane graphs are less privacy sensitive than raw images, secure channels are still required, and privacy and security are not addressed explicitly (Elghazaly et al., 14 Feb 2025).
The future-work trajectory is correspondingly clear: multi-view and multi-sensor fusion, improved association learning beyond RGCN and DistMult, efficient multi-vehicle graph merging and validation, roadside-unit integration with edge/cloud orchestration, and uncertainty propagation from local detection and association into the global HD-map layer (Elghazaly et al., 14 Feb 2025). In that sense, HDMapLaneNet occupies a specific position in the lane-mapping literature: it is a camera-only, graph-centric, V2X-enabled framework that advances association prediction and collaborative lane-layer construction, while leaving large-scale global fusion and operational deployment as open research problems.