---
title: 'HeteroGCLSTM: Graph-LSTM for GNSS Correction'
url: https://www.emergentmind.com/topics/heterogeneous-graph-convlstm-heterogclstm
type: topic
---

# HeteroGCLSTM: Graph-LSTM for GNSS Correction

Searching arXiv for the specified paper and closely related graph-recurrent work to ground the article.
Heterogeneous Graph ConvLSTM (HeteroGCLSTM) is a heterogeneous graph-convolutional recurrent architecture introduced as the core of a dynamic graph regression model for real-time correction of GNSS jamming-induced positioning deviations. In the formulation reported in "Deep Temporal Graph Networks for Real-Time Correction of GNSS Jamming-Induced Deviations" [2509.14000], each GNSS measurement epoch is represented as a heterogeneous receiver–satellite graph rather than as a flat multivariate time series, and a single-layer HeteroGCLSTM processes a short sequence of such graph snapshots to predict the receiver’s 2D horizontal deviation vector. The model is explicitly receiver-centric: it aggregates one-hop spatial context from tracked satellites, models temporal evolution over recent epochs, and outputs a correction term that can be applied online at 1 Hz [2509.14000].

## 1. Definition and problem setting

HeteroGCLSTM is presented within a reframing of GNSS jamming mitigation as dynamic graph regression rather than anomaly detection or absolute position forecasting [2509.14000]. The target is predictive error modeling: the model estimates the receiver’s imminent horizontal position error so that the raw GNSS solution can be corrected online.

The underlying temporal object is a sequence of graph snapshots,
$$
G = (G_1, G_2, \dots, G_T),
$$
where each snapshot at epoch \(t\) is written as
$$
G_t = (V_t, E_t, X_t).
$$
The prediction target is the receiver’s 2D horizontal deviation vector,
$$
Y_t = (\Delta\text{lat}_t, \Delta\text{lon}_t), \quad Y_t \in \mathbb{R}^2.
$$
The paper states the supervised mapping as
$$
\hat{Y}_t = f((G_{t-L+1}, \dots, G_{t-1}); \theta),
$$
while the surrounding prose indicates that an input window of \(L\) snapshots ending at the current epoch is used to predict the target associated with the final snapshot. The text explicitly notes a mild notation inconsistency: the equation lists inputs only up to \(t-1\), whereas the training protocol later makes clear that windows of length 10 are constructed and the temporal model is unrolled over that window, with the final receiver representation mapped to the 2D deviation target [2509.14000].

A central operational consequence of this setup is that one prediction is produced per second, because the GNSS receiver outputs at about 1 Hz. Since the output is the deviation vector itself, the model directly supports online compensation of jamming-induced horizontal error rather than post hoc diagnosis [2509.14000].

## 2. Graph representation and heterogeneity

At each epoch, the receiver environment is encoded as a discrete-time dynamic graph whose topology is explicitly heterogeneous and star-shaped [2509.14000]. Each snapshot contains one receiver node, a variable set of satellite nodes corresponding to currently tracked satellites, and receiver–satellite edges. The receiver occupies the center of the star graph and the satellites form the leaves.

The heterogeneity arises from the distinction between node types. The node types are the receiver node and the satellite nodes. These objects are semantically different and carry different feature sets. The paper explicitly lists the node features as follows: satellite nodes carry \(\text{SNR}\), azimuth, and elevation, while the receiver node carries estimated latitude and estimated longitude [2509.14000].

An edge exists between the receiver node and a satellite node if a signal is being tracked. The topology therefore changes over time as visibility and tracking status vary. The paper does not explicitly state whether the graph is directed or undirected in implementation, does not explicitly define edge weights, and does not describe explicit edge features. What is clear is that the connectivity is receiver–satellite only, the graph is bipartite or star-shaped, and connectivity changes dynamically as satellites appear, disappear, or are lost under jamming [2509.14000].

Time variation enters through both features and topology. Satellite SNR, azimuth, elevation, and receiver latitude/longitude evolve over time, while the set of satellite nodes can also change, so that \(V_t \neq V_{t+1}\). Under jamming, the paper emphasizes that SNR collapses, receiver features become less stable, tracking is intermittently lost, and edges disappear [2509.14000].

This representation is motivated by the variable number of satellites and by the fact that the raw system is naturally relational rather than grid-structured. The paper states that the graph formulation naturally accommodates a variable number of satellites because each snapshot has its own \(V_t\) and \(E_t\). It does not explicitly describe padding, masking, a fixed maximum number of satellites, feature imputation for missing satellites, or the exact normalization procedure. The only explicit statement is that raw time series are transformed into graph snapshots and that outputs are converted back from normalized scale to centimeters for evaluation [2509.14000].

## 3. Architectural design of HeteroGCLSTM

The proposed model is a recurrent GNN whose core is a single-layer HeteroGCLSTM [2509.14000]. The architecture processes a sequence of graph snapshots over a history window, applies recurrent heterogeneous graph convolution, extracts the final hidden state of the receiver node, applies Dropout, and then applies a final Linear layer to produce the 2D deviation prediction.

The model is explicitly receiver-centric. Although the graph contains multiple satellite nodes, only the receiver hidden state is used for final regression. The architectural claim is that HeteroGCLSTM couples spatial aggregation within each graph snapshot with temporal recurrence across snapshots. According to the paper, the layer contains four different HeteroConv sublayers, each assigned to one of the standard LSTM components: input gate, forget gate, output gate, and cell candidate [2509.14000]. Instead of computing gates from ordinary linear transformations of vector inputs, the model first performs heterogeneous graph convolutions and then uses those graph-convolved quantities within LSTM gating dynamics.

The single-layer design is justified by the graph topology. Because each graph snapshot is a star graph, the receiver already reaches all relevant satellites in one hop. Additional graph layers would not expand the receiver’s receptive field and would only repeatedly process the same one-hop neighborhood at additional computational cost [2509.14000]. This one-hop argument is central to the model’s parsimony.

The reported activation functions follow standard LSTM practice: Sigmoid is used for the input, forget, and output gates, and \(\tanh\) is used for the candidate cell state and the hidden-state update. The ablation study and final configuration report one recurrent graph layer, hidden dimensions tested in \(\{16, 32, 64, 128, 256\}\), selected hidden dimension \(256\), selected window length \(10\) s, and output dimension \(2\). The dropout rate is not specified [2509.14000].

## 4. Mathematical formulation and readout

The paper explicitly provides the graph-sequence equation, the deviation-target equation, and the high-level prediction mapping, but it does not print the explicit HeteroGCLSTM gate equations in the provided text [2509.14000]. It states instead that the layer implements the standard LSTM equations and that the outputs of four HeteroConv sublayers are used to compute the gate activations.

A reconstruction given in the source text expresses the update as follows. Let \(X_\tau\) denote node features of graph \(G_\tau\), and let \(H_{\tau-1}\) and \(C_{\tau-1}\) denote the previous hidden and cell states for all nodes. For each time step \(\tau\),
$$
I_\tau = \sigma\!\left(\mathrm{HeteroConv}_i(X_\tau, E_\tau, H_{\tau-1})\right),
$$
$$
F_\tau = \sigma\!\left(\mathrm{HeteroConv}_f(X_\tau, E_\tau, H_{\tau-1})\right),
$$
$$
O_\tau = \sigma\!\left(\mathrm{HeteroConv}_o(X_\tau, E_\tau, H_{\tau-1})\right),
$$
$$
\tilde{C}_\tau = \tanh\!\left(\mathrm{HeteroConv}_c(X_\tau, E_\tau, H_{\tau-1})\right),
$$
followed by the standard LSTM updates
$$
C_\tau = F_\tau \odot C_{\tau-1} + I_\tau \odot \tilde{C}_\tau,
$$
$$
H_\tau = O_\tau \odot \tanh(C_\tau).
$$
The text explicitly characterizes this as a reconstruction faithful to the prose, while also stating that the exact implementation details—such as whether \(X_\tau\) and \(H_{\tau-1}\) are concatenated before graph convolution or transformed separately—are not stated [2509.14000].

The final readout is receiver-specific. If \(h_t^{(r)}\) denotes the hidden state of the receiver node after the final graph in the window, then the readout is described as
$$
z_t = \mathrm{Dropout}(h_t^{(r)}),
$$
$$
\hat{Y}_t = W z_t + b,
$$
with \(W \in \mathbb{R}^{2 \times d_h}\) and \(d_h = 256\). The text notes that this formula is not printed explicitly in the paper but directly follows from the architecture description [2509.14000].

For optimization, the model uses Smooth L1 Loss (Huber Loss) with transition parameter \(\beta = 10^{-2}\). The provided discussion writes the loss as
$$
\mathcal{L}(\hat{Y}, Y) = \mathrm{SmoothL1}_\beta(\hat{Y}-Y),
$$
applied to the two output coordinates. Evaluation uses MAE in latitude and longitude after denormalizing outputs to centimeters, and the results tables report Average Euclidean MAE (cm) [2509.14000].

## 5. Training protocol, datasets, and baselines

The empirical study uses two receiver types: U-blox MAX-M10 (Ublox10), which observes GPS, GLONASS, Galileo, BeiDou, and QZSS, and Ai-Thinker GP01, configured to observe GPS and BeiDou, plus either Galileo or GLONASS [2509.14000].

The jammer profiles differ by receiver. Ublox10 is tested under continuous wave (cw), triple continuous wave (cw3), and frequency modulated (FM). GP01 is tested under cw and cw3. For each active interference mode, six jammer power levels are used: \(-45\), \(-50\), \(-55\), \(-60\), \(-65\), and \(-70\) dBm, yielding 30 scenarios total: 18 for Ublox10 and 12 for GP01 [2509.14000].

Each combination of receiver, interference mode, and jammer power is repeated 50 times. Each repetition consists of 100 s with no interference, 100 s with jammer active, and 80 s of recovery, for a total sequence length of 280 s sampled at approximately 1 Hz [2509.14000]. Additional mixed-mode datasets are also constructed, one per receiver, using 50 measurement sequences at the highest jammer power only and mixed jammer types: cw, cw3, and FM for Ublox10; cw and cw3 for GP01 [2509.14000].

Raw time series are transformed into graph snapshots, and samples are built using a sliding window with window size \(10\) and stride \(10\), so the windows are non-overlapping [2509.14000]. For each of the 30 scenarios, repetitions are split 80:20 into train and test sets, with part of training held out for validation. In a separate low-data analysis, train/test ratios vary from 10:90 to 90:10. Optimization uses Adam with initial learning rate \(0.001\), weight decay for L2 regularization, early stopping based on validation loss, and retention of the best validation checkpoint for final test evaluation. Batch size, maximum number of epochs, exact validation fraction, and dropout probability are not explicitly stated [2509.14000].

The baselines are all time-series models that ignore graph structure. The following baseline configurations are explicitly reported [2509.14000]:

| Baseline | Reported structure | Purpose |
|---|---|---|
| Simple MLP | input window flattened; hidden dense layer with 256 units; ReLU; 2-unit output | test whether a direct non-temporal mapping suffices |
| Uniform CNN | 4 identical Conv1D blocks; kernel size 10; 256 channels; BatchNorm1d; GELU; final flatten + dense(2) | strong modern convolutional time-series baseline |
| Seq2Point-style hierarchical CNN | Conv1D layers with 30 kernels size 10, 30 kernels size 8, 40 kernels size 6, 50 kernels size 5, 50 kernels size 5; then dense 256; then output 2; ReLU throughout | test whether hierarchical temporal feature extraction helps |

The paper’s rationale for expecting graph-temporal modeling to outperform is that GNSS under jamming is a dynamic interaction system rather than a fixed-channel time series. The graph formulation captures a variable number of satellites, explicit receiver–satellite relationships, heterogeneous entities, topology changes due to tracking loss, per-satellite signal and geometry, and the temporal evolution of these relational patterns [2509.14000].

## 6. Empirical performance and data efficiency

The ablation study focuses on the hardest common settings: cw and cw3 at highest power \(-45\) dB for both receivers [2509.14000]. Window sizes tested are
$$
5, 10, 14, 20, 28, 35, 40, 56, 70, 140,
$$
chosen as divisors of the 280-second duration, and hidden dimensions tested are
$$
16, 32, 64, 128, 256.
$$
The main findings are that short windows work best, performance degrades for long windows, hidden size helps up to around 128–256 with diminishing returns, and the best overall regime is 10–20 s windows and 128–256 hidden units. The selected final configuration is window \(= 10\) s and hidden dimension \(= 256\). The specific best ablation points reported are GP01/cw: MAE \(= 2.74\) cm at window 10, hidden 256; GP01/cw3: MAE \(= 7.39\) cm at window 10, hidden 256; Ublox10/cw: MAE \(= 3.97\) cm at window 10, hidden 256; and Ublox10/cw3: MAE \(= 4.59\) cm at window 10, hidden 256 [2509.14000].

Scenario-wise results show that the proposed model consistently attains the lowest mean absolute error among the tested methods [2509.14000]. At \(-45\) dBm, the reported Average Euclidean MAE values include 3.64 cm for GP01/cw, 7.74 cm for GP01/cw3, 4.41 cm for Ublox10/cw, 4.84 cm for Ublox10/cw3, and 4.82 cm for Ublox10/FM. At weaker jamming levels, performance improves, with values in the 1.65–2.08 cm range by \(-60\) to \(-70\) dBm in several settings [2509.14000]. The general trend reported is that errors decrease as jamming weakens and that cw3 is often the hardest case, especially for GP01.

Direct baseline comparisons at strong interference illustrate the margin. At \(-45\) dBm, the reported values are GP01/cw: Proposed \(3.64\), MLP \(5.70\), Seq2Point \(5.03\), CNN \(6.47\); GP01/cw3: Proposed \(7.74\), MLP \(11.40\), Seq2Point \(10.45\), CNN \(11.29\); Ublox10/FM: Proposed \(4.82\), MLP \(9.48\), Seq2Point \(7.13\), CNN \(12.33\); Ublox10/cw: Proposed \(4.41\), MLP \(7.17\), Seq2Point \(6.60\), CNN \(8.82\); Ublox10/cw3: Proposed \(4.84\), MLP \(6.38\), Seq2Point \(5.52\), CNN \(5.99\) [2509.14000]. Within the reported study, these results support the claim that the graph-recurrent model yields the lowest MAE, especially in severe interference.

On mixed datasets pooling all jamming types across powers, the reported MAE is \(4.25 \pm 0.31\) for Ublox10 and \(3.78 \pm 0.27\) for GP01, compared with \(5.82 \pm 0.53\) and \(5.05 \pm 0.87\) for TSMixer/MLP, \(5.12 \pm 0.44\) and \(4.50 \pm 0.19\) for Seq2Point, and \(6.78 \pm 0.82\) and \(8.26 \pm 1.08\) for CNN [2509.14000]. In the low-data split study, the text states that with only 10% training data, the proposed rGNN achieves about 20 cm MAE on GP01, while baselines are in the 36–42 cm range, with the same general trend appearing on Ublox10 [2509.14000]. This suggests that the graph inductive bias is data efficient, although the paper phrases the substantive claim primarily through the observed empirical advantage.

## 7. Interpretation, scope, and limitations

The practical inference loop is described as follows: at each second, construct the current graph snapshot from receiver latitude/longitude, one node per tracked satellite, satellite SNR/azimuth/elevation, and receiver–satellite tracking edges; maintain the latest 10-snapshot history window; unroll the single HeteroGCLSTM layer over the window; take the final receiver hidden state and regress the 2D deviation vector; and use the estimated \((\Delta\text{lat}, \Delta\text{lon})\) to correct the raw GNSS position in real time [2509.14000]. The paper does not provide explicit wall-clock latency or FLOP measurements, but it characterizes the architecture as intentionally lightweight because it uses one recurrent graph layer, one-hop star-graph message passing, and a small output head.

The main contributions claimed in the source are the first framing of GNSS error mitigation as dynamic graph regression, the heterogeneous graph representation of receiver–satellite interactions, a recurrent graph architecture for direct deviation prediction, strong empirical evaluation across two receivers and multiple jammer conditions, and ablation plus comparison against modern time-series baselines [2509.14000].

At the same time, the reported text identifies several omissions. It does not explicitly specify exact HeteroConv message-passing equations, relation types or edge directions, normalization details, batch size, number of epochs, dropout probability, exact validation split, latency benchmarks, or deployment hardware [2509.14000]. The conclusion section is described as empty in the provided text, so future-work claims beyond the body are unavailable.

A common misunderstanding would be to interpret HeteroGCLSTM here as a general-purpose absolute positioning model or a jammer classifier. The paper explicitly positions the task differently: it is not anomaly detection, not absolute position forecasting, and not discrete interference recognition, but direct prediction of the receiver’s horizontal deviation vector for online correction [2509.14000]. Another plausible implication is that the architecture’s single-layer sufficiency is highly contingent on the star topology; the justification given in the source depends on the receiver’s complete one-hop access to all tracked satellites.

Source: https://www.emergentmind.com/topics/heterogeneous-graph-convlstm-heterogclstm