---
title: 'MSPF-Net: Cellular Traffic Forecasting'
url: https://www.emergentmind.com/topics/mspf-net
type: topic
---

# MSPF-Net: Cellular Traffic Forecasting

MSPF-Net is a multimodal cellular traffic forecasting framework introduced in “Multimodal Spatiotemporal-Frequency Fusion with Peak Enhancement for Cellular Traffic Forecasting” [2607.07016]. It is designed for settings in which traffic evolution is governed not only by endogenous temporal and spatial regularities, but also by burst-sensitive local spikes and exogenous urban-event signals. The model integrates four components—a Spatiotemporal-Frequency Traffic Encoder, a Peak Enhancement Module, a News Context Representation Module, and a Dynamic Fusion Prediction Module—to jointly learn from historical traffic, peak-aware variation cues, aligned contextual signals, and spatial relations among cellular regions [2607.07016].

## 1. Identity, scope, and nomenclature

MSPF-Net denotes the framework proposed for **cellular traffic forecasting**, not an edge detector, segmentation refiner, or medical segmentation model. The naming is potentially confusable because several unrelated arXiv works use nearby acronyms: **msmsfnet** is a multi-stream and multi-scale fusion net for edge detection [2404.04856], **MSP** is a multiscale superpixel module for semantic segmentation refinement [2112.01746], and **PMFSNet** is a polarized multi-scale feature self-attention network for lightweight medical image segmentation [2401.07579]. In the 2026 traffic-forecasting paper, however, MSPF-Net is explicitly the proposed multimodal framework for forecasting cellular traffic under bursty and event-driven conditions [2607.07016].

The problem scope is forecasting future traffic over a set of spatial units such as base stations or regions. The paper frames the forecasting challenge as one involving five interacting factors: temporal dependencies, spatial dependencies, spectral or periodic patterns, burst-sensitive variations, and exogenous event effects. A central premise is that many existing methods concentrate on intrinsic traffic dynamics alone, whereas real cellular traffic can be perturbed by public events, emergencies, disruptions, and other urban activities reflected in external information streams [2607.07016].

A common misconception is to interpret MSPF-Net as merely a spatiotemporal graph forecaster with auxiliary features appended at the input. The paper’s formulation is more specific. It separates regular endogenous traffic structure, burst-aware local peak structure, and external context into distinct latent streams, then fuses them adaptively through cross-modal attention and gating. This makes the method multimodal in a structural rather than merely feature-augmented sense [2607.07016].

## 2. Forecasting formulation and representation space

The paper defines a cellular network with \(N\) spatial units and a historical window of length \(L\). Historical traffic is written as
\[
\mathbf{X}\in\mathbb{R}^{N\times L},
\]
aligned news or context features as
\[
\mathbf{C}\in\mathbb{R}^{L\times F_c},
\]
and the spatial graph adjacency as
\[
\mathbf{A}\in\mathbb{R}^{N\times N}.
\]
The forecast horizon is represented by
\[
\hat{\mathbf{Y}}\in\mathbb{R}^{N\times H},
\]
with the overall prediction map defined as
\[
\hat{\mathbf{Y}} = f_{\theta}(\mathbf{X}, \mathbf{C}, \mathbf{A}).
\]
This formulation makes traffic history, exogenous context, and spatial structure co-equal inputs to the forecasting function [2607.07016].

The representation strategy is organized around three latent tensors of shared shape \(N\times L\times d\): the endogenous traffic representation \(\mathbf{H}_{\mathrm{traf}}\), the burst-aware representation \(\mathbf{H}_{\mathrm{peak}}\), and the exogenous context representation \(\mathbf{H}_{\mathrm{news}}\). Their common dimensionality is crucial because it permits later cross-modal attention and position-wise dynamic fusion without additional alignment layers at the fusion stage [2607.07016].

The paper’s framing also clarifies the temporal semantics of external information. Context is aligned at the same hourly resolution as traffic, built only from records with timestamps not later than the current time step, and normalized using training-set statistics only. This explicitly rules out future-context leakage and preserves causal validity in the exogenous branch [2607.07016].

## 3. Core architecture

The full architecture is decomposed into four modules that play distinct roles in the forecasting pipeline [2607.07016].

| Module | Primary input | Role |
|---|---|---|
| Spatiotemporal-Frequency Traffic Encoder | \(\mathbf{X}, \mathbf{A}\) | Captures temporal, spectral, and spatial traffic patterns |
| Peak Enhancement Module | \(\mathbf{X}\) | Extracts burst-aware local spike representations |
| News Context Representation Module | \(\mathbf{C}\) | Encodes aligned exogenous contextual signals |
| Dynamic Fusion Prediction Module | \(\mathbf{H}_{\mathrm{traf}}, \mathbf{H}_{\mathrm{peak}}, \mathbf{H}_{\mathrm{news}}\) | Adaptively integrates heterogeneous signals and produces forecasts |

The **Spatiotemporal-Frequency Traffic Encoder** is the principal endogenous branch. It begins with a learnable projection and positional embedding,
\[
\mathbf{H}^{(0)} = \mathrm{Proj}_{x}(\mathbf{X}) + \mathbf{P}_{x},
\qquad
\mathbf{H}^{(0)} \in \mathbb{R}^{N \times L \times d}.
\]
Temporal modeling is then performed by
\[
\mathbf{H}^{(t)} = \mathrm{TempEnc}\!\left(\mathbf{H}^{(0)}\right),
\]
where \(\mathrm{TempEnc}(\cdot)\) is implemented with temporal self-attention followed by feed-forward transformation [2607.07016].

To preserve periodic and oscillatory structure, the encoder explicitly enters the frequency domain:
\[
\mathbf{F} = \left| \mathrm{FFT}\!\left(\mathbf{H}^{(t)}\right) \right|,
\qquad
\mathbf{H}^{(f)} = \mathrm{FreqEnc}(\mathbf{F}).
\]
Temporal and frequency features are fused residually as
\[
\mathbf{H}^{(tf)} = \mathrm{LayerNorm}\!\left(\mathbf{H}^{(t)} + \mathbf{H}^{(f)}\right).
\]
Spatial propagation then acts on the graph with a normalized adjacency based on \(\mathbf{A}+\mathbf{I}\), followed by
\[
\mathbf{H}_{\mathrm{traf}} = \sigma\!\left( \tilde{\mathbf{A}} \, \mathbf{H}^{(tf)} \mathbf{W}_{s} \right).
\]
This sequence gives the model an explicit decomposition into temporal encoding, spectral encoding, residual temporal-frequency fusion, and graph-based spatial aggregation [2607.07016].

Architecturally, this branch differs from purely temporal sequence models or purely graph-temporal models by making frequency-domain structure first-class rather than implicit. The paper’s position is that periodicity and oscillatory behavior are not adequately represented if spectral cues are left to emerge only through temporal self-attention [2607.07016].

## 4. Peak enhancement and exogenous context encoding

The **Peak Enhancement Module** is motivated by the claim that bursty traffic spikes can be smoothed out by global sequence encoders. It therefore constructs an explicit peak descriptor from raw traffic, temporal differences, and short-window statistics. The first-order temporal difference is defined by \(\Delta \mathbf{X}_{:,1}=\mathbf{0}\) and, for \(t=2,\dots,L\),
\[
\Delta \mathbf{X}_{:,t}=\mathbf{X}_{:,t}-\mathbf{X}_{:,t-1}.
\]
The local peak descriptor is then
\[
\mathbf{R}_{\mathrm{peak}} = \mathrm{Concat}\!\Big(
\mathbf{X}, \Delta \mathbf{X},
\mathrm{MaxPool}_{w}(\mathbf{X}) - \mathrm{AvgPool}_{w}(\mathbf{X})
\Big).
\]
This construction encodes raw local level, abrupt changes, and local extreme-versus-average contrast [2607.07016].

The descriptor is projected and filtered by short-window temporal convolution:
\[
\tilde{\mathbf{R}}_{\mathrm{peak}} = \mathrm{Proj}_{p}(\mathbf{R}_{\mathrm{peak}}),
\qquad
\mathbf{U}_{\mathrm{peak}} = \mathrm{Conv1D}_{w_s}\!\left(\tilde{\mathbf{R}}_{\mathrm{peak}}\right),
\]
and the final burst-aware representation is
\[
\mathbf{H}_{\mathrm{peak}} = \mathrm{ReLU}\!\left(\mathbf{U}_{\mathrm{peak}}\mathbf{W}_{p} + \mathbf{b}_{p}\right).
\]
The paper interprets this branch as emphasizing short-term mutation patterns that can be suppressed in smoother, long-range traffic encoders [2607.07016].

The **News Context Representation Module** converts exogenous information into a time-aligned latent sequence. For each hour \(t\), the context vector is
\[
\mathbf{c}_t = [u_t, s_t, e_t, r_t, m_t] \in \mathbb{R}^{5},
\]
where \(u_t\) is the number of active users, \(s_t\) the number of mentioned cities, \(e_t\) the number of extracted entities, \(r_t\) the number of event types, and \(m_t\) the number of news items observed in interval \(t\). Stacking these yields
\[
\mathbf{C}\in\mathbb{R}^{L\times F_c}, \qquad F_c=5.
\]
If no related record exists, \(\mathbf{c}_t\) is set to a zero vector [2607.07016].

The context sequence is embedded as
\[
\mathbf{E}_{\mathrm{news}} = \mathbf{C}\mathbf{W}_{c} + \mathbf{b}_{c} + \mathbf{P}_{c},
\qquad
\mathbf{E}_{\mathrm{news}} \in \mathbb{R}^{L \times d},
\]
then processed by a Transformer encoder,
\[
\mathbf{H}_{\mathrm{news}}^{(c)} = \mathrm{TransformerEnc}\!\left(\mathbf{E}_{\mathrm{news}}\right),
\]
and finally broadcast to all nodes:
\[
\mathbf{H}_{\mathrm{news}} = \mathbf{1}_{N} \otimes \mathbf{H}_{\mathrm{news}}^{(c)}.
\]
This broadcasting step is explicit: the exogenous branch is globally time-aligned but not node-specific in its original construction [2607.07016].

## 5. Dynamic fusion and forecast generation

The **Dynamic Fusion Prediction Module** projects the three latent streams into a shared space:
\[
\bar{\mathbf{H}}_{\mathrm{traf}} = \mathbf{H}_{\mathrm{traf}}\mathbf{W}_{t}, \qquad
\bar{\mathbf{H}}_{\mathrm{peak}} = \mathbf{H}_{\mathrm{peak}}\mathbf{W}_{p}^{\prime}, \qquad
\bar{\mathbf{H}}_{\mathrm{news}} = \mathbf{H}_{\mathrm{news}}\mathbf{W}_{n}.
\]
Traffic is then used as the query in cross-modal attention against the peak and news branches:
\[
\mathbf{G}_{\mathrm{peak}} = \mathrm{Attn}\!\left(
\bar{\mathbf{H}}_{\mathrm{traf}},
\bar{\mathbf{H}}_{\mathrm{peak}},
\bar{\mathbf{H}}_{\mathrm{peak}}
\right),
\]
\[
\mathbf{G}_{\mathrm{news}} = \mathrm{Attn}\!\left(
\bar{\mathbf{H}}_{\mathrm{traf}},
\bar{\mathbf{H}}_{\mathrm{news}},
\bar{\mathbf{H}}_{\mathrm{news}}
\right).
\]
These operations let the endogenous traffic state selectively retrieve burst-aware and exogenous contextual information rather than receiving them through static concatenation [2607.07016].

Adaptive gating is then computed by
\[
[\alpha,\beta,\gamma] =
\mathrm{Softmax}\!\left(
\mathrm{MLP}\!\left(
\bar{\mathbf{H}}_{\mathrm{traf}}
\,\|\, \mathbf{G}_{\mathrm{peak}}
\,\|\, \mathbf{G}_{\mathrm{news}}
\right)
\right),
\]
where \(\alpha,\beta,\gamma \in \mathbb{R}^{N\times L\times 1}\). The fused representation is
\[
\mathbf{Z} =
\alpha \odot \bar{\mathbf{H}}_{\mathrm{traf}}
+ \beta \odot \mathbf{G}_{\mathrm{peak}}
+ \gamma \odot \mathbf{G}_{\mathrm{news}}.
\]
The weights are therefore normalized and competitive across the three modalities at each spatiotemporal position [2607.07016].

After a temporal pooling operation over the historical horizon, the final prediction head is a two-layer MLP:
\[
\hat{\mathbf{Y}} =
\mathbf{W}_{o}^{(2)} \,\mathrm{ReLU}\!\left(
\mathbf{m}\mathbf{W}_{o}^{(1)} + \mathbf{b}_{o}^{(1)}
\right) + \mathbf{b}_{o}^{(2)}.
\]
A consequential point in the paper is that the benefit is not only multimodality itself, but **dynamic** multimodal weighting. The ablation study identifies removal of dynamic fusion as the most damaging modification among the tested module removals [2607.07016].

## 6. Empirical evaluation, ablation results, and limitations

The paper evaluates MSPF-Net on three datasets—**Milano**, **Trento**, and **LTE traffic**—using **MAE** and **RMSE** as metrics, and compares against **LSTM**, **Transformer**, **FEDformer**, **TimeMixer**, **ST-Tran**, **DDGCRN**, **OpenCity**, **FISTGCN**, and **MSCR** [2607.07016]. On Milano, MSPF-Net reports **MAE 2.2850** and **RMSE 3.7490**; on Trento, **MAE 2.6154** and **RMSE 3.9083**; and on LTE traffic, **MAE 0.3875** and **RMSE 0.5216** [2607.07016].

| Dataset | MSPF-Net | Second-best baseline noted in the paper details |
|---|---|---|
| Milano | MAE 2.2850, RMSE 3.7490 | MSCR: MAE 3.9326, RMSE 5.6841 |
| Trento | MAE 2.6154, RMSE 3.9083 | MSCR: MAE 3.4169, RMSE 5.5361 |
| LTE traffic | MAE 0.3875, RMSE 0.5216 | MSCR: MAE 0.5330, RMSE 0.7152 |

The ablation study removes three components: the News Context Representation Module, the Peak Enhancement Module, and Dynamic Fusion. On Milano, the corresponding results are **2.4786 / 4.0584**, **2.5348 / 4.1267**, and **2.6129 / 4.2842**; on Trento, **2.8567 / 4.3271**, **2.9102 / 4.3985**, and **3.0284 / 4.6156**; on LTE, **0.4218 / 0.5711**, **0.4305 / 0.5806**, and **0.4461 / 0.6060**. These numbers support three claims made by the paper: external context improves forecasting, explicit burst modeling improves forecasting, and adaptive fusion outperforms static aggregation [2607.07016].

Several limitations are also explicit. The news/context vector is only five-dimensional and is broadcast to all nodes equally, so the method does not provide node-specific event localization. The paper does not specify many implementation details, including the exact training loss, optimizer, learning rate, batch size, hidden dimension, number of layers or heads, pooling type, kernel values \(w\) and \(w_s\), or runtime and memory statistics. It also notes a mismatch between the LTE dataset description—which mentions auxiliary and visual context from OpenStreetMap—and the method section, whose formal equations focus on traffic, adjacency, and news-derived context [2607.07016].

These constraints shape the current interpretation of MSPF-Net. Its contribution is clearest as a multimodal forecasting architecture that treats cellular traffic prediction as an adaptive fusion problem over endogenous traffic structure, burst-sensitive local dynamics, and aligned exogenous context. A plausible implication is that future variants would need more spatially localized context modeling and fuller implementation disclosure to support stronger reproducibility and finer-grained causal interpretation.

Source: https://www.emergentmind.com/topics/mspf-net