TrajFusionNet: Transformer Crossing Prediction
- TrajFusionNet is a transformer-based model that fuses future pedestrian trajectory, vehicle speed, and scene images to predict crossing intentions.
- It uses a two-branch design with a Sequence Attention Module (SAM) and a Visual Attention Module (VAM) to effectively integrate temporal and spatial cues.
- Evaluations on PIE and JAAD datasets demonstrate state-of-the-art accuracy and low inference latency with lightweight input modalities.
TrajFusionNet is a transformer-based architecture for pedestrian crossing intention prediction that estimates whether a pedestrian observed by an ego-vehicle will cross the roadway $1$–$2$ s into the future. The model is defined by a two-branch design that fuses predictive sequential cues—future pedestrian trajectory and vehicle speed—with a visual encoding of predicted pedestrian motion overlaid on scene images. In the reported evaluation, it achieves state-of-the-art results on PIE, JAAD, and JAAD, while also attaining the lowest total inference time, including model runtime and data preprocessing, among the compared state-of-the-art approaches (Landry et al., 27 Aug 2025).
1. Task formulation and guiding premise
The task addressed by TrajFusionNet is binary crossing intention prediction: determining whether pedestrians in the scene are likely to cross the road or not. The architectural premise is to use future pedestrian trajectory and vehicle speed predictions as priors for crossing classification, rather than relying on a larger set of heavy input modalities. The model therefore emphasizes a small number of lightweight modalities: pedestrian bounding boxes, vehicle speed, and RGB scene images (Landry et al., 27 Aug 2025).
The design combines two representational views of the same anticipatory signal. One view is sequential, where observed and predicted trajectories and vehicle speed are processed as time-ordered tokens. The other view is visual, where predicted pedestrian bounding boxes are overlaid onto scene frames so that spatial context and scene appearance can be exploited jointly with trajectory priors. A plausible implication is that the architecture is intended to preserve temporal forecast structure while still allowing context-sensitive visual disambiguation, particularly in situations where roadway layout or nearby scene content may modulate the meaning of similar trajectory histories.
The reported benchmark setting uses an observation of $0.53$ s ($16$ frames) and a horizon of $1$–$2$ s ($30$–$60$ frames). Overlapping $2$0-frame sequences are extracted with the stride defined by the Kotseruba et al. benchmark, and the model retains a $2$1-frame future prediction for downstream crossing classification (Landry et al., 27 Aug 2025).
2. Two-branch architecture
TrajFusionNet comprises a Sequence Attention Module (SAM), a Visual Attention Module (VAM), and a late-fusion prediction head (Landry et al., 27 Aug 2025).
The SAM branch has two sub-blocks. The first is a non-autoregressive encoder-decoder transformer that takes as input the past $2$2 frames of pedestrian bounding boxes and vehicle speeds. Its input tensor is $2$3, where the five channels are $2$4. The encoder and decoder each have $2$5 layers, $2$6 attention heads, $2$7, and feed-forward dimension $2$8. The decoder input consists of the observed sequence concatenated with a zero-filled future placeholder. The output is projected back to a sequence of length $2$9, from which only the 0-frame prediction 1 is retained.
The second SAM sub-block is an encoder-only transformer operating on a concatenated sequence 2. This sequence contains the past observed 3, each token tagged with a type-ID scalar 4, and the predicted 5, each token tagged with a type-ID scalar 6. The resulting 7 sequence is processed by 8 transformer encoder layers with 9 heads, 0, and feed-forward dimension 1. A final CLS-like pooling is projected to a 2 embedding 3.
The VAM branch implements two parallel branches of the Visual Attention Network, specifically VAN-B2. The first branch, VAN4, processes the first observed frame 5 augmented with colored rectangles marking the observed pedestrian boxes 6. The second branch, VAN7, processes the last observed frame 8 augmented with rectangles marking the predicted boxes 9. Each VAN uses large-kernel attention (LKA), consisting of a small spatial convolution such as $0.53$0, followed by a dilated convolution such as $0.53$1 atrous to capture long-range context, plus a $0.53$2 channel convolution. The two VAN outputs are concatenated and projected to yield a $0.53$3 visual embedding $0.53$4.
Late fusion is performed at embedding level. The $0.53$5-dimensional sequential embedding $0.53$6 and the $0.53$7-dimensional visual embedding $0.53$8 are concatenated and passed through two fully connected layers with $0.53$9 neurons. The final output $16$0 is obtained through a softmax over the two classes, “cross” and “no-cross.” The paper reports that modality self-attention was tested as an alternative fusion mechanism but did not outperform simple concatenation followed by dense merges (Landry et al., 27 Aug 2025).
3. Input modalities and representation engineering
The sequential input is built from pedestrian bounding boxes and vehicle speed. For each frame, the pedestrian bounding box is represented by $16$1. Sequences of $16$2 past frames are used both by the trajectory predictor and by the sequential encoder, and future boxes are predicted for $16$3 frames. Vehicle speed is represented differently across datasets: continuous m/s for PIE, where it is z-score normalized, and ordinal values $16$4 for JAAD, corresponding to stopped through accelerating (Landry et al., 27 Aug 2025).
The preprocessing pipeline centers trajectories by subtracting coordinates at $16$5 from $16$6. This produces a representation in which pedestrian motion is expressed as offsets from the first observed box. This suggests that the model is encouraged to learn relative motion patterns rather than scene-dependent absolute coordinates, which can reduce sensitivity to camera placement and image scale without introducing an additional geometric module.
The visual input preserves raw appearance while injecting trajectory priors directly into the image tensor. RGB frames are used at full scene resolution. For VAM construction, filled rectangles are drawn in the blue and green channels on $16$7 for observed boxes and on $16$8 for predicted boxes; the red channel remains intact for raw appearance cues. The colors follow the ADE20k segmentation palette to maximize contrast. The resulting representation makes the temporal distinction between observation and prediction explicit in image space, while retaining the unmodified appearance channel (Landry et al., 27 Aug 2025).
A recurrent misconception in this problem setting is that increasing modality count is necessarily the dominant path to accuracy. TrajFusionNet is formulated differently: it relies only on bounding boxes, vehicle speed, and raw images, with no pose and no segmentation, yet the reported results indicate competitive or leading performance together with low end-to-end latency (Landry et al., 27 Aug 2025).
4. Mathematical core and optimization procedure
The attention mechanism in TrajFusionNet is the standard scaled dot-product attention, written as
$16$9
Within SAM, this mechanism is used both for trajectory-and-speed prediction and for the subsequent crossing-oriented encoding of observed and predicted sequence tokens (Landry et al., 27 Aug 2025).
Trajectory and speed prediction are trained with mean squared error:
$1$0
where $1$1 and $1$2 modalities. Crossing classification uses weighted cross-entropy:
$1$3
where $1$4 balances the positive (“cross”) and negative class (Landry et al., 27 Aug 2025).
Training follows a modular, greedy schedule. First, the trajectory encoder-decoder is pretrained on $1$5 for $1$6 epochs with learning rate $1$7 and batch size $1$8. Second, that block is frozen and the $1$9-layer transformer encoder plus its dense head are trained on $2$0 for $2$1 epochs with learning rate $2$2 and batch size $2$3. Third, each VAN branch is pretrained independently for $2$4 for $2$5 epochs with learning rate $2$6 and batch size $2$7. Fourth, the full model is assembled, SAM and VAM backbones are frozen, and the final projection plus MLP head are trained for $2$8 epochs with learning rate $2$9 and batch size $30$0; during the first $30$1 epochs, the VAM projection is frozen to focus learning on SAM. Optimization uses AdamW with linear warm-up followed by linear decay, and regularization is provided by early stopping via validation loss together with AdamW weight decay (Landry et al., 27 Aug 2025).
For PIE, the reported training data comprise approximately $30$2 train sequences for the trajectory predictor and approximately $30$3 for the crossing classifier. JAAD$30$4 and JAAD$30$5 use the same splits as Kotseruba et al. (Landry et al., 27 Aug 2025).
5. Quantitative results
On PIE, the reported results are: PCPA (2021), Accuracy $30$6, AUC $30$7, F1 $30$8; Song et al. (2022), Accuracy $30$9, AUC $60$0, F1 $60$1; TrajFusionNet-Small, Accuracy $60$2, AUC $60$3, F1 $60$4; and TrajFusionNet, Accuracy $60$5, AUC $60$6, F1 $60$7. On JAAD$60$8, the reported results are: PCPA, Accuracy $60$9, AUC $2$00, F1 $2$01; PedAST-GCN, Accuracy $2$02, AUC $2$03, F1 $2$04; TrajFusionNet-Small, Accuracy $2$05, AUC $2$06, F1 $2$07; and TrajFusionNet, Accuracy $2$08, AUC $2$09, F1 $2$10. On JAAD$2$11, the reported results are: PCPA, Accuracy $2$12, AUC $2$13, F1 $2$14; PedAST-GCN, Accuracy $2$15, AUC $2$16, F1 $2$17; TrajFusionNet-Small, Accuracy $2$18, AUC $2$19, F1 $2$20; and TrajFusionNet, Accuracy $2$21, AUC $2$22, F1 $2$23 (Landry et al., 27 Aug 2025).
The paper summarizes these results by stating that TrajFusionNet attains the highest accuracy and F1 on PIE and JAAD$2$24 and remains among the top on JAAD$2$25, demonstrating strong generalization. Because the architecture uses future trajectory and speed predictions as priors, a plausible reading of these results is that explicit anticipation can be beneficial even when the final task is binary classification rather than sequence forecasting.
Efficiency is a central reported property. Measured on an NVIDIA RTX 3060 with model runtime and preprocessing included, the inference times and parameter counts are: PCPA, $2$26 ms and $2$27 M parameters; PedestrianGraph$2$28, $2$29 ms and $2$30 M; TrajFusionNet-Small, $2$31 ms and $2$32 M; and TrajFusionNet, $2$33 ms and $2$34 M. The paper therefore attributes the low end-to-end latency to the reliance on bounding boxes, vehicle speed, and raw images, without pose or segmentation (Landry et al., 27 Aug 2025).
6. Ablation findings and broader interpretation
The ablation study evaluates six modifications, reported as PIE/JAAD$2$35/JAAD$2$36 accuracy. Replacing dense fusion with modality self-attention yields $2$37, described as providing no gain. Using a single VAN with both observed and predicted overlays yields $2$38, with a drop on JAAD. Removing sequence-type identifiers in SAM yields $2$39, characterized as a modest drop. Omitting vehicle speed yields $2$40, producing a large drop on PIE and JAAD$2$41. Disabling trajectory prediction in SAM yields $2$42, described as a significant drop. Disabling predicted boxes in VAM yields $2$43, with a JAAD drop (Landry et al., 27 Aug 2025).
The paper draws four direct conclusions from these scenarios. Predicting future trajectories and speeds is critical. Keeping observed and predicted overlays in two separate VANs helps capture before/after context. Explicit sequence-type IDs improve attention focus on past versus future. Vehicle speed materially boosts predictive power on more dynamic datasets (Landry et al., 27 Aug 2025).
These findings clarify the role of TrajFusionNet within pedestrian intention prediction research. Its contribution is not simply the addition of a transformer backbone, but a specific fusion of lightweight trajectory priors, long-sequence attention, and scene-conditioned visual overlays. The ablations also counter a straightforward assumption that more sophisticated fusion is automatically preferable: in this case, embedding-level late fusion with a small MLP outperforms modality self-attention. This suggests that, for this task, representational complementarity between SAM and VAM is sufficiently strong that simple fusion is adequate once each branch has already encoded its modality-specific structure.
Taken together, the reported results position TrajFusionNet as a model in which anticipatory sequence prediction is directly harnessed for intention classification, and in which efficiency is treated as a first-order design constraint rather than a secondary implementation concern (Landry et al., 27 Aug 2025).