Accurate and resource-efficient trajectory prediction for surrounding vehicles is essential for autonomous driving. Existing Graph Convolutional Network (GCN) and attention-based methods often consume excessive computational resources by processing all historical trajectories.
To address this challenge, we introduce GP-STFF (Graph Positional and selected Spatial-Temporal Feature Fusion)[cite: 1]:
- Positional Graph (GP): Combines relative distance graph structures with learnable Position Encoding (PE) to capture both relative interactions and absolute spatial positions[cite: 1].
-
Selected Spatial-Temporal Feature Fusion (selected-STFF): Selects spatial interactions at specific key time steps (e.g.,
$t=1$ ) and fuses them with temporal features via a Multi-Head Self-Attention mechanism, drastically reducing computational overhead while retaining accuracy[cite: 1]. - Performance: Achieves a 9% reduction in average RMSE compared to state-of-the-art methods on the public NGSIM dataset[cite: 1].
The GP-STFF model consists of:
-
Spatial Dimension Module: Constructs the relative distance relationship graph
$G_w$ and adds learnable Position Encoding$PE$ to build$G_{GP} = G_w + PE$ [cite: 1]. -
Temporal Dimension Module & Selected-STFF: Processes historical trajectory
$X$ of the target vehicle and fuses it with selective$G_{GP}$ using an LSTM and Multi-Head Self-Attention mechanism[cite: 1]. -
Feature Aggregation & Decoder: Concatenates spatial-temporal features with surrounding vehicle trajectory representations (
$F_n$ ), applies an attention bottleneck, and outputs future trajectory points ($Y$ ) via an LSTM decoder[cite: 1].
- Efficient Spatial-Temporal Integration: Seamlessly integrates spatial factors directly into time-series trajectory learning[cite: 1].
- Positional Graph Representation: Captures both relative inter-vehicle distances and absolute positional importance through learnable position encodings[cite: 1].
- Interpretability: Provides qualitative insights into driver habits and attention distributions across different vehicle types (e.g., Automobiles vs. Heavy Trucks)[cite: 1].
Evaluated on the NGSIM highway dataset (US-101 and I-80) using Root-Mean-Square Error (RMSE) across 5 prediction time steps (1.0 second horizon at 0.2s/step)[cite: 1]:
| Model | 1st Step | 2nd Step | 3rd Step | 4th Step | 5th Step |
|---|---|---|---|---|---|
| Naive LSTM[cite: 1] | 0.1012 | 0.2093 | 0.3384 | 0.4830 | 0.6406 |
| CS-LSTM (CVPRW '18)[cite: 1] | 0.1029 | 0.2023 | 0.3146 | 0.4364 | 0.5674 |
| WSiP (AAAI '23)[cite: 1] | 0.1030 | 0.1998 | 0.3114 | 0.4340 | 0.5631 |
| STA-LSTM (IEEE T-ITS '22)[cite: 1] | 0.0995 | 0.2002 | 0.3130 | 0.4348 | 0.5615 |
| STA-LSTM + Graph[cite: 1] | 0.0827 | 0.1825 | 0.2965 | 0.4183 | 0.5462 |
| GP-STFF (Ours)[cite: 1] | 0.0769 | 0.1727 | 0.2851 | 0.4057 | 0.5311 |
The overall execution pipeline of the GP-STFF trajectory prediction framework follows a systematic, end-to-end workflow:
To ensure consistent temporal tracking and multi-agent spatial interaction modeling, each vehicle within the ROI (Region of Interest) is assigned a unique identifier sequence (ID tracklet) across all consecutive observation frames.
Figure 0: Unique ID sequence assignment for target and surrounding traffic agents across temporal frames.
First, historical trajectory data of the ego vehicle (target vehicle) are extracted across the highway segment over the predefined observation horizon (
On-board sensors (e.g., radar, LiDAR) collect spatial context and historical motion trajectories from surrounding traffic agents within the local grid interaction area.
Figure 2: Perception field capturing surrounding vehicles' historical movement patterns via vehicle radar.
The collected spatial-temporal sequences are processed using the trained GP-STFF architecture[cite: 1]. The positional graph (
The decoder outputs sequence-level future trajectory coordinates (
To quantify prediction uncertainty and multi-modal trajectory distributions, fine-grained spatio-temporal features are analyzed using spatial heatmaps and positional probability density point maps.
Figure 5: Spatial trajectory uncertainty heatmap.
Figure 6: Predicted trajectory probability density distribution map.
# Clone repository
git clone [https://github.com/victor1469/GP-STFF.git](https://github.com/victor1469/GP-STFF.git)
cd GP-STFF
# Install requirements
pip install -r requirements.txt


