Frequency-Domain Multi-Modality Transportation Modeling Jiewen Deng
Hangchen Liu
Junchen Li
[email protected] Southern University of Science and Technology Shenzhen, China
[email protected] The University of Tokyo Tokyo, Japan
[email protected] Southern University of Science and Technology Shenzhen, China
Boyuan Zhang
Renhe Jiang∗
[email protected] The University of Tokyo Tokyo, Japan
[email protected] The University of Tokyo Tokyo, Japan
Multi-modality transportation refers to urban systems composed of multiple transportation modes, such as traffic flow and public transit, whose dynamics are coupled by shared temporal patterns. Accurate multi-modality transportation forecasting remains challenging because (1) different modalities exhibit distinct spectral characteristics and (2) interact unevenly across frequencies, whereas most existing methods operate primarily in the time domain or rely on coarse feature fusion. To address these limitations, we propose a lightweight yet effective Frequency-Domain MultiModality modeling (FreMo) that explicitly exploits the frequency domain to enable adaptive and selective cross-modality synergy. FreMo disentangles modality-wise spectral refinement from crossmodality synergy and supports plug-and-play integration with general time series backbones. Specifically, FreMo introduces a Modality-Wise Frequency Filter (MFF) to adaptively refine spectral components within each modality, emphasizing informative frequencies while suppressing noise. FreMo further incorporates a Frequency-Guided Synergy Integrator (FSI) that selectively aggregates information across modalities based on their relative contribution at each frequency, facilitating effective cross-modality knowledge sharing while mitigating negative transfer. Extensive experiments on real-world datasets show that FreMo consistently outperforms state-of-the-art baselines, with superior performance and generalization across diverse forecasting scenarios. The code is available at https://github.com/beginner-sketch/FreMo.
ACM Reference Format: Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang. 2026. Frequency-Domain Multi-Modality Transportation Modeling. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD ’26), August 09–13, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/3770855.3818022
Aligned
Bike Outflow Taxi Outflow
Bike Outflow Taxi Outflow
Unaligned
Time
Frequency (Low → High)
(a) Temporal patterns
(b) Frequency spectra 1.0
High consensus
Low consensus Frequency (Low → High)
0.5 0.0
Coherence
arXiv:2607.08475v1 [cs.LG] 9 Jul 2026
Abstract
(c) Spectral coherence between Bike and Taxi Outflow
Figure 1: Multi-modality transportation data with representative modalities (Bike and Taxi Outflow). (a) Temporal patterns exhibit aligned trends but mismatched details. (b) Frequency spectra differ across modalities. (c) Spectral coherence is high at low frequencies but drops at high frequencies.
CCS Concepts
1
• Information systems → Spatial-temporal systems; • Computing methodologies → Artificial intelligence.
Multi-modality transportation data are structured time series that record traffic dynamics over time across multiple spatial units and transportation modes, such as public transit and bike-sharing systems. At each timestamp, observations from different modalities (e.g., Bike and Taxi flows [28, 70, 71]) describe complementary aspects of traffic state. Recently, time series analysis has received increasing attention due to its vital role in urban computing for various downstream tasks [4–6, 24, 34, 38, 40, 45, 49, 66, 74]. Within this field, multi-modality time series modeling has become increasingly important for its ability to integrate diverse information sources and improve predictive performance. It has been widely applied in realworld scenarios such as transportation forecasting [8–10, 52, 79], crime analysis [22, 23, 58], and air quality monitoring [8–10, 20]. However, effectively coordinating multiple modalities for reliable forecasting remains challenging. A central goal of multi-modality
Keywords frequency-domain, multi-modality transportation modeling ∗ Corresponding author.
This work is licensed under a Creative Commons Attribution 4.0 International License. KDD ’26, Jeju Island, Republic of Korea © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2259-2/2026/08 https://doi.org/10.1145/3770855.3818022
Introduction
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
learning is to capture shared patterns across modalities while preserving modality-specific characteristics, enabling complementary information exchange without introducing interference. Real-world data exhibit complex and diverse temporal dynamics across modalities [8, 10]. Different modalities are informative at different temporal scales: some mainly reflect long-term trends, while others are more sensitive to short-term fluctuations. These limitations further expose three key challenges across temporal and frequency dimensions, as illustrated in Figure 1. First, rigid temporal alignment fails to capture fine-grained cross-modality differences. In the time domain (Figure 1(a)), both modalities exhibit aligned trends driven by daily commuting patterns (e.g., morning and evening peaks). At finer time resolutions, clear unaligned details emerge. Bike Outflow shows sharper fluctuations and intermittent drops that do not align well with Taxi Outflow. Most existing methods [15, 22, 55, 61] primarily fuse modalities in the time domain and apply a uniform fusion rule across time steps. This scale-agnostic design may enforce coarse cross-modal consistency, but it cannot separate scale-dependent dynamics or identify when each modality should dominate, often leading to misaligned interactions, noisy information exchange, and limited gains. To explicitly separate temporal scales, a natural way is to move from the time domain to the frequency domain, where different scales can be characterized by distinct frequency components. Second, distinct spectral compositions across modalities challenge uniform modeling. Figure 1(b) reveals clearer modality differences in the frequency domain. At low frequencies, modalities exhibit strong spectral energy, indicating consistent daily patterns. At higher frequencies, their spectra diverge with modality-specific peaks and distinct decay behaviors, suggesting that modality contributions vary across frequency bands. However, existing frequencydomain methods [16, 44, 62, 65, 68, 77] are modality-agnostic, leaving the modality-frequency relationship (i.e., which spectral components are informative for each modality) underexplored. Third, indiscriminate collaboration obscures reliable crossmodality consensus signals. Spectral coherence (Figure 1(c)) measures dependency between modalities at different frequencies. High Consensus appears in low-frequency regions, indicating reliable shared dynamics. In contrast, coherence drops at high frequencies, marking a Low Consensus area dominated by modality-specific variations. This suggests that cross-modality collaboration should be selective and conditioned on frequency-dependent reliability, rather than uniformly enforced across components. We refer to selective and reliability-aware collaboration as synergy across modalities. To address these challenges, we propose a novel FrequencyDomain Multi-Modality modeling framework (FreMo) for transportation forecasting. The central idea of FreMo is to explicitly achieve adaptive and selective cross-modality synergy in the frequency domain. First, to handle diverse spectral characteristics within each modality, we introduce a Modality-Wise Frequency Filter (MFF), which adaptively refines frequency components to emphasize informative bands while suppress noise. MFF enables flexible modality-wise frequency refinement without altering temporal structure. Second, to enable selective cross-modality synergy, we design a Frequency-Guided Synergy Integrator (FSI), which constructs a synergy consensus based on the relative contribution of
Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang
each modality at every frequency. FSI strengthens synergy in frequency components that capture shared cross-modality patterns, while limiting interference in components dominated by modalityspecific variations. FreMo is architecture-agnostic and can be integrated into existing time series models as a plug-and-play module. Our main contributions are summarized as follows: • To the best of our knowledge, FreMo is the first work to systematically formulate frequency-domain multi-modality modeling for transportation forecasting, uncovering the effectiveness of selective and frequency-dependent synergy mechanism. • We propose the MFF that performs adaptive, modality-wise frequency filtering, enabling each modality to emphasize informative components while suppressing noise. • We propose the FSI that identifies and aggregates high-consensus modalities at each frequency, allowing selective and reliable crossmodality synergy. • Extensive experiments demonstrate the superior performance and strong generalization of FreMo across diverse scenarios.
2
Related Works
Uni-Modality Modeling. Uni-modality time series forecasting has been widely studied. Graph neural networks model spatial relations via diffusion-based mechanisms [33] , fully-connected temporal structures [42], and cross-scale [25], while convolutional operators capture local dependencies through spatial [11], temporal [56, 57], spatio-temporal [18, 19, 59], and adaptive convolutions [43]. [29] provides a systematic benchmark, and [80] conducts micro-cluster analysis. Attention mechanisms [48] are adopted for long-range dependencies, including spatial attention [13, 75], temporal attention [54, 76], adaptive embedding strategies [35], with further evidence on capturing global correlations [27, 31]. Recent architectures integrate Transformer-based global modeling with local pattern [14, 39, 41, 53, 73]. Self-supervised learning and large language model have been introduced to enhance time series representations [17, 26, 30, 37, 46, 72, 78]. Further extensions model distribution shift [12], heterogeneity [45], lead-lag under local stationarity [74], and multi-period multi-scale transformations [50, 51]. Frequency-aware modeling has attracted attention for non-stationarity and distribution shifts, including Transformer variants with frequency attention [3, 77], filtering [63, 68], debiasing [44], frequency-domain MLPs[65] and hypervariate graph [64]. Related robustness studies explore frequency-adaptive normalization [62], frequency-masked inference [16], and time-frequency consistency pre-training [72]. Despite their effectiveness, most existing methods focus on temporal/spatial or frequency modeling within a single modality, overlooking cross-modality interactions essential for multi-modality forecasting task. Multi-Modality Modeling. Early works model cross-domain correlations using transfer learning or heterogeneous recurrent architectures across cities or demand types [60, 61]. With the rise of graph neural networks, Zhang et al. [69] combine sparse graph attention with temporal convolution, while MDTP [15] integrates GCNs and LSTMs to fuse multi-source trajectory data. Transformerbased models further capture long-range and cross-modality dependencies. MiST [22] adopts a multi-view Transformer for complex interaction modeling, whereas DMSTGCN [21] fuses flow and speed
Frequency-Domain Multi-Modality Transportation Modeling
representations through multi-dimensional interaction modules. Hierarchical and cross-modality Transformers explicitly model semantic dependencies across data sources. STtrans [55] introduces cross-modality attention. Several studies extend multi-modality forecasting to multi-task and multi-scenario settings by integrating heterogeneous demand signals or jointly optimizing multiple objectives [36, 52, 67]. Recent methods further address heterogeneity in multi-modality data through normalization [9], Gaussian mixture representation learning [8], and self-supervised learning [10]. Despite these advances, existing methods often rely on shallow fusion or limited cross-attention, which is insufficient for learning robust representations from noisy and sparse multi-modality data. Our method explicitly models modality reliability and selectively aggregates information from frequency domain to enable more effective cross-modality synergy.
3
Problem Definition
Let 𝑁 denote the set of spatial nodes and 𝑀 the number of transportation modalities. At time step 𝑡, the multi-modality transportation state is represented as a matrix 𝑋 :,:,𝑡 ∈ R𝑀 ×𝑁 , where each entry corresponds to the traffic demand of a specific modality at a given node. Given a historical window of length 𝑇 , the observations form a multi-modality transportation tensor, denoted as 𝑋 = (𝑋 :,:,1, . . . , 𝑋 :,:,𝑇 ) ∈ R𝑀 ×𝑁 ×𝑇 . Given the historical tensor 𝑋 , the goal of multi-modality transportation forecasting is to predict traffic states for the subsequent 𝑂 time steps, formulated as 𝑌ˆ = (𝑋ˆ :,:,𝑇 +1, . . . , 𝑋ˆ :,:,𝑇 +𝑂 ) ∈ R𝑀 ×𝑁 ×𝑂 .
4
Methodology
Figure 2 depicts an overview of the proposed FreMo. A spatiotemporal encoder initializes the latent representation from the multi-modality transportation tensor 𝑋 . FreMo is then inserted as a plug-and-play refinement module that operates in the frequency domain and comprises two sequential components: (i) ModalityWise Frequency Filter (MFF), which learns modality-wise frequency weights to perform modality-adaptive, phase-preserving filtering; and (ii) Frequency-Guided Synergy Integrator (FSI), which computes frequency-wise synergy weights across modalities, forms a shared synergy consensus, and injects it back through residual feedback. The calibrated representation is transformed back to the temporal domain and passed to a predictor to generate future forecasts.
4.1
Spatio-Temporal Encoder
In this study, we design a spatio-temporal encoder to jointly preserve spatio-temporal contextual information in multi-modality transportation data. The encoder aims to unifiedly capture sequential patterns across time periods, spatial dependencies among nodes, and shared representations across modalities. For notational simplicity, a nonlinear projection function is defined as 𝑓 (𝑥) = 𝑅𝑒𝐿𝑈 (𝑥 · 𝑊 + 𝑏), where 𝑊 and 𝑏 are trainable parameters. The Encoder takes the projected representation 𝐻 (in) = 𝑓 (𝑋 ) ∈ R𝑀 ×𝑁 ×𝑇 ×𝑑 as input. Temporal Encoder. We adopt a Temporal Convolution Network (TCN) [47] as the Temporal Encoder to fuse high-dimensional temporal features. The architecture comprises two dilated causal convolutions to enlarge the receptive field for learned representation,
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
a gated linear unit for incorporating nonlinearity into the network, and a 1 × 1 convolution operation for refining features. The Temporal Encoder is defined as follows: 𝐻 (te) = Conv tanh(𝐻 (in) ∗ 𝑊f ) ⊙ 𝜎 (𝐻 (in) ∗ 𝑊g ) , (1) where ∗ denotes the dilated causal convolution, 𝑊f and 𝑊g are learnable weights for the filter and gate, respectively, and ⊙ indicates element-wise multiplication. The operator Conv(·) represents a 1×1 convolution. Spatial Encoder. To capture spatial dependencies among nodes while circumventing the quadratic complexity of standard selfattention (i.e., O (𝑁 2 )), we introduce a latent-based spatial encoding mechanism motivated by [32]. We initialize a set of learnable latent variables 𝑍 ∈ R𝐿×𝑑 (𝐿 ≪ 𝑑) to act as a global information bottleneck. The Spatial Encoder consists of two cross-attention stages with Layer Normalization [1] followed by a feed-forward network (FFN). First, the latent variables query the node-level information to aggregate global contexts: 𝑍˜ = LN 𝑍 + CrossAttn(𝑍, 𝐻 (te) , 𝐻 (te) ) , (2) where LN(·) denotes Layer Normalization, and CrossAttn(𝑄, 𝐾, 𝑉 ) represents the cross-attention mechanism. Then, the updated latents 𝑍˜ act as keys and values to refine the node representations, distributing the global information back to the spatial units: ˜ 𝑍˜ ) , 𝐻 (se) = LN 𝐻 (te) + CrossAttn(𝐻 (te) , 𝑍, (3) 𝐻 (se) ← 𝐻 (se) + FFN(𝐻 (se) ). The resulting representation 𝐻 (se) preserves the original multimodality spatio-temporal structure while incorporating latent-guided spatial interactions with a linear complexity O (𝐿𝑁 ). A stacked architecture comprising multiple layers is employed to model spatio-temporal interactions across modalities effectively. In each layer, the input undergoes a Temporal Encoder followed by a Spatial Encoder via residual learning, transforming the raw data into a comprehensive latent representation 𝐻 ∈ R𝑀 ×𝑁 ×𝑇 ×𝑑 .
4.2
Modality-Wise Frequency Filter
Multi-modality transportation signals exhibit frequency-dependent structure, where informative components and noise distribute differently across frequencies. Modalities also have distinct spectral profiles: some are dominated by low-frequency periodic trends, whereas others contain pronounced high-frequency fluctuations. Moreover, informative frequencies may vary across spatial units. These factors make a single shared frequency filter suboptimal: it may over-suppress useful components for certain modalities/nodes, or retain excessive noise for others. To address this challenge, we propose the Modality-Wise Frequency Filter (MFF), which performs adaptive and independent frequency filtering for each modality and node. MFF derives modality-wise frequency gates from spectral amplitude profiles and applies them in a phase-preserving manner, so that informative frequency components are emphasized while noise is suppressed without distorting temporal alignment. Frequency-Domain Transformation. We first expose frequency components explicitly by transforming temporal representations
FreMo: Frequency-Domain Multi-Modality Transportation Modeling
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Input 𝑿𝑿
Value
t … t N Modality 1 Modality M
T
Frequency-Domain Multi-Modality Modeling (FreMo)
Spatio-Temporal Encoder 𝑯𝑯
Frequency-Domain Multi-Modality Modeling
𝑯𝑯
M
𝑯𝑯
Modality 1
𝑭𝑭
𝑯𝑯𝑭𝑭 𝟏𝟏,:,:
Spatial Emb. 𝑬𝑬
Frequency
Amp
⦷
𝑭𝑭
𝑯𝑯 𝒎𝒎,:,:
Frequency
Get Weight Amp. Gen. 𝒈𝒈𝑚𝑚 Phase Preserving
Frequency-Guided Synergy Integrator
Modality-Wise Frequency Filter
𝓕𝓕 −𝟏𝟏
� 𝑯𝑯
Frequency-Guided Synergy Integrator (FSI)
𝒘𝒘𝒎𝒎
� 𝑭𝑭 𝒎𝒎,:,: 𝑯𝑯
Filtered Frequency
Frequency Stack Score Stack
Synergy Consensus
Get Amp. Modality as Score Softmax
…
O
𝑯𝑯𝟏𝟏,:,:
𝑯𝑯𝑭𝑭 𝑴𝑴,:,:
𝑯𝑯𝑭𝑭 𝟐𝟐,:,:
…
Value
t … t N Modality 1 Modality M
𝓕𝓕
Per-Modality
� 𝑯𝑯
� Output 𝒀𝒀
𝑯𝑯𝟐𝟐,:,:
Modality-Wise Frequency Filter (MFF)
Predictor Value
𝑯𝑯𝑴𝑴,:,:
M
Amp
Value
Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang
𝜶𝜶𝟏𝟏 𝜶𝜶𝟐𝟐
𝑯𝑯𝑭𝑭 𝓕𝓕 −𝟏𝟏 � 𝑯𝑯 Feedback
𝜸𝜸 Modality-Shared Scalar
𝜶𝜶𝒎𝒎
Figure 2: An overview of the proposed FreMo. Left: As a plug-and-play component, FreMo can be integrated into time series frameworks to enable cross-modality synergy. Right: Detailed design of FreMo. Modality-Wise Frequency Filter (MFF) adaptively learns modality-wise frequency weights 𝑤 from spectral amplitudes by incorporating spatial embedding 𝐸, and applies them as soft gates to preserve informative frequency components while suppressing noise in a phase-preserving manner. FrequencyGuided Synergy Integrator (FSI) computes frequency-wise synergy weights 𝛼 across modalities, aggregates the filtered spectra into a unified synergy consensus, and injects it back through a learnable scalar 𝛾 for residual calibration. into the frequency domain via real-valued FFT (rFFT): F 𝐻𝑚,𝑛,: = F (𝐻𝑚,𝑛,: ) ∈ C𝐹 ×𝑑 , (4) 𝑇 where F (·) denotes the rFFT operator and 𝐹 = 2 +1 is the number of frequency bins. We then compute the amplitude spectrum as: F 𝐴𝑚,𝑛,: = |𝐻𝑚,𝑛,: | ∈ R𝐹 ×𝑑 ,
(5)
where | · | denotes the modulus of the complex number. Modality-Wise Frequency Weight Generation. The amplitude spectrum provides a direct cue of how strong each frequency component is. Based on 𝐴𝑚,𝑛,: , we learn a frequency-wise gating vector that highlights informative spectral bands while attenuating noise. To capture both modality differences and spatial variation, we use an independent weight generator for each modality and incorporate a learnable node embedding as local context. Specifically, for the 𝑚 th modality, we first aggregate the amplitude spectrum along the channel dimension to obtain 𝐴¯𝑚,𝑛,: ∈ R𝐹 , and then concatenate it with the node embedding 𝐸𝑛 ∈ R𝐿 to form the context 𝜙𝑚,𝑛,: ∈ R𝐹 +𝐿 : 𝐴¯𝑚,𝑛,: = Pool𝑑 (𝐴𝑚,𝑛,: ), (6) 𝜙𝑚,𝑛,: = [𝐴¯𝑚,𝑛,: ∥ 𝐸𝑛 ], where Pool𝑑 (·) denotes average pooling over the hidden dimension and ∥ represents concatenation. The modality-wise weight generator G𝑚 (·) maps the context to a frequency-bin gate: 𝑤𝑚,𝑛,: = 𝜎 G𝑚 (𝜙𝑚,𝑛,: ) ∈ [0, 1] 𝐹 , (7) where 𝜎 (·) is the Sigmoid activation function. The resulting 𝑤𝑚,𝑛,: is a modality-wise frequency weight for modality 𝑚 at node 𝑛, which assigns a retention score to each frequency bin: larger values keep the corresponding components, while smaller values attenuate them, thereby serving as a soft frequency gate within the modality.
Phase-Preserving Frequency Filtering. Based on the learned gate 𝑤𝑚,𝑛,: , we then modulate the complex spectrum to perform modality-wise frequency filtering: F F 𝐻˜ 𝑚,𝑛,: = 𝐻𝑚,𝑛,: ⊙ 𝑤𝑚,𝑛,:, (8) where ⊙ denotes the element-wise Hadamard product with broadcasting over the hidden dimension. Since 𝑤𝑚,𝑛,: is real-valued, this operation rescales amplitudes while preserving phase, thereby suppressing noise without disrupting temporal alignment. After obtaining filtered frequency representations of each modality independently, we gather the results as 𝐻˜ F ∈ C𝑀 ×𝑁 ×𝐹 ×𝑑 .
4.3
Frequency-Guided Synergy Integrator
While MFF independently refines spectral components for each modality, effective forecasting also depends on cross-modality synergy. Synergy is inherently frequency-dependent: low-frequency trends (e.g., daily commuting) often consistent across modalities, whereas high-frequency components are more likely to be modalityspecific details. Existing fusion methods [8, 15, 21, 22, 69] typically operate in the time domain or adopt shallow interactions, and thus do not explicitly model frequency-wise modality consensus. To bridge this gap, we propose the Frequency-Guided Synergy Integrator (FSI). FSI constructs a frequency-wise synergy consensus by assigning a relative modality weight at each frequency, and then uses consensus to calibrate modality-wise spectral representations. Frequency-Wise Synergy Weighting. To quantify how much each modality should contribute at each frequency bin, we compute F a frequency-wise strength score from the filtered spectrum 𝐻˜ 𝑚,𝑛,: : F 𝑆𝑚,𝑛,: = Pool𝑑 (|𝐻˜ 𝑚,𝑛,: |),
(9)
∈ R𝐹 serves as a proxy of spectral strength for modal-
where 𝑆𝑚,𝑛,: ity 𝑚 at node 𝑛, and larger values indicate stronger responses at
Frequency-Domain Multi-Modality Transportation Modeling
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
the corresponding frequencies. We then apply the Softmax over modalities to obtain the frequency-wise synergy weights: exp(𝑆𝑚,𝑛,: ) 𝛼𝑚,𝑛,: = Í , exp(𝑆𝑚 ′ ,𝑛,: )
(10)
𝑚 ′ ∈𝑀
where 𝛼𝑚,𝑛,: ∈ R𝐹 gives the relative contribution of 𝑚 th modality Í at each frequency bin (with 𝑚 𝛼𝑚,𝑛,𝑓 = 1), enabling frequencygranular competition and supporting subsequent synergy consensus construction. Large 𝛼𝑚,𝑛,𝑓 indicates that modality 𝑚 is more reliable for forming the cross-modality consensus at frequency 𝑓 . Synergy Consensus Construction. Given the learned frequencywise synergy weights 𝛼𝑚,𝑛,: , we form a unified synergy consensus by aggregating the filtered spectra of all modalities with frequencygranular weighting: C𝑛,: =
𝑀 ∑︁
F 𝛼𝑚,𝑛,: ⊙ 𝐻˜ 𝑚,𝑛,: ,
(11)
𝑚=1
where ⊙ denotes element-wise multiplication with broadcasting over the hidden dimension. The resulting C𝑛,: ∈ C𝐹 ×𝑑 summarizes the cross-modality consensus at each frequency bin, serving as a shared frequency-domain reference for all modalities. Feedback Injection and Reconstruction. The synergy consensus is then fed back to each modality via residual injection to calibrate its original spectrum: F F 𝐻ˆ𝑚,𝑛,: = 𝐻𝑚,𝑛,: + 𝛾 · C𝑛,:,
Model Optimization
Given the refined multi-modality representation obtained as described above, we employ a Predictor to project it into the future horizon. Specifically, our Predictor contains a gated linear unit and a linear projection layer, which is formulated as follows: 𝑌ˆ𝑚,𝑛,: = (𝐻ˆ𝑚,𝑛,: ∗ 𝑊vc ) ⊙ 𝜎 (𝐻ˆ𝑚,𝑛,: ∗ 𝑊gc ) ∗ 𝑊o, (14) where𝑊vc,𝑊gc ∈ R𝑇 ×𝑑 ×𝑑 denote temporal convolution kernels, and 𝑊o ∈ R𝑑 ×𝑂 projects the gated features to the next 𝑂-step horizon. The model is trained in an end-to-end manner by minimizing the Mean Absolute Error (MAE) between the predicted multi-modality transportation tensor 𝑌ˆ ∈ R𝑀 ×𝑁 ×𝑂 and the ground truth 𝑌 : L=
𝑀 ∑︁ 𝑁 ∑︁ 𝑂 ∑︁ 1 |𝑌ˆ𝑚,𝑛,𝑡 − 𝑌𝑚,𝑛,𝑡 |. 𝑀 × 𝑁 × 𝑂 𝑚=1 𝑛=1 𝑡 =1
5 Experiment 5.1 Experiment Setup Table 1: Summary of multi-modality transportation datasets.
(13)
where F −1 (·) denotes inverse rFFT, yielding the augmented temporal representation 𝐻ˆ ∈ R𝑀 ×𝑁 ×𝑇 ×𝑑 . The complete workflow of FreMo is summarized in Algorithm 1.
4.4
Require: Input hidden representation H ∈ R𝑀 ×𝑁 ×𝑇 ×𝑑 ; Learnable node embeddings E ∈ R𝑁 ×𝑙 ; 𝑀 ; Modality-wise weight generators {𝑔𝑚 (·)}𝑚=1 Modality-shared learnable scalar 𝛾. Ensure: Refined hidden representation Ĥ. 1: // Step 1: Frequency Domain Transformation 2: H𝐹 ← rFFT(H, dim = 𝑇 ) ▷ Transform to frequency domain 3: A ← |H𝐹 | ▷ Compute spectral amplitude 4: // Step 2: Modality-Wise Frequency Filter (MFF) 5: for 𝑚 = 1 to 𝑀 do 𝐹 for modality 6: Extract amplitude A𝑚 and complex features H𝑚 𝑚 7: Ā𝑚 ← Mean(A𝑚 , dim = 𝑑) ▷ Aggregate hidden dimension 8: 𝜙𝑚 ← Concat( Ā𝑚 , E) ▷ Inject node embedding 9: w𝑚 ← 𝑔𝑚 (𝜙𝑚 ) ▷ Generate frequency weights ∈ [0, 1] 𝐹 ← H𝐹 ⊙ w 10: H̃𝑚 ▷ Phase-preserving filtering 𝑚 𝑚 𝐹 |, dim = 𝑑) 11: S𝑚 ← Mean(| H̃𝑚 ▷ Compute reliability score 12: end for 13: // Step 3: Frequency-Guided Synergy Integrator (FSI) 14: 𝜶 ← Softmax(Stack([S1 , . . . , S𝑀 ]), dim = 𝑀) ▷ Calculate frequency-wise synergy weight Í𝑀 𝐹 15: C 𝐹 ← 𝑚=1 𝛼𝑚 ⊙ H̃𝑚 ▷ Construct synergy consensus 16: Ĥ𝐹 ← H𝐹 + 𝛾 · C 𝐹 ▷ Inject consensus via residual connection 17: Ĥ ← irFFT( Ĥ𝐹 , dim = 𝐹 ) ▷ Reconstruct to time domain 18: return Ĥ
(12)
where 𝛾 ∈ R is a learnable scalar initialized to zero and shared across modalities. This global scaling factor regulates the strength of synergy feedback, balancing modality-specific spectra with the shared consensus. Finally, the calibrated frequency representation is transformed back to the temporal domain: F 𝐻ˆ𝑚,𝑛,: = F −1 (𝐻ˆ𝑚,𝑛,: ).
Algorithm 1 Pseudocode of the proposed FreMo
(15)
Dataset
Time Period
Nodes
NYC
2016/4/1 ∼ 2016/6/30 (0.5h)
98
DC
2015/10/24 ∼ 2016/1/31 (1h)
108
Chicago
2016/4/1 ∼ 2016/6/30 (0.5h)
510
Modalities
Horizons
Bike Inflow, Bike Outflow, Taxi Inflow, Taxi Outflow
Input: 16 ↓ Out: 1/2/3
5.1.1 Datasets. To fully evaluate the proposed FreMo under different scenarios, we conduct experiments on three real-world multimodality transportation datasets, as presented in Table 1. These city-level datasets, collected from New York City (NYC), Washington D.C. (DC), and Chicago, span distinct spatio-temporal scales and coverage. Specifically, multi-modality is represented by four distinct transport modes, including Bike Inflow, Bike Outflow, Taxi Inflow, and Taxi Outflow. 5.1.2 Implementation Details. For the model, the spatio-temporal encoder consists of four layers with a hidden dimension of 64. The temporal encoder adopts a kernel size of 3, while the spatial encoder employs a latent dimension of 64. The node embedding dimension in FreMo is set the same as the latent dimension. The training
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang
Table 2: Performance on multi-modality transportation datasets for the first/second/third horizon. The best accuracy is highlighted in bold, and the second-best performance is underlined. Bike Inflow
NYC MAE
Bike Outflow RMSE
MAE
RMSE
Taxi Inflow MAE
Taxi Outflow RMSE
(GWN) [57] 2.37 / 2.66 / 2.91 6.89 / 7.63 / 8.23 2.57 / 2.91 / 3.17 7.42 / 8.26 / 8.85 7.00 / 8.23 / 9.59 13.30 / 16.11 / 18.99 AGCRN [2] 2.40 / 2.70 / 3.02 6.96 / 7.69 / 8.41 2.60 / 2.97 / 3.31 7.39 / 8.32 / 9.15 7.08 / 8.10 / 9.16 13.42 / 15.71 / 18.16 2.28 / 2.44 / 2.69 6.68 / 7.05 / 7.62 2.45 / 2.60 / 2.83 7.07 / 7.46 / 8.02 7.24 / 8.04 / 8.78 13.82 / 15.82 / 17.71 MTGNN [56] TimesNet [53] 2.65 / 2.88 / 3.21 4.61 / 5.11 / 5.88 2.81 / 3.10 / 3.45 5.03 / 5.65 / 6.34 7.54 / 8.44 / 9.25 13.61 / 15.98 / 18.04 iTransformer [39] 3.05 / 3.75 / 4.52 5.61 / 7.22 / 8.85 3.31 / 4.10 / 4.80 6.20 / 7.99 / 9.44 8.11 / 10.32 / 12.54 14.40 / 19.16 / 23.98 STAEformer [35] 3.57 / 3.67 / 3.81 7.13 / 7.42 / 7.78 3.44 / 3.57 / 3.73 7.10 / 7.40 / 7.74 8.76 / 10.33 / 11.89 15.40 / 19.51 / 24.06 2.32 / 2.52 / 2.76 6.78 / 7.26 / 7.84 2.57 / 2.83 / 3.01 7.48 / 8.10 / 8.54 7.93 / 8.73 / 9.61 14.10 / 16.51 / 19.06 MiST [22] STtrans [55] 2.72 / 3.06 / 3.38 7.74 / 8.54 / 9.35 2.96 / 3.27 / 3.56 8.33 / 8.93 / 9.67 8.46 / 9.62 / 11.03 15.67 / 17.74 / 20.42 COCOA [7] 2.82 / 3.11 / 3.48 5.06 / 5.82 / 6.53 2.93 / 3.28 / 3.60 5.43 / 6.23 / 6.85 8.05 / 9.81 / 11.95 13.98 / 17.85 / 22.18 MoSSL [10] 2.26 / 2.39 / 2.58 4.10 / 4.28 / 4.51 2.43 / 2.55 / 2.67 4.42 / 4.65 / 4.86 6.73 / 7.53 / 8.38 11.89 / 14.13 / 16.22 FreMo (Ours) 2.22 / 2.33 / 2.42 4.06 / 4.21 / 4.39 2.34 / 2.45 / 2.55 4.40 / 4.61 / 4.80 6.43 / 7.01 / 7.54 11.83 / 13.85 / 15.78 Bike Inflow
DC MAE
Bike Outflow RMSE
MAE
RMSE
(GWN) [57] 0.87 / 0.93 / 0.99 1.51 / 1.60 / 1.66 0.88 / 0.94 / 1.00 1.53 / 1.62 / 1.67 AGCRN [2] 0.99 / 1.10 / 1.24 1.49 / 1.63 / 1.79 1.00 / 1.11 / 1.25 1.51 / 1.65 / 1.81 MTGNN [56] 0.67 / 0.75 / 0.81 1.23 / 1.38 / 1.48 0.69 / 0.77 / 0.82 1.27 / 1.40 / 1.50 TimesNet [53] 0.54 / 0.59 / 0.62 1.17 / 1.27 / 1.32 0.56 / 0.60 / 0.63 1.20 / 1.30 / 1.35 iTransformer [39] 0.63 / 0.71 / 0.77 1.38 / 1.55 / 1.65 0.64 / 0.71 / 0.77 1.41 / 1.55 / 1.66 STAEformer [35] 0.64 / 0.71 / 0.73 1.23 / 1.31 / 1.35 0.65 / 0.71 / 0.75 1.28 / 1.34 / 1.38 MiST [22] 0.72 / 0.80 / 0.91 1.37 / 1.55 / 1.68 0.74 / 0.81 / 0.92 1.43 / 1.59 / 1.71 STtrans [55] 0.68 / 0.81 / 0.87 2.09 / 2.42 / 2.67 0.74 / 0.83 / 0.89 2.26 / 2.43 / 2.63 COCOA [7] 0.70 / 0.71 / 0.79 1.28 / 1.36 / 1.45 0.71 / 0.74 / 0.80 1.30 / 1.38 / 1.47 MoSSL [10] 0.56 / 0.63 / 0.66 1.11 / 1.21 / 1.29 0.61 / 0.65 / 0.67 1.20 / 1.27 / 1.32 FreMo (Ours) 0.46 / 0.51 / 0.54 1.07 / 1.18 / 1.21 0.50 / 0.52 / 0.55 1.16 / 1.20 / 1.26 Bike Inflow
Chicago MAE
MAE
RMSE
(GWN) [57] 0.43 / 0.45 / 0.48 1.79 / 1.90 / 2.03 0.43 / 0.46 / 0.49 1.88 / 1.99 / 2.08 AGCRN [2] 0.51 / 0.52 / 0.60 1.35 / 1.52 / 1.75 0.52 / 0.53 / 0.60 1.45 / 1.57 / 1.73 MTGNN [56] 0.32 / 0.34 / 0.35 1.19 / 1.27 / 1.37 0.33 / 0.34 / 0.35 1.28 / 1.34 / 1.42 TimesNet [53] 0.34 / 0.36 / 0.39 1.25 / 1.38 / 1.56 0.35 / 0.37 / 0.39 1.35 / 1.43 / 1.54 iTransformer [39] 0.38 / 0.43 / 0.49 1.46 / 1.76 / 2.07 0.38 / 0.41 / 0.45 1.54 / 1.72 / 1.87 STAEformer [35] 0.35 / 0.35 / 0.37 1.91 / 1.92 / 1.96 0.37 / 0.38 / 0.40 1.98 / 2.01 / 2.07 MiST [22] 0.34 / 0.38 / 0.43 1.35 / 1.55 / 1.76 0.35 / 0.39 / 0.43 1.42 / 1.57 / 1.74 STtrans [55] 0.37 / 0.41 / 0.45 2.26 / 2.54 / 2.74 0.39 / 0.42 / 0.45 2.28 / 2.35 / 2.73 COCOA [7] 0.34 / 0.35 / 0.38 1.30 / 1.45 / 1.63 0.34 / 0.36 / 0.38 1.35 / 1.45 / 1.58 0.31 / 0.32 / 0.34 1.20 / 1.26 / 1.31 0.32 / 0.33 / 0.34 1.21 / 1.24 / 1.31 MoSSL [10] FreMo (Ours) 0.28 / 0.29 / 0.30 1.15 / 1.20 / 1.25 0.28 / 0.29 / 0.30 1.17 / 1.20 / 1.24
phase is performed using the Adam optimizer, and the batch size is 16. The inputs are normalized by Z-Score. We implement FreMo using PyTorch, and conduct all experiments on a GPU server with NVIDIA GeForce GTX 2080 Ti graphic cards. For evaluation, we adopt Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) as evaluation metrics. 5.1.3 Baselines. We choose ten state-of-the-art models as baselines, which can be broadly categorized into two groups. (1) Unimodality forecasting models: well-established general time series forecasting models, including Graph WaveNet (GWN) [57], AGCRN [2], MTGNN [56], TimesNet [53], iTransformer [39], and STAEformer [35]. (2) Multi-modality forecasting models: representative methods explicitly designed for multi-modality modeling, including MiST [22], STtrans [55], COCOA [7], and MoSSL [10].
5.2
Overall Performance
The average results of three repeated experiments are listed in Table 2, with the best in bold and the second underlined. Among the
RMSE
7.72 / 9.05 / 10.29 14.98 / 17.48 / 19.69 7.81 / 9.03 / 10.19 15.03 / 17.41 / 19.29 7.97 / 8.97 / 9.96 15.47 / 17.59 / 19.26 8.45 / 9.52 / 10.57 15.94 / 18.15 / 20.37 8.69 / 10.85 / 12.80 16.15 / 20.60 / 24.37 9.68 / 10.96 / 12.14 16.74 / 19.34 / 21.75 8.44 / 9.19 / 10.82 15.56 / 17.93 / 19.33 8.78 / 10.04 / 11.28 16.66 / 19.12 / 21.39 8.80 / 10.69 / 12.83 16.20 / 20.15 / 24.02 7.37 / 8.18 / 8.90 13.80 / 15.61 / 17.23 7.07 / 7.67 / 8.25 13.49 / 14.71 / 15.86
Taxi Inflow
Taxi Outflow
MAE
RMSE
MAE
RMSE
3.59 / 3.78 / 4.00 2.76 / 3.00 / 3.29 2.78 / 3.12 / 3.38 2.55 / 2.78 / 2.94 3.10 / 3.70 / 4.24 2.94 / 3.31 / 3.70 3.05 / 3.64 / 4.09 3.13 / 3.51 / 3.81 2.68 / 3.03 / 3.31 2.49 / 2.75 / 2.89 2.39 / 2.60 / 2.82
6.57 / 6.85 / 7.21 4.69 / 5.22 / 5.68 4.90 / 5.73 / 6.21 4.38 / 4.87 / 5.58 5.35 / 6.60 / 7.58 5.36 / 6.23 / 6.95 5.39 / 6.61 / 7.48 5.85 / 6.54 / 6.95 4.77 / 5.64 / 6.23 4.45 / 5.21 / 5.51 4.33 / 4.94 / 5.40
4.03 / 4.22 / 4.39 2.95 / 3.29 / 3.61 2.95 / 3.41 / 3.68 2.79 / 3.04 / 3.22 3.26 / 4.03 / 4.60 3.36 / 3.52 / 3.81 3.25 / 3.87 / 4.37 3.21 / 3.66 / 4.26 2.88 / 3.37 / 3.70 2.69 / 2.98 / 3.36 2.57 / 2.90 / 3.11
8.55 / 8.87 / 9.19 5.73 / 6.65 / 7.32 6.04 / 7.23 / 7.83 5.41 / 5.92 / 7.26 6.37 / 8.08 / 9.33 7.71 / 7.94 / 8.51 7.00 / 8.49 / 9.56 6.56 / 7.71 / 8.51 5.89 / 7.08 / 7.94 5.44 / 6.63 / 6.87 5.33 / 6.29 / 6.61
Bike Outflow RMSE
MAE
Taxi Inflow
Taxi Outflow
MAE
RMSE
MAE
RMSE
0.75 / 0.81 / 0.86 0.76 / 0.79 / 0.90 0.56 / 0.60 / 0.65 0.59 / 0.64 / 0.69 0.61 / 0.72 / 0.84 0.74 / 0.78 / 0.83 0.57 / 0.64 / 0.72 0.68 / 0.72 / 0.78 0.60 / 0.67 / 0.77 0.54 / 0.56 / 0.60 0.50 / 0.52 / 0.56
3.66 / 3.97 / 4.31 2.51 / 2.74 / 3.18 2.33 / 2.54 / 2.89 2.56 / 2.88 / 3.26 2.60 / 3.16 / 3.90 4.65 / 4.93 / 5.20 2.42 / 2.75 / 3.28 3.02 / 3.07 / 3.37 2.60 / 2.94 / 3.57 2.29 / 2.45 / 2.72 2.21 / 2.37 / 2.63
0.68 / 0.74 / 0.79 0.69 / 0.73 / 0.85 0.50 / 0.54 / 0.61 0.52 / 0.57 / 0.63 0.55 / 0.65 / 0.79 0.95 / 0.99 / 1.03 0.51 / 0.57 / 0.65 0.56 / 0.60 / 0.70 0.54 / 0.63 / 0.74 0.47 / 0.49 / 0.55 0.44 / 0.46 / 0.50
3.73 / 4.11 / 4.48 2.53 / 2.87 / 3.44 2.40 / 2.73 / 3.29 2.66 / 2.99 / 3.54 2.63 / 3.29 / 4.22 4.68 / 4.87 / 5.06 2.45 / 2.79 / 3.32 2.80 / 3.19 / 3.78 2.69 / 3.29 / 4.16 2.32 / 2.61 / 3.05 2.22 / 2.42 / 2.70
uni-modality baselines, MTGNN and TimesNet consistently outperform other models. MTGNN benefits from graph-based temporal dependency modeling that partially captures structured spatial correlations within each modality. TimesNet further demonstrates competitive performance, achieving second-best results on a substantial portion of metrics in the Washington DC dataset, owing to its frequency-aware period discovery mechanism. Nevertheless, these models remain inherently limited in multi-modality settings, as they process each modality independently and fail to exploit complementary cross-modality information, resulting in suboptimal performance when inter-modality dynamics are critical. In contrast, most multi-modality baselines (i.e., MiST, STtrans, and COCOA) even underperform strong uni-modality models, primarily due to coarse or static fusion strategies that indiscriminately aggregate modalities without considering their varying reliability across temporal patterns. MoSSL stands out by achieving second-best performance across all datasets, benefiting from self-supervised objectives that promote selective cross-modality alignment and alleviate negative transfer. However, MoSSL operates mainly in
Frequency-Domain Multi-Modality Transportation Modeling
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Table 3: Ablation studies on the NYC. Results are taken from the third horizon.
Variant
Bike Inflow Bike Outflow Taxi Inflow Taxi Outflow
Table 4: Applying FreMo to representative time series models on the NYC. Results are taken from the third horizon. Model
MAE RMSE MAE RMSE MAE RMSE MAE RMSE w/o FreMo 2.64 Time Domain 2.52 w/o Synergy 2.45 Shared Weight 2.46 w/o Node Emb 2.52 FreMo 2.42 Removed
4.91 4.61 4.41 4.50 4.58 4.39
Modality-Wise
Bike
2.67 2.59 2.55 2.59 2.59 2.55
5.19 4.93 4.81 4.82 4.87 4.80
Dynamic Gating
Taxi
5.1
8.80 17.60 9.15 7.96 16.21 8.37 7.72 15.95 8.39 8.07 16.62 8.57 8.44 17.08 8.77 7.54 15.78 8.25
Modality-Shared (FreMo)
Bike
Taxi
16.8
4.5
16.4 16.0
Inflow Outflow
2.5
Inflow Outflow
Taxi Inflow
Taxi Outflow
AGCRN [2] +FreMo Δ(%)
3.02 8.41 3.31 9.15 9.16 18.16 10.19 19.29 2.89 5.78 3.27 6.18 8.84 16.69 9.75 18.07 4.30% 31.27% 1.21% 32.46% 3.49% 8.09% 4.32% 6.32%
TimesNet [53] +FreMo Δ(%)
3.21 5.88 3.45 6.34 9.25 18.04 10.57 20.37 3.15 5.78 3.41 6.26 9.22 17.95 10.35 20.02 1.87% 1.70% 1.16% 1.26% 0.32% 0.50% 2.08% 1.72%
iTransformer [39] 4.52 8.85 4.80 9.44 12.54 23.98 12.80 24.37 +FreMo 3.77 6.88 3.94 7.58 10.66 19.81 11.38 20.75 Δ(%) 16.59% 22.26% 17.92% 19.70% 14.99% 17.39% 11.09% 14.85% STAEformer [35] 3.81 7.78 3.73 7.74 11.89 24.06 12.14 21.75 +FreMo 3.10 6.32 3.24 7.07 9.26 19.02 9.46 18.48 Δ(%) 18.64% 18.77% 13.14% 8.66% 22.12% 20.95% 22.08% 15.03%
8.0
Inflow Outflow
Inflow Outflow
Figure 3: Performance of different synergy injection strategies (parameterization of 𝛾) on the NYC. the temporal domain and lacks explicit mechanisms to regulate modality contributions at finer structural levels. FreMo consistently achieves the best performance by explicitly modeling modality-wise spectral patterns and coordinating cross-modality collaboration in the frequency domain, enabling fine-grained signal selection and robust information sharing.
5.3
Bike Outflow
7.5
2.4
15.6
4.2
8.5
MAE
4.8
MAE
RMSE
2.6
RMSE
17.71 16.15 16.19 16.62 17.09 15.86
Bike Inflow
MAE RMSE MAE RMSE MAE RMSE MAE RMSE
Ablation Study
5.3.1 Effect of Each Component. To evaluate the contribution of each component in FreMo, ablation studies are conducted on the NYC dataset (Table 3) with the following variants: (1) w/o FreMo: removes the entire FreMo module and uses the spatio-temporal encoder as the backbone. (2) Time Domain: replaces the frequency-domain operations in FreMo with equivalent time-domain processing blocks. (3) w/o Synergy: removes the FSI and keeps only the modalitywise refinement (MFF). (4) Shared Weight: replaces the modality-wise weight generators with a single shared generator across modalities. (5) w/o Node Emb: removes the node embedding used in MFF. Table 3 yields the following observations: (1) Removing FreMo (w/o FreMo) degrades all modalities, indicating that the frequencydomain refinement and synergy modeling in FreMo provide substantial gains beyond the encoder alone. (2) Time Domain underperforms FreMo, suggesting that explicitly operating in the frequency domain better captures scale-dependent structures than purely timedomain modeling. (3) Removing FSI (w/o Synergy) causes minor changes on Bike flows but larger drops on Taxi flows, implying that frequency-wise synergy consensus is particularly helpful for modalities that benefit more from cross-modality consistent components. (4) Shared Weight consistently degrades performance (especially on Taxi), showing that using separate modality-wise generators is important to accommodate modality-dependent spectral profiles,
rather than enforcing a single shared parameterization. (5) w/o Node Emb performs worse, highlighting that incorporating node embeddings helps MFF adapt its frequency gates to different spatial units, which improves the effectiveness of frequency refinement. Overall, the gains of FreMo come from the combined contributions of MFF and FSI, rather than any single design in isolation. 5.3.2 Impact of Synergy Scalar 𝛾. Figure 3 examines how the residual scaling factor 𝛾 of the synergy consensus affects the feedback injection into modality-wise representations. Removed sets 𝛾 = 1 (i.e., injecting C without learnable scaling), Modality-Wise learns an independent scalar for each modality, Dynamic Gating uses a two-layer MLP to generate the gating coefficient, and ModalityShared (FreMo) uses a single learnable scalar shared across modalities. (Removed) consistently degrades performance, indicating that a learnable coefficient is important for balancing the original features and the injected consensus. Both Modality-Wise and Dynamic Gating lead to worse results across metrics, and Dynamic Gating is the most unstable choice. This suggests that introducing extra flexibility in the feedback (either per-modality parameters or dynamically generated gates) brings limited benefit and can make optimization harder. In contrast, (Modality-Shared) achieves the best performance, providing a simple and effective global control of synergy injection with minimal parameter overhead. 5.3.3 Plug-and-Play Capacity. To evaluate the plug-and-play capability, we integrate FreMo into four representative backbones, including AGCRN, TimesNet, iTransformer, and STAEformer. FreMo is inserted as an auxiliary enhancement module without modifying the original backbone architecture. Full results are provided in Appendix B. Table 4 shows that FreMo consistently improves performance across all backbones, modalities, and metrics. We make the following observations: (1) FreMo brings substantial gains on Transformer-based backbones. For iTransformer, FreMo achieves double-digit improvements, reducing MAE by 11.09% − 17.92% and RMSE by 14.85% − 22.26%. For STAEformer, the gains are more pronounced on Taxi flows, with about 22.1% MAE reduction and 15.03% − 20.95% RMSE reduction. (2) FreMo improves AGCRN notably in RMSE. MAE reductions are modest (around
Case Study 02:inter-modality -> 在同一频率上:哪个模态更值得信任? 不同模态在不同频率上的相对重要性(Softmax 后的score) 回答以下问题: (1)模型如何决策出协同共识? (2)协同共识的效果体现在哪里?
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Inflow
Outflow
Best Inflow
Taxi
Bike
Bike RMSE
Best Outflow Taxi
4.8
16.0
5.1
Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang
4.8
15.2
4.6
4.5
14.4
4.4
15.0 14.5 14.0
4.2
4.2
13.6
13.5
3.9
8
16
32
64 128
8
Hidden Dim.
16
32
64 128
16
Hidden Dim.
32
48
64
80
16
Latent Dim.
32
48
64
80
Latent Dim.
0.2 0
0.0
Frequency (Low → High) Spectral Energy
Taxi Inflow 1.0 0.8
4
0.6 2
0.4 0.2
0 1
2
3
4
5
6
7
Frequency (Low → High)
8
0.0
1.0 0.8
6
0.6
4
0.4
2
0.2
0
0.0
Frequency (Low → High) Taxi Outflow
1.0
6
0.8 4
0.6 0.4
2
0.2 0 0
1
2
3
4
5
6
7
Frequency (Low → High)
8
0.0
Frequency Weights (w)
0.4
2
Bike Outflow
8
Frequency Weights (w)
0.6
Spectral Energy
0.8
4
Spectral Energy
1.0
Frequency Weights (w)
Spectral Energy
Bike Inflow 6
Frequency Weights (w)
Figure 4: Hyperparameter sensitivity on the NYC.
0
Figure 5: The learned modality-wise frequency weights on the NYC. Gray shaded areas denote the raw spectral amplitude 𝐴 (left axis), and the colored curves show the frequency weights 𝑤 (right axis).
1.21% − 4.30%), while RMSE on Bike flows drops by 31.27% − 32.46%, suggesting fewer large-error cases. (3) FreMo remains effective with a frequency-based backbone. Although TimesNet incorporates frequency-related modeling, FreMo still provides consistent gains, typically around 1% − 2% on both MAE and RMSE, indicating complementary benefits. These results are consistent with the motivation of FreMo: common backbones often treat modalities as independent channels or simply concatenated inputs, while FreMo explicitly refines frequency components for each modality and enables selective cross-modality synergy. Overall, FreMo serves as an architecture-agnostic enhancer that can be plugged into diverse backbones to improve multi-modality forecasting. 5.3.4 Hyperparameter Sensitivity. As shown in Figure 4, the performance of FreMo on the NYC dataset is influenced by two considered hyperparameters, i.e., the hidden dimension 𝑑 and the latent dimension 𝐿. As hidden dimension increases, the performance first decreases and then rises, consistent with the underfitting-overfitting trade-off. The best result is achieved at 𝑑 = 64. For the latent dimension, Taxi flows exhibit a similar pattern with the optimum at 𝐿 = 64, whereas Bike flows are relatively stable across 𝐿 ∈ {32, 48, 64}. Complete results on all datasets are provided in Appendix C.
5.4
Figure a: 低频区域T 统治;在高 有模态的信 证明协同机 的权重是频 如Taxi Out 节(高频)
Case Study
5.4.1 Modality-wise Frequency Weights. To examine whether MFF can identify informative frequency components within each modality, we visualize the learned frequency weights 𝑤 in Eq. (7) on a representative node, as shown in Figure 5. The gray shaded area
Large Calibration
Figure 6: The frequency-wise synergy weight and selective temporal calibration on the NYC. (a) Synergy weight 𝛼 indicates each modality’s relative contribution at each frequency bin. (b) Temporal comparison between original (𝐻 ) and synergy-calibrated (𝐻ˆ ) Taxi Outflow representations.
denotes the spectral amplitude, while the colored curve shows the frequency weights (right axis), indicating how strongly each frequency bin is retained after filtering. Across modalities, the raw spectral amplitude exhibits a similar low-to-high decay trend, reflecting the low-frequency dominance of traffic signals. However, the learned frequency responses differ clearly across modalities. Bike Inflow shows a relatively balanced response across frequencies, whereas Bike Outflow places more emphasis on mid-frequency components. Taxi Inflow exhibits strong low-frequency responses with an additional mid-frequency emphasis, followed by attenuation at high frequencies. Taxi Outflow is highly concentrated on the lowest frequencies and decays rapidly thereafter. These results suggest that spectral amplitude alone is not sufficient to determine frequency importance. If filtering were solely driven by amplitude magnitude, the model would tend to overemphasize low-frequency components and under-utilize informative mid/high-frequency patterns. Instead, even under a similar amplitude decay trend, FreMo learns distinct modality-wise frequency gates, supporting the motivation that different modalities benefit from different frequency-band selections. 5.4.2 Frequency-wise Synergy Weights and Selective Temporal Calibration. This case study examines how FSI assigns frequency-wise synergy weights across modalities and how the resulting feedback leads to selective temporal calibration. We visualize the Softmaxnormalized synergy weights 𝛼 in Eq. (10) over frequency bins, and then compare the original temporal feature 𝐻 with the synergycalibrated feature 𝐻ˆ for Taxi Outflow. Figure 6(a) shows the frequency-wise synergy weight map on a representative node, where each entry 𝛼 indicates the relative contribution of a modality at a specific frequency bin. At the lowestfrequency bin, the weights are highly concentrated on Taxi Outflow, suggesting that the synergy consensus in this regime is mainly driven by Taxi-related signals. As frequency increases, the weights become less concentrated and spread more evenly across modalities, indicating that no single modality consistently dominates in higher-frequency components. Figure 6(b) further illustrates the
Figure b: 在 Time St 红线和灰线 证明了模型 或不确定性 原始特征。
t-SNE Dimension 2
Frequency-Domain Multi-Modality Transportation Modeling
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
References
Rep.C
Rep.B
Group A Group B Group C
Rep.A
Synergy Weight (α)
t-SNE Dimension 1
(a) Node clustering by frequency weight (MFF) Bike Inflow
Bike Outflow
Taxi Inflow
0.30
0.30
0.25
0.25
Taxi Outflow
0.5
0.0
Rep.A
Rep.B
Rep.C
Low-fre (0-2)
0.20
Rep.A
Rep.B
Rep.C
Mid-fre (3-5)
0.20
Rep.A
Rep.B
Rep.C
High-fre (6-8)
(b) Node-wise synergy allocation across frequency (FSI)
Figure 7: Node-level frequency gating and synergy allocation. (a) Node clustering by frequency weight 𝑤. (b) Synergy weights 𝛼 of representative nodes across frequencies.
temporal effect of synergy feedback on Taxi Outflow. Synergy feedback does not alter the sequence uniformly. Larger calibrations appear in the earlier segment, while the later segment shows much smaller calibrations and the two curves stay close. This indicates that FreMo performs selective temporal calibration rather than applying a global smoothing effect, strengthening the representation only when it is needed. Overall, FreMo learns frequency-wise synergy weights and translates them into interpretable, adaptive refinement in the temporal domain. 5.4.3 Node-level Frequency Gating and Synergy Allocation. To examine whether FreMo exhibits structured node-wise differences in frequency filtering, we cluster nodes using the modality-wise frequency weight 𝑤 and visualize them with t-SNE, as shown in Figure 7(a). The separable clustering trend suggests that MFF learns node-dependent frequency-gating profiles rather than applying a uniform gate. We further inspect frequency-wise synergy weights (𝛼) on representative nodes (Rep.A/Rep.B/Rep.C) selected from different clusters, as shown in Figure 7(b). The nodes exhibit frequency-dependent variations in modality allocation, especially in the mid/high-frequency bands. This highlights that FreMo learns node-dependent frequency gating, and the resulting node diversity manifests as frequency-dependent variations in synergy allocation.
6
Conclusion
In this paper, we propose a frequency-domain modeling method (FreMo) for multi-modality transportation forecasting, designed to achieve outstanding performance with a simple architecture. FreMo decouples the modeling pipeline into two complementary paradigms by leveraging frequency-wise discriminative properties: a Modality-Wise Frequency Filter that accentuates informative spectral patterns within each modality without compromising temporal alignment, and a Frequency-Guided Synergy Integrator that coordinates cross-modal synergy by dynamically weighting reliable modality contributions at each frequency. Experiments on several multi-modality transportation datasets demonstrate that FreMo achieves state-of-the-art forecasting performance with robust generalization across diverse scenarios.
[1] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450 (2016). [2] Lei Bai, Lina Yao, Can Li, Xianzhi Wang, and Can Wang. 2020. Adaptive graph convolutional recurrent network for traffic forecasting. Advances in Neural Information Processing Systems 33 (2020), 17804–17815. [3] Wanlin Cai, Yuxuan Liang, Xianggen Liu, Jianshuai Feng, and Yuankai Wu. 2024. Msgnet: Learning multi-scale inter-series correlations for multivariate time series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 11141–11149. [4] Peng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu, Yihang Wang, Qingsong Wen, Bin Yang, and Chenjuan Guo. 2024. Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series Forecasting. In International Conference on Learning Representations. [5] Mingyue Cheng, Qi Liu, Zhiding Liu, Zhi Li, Yucong Luo, and Enhong Chen. 2023. Formertime: Hierarchical multi-scale representations for multivariate time series classification. In Proceedings of the ACM web conference. 1437–1445. [6] Tao Dai, Beiliang Wu, Peiyuan Liu, Naiqi Li, Jigang Bao, Yong Jiang, and Shu-Tao Xia. 2024. Periodicity decoupling framework for long-term series forecasting. In International Conference on Learning Representations. [7] Shohreh Deldari, Hao Xue, Aaqib Saeed, Daniel V Smith, and Flora D Salim. 2022. Cocoa: Cross modality contrastive learning for sensor data. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 3 (2022), 1–28. [8] Jiewen Deng, Jinliang Deng, Renhe Jiang, and Xuan Song. 2023. Learning Gaussian Mixture Representations for Tensor Time Series Forecasting. In Proceedings of the International Joint Conference on Artificial Intelligence. 2077–2085. [9] Jiewen Deng, Jinliang Deng, Du Yin, Renhe Jiang, and Xuan Song. 2023. Tts-norm: Forecasting tensor time series via multi-way normalization. ACM Transactions on Knowledge Discovery from Data 18, 1 (2023), 1–25. [10] Jiewen Deng, Renhe Jiang, Jiaqi Zhang, and Xuan Song. 2024. Multi-Modality Spatio-Temporal Forecasting via Self-Supervised Learning. In Proceedings of the International Joint Conference on Artificial Intelligence. 2018–2026. [11] Leyan Deng, Defu Lian, Zhenya Huang, and Enhong Chen. 2022. Graph convolutional adversarial networks for spatiotemporal anomaly detection. IEEE Transactions on Neural Networks and Learning Systems 33, 6 (2022), 2416–2428. [12] Wei Fan, Pengyang Wang, Dongkun Wang, Dongjie Wang, Yuanchun Zhou, and Yanjie Fu. 2023. Dish-ts: a general paradigm for alleviating distribution shift in time series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 7522–7529. [13] Shen Fang, Qi Zhang, Gaofeng Meng, Shiming Xiang, and Chunhong Pan. 2019. GSTNet: Global Spatial-Temporal Network for Traffic Flow Prediction.. In IJCAI. 2286–2293. [14] Yuchen Fang, Yuxuan Liang, Bo Hui, Zezhi Shao, Liwei Deng, Xu Liu, Xinke Jiang, and Kai Zheng. 2025. Efficient large-scale traffic forecasting with transformers: A spatial data management perspective. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 307–317. [15] Ziquan Fang, Lu Pan, Lu Chen, Yuntao Du, and Yunjun Gao. 2021. MDTP: a multi-source deep traffic prediction framework over spatio-temporal trajectory data. Proceedings of the VLDB Endowment 14, 8 (2021), 1289–1297. [16] En Fu and Yanyan Hu. 2025. Frequency-Masked Embedding Inference: A NonContrastive Approach for Time Series Representation Learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 16639–16647. [17] Kan Guo, Yongli Hu, Yanfeng Sun, Sean Qian, Junbin Gao, and Baocai Yin. 2021. Hierarchical Graph Convolution Network for Traffic Forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 151–159. [18] Shengnan Guo, Youfang Lin, Ning Feng, Chao Song, and Huaiyu Wan. 2019. Attention based spatial-temporal graph convolutional networks for traffic flow forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 922–929. [19] Shengnan Guo, Youfang Lin, Shijie Li, Zhaoming Chen, and Huaiyu Wan. 2019. Deep spatial–temporal 3D convolutional neural networks for traffic data forecasting. IEEE Transactions on Intelligent Transportation Systems 20, 10 (2019), 3913–3926. [20] Jindong Han, Hao Liu, Hengshu Zhu, Hui Xiong, and Dejing Dou. 2021. Joint air quality and weather prediction based on multi-adversarial spatiotemporal networks. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 4081–4089. [21] Liangzhe Han, Bowen Du, Leilei Sun, Yanjie Fu, Yisheng Lv, and Hui Xiong. 2021. Dynamic and multi-faceted spatio-temporal deep learning for traffic speed forecasting. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 547–555. [22] Chao Huang, Chuxu Zhang, Jiashu Zhao, Xian Wu, Dawei Yin, and Nitesh Chawla. 2019. Mist: A multiview and multimodal spatial-temporal learning framework for citywide abnormal event forecasting. In The World Wide Web conference. 717–728.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
[23] Chao Huang, Junbo Zhang, Yu Zheng, and Nitesh V. Chawla. 2018. DeepCrime: Attentive Hierarchical Recurrent Networks for Crime Prediction. In Proceedings of the ACM International Conference on Information and Knowledge Management. 1423–1432. [24] Qihe Huang, Lei Shen, Ruixin Zhang, Jiahuan Cheng, Shouhong Ding, Zhengyang Zhou, and Yang Wang. 2024. Hdmixer: Hierarchical dependency with extendable patch for multivariate time series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 12608–12616. [25] Qihe Huang, Lei Shen, Ruixin Zhang, Shouhong Ding, Binwu Wang, Zhengyang Zhou, and Yang Wang. 2023. Crossgnn: Confronting noisy multivariate time series via cross interaction refinement. Advances in Neural Information Processing Systems 36 (2023), 46885–46902. [26] Jiahao Ji, Jingyuan Wang, Chao Huang, Junjie Wu, Boren Xu, Zhenhe Wu, Junbo Zhang, and Yu Zheng. 2023. Spatio-temporal self-supervised learning for traffic flow prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 4356–4364. [27] Jiawei Jiang, Chengkai Han, Wayne Xin Zhao, and Jingyuan Wang. 2023. Pdformer: Propagation delay-aware dynamic long-range transformer for traffic flow prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 4365–4373. [28] Renhe Jiang, Zekun Cai, Zhaonan Wang, Chuang Yang, Zipei Fan, Quanjun Chen, Kota Tsubouchi, Xuan Song, and Ryosuke Shibasaki. 2021. DeepCrowd: A deep model for large-scale citywide crowd density and flow prediction. IEEE Transactions on Knowledge and Data Engineering 35, 1 (2021), 276–290. [29] Renhe Jiang, Du Yin, Zhaonan Wang, Yizhuo Wang, Jiewen Deng, Hangchen Liu, Zekun Cai, Jinliang Deng, Xuan Song, and Ryosuke Shibasaki. 2021. Dl-traff: Survey and benchmark of deep learning models for urban traffic prediction. In Proceedings of the ACM International Conference on Information and Knowledge Management. 4515–4525. [30] Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen. 2024. Time-LLM: Time series forecasting by reprogramming large language models. In International Conference on Learning Representations. [31] Hyunwook Lee and Sungahn Ko. 2024. TESTAM: A Time-Enhanced SpatioTemporal Attention Model with Mixture of Experts. In International Conference on Learning Representations. [32] Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks. In International Conference on Machine Learning. 3744–3753. [33] Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In International Conference on Learning Representations. [34] Shengsheng Lin, Weiwei Lin, Xinyi Hu, Wentai Wu, Ruichao Mo, and Haocheng Zhong. 2024. Cyclenet: Enhancing time series forecasting through modeling periodic patterns. Advances in Neural Information Processing Systems 37 (2024), 106315–106345. [35] Hangchen Liu, Zheng Dong, Renhe Jiang, Jiewen Deng, Jinliang Deng, Quanjun Chen, and Xuan Song. 2023. Spatio-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting. In Proceedings of the ACM International Conference on Information and Knowledge management. 4125–4129. [36] Hao Liu, Qiyu Wu, Fuzhen Zhuang, Xinjiang Lu, Dejing Dou, and Hui Xiong. 2021. Community-aware multi-task transportation demand prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 320–327. [37] Peiyuan Liu, Hang Guo, Tao Dai, Naiqi Li, Jigang Bao, Xudong Ren, Yong Jiang, and Shu-Tao Xia. 2025. Calf: Aligning llms for time series forecasting via crossmodal fine-tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 18915–18923. [38] Qinghua Liu and John Paparrizos. 2024. The elephant in the room: Towards a reliable time-series anomaly detection benchmark. Advances in Neural Information Processing Systems 37 (2024), 108231–108261. [39] Yong Liu, Tengge Hu, Haoran Zhang, Haixu Wu, Shiyu Wang, Lintao Ma, and Mingsheng Long. 2024. iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. In International Conference on Learning Representations. [40] Wang Lu, Jindong Wang, Xinwei Sun, Yiqiang Chen, and Xing Xie. 2023. Out-ofdistribution Representation Learning for Time Series Classification. In International Conference on Learning Representations. [41] Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In International Conference on Learning Representations. [42] Boris N Oreshkin, Arezou Amini, Lucy Coyle, and Mark Coates. 2021. FC-GAGA: Fully connected gated graph architecture for spatio-temporal traffic forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 9233–9241. [43] Zheyi Pan, Yuxuan Liang, Weifeng Wang, Yong Yu, Yu Zheng, and Junbo Zhang. 2019. Urban traffic prediction from spatio-temporal data using deep meta learning. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1720–1730.
Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang
[44] Xihao Piao, Zheng Chen, Taichi Murayama, Yasuko Matsubara, and Yasushi Sakurai. 2024. Fredformer: Frequency debiased transformer for time series forecasting. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2400–2410. [45] Xiangfei Qiu, Xingjian Wu, Yan Lin, Chenjuan Guo, Jilin Hu, and Bin Yang. 2025. Duet: Dual clustering enhanced multivariate time series forecasting. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1185–1196. [46] Zezhi Shao, Zhao Zhang, Fei Wang, and Yongjun Xu. 2022. Pre-training enhanced spatial-temporal graph neural network for multivariate time series forecasting. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1567–1577. [47] Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016. WaveNet: A Generative Model for Raw Audio. In Proceedings of the ISCA Workshop on Speech Synthesis Workshop. [48] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017). [49] Chengsen Wang, Zirui Zhuang, Qi Qi, Jingyu Wang, Xingyu Wang, Haifeng Sun, and Jianxin Liao. 2023. Drift doesn’t matter: Dynamic decomposition with diffusion reconstruction for unstable multivariate time series anomaly detection. Advances in Neural Information Processing Systems 36 (2023), 10758–10774. [50] Shiyu Wang, Jiawei Li, Xiaoming Shi, Zhou Ye, Baichuan Mo, Wenze Lin, Shengtong Ju, Zhixuan Chu, and Ming Jin. 2025. TimeMixer++: A General Time Series Pattern Machine for Universal Predictive Analysis. In International Conference on Learning Representations. [51] Shiyu Wang, Haixu Wu, Xiaoming Shi, Tengge Hu, Huakun Luo, Lintao Ma, James Y Zhang, and JUN ZHOU. 2024. TimeMixer: Decomposable Multiscale Mixing for Time Series Forecasting. In International Conference on Learning Representations. [52] Zhaonan Wang, Renhe Jiang, Hao Xue, Flora D Salim, Xuan Song, and Ryosuke Shibasaki. 2022. Event-aware multimodal mobility nowcasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 4228–4236. [53] Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. 2023. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In International Conference on Learning Representations. [54] Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. 2021. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems 34 (2021), 22419–22430. [55] Xian Wu, Chao Huang, Chuxu Zhang, and Nitesh V Chawla. 2020. Hierarchically structured transformer networks for fine-grained spatial event forecasting. In Proceedings of the Web Conference. 2320–2330. [56] Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, Xiaojun Chang, and Chengqi Zhang. 2020. Connecting the dots: Multivariate time series forecasting with graph neural networks. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 753–763. [57] Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, and Chengqi Zhang. 2019. Graph wavenet for deep spatial-temporal graph modeling. In Proceedings of the International Joint Conference on Artificial Intelligence. 1907–1913. [58] Lianghao Xia, Chao Huang, Yong Xu, Peng Dai, Liefeng Bo, Xiyue Zhang, and Tianyi Chen. 2021. Spatial-Temporal Sequential Hypergraph Network for Crime Prediction with Dynamic Multiplex Relation Learning. In Proceedings of the International Joint Conference on Artificial Intelligence. 1631–1637. [59] Song Yang, Jiamou Liu, and Kaiqi Zhao. 2021. Space Meets Time: Local Spacetime Neural Network For Traffic Flow Forecasting. In IEEE International Conference on Data Mining. 817–826. [60] Huaxiu Yao, Yiding Liu, Ying Wei, Xianfeng Tang, and Zhenhui Li. 2019. Learning from multiple cities: A meta-learning approach for spatial-temporal prediction. In The World Wide Web conference. 2181–2191. [61] Junchen Ye, Leilei Sun, Bowen Du, Yanjie Fu, Xinran Tong, and Hui Xiong. 2019. Co-prediction of multiple transportation demands based on deep spatio-temporal neural network. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 305–313. [62] Weiwei Ye, Songgaojun Deng, Qiaosha Zou, and Ning Gui. 2024. Frequency adaptive normalization for non-stationary time series forecasting. Advances in Neural Information Processing Systems 37 (2024), 31350–31379. [63] Kun Yi, Jingru Fei, Qi Zhang, Hui He, Shufeng Hao, Defu Lian, and Wei Fan. 2024. Filternet: Harnessing frequency filters for time series forecasting. Advances in Neural Information Processing Systems 37 (2024), 55115–55140. [64] Kun Yi, Qi Zhang, Wei Fan, Hui He, Liang Hu, Pengyang Wang, Ning An, Longbing Cao, and Zhendong Niu. 2023. FourierGNN: Rethinking multivariate time series forecasting from a pure graph perspective. Advances in Neural Information Processing Systems 36 (2023), 69638–69660. [65] Kun Yi, Qi Zhang, Wei Fan, Shoujin Wang, Pengyang Wang, Hui He, Ning An, Defu Lian, Longbing Cao, and Zhendong Niu. 2023. Frequency-domain MLPs are more effective learners in time series forecasting. Advances in Neural Information Processing Systems 36 (2023), 76656–76679.
Frequency-Domain Multi-Modality Transportation Modeling
[66] Chengqing Yu, Fei Wang, Zezhi Shao, Tangwen Qian, Zhao Zhang, Wei Wei, and Yongjun Xu. 2024. Ginar: An end-to-end multivariate time series forecasting model suitable for variable missing. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3989–4000. [67] Haitao Yuan, Guoliang Li, Zhifeng Bao, and Ling Feng. 2021. An effective joint prediction model for travel demands and traffic flows. In IEEE International Conference on Data Engineering. 348–359. [68] Wenzhen Yue, Yong Liu, Xianghua Ying, Bowei Xing, Ruohao Guo, and Ji Shi. 2025. FreEformer: frequency enhanced transformer for multivariate time series forecasting. In Proceedings of the International Joint Conference on Artificial Intelligence. 3606–3614. [69] Dongran Zhang, Jiangnan Yan, Kemal Polat, Adi Alhudhaif, and Jun Li. 2024. Multimodal joint prediction of traffic spatial-temporal data with graph sparse attention mechanism and bidirectional temporal convolutional network. Advanced Engineering Informatics 62 (2024), 102533. [70] Junbo Zhang, Yu Zheng, and Dekang Qi. 2017. Deep spatio-temporal residual networks for citywide crowd flows prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 31. 1655–1661. [71] Qianru Zhang, Chao Huang, Lianghao Xia, Zheng Wang, Zhonghang Li, and Siuming Yiu. 2023. Automated Spatio-Temporal Graph Contrastive Learning. In Proceedings of the ACM Web Conference. 295–305. [72] Xiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, and Marinka Zitnik. 2022. Self-supervised contrastive pre-training for time series via time-frequency consistency. Advances in Neural Information Processing Systems 35 (2022), 3988–4003. [73] Yunhao Zhang and Junchi Yan. 2023. Crossformer: Transformer Utilizing CrossDimension Dependency for Multivariate Time Series Forecasting. In International Conference on Learning Representations. [74] Lifan Zhao and Yanyan Shen. 2024. Rethinking Channel Dependence for Multivariate Time Series Forecasting: Learning from Leading Indicators. In International Conference on Learning Representations. [75] Chuanpan Zheng, Xiaoliang Fan, Cheng Wang, and Jianzhong Qi. 2020. Gman: A graph multi-attention network for traffic prediction. In Proceedings of the AAAI conference on artificial intelligence, Vol. 34. 1234–1241. [76] Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 11106–11115. [77] Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. 2022. Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. In International Conference on Machine Learning. 27268–27286. [78] Tian Zhou, Peisong Niu, Liang Sun, Rong Jin, et al. 2023. One fits all: Power general time series analysis by pretrained lm. Advances in neural information processing systems 36 (2023), 43322–43355. [79] Xian Zhou, Yanyan Shen, Yanmin Zhu, and Linpeng Huang. 2018. Predicting Multi-step Citywide Passenger Demands Using Attention-based Neural Networks. In Proceedings of the ACM International Conference on Web Search and Data Mining. 736–744. [80] Zhibo Zhu, Ziqi Liu, Ge Jin, Zhiqiang Zhang, Lei Chen, Jun Zhou, and Jianyong Zhou. 2021. MixSeq: Connecting Macroscopic Time Series Forecasting with Microscopic Time Series Data. Advances in Neural Information Processing Systems 34 (2021), 12904–12916.
A
Complexity and Efficiency Analysis
Theoretical Complexity. FreMo consists of three main costs. (1) rFFT and irFFT along the temporal dimension cost O (𝑀𝑁𝑑·𝑇 log𝑇 ), which is the only term super-linear in 𝑇 . (2) MFF performs amplitude computation, channel pooling, frequency-wise gating, and gate generation, all linear in 𝐹 , with cost O (𝑀𝑁𝑑 · 𝐹 ). (3) FSI performs frequency-wise scoring, modality Softmax, consensus aggregation, and residual feedback, with cost O (𝑀𝑁𝑑 · 𝐹 + 𝑀𝑁 𝐹 ), where O (𝑀𝑁 𝐹 ) comes from the Softmax. Thus, the total time complexity is O (𝑀𝑁𝑑 · 𝑇 log𝑇 + 𝑀𝑁𝑑 · 𝐹 ), which asymptotically reduces to O (𝑀𝑁𝑑 · 𝑇 log𝑇 ). Practical Efficiency. FreMo processes modalities and nodes in parallel, with costs scaling linearly with 𝑀 and 𝑁 . Its additional memory footprint is O (𝑀𝑁 𝐹𝑑) due to intermediate frequencydomain tensors. As shown in Table 5, FreMo adds only +0.028M parameters on NYC/DC and +0.054M on Chicago, with moderate training-time overhead ranging from +1.2s to +13.9s per epoch.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
This confirms that FreMo is lightweight, highly parallelizable, and can be efficiently integrated into diverse forecasting backbones for large-scale multi-modality transportation forecasting. Table 5: Overhead of Params / Training Time. Model
NYC
DC
Chicago
AGCRN [2] +FreMo
0.597M / 48.9s 0.597M / 27.3s 0.600M / 90.1s +0.028M / +2.4s +0.028M / +10.6s +0.054M / +13.9s
TimesNet [53] +FreMo
4.718M / 28.3s 4.720M / 15.8s 4.824M / 29.8s +0.028M / +3.1s +0.028M / +7.9s +0.054M / +12.8s
iTransformer [39] 0.052M / 7.6s 0.052M / 4.5s 0.052M / 21.8s +FreMo +0.028M / +4.0s +0.028M / +3.4s +0.054M / +12.3s STAEformer [35] 1.187M / 16.2s 1.199M / 23.3s 1.714M / 101.2s +FreMo +0.028M / ++1.2s +0.028M / +5.8s +0.054M / +12.6s
B
Full Results of Plug-and-Play Capacity
Table 6 reports the full generality results of integrating FreMo into four representative forecasting backbones (AGCRN, TimesNet, iTransformer, and STAEformer) across all datasets. Consistent with the observations in the main text, FreMo yields universal performance improvements across diverse datasets, backbone architectures, and transportation modalities, indicating FreMo’s adaptability to diverse urban mobility patterns and spatial configurations, regardless of the scale or density of the road network. This universality implies that explicit spectral disentanglement and synergy modeling capture fundamental cross-modality correlations that remain under-exploited by current time series modeling paradigms. These extensive results further validate the robustness and effectiveness of FreMo as a general, plug-and-play enhancement module for multi-modality transportation forecasting.
C
Full Results of Hyperparameter Sensitivity
To investigate the impact of hyperparameter settings on FreMo, we conduct a sensitivity analysis on two key parameters: the hidden dimension (𝑑) and the latent dimension (𝐿). We evaluate the RMSE metric across three datasets by varying 𝑑 within {8, 16, 32, 64, 128} and 𝐿 within {16, 32, 48, 64, 80}. The results are visualized in Figure 8. Regarding the hidden dimension 𝑑, we observe a consistent trend where performance improves with increased model capacity but degrades at higher values (e.g., 128) due to overfitting. NYC and Washington DC achieve optimal results at 𝑑 = 64, maximizing spectral feature encoding. In contrast, the Chicago dataset favors a more compact representation, peaking at 𝑑 = 32. For the latent dimension 𝐿, the results indicate that a moderate embedding size is sufficient to capture spatial heterogeneity without introducing redundancy. Performance generally stabilizes around 𝐿 = 64 for NYC. However, Washington DC and Chicago exhibit a preference for a lower dimension, achieving its lowest error at 𝐿 = 32. Based on these empirical observations, we adopt the optimal (𝑑, 𝐿) configurations of (64, 64) for NYC, (64, 32) for Washington DC, and (32, 32) for Chicago in our main experiments (Table 2).
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang
Table 6: Full results of applying FreMo to representative time series models. Bike Inflow
NYC MAE
Bike Outflow RMSE
MAE
Taxi Inflow
RMSE
Taxi Outflow
MAE
RMSE
MAE
RMSE
7.81 / 9.03 / 10.19 7.81 / 8.75 / 9.75
15.03 / 17.41 / 19.29 14.63 / 16.34 / 18.07
AGCRN [2] + FreMo
2.40 / 2.70 / 3.02 6.96 / 7.69 / 8.41 2.60 / 2.97 / 3.31 7.39 / 8.32 / 9.15 2.37 / 2.63 / 2.89 4.91 / 5.31 / 5.78 2.54 / 2.87 / 3.27 5.21 / 5.75 / 6.18
7.08 / 8.10 / 9.16 6.86 / 7.92 / 8.84
13.42 / 15.71 / 18.16 13.13 / 14.72 / 16.69
TimesNet [53] + FreMo
2.65 / 2.88 / 3.21 4.61 / 5.11 / 5.88 2.81 / 3.10 / 3.45 5.03 / 5.65 / 6.34 2.65 / 2.86 / 3.15 4.55 / 4.98 / 5.78 2.79 / 3.07 / 3.41 4.94 / 5.54 / 6.26
7.54 / 8.44 / 9.25 7.50 / 8.40 / 9.22
13.61 / 15.98 / 18.04 8.45 / 9.52 / 10.57 15.94 / 18.15 / 20.37 13.57 / 15.65 / 17.95 8.42 / 9.43 / 10.35 15.93 / 18.06 / 20.02
iTransformer [39] 3.05 / 3.75 / 4.52 5.61 / 7.22 / 8.85 3.31 / 4.10 / 4.80 6.20 / 7.99 / 9.44 8.11 / 10.32 / 12.54 14.40 / 19.16 / 23.98 8.69 / 10.85 / 12.80 16.15 / 20.60 / 24.37 + FreMo 2.93 / 3.36 / 3.77 5.20 / 6.06 / 6.88 3.09 / 3.54 / 3.94 5.79 / 6.73 / 7.58 7.71 / 9.21 / 10.66 13.47 / 16.73 / 19.81 8.44 / 9.91 / 11.38 15.46 / 18.33 / 20.75 STAEformer [35] 3.57 / 3.67 / 3.81 7.13 / 7.42 / 7.78 3.44 / 3.57 / 3.73 7.10 / 7.40 / 7.74 8.76 / 10.33 / 11.89 15.40 / 19.51 / 24.06 9.68 / 10.96 / 12.14 16.74 / 19.34 / 21.75 + FreMo 2.62 / 2.83 / 3.10 5.01 / 5.58 / 6.32 2.72 / 2.99 / 3.24 5.48 / 6.29 / 7.07 7.24 / 8.19 / 9.26 13.92 / 16.07 / 19.02 7.69 / 8.60 / 9.46 14.77 / 16.76 / 18.48 Bike Inflow
DC MAE
Bike Outflow RMSE
MAE
Taxi Inflow
RMSE
Taxi Outflow
MAE
RMSE
MAE
RMSE
AGCRN [2] + FreMo
0.99 / 1.10 / 1.24 1.49 / 1.63 / 1.79 1.00 / 1.11 / 1.25 1.51 / 1.65 / 1.81 0.95 / 1.07 / 1.20 1.43 / 1.58 / 1.63 0.94 / 0.99 / 1.21 1.42 / 1.46 / 1.76
2.76 / 3.00 / 3.29 2.62 / 2.91 / 3.14
4.69 / 5.22 / 5.68 4.17 / 4.79 / 5.45
2.95 / 3.29 / 3.61 2.80 / 3.08 / 3.26
5.73 / 6.65 / 7.32 5.43 / 6.25 / 7.13
TimesNet [53] + FreMo
0.54 / 0.59 / 0.62 1.17 / 1.27 / 1.32 0.56 / 0.60 / 0.63 1.20 / 1.30 / 1.35 0.52 / 0.58 / 0.61 1.15 / 1.23 / 1.28 0.55 / 0.58 / 0.60 1.17 / 1.26 / 1.32
2.55 / 2.78 / 2.94 2.50 / 2.74 / 2.84
4.38 / 4.87 / 5.58 4.34 / 4.80 / 5.06
2.79 / 3.04 / 3.22 2.63 / 2.89 / 3.02
5.41 / 5.92 / 7.26 5.10 / 5.63 / 5.89
iTransformer [39] 0.63 / 0.71 / 0.77 1.38 / 1.55 / 1.65 0.64 / 0.71 / 0.77 1.41 / 1.55 / 1.66 0.62 / 0.67 / 0.72 1.37 / 1.50 / 1.59 0.63 / 0.68 / 0.72 1.35 / 1.44 / 1.60 + FreMo
3.10 / 3.70 / 4.24 3.08 / 3.61 / 4.10
5.35 / 6.60 / 7.58 5.30 / 6.49 / 7.36
3.26 / 4.03 / 4.60 3.21 / 3.79 / 4.01
6.37 / 8.08 / 9.33 6.31 / 7.88 / 9.12
STAEformer [35] 0.64 / 0.71 / 0.75 1.23 / 1.31 / 1.35 0.65 / 0.71 / 0.78 1.28 / 1.36 / 1.38 0.62 / 0.67 / 0.73 1.19 / 1.29 / 1.31 0.62 / 0.68 / 0.75 1.26 / 1.34 / 1.35 + FreMo
2.94 / 3.31 / 3.70 2.75 / 3.13 / 3.38
5.36 / 6.23 / 6.95 5.12 / 5.97 / 6.56
3.36 / 3.52 / 3.81 2.96 / 3.33 / 3.57
7.71 / 7.94 / 8.51 6.76 / 7.67 / 8.23
Bike Inflow
Chicago MAE
Bike Outflow RMSE
MAE
Taxi Inflow
RMSE
Taxi Outflow
MAE
RMSE
MAE
RMSE
AGCRN [2] + FreMo
0.51 / 0.52 / 0.60 1.35 / 1.52 / 1.75 0.52 / 0.53 / 0.60 1.45 / 1.57 / 1.73 0.40 / 0.42 / 0.45 1.31 / 1.41 / 1.56 0.39 / 0.42 / 0.44 1.35 / 1.46 / 1.56
0.76 / 0.79 / 0.90 0.64 / 0.68 / 0.75
2.51 / 2.74 / 3.18 2.35 / 2.55 / 2.88
0.69 / 0.73 / 0.85 0.58 / 0.62 / 0.70
2.53 / 2.87 / 3.44 2.41 / 2.70 / 3.15
TimesNet [53] + FreMo
0.34 / 0.36 / 0.39 1.25 / 1.38 / 1.56 0.35 / 0.37 / 0.39 1.35 / 1.43 / 1.54 0.34 / 0.36 / 0.39 1.23 / 1.36 / 1.51 0.34 / 0.36 / 0.37 1.34 / 1.41 / 1.52
0.59 / 0.64 / 0.69 0.57 / 0.62 / 0.66
2.56 / 2.88 / 3.26 2.52 / 2.78 / 3.10
0.52 / 0.57 / 0.63 0.50 / 0.55 / 0.61
2.66 / 2.99 / 3.54 2.61 / 2.87 / 3.40
iTransformer [39] 0.38 / 0.43 / 0.49 1.46 / 1.76 / 2.07 0.38 / 0.41 / 0.45 1.54 / 1.72 / 1.87 + FreMo 0.37 / 0.41 / 0.47 1.40 / 1.63 / 1.91 0.37 / 0.39 / 0.41 1.44 / 1.58 / 1.75
0.61 / 0.72 / 0.84 0.58 / 0.68 / 0.74
2.60 / 3.16 / 3.90 2.57 / 2.98 / 3.60
0.55 / 0.65 / 0.79 0.54 / 0.60 / 0.71
2.63 / 3.29 / 4.22 2.59 / 3.06 / 3.80
STAEformer [35] 0.35 / 0.38 / 0.31 1.91 / 1.92 / 1.96 0.37 / 0.38 / 0.42 1.98 / 2.01 / 2.27 0.34 / 0.35 / 0.37 1.68 / 1.88 / 1.91 0.35 / 0.38 / 0.40 1.77 / 1.91 / 2.07 + FreMo
0.74 / 0.78 / 0.83 0.60 / 0.66 / 0.74
4.65 / 4.93 / 5.20 2.99 / 3.46 / 4.27
0.95 / 0.99 / 1.03 0.54 / 0.58 / 0.66
4.68 / 4.87 / 5.06 3.07 / 3.25 / 3.79
NYC
Bike · Hidden Dim. 5.00 4.80 4.60 4.40 4.20
DC
4.65
15.00
4.50
32
64
128
4.20 8
16
32
64
128
7.00 6.50 6.00 5.50 5.00
1.26 1.23 1.20 1.17 1.14 8
16
32
64
128
1.25 1.23 1.20 8
16
32
64
48
64
80
1.16 16
32
64
128
16
32
64
128
Outflow
1.22 1.22 1.21 1.20 1.19
16
32
48
64
80
16
32
48
64
80
16
32
48
64
80
6.50 6.00 5.50 5.00
1.20
8
Inflow
32
1.24
8
128
16 1.28
2.64 2.58 2.52 2.46 2.40
1.28
Taxi · Latent Dim. 15.00 14.70 14.40 14.10 13.80
4.35
14.00 16
Bike · Latent Dim.
15.50 14.50
8
Chicago
Taxi · Hidden Dim.
16
32
48
64
80 2.48 2.45 2.42 2.40
16
32
Best Inflow
48
64
80
Best Outflow
Figure 8: Hyperparameter sensitivity studies of hidden dimension 𝑑 and latent dimension 𝐿 on three datasets.