FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction Fermin Orozco
Man Luo
Johan Wahlström
University of Exeter Exeter, United Kingdom [email protected]
University of Exeter Exeter, United Kingdom [email protected]
University of Exeter Exeter, United Kingdom [email protected]
arXiv:2609.20026v1 [cs.AI] 17 Sep 2026
ABSTRACT Urban traffic forecasting often relies on information distributed across stakeholders who may be unable to share raw data due to privacy or commercial constraints, motivating federated spatialtemporal approaches. In such federated settings, each client observes traffic over a distinct sensor subgraph with its own spatial topology and temporal dynamics, leading to significant heterogeneity across clients. Existing federated spatial-temporal methods typically rely on model parameter aggregation and provide limited mechanisms for recovering spatial dependencies across client boundaries. This introduces two key limitations. Specifically, parameter aggregation across heterogeneous graph domains tends to dilute client-specific representations, while road network partitioning breaks the propagation of traffic dynamics across client boundaries. To address these challenges, we propose FedeRICo, a federated traffic forecasting framework that combines gradientlevel collaboration with boundary-aware residual communication. FedeRICo employs a dual-branch forecasting architecture in which a globally guided branch captures transferable forecasting structure, while a private residual branch preserves client-specific corrections and incorporates boundary residual signals. The global branch is coordinated through gradient alignment across all clients, enabling collaborative optimisation without destructive parameter interference. To recover cross-client spatial dependencies, boundary messages are extracted through a trend-residual decomposition that suppresses periodic structure and communicates only transient spatial-temporal residual signals between physically adjacent clients. Experiments across four real-world traffic forecasting benchmarks demonstrate that FedeRICo consistently outperforms stateof-the-art federated spatial-temporal baselines while maintaining competitive training runtime.
CCS CONCEPTS • Information systems → Spatial-temporal systems; • Computing methodologies → Distributed computing methodologies.
KEYWORDS Federated Learning, Traffic Flow Prediction, Spatial-Temporal Systems
1
INTRODUCTION
Spatial-temporal forecasting is an essential part of modern Intelligent Transport Systems; the ubiquitous deployment of infrastructure and sensors have led to large-scale data that can be leveraged to optimise urban services [8, 14, 29] and enhance safety [5, 27, 31]. To model the inherent interdependencies within this data, deep
spatial-temporal models implement a combination of GNNs [19], and RNN [7] or attention [22] based architectures to model the spatial and temporal relationships, respectively [2, 4, 11, 26, 30]. Yet these approaches assume a centralised setting, where all the data is centrally available on a single system. In real-world settings, there may be multiple organisations, each with their own independent dataset, that may not be able to centrally aggregate their data due to privacy or commercial limitations. This motivates decentralised data-driven methods for optimising distributed traffic forecasting models. Federated Learning (FL) [16] enables training under this constraint since it does not require centralisation of data. Organisations may collaborate in a decentralised manner to train optimal models without needing to share their raw sensitive data. Several Federated methods have been proposed for traffic flow prediction [13, 17, 28, 32, 33]. Such approaches collaborate primarily through parameter aggregation on shared model components [13, 17, 32, 33], while others learn cross-client interactions through latent representations [28]. The spatial and temporal heterogeneity across client datasets poses difficulties for FL. First, naive parameter aggregation across clients can dilute client-specific representations built during local training. Decoupling spatial and temporal parameters as an alternative, risks losing the integrated spatial-temporal modelling that gives these architectures their predictive power. Second, partitioning a connected road network across clients disrupts spatial dependencies at partition boundaries, limiting the cross-client information flow that boundary nodes may otherwise contribute. Figure 1 illustrates the challenge of FL under real-world urban traffic constraints wherein organisations or municipalities may have ownership of spatially distributed infrastructure, and hence would have a partitioned view of the full network. Furthermore, they will possess different spatial distributions, and as a result, will learn spatial-temporal models over different underlying graph structures. The proposed Federated Region-Influenced Coupling (FedeRICo) framework addresses both of these difficulties through two complementary mechanisms. First, gradient-based training collaboration is implemented, replacing parameter averaging with a learningbased alignment that preserves client specific representations. Second, a boundary messaging protocol exchanges transient, eventinformative signals between physically adjacent clients to recover some of the cross-client information flow at partition boundaries. These mechanisms are realised through a dual-branch architecture, wherein a global branch captures shared dynamics through gradient alignment among all clients, while a local branch models client-specific patterns and the residual error. Figure 1(b) shows
Conference’17, July 2017, Washington, DC, USA
Orozco et al.
2 RELATED WORKS 2.1 Spatial-Temporal Modelling
FedAvg
11
3.4
3.22 3.01
7 6.35 6.07
6
8.93
9 8.12
7
5
(a)
10
8
3 2.8
MAPE (%)
7.15
3.6 3.2
12 11.05
3.69
RMSE
MAE
3.8
FedeRICo (𝑘 = 4)
Centralised 8
4
(b)
(c)
Figure 1: Performance and spatial partitioning summary on METRLA. Top: Spatial partitioning into 𝑘 = 4 localized clusters. Bottom (a–c): Evaluation of our proposed framework against federated and centralised baselines, demonstrating superior predictive accuracy against the federated setting, approaching the centralised model performance.
that this approach substantially outperforms naive parameter averaging (FedAvg) and approaches the centralised STDN model [4] upper bound, in which all data is idealistically aggregated to a single system. These results indicate that gradient-based alignment is a more effective collaboration mechanism than parameter aggregation, and that the boundary messaging method recovers some of the information lost when a connected network is partitioned across clients. The major contributions are summarised as follows: • A client-to-client boundary message-passing mechanism is proposed to exchange spatial-temporal-aware residual signals between adjacent clients without sharing raw traffic observations. • A dual-branch forecasting architecture is introduced to separate shared forecasting structure from client-specific residual corrections. Instead of relying on parameter averaging, clients are coordinated through branch-wise gradientdirection updates: shared-branch gradients are aligned to a population reference, while local-branch gradients are discouraged from collapsing to the same direction. • Extensive experiments are conducted on four real-world benchmark traffic datasets, demonstrating improved forecasting performance with competitive training cost compared with existing federated spatial-temporal forecasting baselines.
Spatial-temporal modelling aims to capture spatial and temporal dependencies within data, making it central to tasks such as traffic prediction. Early GNNs focused on exploiting static explicit graph structured data [11, 30]. Adaptive approaches later aimed to learn a representation of the graph structure from the data itself, such as AGCRN [2] which foregoes any graph structure prior and aims to learn a symmetric graph, and Graph Wavenet (GWNet) [26] which retains the static graph as a support, and learns bi-directional node embeddings as a correction. A parallel line of work applies seasonal-trend decomposition to disentangle periodic, trend, and high-frequency residual components for improved spatial-temporal modelling [20, 23]. Most recently, Cao et al. [4] reframe seasonal-trend decomposition from a spatial-temporal perspective, extracting the trend component conditional on each node’s spatial location and temporal context rather than per-node or purely temporally. Such approaches rely on a centralised data setting and do not address distributed real-world systems.
2.2
Federated Learning
FL was developed to enable collaborative client-local training on distributed data [16], and has been increasingly explored for spatialtemporal modelling. MFVSTGNN [12] introduces a multilevel federated spatial-temporal graph framework that combines local traffic modelling with federated knowledge sharing across clients, while FedGTP [28] models inter-client spatial dependencies via a polynomial decomposition over client-encoded representations. Furthermore, pFedCTP [32] performs personalised knowledge transfer between target areas through adaptive parameter aggregation, and FedDis [33] shares a global traffic pattern bank, while disentangling and capturing local client-specific patterns. These methods collaborate through parameter aggregation on shared components or through the aggregation of learned spatial-temporal representations. However, they do not intervene directly in the underlying optimisation dynamics or exchange targeted, decomposition-aware signals between physically adjacent clients. A separate line of work intervenes in training dynamics directly; SCAFFOLD [9] and FedDyn [1] use gradient correction and dynamic regularisation to align local and global optimisation, and Lu et al. [15] propose conflict-averse gradient aggregation across heterogeneous clients. In the federated graph learning literature, FedIGL [24] addresses graph-level classification across clients holding disjoint graphs from molecular, social, and biological domains, using a bi-gradient regularisation strategy to disentangle invariant from domain-specific substructures. None of these methods coordinate client collaboration through gradient-direction alignment or exchange targeted, decomposition-aware residual signals between physically adjacent clients.
3
PROBLEM FORMULATION
Definition 1 (Spatial-Temporal Prediction): Spatial-temporal traffic forecasting predicts future traffic measurements from a window of past observations over a graph network. The network is denoted
FedeRICo
Conference’17, July 2017, Washington, DC, USA
Figure 2: FedeRICo framework. The left panel shows the dual-branch architecture with shared and local branches using the same ST-Block design but separate parameters. The shared branch predicts the main signal, while the local branch predicts residual corrections informed by boundary messages, and both outputs are added for the final forecast. The top-right panel shows shared-gradient alignment and local-gradient diversification, and the bottom-right panel shows sender-local boundary-message composition.
as G = (V, A) with 𝑁 = |V | sensors and adjacency A ∈ R𝑁 ×𝑁 describing pairwise spatial or semantic relationships. Each sensor 𝑣𝑖 𝑇total produces a time series X𝑖 = {𝑥𝑡 }𝑡 =1 , where 𝑥𝑡 ∈ R𝐹 denotes the traffic state at time 𝑡. Given a window of 𝑇 past observations X𝑡 −𝑇 +1:𝑡 , a forecasting model F produces predictions Ŷ over the next 𝑇 ′ steps: Ŷ = F (X𝑡 −𝑇 +1:𝑡 , A; 𝜃 ),
(1)
′ where F (·) is a learnable model with parameters 𝜃 and Ŷ ∈ R𝑁 ×𝑇 ×𝐹 .
Definition 2 (Federated Spatial-Temporal Prediction): In the federated setting, the sensor set is split across 𝑀 clients. Client 𝑚 holds a local subgraph G (𝑚) = (V (𝑚) , A (𝑚) ) together with its measurements X (𝑚) , where the client partitions are disjoint, V (𝑚) ∩ V (𝑛) = ∅ for 𝑚 ≠ 𝑛, and jointly cover the full network. This partitioning severs the boundary edges E𝑚,𝑛 that originally connected adjacent client subgraphs in the original network, removing a source of spatial information that traditional federated approaches do not recover. The aim is for clients to collaboratively improve their forecasting performance without aggregating their raw observations. This is formulated as a joint minimisation over per-client parameters: 𝑀 ∑︁ (𝑚) (𝑚) min L F (𝑚) X𝑡(𝑚) ; 𝜃 , Y (2) −𝑇 +1:𝑡 𝑀 {𝜃 (𝑚) }𝑚=1 𝑚=1
where L (·) is the local forecasting loss, Y (𝑚) denotes the future traffic states at client 𝑚, and each 𝜃 (𝑚) remains local to its client. Our framework aligns client collaboration through complementary mechanisms that operate on gradients and on boundary information, as defined in Section 4.
4
METHODOLOGY
Figure 2 provides an overview of the proposed FedeRICo framework. This approach is designed for federated spatial-temporal forecasting under graph partitioning, where each client owns a different road subgraph and therefore learns spatial operators that are not directly interchangeable. The framework addresses two complementary challenges. First, clients should benefit from each
other’s optimisation dynamics without requiring full parameter aggregation across heterogeneous graphs. Second, adjacent clients should be able to exchange limited boundary information so that disturbances across partition boundaries are not entirely removed by the federated split. This section describes the components in the framework which address these challenges. Section 4.1 introduces the dual-branch forecasting model and its residual training objective. Section 4.2 describes the gradient alignment and divergence regularisation method. Section 4.3 presents the construction and injection of boundary residual messages.
4.1
Dual-Branch Architecture
FedeRICo uses a dual-branch forecasting architecture in which both branches are built from spatial-temporal (ST) Blocks based on Graph WaveNet [26]. Each ST-Block combines gated temporal convolutions with graph diffusion convolution. For client 𝑚, let 𝑁𝑚 = |V (𝑚) |. The input traffic window X𝑡(𝑚) −𝑇 +1:𝑡 is first projected to a hidden representation. Denote the input to the ℓ-th ST-Block by Hℓ(𝑚) . The temporal module applies a gated temporal convolution: 𝑓 Uℓ(𝑚) = tanh Θℓ ★ Hℓ(𝑚) , 𝑔 Rℓ(𝑚) = 𝜎 Θℓ ★ Hℓ(𝑚) , Zℓ(𝑚) = Uℓ(𝑚) ⊙ Rℓ(𝑚) .
(3)
where Uℓ(𝑚) is the candidate temporal response, Rℓ(𝑚) is the temporal 𝑓 𝑔 gate, Θℓ and Θℓ are the filter and gate temporal convolution kernels, ★ denotes temporal convolution, 𝜎 (·) is the sigmoid function, and ⊙ denotes element-wise multiplication. Spatial dependencies are then modelled through graph diffusion convolution over physical and learned graph supports. Following bi-directional diffusion convolution [11, 26], the physical supports are the forward random-walk transition matrix induced by A (𝑚) ⊤ and the reverse transition matrix induced by A (𝑚) . In addition,
Conference’17, July 2017, Washington, DC, USA
Orozco et al.
each client learns an adaptive graph support from trainable node embeddings: b (𝑚) = softmax ReLU E (𝑚) E (𝑚) , A (4) 1 2 adp where E1(𝑚) ∈ R𝑁𝑚 ×𝑑 and E2(𝑚) ∈ R𝑑 ×𝑁𝑚 are learnable clientspecific node embeddings, and the softmax is applied row-wise. Because this support is produced from two unconstrained embedding matrices it can represent asymmetric spatial influence. This is compatible with boundary-message inputs since information introduced at boundary nodes can be propagated through a directed spatial operator. The resulting adaptive support is used together with the physical supports, so each client’s spatial propagation depends on both the observed road topology and a learned clientspecific graph operator. The shared and local branches use the same ST-Block backbone but maintain separate parameters. The shared branch receives only the local traffic observations, while the local branch receives both the local observations and the boundary message representation from adjacent clients. Boundary messages are projected to the model hidden dimension and added to the local branch hidden representation through a learned gate, after the initial input projection. The final prediction is composed additively: b (𝑚) = Y b (𝑚) + Y b (𝑚) , Y sh loc
(5)
b (𝑚) and Y b (𝑚) denote the shared-branch and local-branch where Y sh loc
outputs, respectively. During training, the shared branch is supervised against the forecasting target, while the local branch is supervised against the residual left by the shared branch. This decomposition encourages the shared branch to capture the dominant local forecasting signal, while the local branch focuses on residual corrections informed by boundary messages.
4.2
Bi-Gradient Regularisation
FedeRICo coordinates clients through their gradients rather than their parameters. As shown in Table 2, parameter averaging performs poorly under heterogeneous client partitions, suggesting that direct aggregation can dilute client-specific representations. This approach instead draws on the bi-gradient principle of FedIGL [24], in which gradients associated with a shared branch are encouraged to align across clients, while gradients associated with a client-specific branch are discouraged from collapsing to the same direction. This framework adapts this principle to spatial-temporal forecasting by applying coordination directly to gradient directions after backpropagation and before the optimiser step, preserving each client’s gradient norm. Both alignment and diversification operate on cosine similarity rather than 𝐿2 distance, since gradient magnitudes can vary significantly across clients due to different underlying spatial operators. Cosine similarity isolates directional agreement, providing a coordination signal that is more robust to magnitude differences that may arise from heterogeneous graph spectra and local traffic dynamics [18, 25]. (𝑚) (𝑚) Let gsh and gloc denote the current gradients of the shared and local branches of client 𝑚 for non-node-specific parameters. At the end of each communication round, FedeRICo estimates allclient reference directions from the accumulated client gradients.
(𝑚) The shared-branch reference for client 𝑚 is denoted ḡsh , and the
(𝑚) local-branch reference is denoted ḡloc . Let u(g) = g/∥g∥ 2 denote the unit direction of a non-zero gradient. For the shared branch, the client gradient direction is merged with the population reference direction, (𝑚) (𝑚) (𝑚) ash = (1 − 𝜆align )u gsh + 𝜆align u ḡsh , (𝑚) (𝑚) (𝑚) e gsh = ∥gsh ∥ 2 u ash . (6)
For the local branch, this approach only modifies the gradient when it is too aligned with the population local-branch reference: (𝑚) (𝑚) 𝑐𝑚 = cos gloc , ḡloc , (𝑚) (𝑚) (𝑚) aloc = u gloc − 𝜆div [𝑐𝑚 − 𝛿] + u ḡloc , (𝑚) (𝑚) (𝑚) e gloc = ∥gloc ∥ 2 u aloc . (7) where [𝑥] + = max(𝑥, 0) denotes the positive-part operator, 𝜆div controls the strength of the diversification update, and 𝛿 is the local-branch diversity margin. The modified gradients are then used for the optimiser step. FedeRICo therefore couples clients through branch-level gradient directions rather than through direct parameter averaging. Shared branches are guided toward population-level forecasting structure, while the local branch remains free to preserve personalised residual corrections, including those induced by boundary messages.
4.3
Boundary Message Composition
Federated partitioning removes physical edges that originally connected neighbouring road regions. These removed edges are important for traffic forecasting because congestion, incidents, and other short-term disturbances often propagate across the partition boundary. The proposed framework therefore introduces boundary message composition to integrate an encoded view of cross-client spatial context. This approach does not communicate full raw trajectories or low-frequency traffic structure; such signals may encode persistent demand levels, commuting periodicity, and other commercially sensitive regional patterns. Instead, the communicated message is designed to represent boundary-local residual behaviour, such as non-periodic short-term deviations near the boundary after smooth temporal structure has been suppressed. For a neighbouring client pair (𝑛, 𝑚), let B𝑚←𝑛 denote the receiverside boundary nodes in client 𝑚 adjacent to client 𝑛. For each receiver boundary node, FedeRICo selects the top-𝑘 sender nodes in client 𝑛 closest to the cross-client interface. The sender-side boundary packet is h i (𝑣𝑘 ) (𝑢,𝑡 ) 1) B𝑛→𝑚 = x𝑡(𝑣−𝑊 (8) +1:𝑡 , . . . , x𝑡 −𝑊 +1:𝑡 , where 𝑢 ∈ B𝑚←𝑛 is the receiver boundary node, {𝑣 1, . . . , 𝑣𝑘 } are the selected sender nodes, and 𝑊 is the boundary history window. The packet is processed by a lightweight boundary spatial-temporal module. This module first encodes the sender boundary packet using gated temporal convolutions and a learned mixing operator over (𝑢,𝑡 ) the selected sender nodes. Let Z𝑛→𝑚 denote the encoded boundary
FedeRICo
Conference’17, July 2017, Washington, DC, USA
representation produced by this first boundary ST module. Following the motivation of spatial-temporal decomposition methods such as STDN [4], this approach uses temporal embeddings such as time-of-day and day-of-week to estimate and suppress periodic structure in the encoded boundary representation: (𝑢,𝑡 ) (𝑢,𝑡 ) (𝑢,𝑡 ) R𝑛→𝑚 = Z𝑛→𝑚 − Z𝑛→𝑚 ⊙ 𝜎 𝑞𝜓 (T𝑡 ) ,
(9)
where T𝑡 denotes the temporal embedding, 𝑞𝜓 (·) projects the temporal embedding to the boundary representation dimension, 𝜎 (·) is the sigmoid function, and ⊙ denotes element-wise multiplication. The (𝑢,𝑡 ) resulting residual representation R𝑛→𝑚 suppresses time-periodic structure while retaining short-term boundary deviations. A second boundary ST module processes this residual representation to form the transmitted message: (𝑢,𝑡 ) (𝑢,𝑡 ) m𝑛→𝑚 = ℎ𝜙 R𝑛→𝑚 , (10) where ℎ𝜙 (·) denotes the residual boundary-message module. The (𝑢,𝑡 ) output m𝑛→𝑚 is therefore a learned residual boundary message rather than a raw boundary trajectory. For client 𝑚, all incoming messages from adjacent clients are assembled into a boundary context tensor M (𝑚) . This context is provided only to the local branch, together with the client’s own input window and graph supports. The shared branch does not receive boundary messages. This confines cross-client information to the personalised residual-correction pathway while keeping the shared branch focused on the client’s own spatial-temporal structure.
4.4
Federated Framework
The proposed framework has two stages. First, each client trains a small sender-local boundary-message encoder on its own boundary nodes, where boundary nodes are determined by cross-client proximity. This pretraining uses only local sender-side traffic input (𝑢,𝑡 ) data B𝑛→𝑚 , and sender-side forecasting targets. Hence, there is no cross-client communication or collaboration in pre-training this client-specific module. In our final configuration, these encoders are trained for 20 local epochs and then frozen, so during federated training they produce compact boundary messages for neighbouring receiver clients. In the main federated stage, each client maintains a personalised spatial-temporal model with shared and local branches. At each communication round, clients perform local optimisation on their own subgraphs. For each batch, incoming boundary messages are constructed from adjacent sender clients, aligned by time, and injected into the receiver model. The server coordinates learning using normalised gradients from model branches. Shared branches are encouraged to follow the reference directions, while private branches are discouraged from collapsing to the same direction. Thus, clients exchange model-update information through server and boundary-context messages across adjacent clients, but raw node observations remain local. Clients do not transmit raw traffic observations, boundary messages, or per-batch gradients to the server. Each client only communicates round-level branch gradients used to form reference directions for the next communication round. The server-side mechanism is detailed in Algorithm 1.
Algorithm 1 FedeRICo Federated Training Require: Clients C, datasets {D𝑘 }, boundary encoders {𝜓𝑘 } Require: Rounds 𝑅, local epochs 𝐸 Ensure: Personalised client models {𝜃 𝑘 } 1: Initialise client models {𝜃 𝑘 }. 2: Set server reference directions as unavailable. 3: for federated round 𝑟 = 1, . . . , 𝑅 do 4: if 𝑟 > 1 then 5: Server sends each client its reference directions. 6: end if 7: for each client 𝑘 in parallel do 8: Initialise branch-gradient accumulators. 9: for local epoch 𝑒 = 1, . . . , 𝐸 do 10: for each mini-batch do (𝑢,𝑡 ) 11: Encode packets B𝑛→𝑘 into context 𝐶𝑘 (𝑡). 12: Predict 𝑌ˆ𝑘 = 𝑓𝜃𝑘 (𝑋𝑘 , 𝐶𝑘 (𝑡)). 13: Compute L𝑘 using shared and local-residual targets. 14: Backpropagate ∇𝜃𝑘 L𝑘 . 15: Accumulate shared and local branch gradients. 16: if 𝑟 > 1 then 17: Align shared gradients to reference from Eq. (6). 18: Diverge local gradients from reference from Eq. (7). 19: end if 20: Update 𝜃 𝑘 . 21: end for 22: end for 23: Send round-level gradient summaries to the server. 24: end for 25: Server forms references for the next round. 26: Validate clients and retain best checkpoints. 27: end for
5 EXPERIMENTS 5.1 Datasets Table 1: Statistics of traffic datasets.
Datasets METR-LA PEMS-BAY PEMS03 PEMS07 (M)
# Sensors 207 325 358 228
# Samples 34,272 52,116 26,208 12,672
Target Feature Speed Speed Traffic Flow Speed
FedeRICo is evaluated on four real-world traffic forecasting datasets commonly used in spatial-temporal prediction: METR-LA, PEMSBAY [11], PEMS03 [21], and PEMS07 (M) [30]. All datasets record data at 5-minute intervals. METR-LA and PEMS-BAY contain traffic speed measurements from freeway sensors in the Los Angeles County and the Bay Area, respectively. PEMS03 and PEMS07 (M) are collected from California Performance Measurement System (PEMS) districts, with PEMS03 using traffic flow and PEMS07 (M) using traffic speed as the prediction target.
Conference’17, July 2017, Washington, DC, USA
5.2
Baseline Methods
To assess the performance of FedeRICo, several baseline centralised spatial-temporal forecasting models and federated traffic forecasting methods are used. The centralised models are trained with centralised, full road network datasets to provide reference performance for the non-federated setting. The federated baselines operate under the same partitioning protocol as FedeRICo. • STDN [4]: Decomposes traffic signals into trend and seasonality components via spatial-temporal embeddings, encoding each stream with separate GRUs before decoding with multi-head attention. • GWNet [26]: Combines stacked layers of dilated temporal convolutions with graph convolutions over both predefined and learned adaptive graph supports. • AGCRN [2]: Learns node-specific embeddings to construct an adaptive graph for recurrent spatial-temporal modelling. For the federated traffic prediction methods, the following baselines are implemented: • FedAvg [16]: Trains local client models and periodically averages model parameters across clients using GWNet as the spatialtemporal model. • MFVSTGNN [12]: Federates a multi-view spatial-temporal graph neural network, using separate local and collaborative views to model client-specific traffic dynamics and shared spatial-temporal dependencies. • FedGTP [28]: Performs federated traffic prediction on an AGCRNbased model by exchanging learned spatial-temporal representations between clients through a graph-based collaborative training mechanism. • pFedCTP [32]: Coordinates FL through a shared temporal module and a private spatial module, with personalisation achieved by adaptively aggregating shared parameters. • FedDis [33]: Implements an AGCRN-based spatial temporal model with separate shared and personalised representations using client similarity to guide federated parameter aggregation.
5.3
Experiment Settings
All experiments are implemented in PyTorch and conducted on NVIDIA RTX 4090 GPUs. Each dataset is split into training, validation, and test sets using a 60:20:20 ratio. Inputs are processed into 𝑇 = 12 window time steps, with the prediction horizon 𝑇 ′ = 12. All models are trained using the Adam optimiser [10] with a learning rate of 1𝑒 −3 , and batch size 128. Federated methods are trained for 100 global rounds and 1 local round of training. Source code is made available at https://anonymous.4open.science/r/FedeRICo-CA6B.
6 RESULTS 6.1 Performance Comparison Table 2 reports the forecasting performance across four traffic benchmarks under centralised and federated settings. FedeRICo
Orozco et al.
achieves the best federated performance on every dataset and metric, consistently improving over both parameter-averaging baselines and recent personalised and collaborative FL methods. Compared with FedAvg, FedeRICo reduces MAE by 18.8% on METR-LA, 10.2% on PEMS-BAY, 24.6% on PEMS-D3, and 10.6% on PEMS-D7. The comparison also highlights a gap between existing federated traffic prediction methods and centralised methods. The performance of FedGTP, MFVSTGNN, pFedCTP, and FedDis remain substantially below all centralised models on most datasets, suggesting that direct parameter averaging or representation-level collaboration may be insufficient when traffic networks are partitioned into heterogeneous client subgraphs. FedeRICo narrows this gap most consistently: its MAE remains within 3.2% − 7.3% of centralised GWNet across datasets. Moreover, FedeRICo matches or outperforms centralised AGCRN across all datasets, despite operating without centralising client traffic datasets. Relative to the strongest federated baseline, FedDis, FedeRICo lowers MAE by 0.32 on METR-LA (9.1%), 0.12 on PEMS-BAY (6.7%), 0.65 on PEMS-D3 (3.9%), and 0.26 on PEMS-D7 (8.8%). These results support the main design motivation of FedeRICo. Boundary messages provide cross-client spatial context without merging raw subgraphs, while branch-gradient coordination avoids the detrimental effect caused by averaging parameters.
6.2
Ablation Studies
6.2.1 FL Framework Ablation. Figure 3 isolates the contribution of each component in FedeRICo. w/o Coll trains the dual-branch model locally without federated collaboration, w/o BM removes cross-client boundary messages while retaining gradient alignment, w/o GA replaces gradient alignment with standard parameter averaging, w OS uses operator-similarity-based client selection, and FedeRICo denotes the full proposed framework. The local-only dual-branch variant (w/o Coll) already provides a competitive personalised forecasting model, but it is consistently worse than the collaborative variants on both METR-LA and PEMSBAY. This shows that the gains of this approach do not come solely from increasing model capacity through the dual-branch design; cross-client coordination is necessary to recover spatial dependencies that are broken by graph partitioning. Removing boundary messages (w/o BM) degrades performance relative to the full model across both datasets. On METR-LA, MAE increases from 3.1878 to 3.2845, while RMSE increases from 6.3095 to 6.5712. These drops indicate that boundary messages provide useful cross-client residual context beyond what can be recovered through gradient coordination alone. The w/o GA variant evaluates whether standard federated parameter averaging can replace the proposed branch-gradient coordination. This ablation is particularly revealing on METR-LA, where MAE rises to 3.4856 and MAPE increases sharply to 10.76. The degradation supports the premise that directly averaging parameters across heterogeneous traffic subgraphs can interfere with client-specific representations. Gradient-space coordination provides a softer coupling mechanism, aligning the shared branch while preserving local residual specialisation. For the W OS variant, client compatibility is estimated from the graph operators learned by the forecasting backbone. In GWNet [26],
FedeRICo
Conference’17, July 2017, Washington, DC, USA
Table 2: Results across METR-LA, PEMS-BAY, PEMS-D3, and PEMS-D7 datasets.
Framework
Model
METR-LA
PEMS-BAY
PEMS-D3
PEMS-D7
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
Centralised Centralised Centralised
STDN GWNet AGCRN
3.01 3.09 3.26
6.07 6.20 6.52
8.12 8.38 9.06
1.65 1.61 1.67
3.74 3.58 3.79
3.69 3.62 3.80
15.04 14.81 15.90
26.36 25.38 27.46
15.20 14.74 16.11
2.55 2.60 2.77
5.09 5.09 5.34
6.36 6.61 6.94
Federated Federated Federated Federated Federated
FedAvg FedGTP MFVSTGNN pFedCTP FedDis
3.93 3.92 3.79 3.76 3.51
7.49 7.61 7.43 7.01 6.98
11.98 11.82 10.82 10.39 10.02
1.86 1.97 1.83 2.05 1.79
4.04 4.37 4.09 4.61 3.95
4.30 4.62 4.21 4.88 4.09
21.07 18.49 17.30 20.83 16.54
33.48 30.38 28.88 33.85 28.43
25.44 20.05 17.64 22.68 16.42
3.03 3.01 3.10 3.38 2.97
5.71 5.93 5.89 6.33 5.70
7.82 7.79 7.95 8.71 7.42
Federated
FedeRICo
3.19
6.31
8.87
1.67
3.67
3.72
15.89
26.83
16.03
2.71
5.24
6.76
w/o BM
w/o Coll
3.1
(a)
FedeRICo
10.8
3.82
10.4
1.72
3.78
1.7
3.74
3.84 3.8
6.45
9.6 9.2 8.8
6.25
8.4
6.15
8
(b)
1.68
3.7
1.66
3.66
1.64
3.62
(c)
MAPE
10
6.55
6.35 3.2
W OS
1.74
RMSE
3.3
w/o GA
MAE
RMSE
MAE
3.4
6.65
MAPE
3.5
6.75
3.76 3.72 3.68
(d)
(f)
(e)
Figure 3: Ablation study. Subfigures (a), (b), and (c) are the results of the METR-LA ablation experiments, while subfigures (d), (e), and (f) present the results of the PEMS-BAY ablation experiments. 3.28
RMSE
MAE
6.4
3.24 3.22 3.2
6.3 6.25 8 (a)
4
12
4
3.22
8 (b)
12
3 (d)
5
6.32
RMSE
3.21 MAE
6.35
3.18
3.2 3.19 3.18
6.3 6.28 6.26
1
3 (c)
5
1
3.22
6.33
3.21
6.32 RMSE
6.2.2 Hyperparameter Ablation. We study the sensitivity of FedeRICo to two groups of hyperparameters on METR-LA: boundarymessage configuration and gradient-controller strength. For boundary messages, we vary the message hidden dimension 𝑑𝑚 , the temporal decorrelation coefficient 𝜆𝑡 , and the number of received sender contexts top-𝑘. For the gradient controller, we jointly vary 𝜆align , 𝜆div , and 𝛿 using a shared value, testing whether the strength of branch-gradient coordination requires narrow tuning. Figure 4 shows that FedeRICo is not highly sensitive to either boundary-message or gradient-controller hyperparameters. Increasing the message dimension from 𝑑𝑚 = 4 to 𝑑𝑚 = 8 improves both MAE and RMSE, while increasing to 𝑑𝑚 = 12 gives only marginal
6.45
3.26
MAE
each client learns an adaptive graph support in addition to the physical road network. We construct structural signatures from these learned graph operators using GraphWave [6], which characterises nodes by heat-kernel diffusion responses and therefore captures structural roles rather than node identities. Since clients contain disjoint node sets, we compare the resulting signature distributions using sliced Wasserstein distance [3]. The operator-similarity variant (W OS) performs competitively, but does not consistently improve over all-client coordination. This suggests that the learned gradient references are sufficiently stable to benefit from broader all-client coordination. Overall, the ablation results show that the two main mechanisms are complementary. Boundary messages restore missing spatial context at partition boundaries, while branch-gradient coordination prevents this collaboration from collapsing into destructive parameter averaging.
3.2
6.31
3.19
6.3
3.18
6.29 0.05
0.10
0.20
(e)
0.05
0.10
0.20
(f)
Figure 4: Hyperparameter ablation on METR-LA. Subfigures (a) and (b) vary the boundary-message dimension 𝑑𝑚 under two temporal decorrelation strengths. Subfigures (c) and (d) vary the number of received sender contexts top-𝑘 using the default 𝜆𝑡 = 0.10. Subfigures (e) and (f) vary the gradient-controller strength by jointly setting 𝜆align = 𝜆div = 𝛿.
Conference’17, July 2017, Washington, DC, USA
Orozco et al.
Table 3: Federated model results on METR-LA across different numbers of clients.
# Clients
Metric
pFedCTP
FedDis
FedeRICo
4
MAE RMSE MAPE
3.75 7.27 10.74
3.47 6.86 9.74
3.22 6.35 8.93
8
MAE RMSE MAPE
3.80 7.27 10.85
3.49 6.91 9.88
3.16 6.24 8.76
12
MAE RMSE MAPE
3.76 7.01 10.39
3.51 6.97 9.94
3.19 6.31 8.87
16
MAE RMSE MAPE
3.79 7.21 11.02
3.55 7.04 10.08
3.26 6.45 9.11
20
MAE RMSE MAPE
3.82 7.27 10.97
3.61 7.17 10.28
3.26 6.48 9.13
Table 4: Computational cost comparison on METR-LA with 12 clients. Parameter counts are reported as mean trainable parameters per client. Training time is the median logged wall-clock time per federated round. For FedeRICo, the two times denote sender-local encoder fitting and federated forecasting training, respectively.
Method
Param Count
Training Time/Round
MFVSTGNN FedGTP pFedCTP FedDis FedeRICo
11.4K 313.9K 69.0K 1.52M 165.9K + 1.3K/encoder
40.6s 1186.5s 58.6s 148.4s 21.3s / 40.7s
additional benefit. The two 𝜆𝑡 curves remain close, indicating that performance is not tied to a narrowly tuned temporal decorrelation strength. The top-𝑘 sweep is similarly stable, with one, three, or five sender contexts producing small changes. The gradientcontroller sweep also remains within a narrow range, suggesting that branch-gradient coordination does not depend on a precisely tuned coupling strength.
6.3
Client Scalability Comparison
Table 3 evaluates robustness to different numbers of METR-LA clients. FedeRICo consistently outperforms pFedCTP and FedDis across all partition sizes, showing that its boundary-message and gradient-coordination mechanisms remain effective as the graph is partitioned into more subsets for each client. This also demonstrates that FedeRICo scales better than the competing federated baselines under increasing client partitioning.
6.4
Computational Costs Comparison
Table 4 compares the computational cost of the federated methods on METR-LA. FedeRICo uses more parameters than lightweight personalised baselines such as pFedCTP, but remains substantially smaller than FedDis and avoids the high per-round runtime of
Table 5: Compromised-receiver transfer reconstruction audit on METR-LA, 𝐾 = 12 Voronoi. Here, metadata denotes the temporal features, receiver boundary-node identity, and sender rank.
Input
MAE RMSE Corr.
Packet + metadata 25.22 Metadata only 15.80
37.74 23.57
𝑅2
0.200 -3.123 0.084 -0.190
FedGTP. The additional sender-local encoder cost is separated from the main federated training loop, and the forecasting round time remains competitive with the other federated baselines. This indicates that the performance gains of FedeRICo do not come from simply scaling model size or incurring substantially higher federated training cost.
6.5
Boundary Messaging Privacy Consideration
We audit a compromised-receiver attack in which one client trains an inversion model using boundary-message examples for which it has corresponding local raw observations, and then transfers this model to incoming boundary messages from neighbouring clients. The attacker is implemented as a contextual MLP regressor, where the boundary packet is encoded by a two-layer MLP, temporal covariates are projected, receiver boundary-node identity and sender rank are embedded, and the fused representation is decoded to the raw boundary signal using an ℓ1 reconstruction loss. The packet-conditioned attacker does not improve over the metadata-only baseline in this transfer setting. This suggests that sender-local boundary packets do not provide a stable inversion signal to a receiver trained only on its own locally observable examples.
7
CONCLUSION
This paper introduced FedeRICo, a federated spatial-temporal forecasting framework for traffic networks partitioned across heterogeneous client subgraphs. This approach separates collaboration from parameter sharing by coordinating clients through gradientdirection alignment while preserving client-specific forecasting parameters. It further restores cross-partition spatial context through boundary-aware residual messages exchanged between physically adjacent clients without sharing raw traffic observations. Experiments on four real-world traffic benchmarks show that FedeRICo consistently outperforms existing federated spatial-temporal baselines and narrows the gap to centralised forecasting. Ablation, scalability, hyperparameter, and reconstruction-audit results further indicate that boundary residual messaging and gradient-level coordination are complementary, robust across client partitions, and do not provide a stable inversion signal under the compromisedreceiver attack considered. Future work may aim extend the framework for asynchronous participation and stronger formal privacy mechanisms for boundarymessage exchange.
GENAI DISCLOSURE STATEMENT Use of Generative AI for this work was solely for language refinement, formatting assistance, and grammatical corrections.
FedeRICo
REFERENCES [1] Durmus Alp Emre Acar, Yue Zhao, Ramon Matas, Matthew Mattina, Paul Whatmough, and Venkatesh Saligrama. 2021. Federated Learning Based on Dynamic Regularization. In International Conference on Learning Representations. [2] LEI BAI, Lina Yao, Can Li, Xianzhi Wang, and Can Wang. 2020. Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 17804–17815. [3] Nicolas Bonneel, Julien Rabin, Gabriel Peyré, and Hanspeter Pfister. 2015. Sliced and Radon Wasserstein Barycenters of Measures. Journal of Mathematical Imaging and Vision 51, 1 (2015), 22–45. [4] Lingxiao Cao, Bin Wang, Guiyuan Jiang, Yanwei Yu, and Junyu Dong. 2025. Spatiotemporal-aware trend-seasonality decomposition network for traffic flow forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 11463–11471. [5] Minxiao Chen, Haitao Yuan, Nan Jiang, Zhifeng Bao, and Shangguang Wang. 2024. Urban Traffic Accident Risk Prediction Revisited: Regionality, Proximity, Similarity and Sparsity. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (Boise, ID, USA) (CIKM ’24). Association for Computing Machinery, New York, NY, USA, 281–290. https://doi.org/10. 1145/3627673.3679567 [6] Claire Donnat, Marinka Zitnik, David Hallac, and Jure Leskovec. 2018. Learning Structural Node Embeddings via Diffusion Wavelets. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1320–1329. [7] Jeffrey L. Elman. 1990. Finding Structure in Time. Cognitive Science 14, 2 (1990), 179–211. https://doi.org/10.1207/s15516709cog1402_1 [8] Yang Hu, Shaobo Li, Dawen Xia, Zhiheng Zhou, Wenyong Zhang, Huaqing Li, Xingxing Zhang, and Senzhang Wang. 2025. Mixture of Semantic and Spatial Experts for Explainable Traffic Prediction. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management (Seoul, Republic of Korea) (CIKM ’25). Association for Computing Machinery, New York, NY, USA, 908–917. https://doi.org/10.1145/3746252.3761412 [9] Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh. 2020. SCAFFOLD: Stochastic Controlled Averaging for Federated Learning. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119), Hal Daumé III and Aarti Singh (Eds.). PMLR, 5132–5143. [10] Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). [11] Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In International Conference on Learning Representations. [12] Lei Liu, Yuxing Tian, Chinmay Chakraborty, Jie Feng, Qingqi Pei, Li Zhen, and Keping Yu. 2023. Multilevel Federated Learning-Based Intelligent Traffic Flow Forecasting for Transportation Network Management. IEEE Transactions on Network and Service Management 20, 2 (2023), 1446–1458. https://doi.org/10. 1109/TNSM.2023.3280515 [13] Yi Liu, James J. Q. Yu, Jiawen Kang, Dusit Niyato, and Shuyu Zhang. 2020. PrivacyPreserving Traffic Flow Prediction: A Federated Learning Approach. IEEE Internet of Things Journal 7, 8 (2020), 7751–7763. https://doi.org/10.1109/JIOT.2020. 2991401 [14] Zhanyu Liu, Guanjie Zheng, and Yanwei Yu. 2023. Cross-city Few-Shot Traffic Forecasting via Traffic Pattern Bank. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (Birmingham, United Kingdom) (CIKM ’23). Association for Computing Machinery, New York, NY, USA, 1451–1460. https://doi.org/10.1145/3583780.3614829 [15] Yuxiang Lu, Suizhi Huang, Yuwen Yang, Shalayiding Sirejiding, Yue Ding, and Hongtao Lu. 2024. Fedhca2: Towards Hetero-Client Federated Multi-Task Learning . In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, Los Alamitos, CA, USA, 5599–5609. https://doi.org/10.1109/CVPR52733.2024.00535 [16] Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017. Communication-Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 54), Aarti Singh and Jerry Zhu (Eds.). PMLR, 1273–1282. https://proceedings. mlr.press/v54/mcmahan17a.html [17] Chuizheng Meng, Sirisha Rambhatla, and Yan Liu. 2021. Cross-Node Federated Graph Neural Network for Spatio-Temporal Data Modeling. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (Virtual Event, Singapore) (KDD ’21). Association for Computing Machinery, New York, NY, USA, 1202–1211. https://doi.org/10.1145/3447548.3467371 [18] Felix Sattler, Klaus-Robert Müller, and Wojciech Samek. 2021. Clustered Federated Learning: Model-Agnostic Distributed Multitask Optimization under Privacy
Conference’17, July 2017, Washington, DC, USA
Constraints. IEEE Transactions on Neural Networks and Learning Systems 32, 8 (2021), 3710–3722. [19] Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. 2009. The Graph Neural Network Model. IEEE Transactions on Neural Networks 20, 1 (2009), 61–80. https://doi.org/10.1109/TNN.2008.2005605 [20] Zezhi Shao, Zhao Zhang, Wei Wei, Fei Wang, Yongjun Xu, Xin Cao, and Christian S. Jensen. 2022. Decoupled dynamic spatial-temporal graph neural network for traffic forecasting. Proc. VLDB Endow. 15, 11 (July 2022), 2733–2746. https://doi.org/10.14778/3551793.3551827 [21] Chao Song, Youfang Lin, Shengnan Guo, and Huaiyu Wan. 2020. Spatial-Temporal Synchronous Graph Convolutional Networks: A new framework for spatialtemporal network data forecasting. Proc. Conf. AAAI Artif. Intell. 34, 01 (April 2020), 914–921. [22] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/ 2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf [23] Binwu Wang, Pengkun Wang, Yudong Zhang, Xu Wang, Zhengyang Zhou, Lei Bai, and Yang Wang. 2024. Towards dynamic spatial-temporal graph learning: A decoupled perspective. 38, 8 (March 2024), 9089–9097. [24] Lingren Wang, Wenxuan Tu, Jiaxin Wang, Xiong Wang, Jieren Cheng, and Jingxin Liu. 2026. FedIGL: Federated Invariant Graph Learning for Non-IID Graphs. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=H9BdN4f2vz [25] Zheng Wang, Xiaoliang Fan, Jianzhong Qi, Chenglu Wen, Cheng Wang, and Rongshan Yu. 2021. Federated Learning with Fair Averaging. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence. 1615–1623. [26] Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, and Chengqi Zhang. 2019. Graph wavenet for deep spatial-temporal graph modeling (IJCAI’19). AAAI Press, 1907–1913. [27] Guang Yang, Yuequn Zhang, Jinquan Hang, Xinyue Feng, Zejun Xie, Desheng Zhang, and Yu Yang. 2023. CARPG: Cross-City Knowledge Transfer for Traffic Accident Prediction via Attentive Region-Level Parameter Generation. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (Birmingham, United Kingdom) (CIKM ’23). Association for Computing Machinery, New York, NY, USA, 2939–2948. https: //doi.org/10.1145/3583780.3614802 [28] Linghua Yang, Wantong Chen, Xiaoxi He, Shuyue Wei, Yi Xu, Zimu Zhou, and Yongxin Tong. 2024. FedGTP: Exploiting Inter-Client Spatial Dependency in Federated Graph-based Traffic Prediction. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain) (KDD ’24). Association for Computing Machinery, New York, NY, USA, 6105–6116. https://doi.org/10.1145/3637528.3671613 [29] Xiuwen Yi, Zhewen Duan, Ting Li, Tianrui Li, Junbo Zhang, and Yu Zheng. 2019. CityTraffic: Modeling Citywide Traffic via Neural Memorization and Generalization Approach. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (Beijing, China) (CIKM ’19). Association for Computing Machinery, New York, NY, USA, 2665–2671. https://doi.org/10.1145/3357384.3357822 [30] Bing Yu, Haoteng Yin, and Zhanxing Zhu. 2018. Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18. International Joint Conferences on Artificial Intelligence Organization, 3634–3640. https://doi.org/10.24963/ijcai.2018/505 [31] Yupu Zhang, Lei Jia, Hao Miao, Weizhu Qian, Yan Zhao, and Kai Zheng. 2025. Traffic Safety Evaluation Based on Macroscopic Traffic Features in Road Tunnels. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management (Seoul, Republic of Korea) (CIKM ’25). Association for Computing Machinery, New York, NY, USA, 4253–4262. https://doi.org/10.1145/ 3746252.3761296 [32] Yu Zhang, Hua Lu, Ning Liu, Yonghui Xu, Qingzhong Li, and Lizhen Cui. 2024. Personalized Federated Learning for Cross-City Traffic Prediction. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI24, Kate Larson (Ed.). International Joint Conferences on Artificial Intelligence Organization, 5526–5534. https://doi.org/10.24963/ijcai.2024/611 Main Track. [33] Chengyang Zhou, Zijian Zhang, Chunxu Zhang, Hao Miao, Yulin Zhang, Kedi Lyu, and Juncheng Hu. 2026. FedDis: A Causal Disentanglement Framework for Federated Traffic Prediction (WWW ’26). Association for Computing Machinery, New York, NY, USA, 7541–7551. https://doi.org/10.1145/3774904.3792663