Falcon-X: A Time Series Foundation Model
for Heterogeneous Multivariate Modeling
Yiding Liu∗ , Yifan Hu∗ , Hongjie Xia∗ , Peiyuan Liu∗ , Hongzhou Chen , Xilin Dai , Zewei Dong† , Jiangming Yang
arXiv:2605.27286v1 [cs.LG] 26 May 2026
∗ Equal Contribution, † Corresponding Author, # {yiding.lyd, zewei.dong}@ant-intl.com
Time series foundation models (TSFMs) are transforming the forecasting paradigm through large-scale cross-domain pretraining. However, most existing TSFMs remain univariate, and recent efforts to enable cross-variate modeling still operate directly within the raw variate space. This design introduces fundamental limitations in semantic alignment and relational expressivity. Specifically, raw-space group mixing lacks a dedicated mechanism to align heterogeneous physical quantities, while standard non-negative attention fails to capture the complex synergistic and antagonistic interactions ubiquitous in real-world systems. To address these challenges, we propose Falcon-X, decouples variates from the raw space and maps them into a unified latent prototype space. Falcon-X employs a Unified Prototype Diff-Attention mechanism that explicitly evaluates both positive and negative semantic affinities to explicitly align heterogeneous variates. Cross-variate interactions are then efficiently performed within this shared space via Latent Entity Attention, naturally facilitating zero-shot structural transfer. Finally, a Variate Reassembly Router robustly reconstructs variate-specific trajectories via a request-and-dispatch mechanism. Extensive evaluations on the GIFT-Eval and fev-bench benchmarks demonstrate that Falcon-X achieves state-ofthe-art forecasting performance, offering a principled and scalable paradigm for complex multivariate environments. Falcon-X is publicly released to support future research. Date: May 27, 2026 Code: https://github.com/ant-intl/Falcon-TST Model: Coming soon.
1
Introduction
Time series forecasting is a fundamental task for understanding dynamic systems and supporting future-oriented decision making. Conventional deep forecasting models are typically trained at the dataset level (Kong et al., 2025), making them difficult to reuse across domains, sampling frequencies, and variate structures that vary in the real world. Time series foundation models (TSFM) are reshaping this paradigm by pretraining on large-scale cross-domain data and transferring directly to new forecasting tasks (Kottapalli et al., 2025), thereby substantially reducing the cost of repeated training and tuning. However, most existing models still take univariate series as the basic modeling unit, extrapolating the future solely from each series’ own history. This formulation disconnects the co-evolving relationships that are ubiquitous in real systems, limiting the ability of TSFMs to capture complex multivariate dynamics. To enable cross-variate modeling, recent TSFMs explore how foundation models can accommodate varying numbers of variates. As shown in Table1, Moirai-1.0 (Woo et al., 2024) flattens multivariate series into a single sequence, enabling joint attention across time and variates. While straightfor1
Follow-up process ... 7 variates
jena_weather/10T (21 variates)
Variate Reassembly Router
Last-Layer
bizitobs_12c/5T
Latent Entity Attention bizitobs_12c/5T jena_weather/10T (21 variates) ett1/15T
7 variates
(a) Inputs with highly distinct temporal patterns.
sim=0.995
Shared Prototypes
📈📉
...
Unified Prototype Diff-Attention
Fixed prototype dim.=6
ett1/15T
sim=0.719
Heterogeneous Series
(b) Chronos-2
(c) Prototype Routing
(d) Falcon-X
Figure 1 Comparison of multivariate modeling paradigms. (a) Heterogeneous inputs with highly distinct temporal patterns. (b) Group attention produces almost identical attention maps for completely dissimilar inputs, exposing the severe semantic collapse and over-smoothing in the raw variate space. (c-d) In contrast, Falcon-X projects variates into a latent prototype space, yielding highly discriminative attention maps that accurately capture the underlying dynamics.
ward, this design scales poorly with the number of variates, as attention cost increases rapidly in high-dimensional settings. A more effective alternative is group attention (Cohen et al., 2025; Ansari et al., 2025), which organizes variates into dataset-, entity-, or task-level groups and confines attention to variates within the same group. This avoids indiscriminate mixing across unrelated samples while allowing a shared backbone to process multivariate inputs with varying dimensionalities. By replacing global all-to-all interaction with structured within-group communication, group attention marks an important step toward scalable multivariate TSFMs. However, group attention still exhibits fundamental limitations in semantic alignment and relational expressivity. ❶ It operates directly in the raw variate space, controlling which variates can interact but not how their heterogeneous semantics are aligned. In high-dimensional systems, only a small subset of variates typically exhibits strong dependencies, while many others are weakly related or noisy. As shown in Figure 1, Dense attention over raw variates can therefore dilute meaningful signals and promote dataset-specific correlations as if they were transferable patterns. Moreover, different datasets share the same Transformer backbone, yet their variates often correspond to entirely different physical quantities and dynamics. Without an dedicated alignment space, the model must absorb such heterogeneity implicitly in its parameters, making cross-domain transfer a by-product of parameter sharing rather than an principled process of organizing reusable temporal structures. ❷ Existing attention mechanisms have limited relational expressivity. In real-world systems, cross-variate dependencies often involve both synergistic and antagonistic interactions. However, current attention formulations primarily capture aggregative effects, lacking the ability to represent opposing dynamics and more complex interaction patterns. In this paper, we address these limitations by decoupling physical variates from the latent space used for cross-variate interaction. Instead of mixing raw variates, we map them into a shared fixed-dimensional latent prototype space, where interactions are mediated by pairs of learnable prototypes that explicitly capture both positive and negative semantic affinities. These prototypes absorb recurring temporal structures across datasets and filter out variate-specific noise, enabling the model to learn reusable cross-domain patterns in an unified aligned space. This alignment establishes a common semantic coordinate system for heterogeneous variates, making cross-domain interactions more structured, transferable, and scalable. It also replaces dense variate-level attention with lightweight variate-to-prototype interaction, substantially improving efficiency while preserving cross-variate modeling capacity. 2
Technically, we propose Falcon-X, a novel encoder-only TSFM with 591 million parameters materializing this latent prototype paradigm for heterogeneous multivariate forecasting. At its core, the Unified Prototype Diff-Attention decouples heterogeneous variates into a fixed semantic space, utilizing a differential mechanism to explicitly capture both synergistic and antagonistic variate relationships. Once aligned, the Latent Entity Attention performs global cross-variate interactions entirely within this unified space, naturally facilitating seamless cross-domain transfer without coupling to raw variate dimensions. To project back to the original physical space, a dynamic Variate Reassembly Router robustly reconstructs variate-specific trajectories via a request-and-dispatch mechanism and gated residual connections. Furthermore, Falcon-X integrates essential instance-wise normalization, patching, and a probabilistic forecasting head to ensure end-to-end stability. Extensive experiments on the GIFT-Eval (Aksu et al., 2024) and fev-bench (Shchur et al., 2025) validate that Falcon-X advances the state-of-the-art in scalable and zero-shot multivariate forecasting. In a nutshell, our main contribution can be summarized as: • A Novel Heterogeneous Modeling Paradigm. We present a paradigm shift for multivariate time series foundation models, moving from raw-space mixing to a unified latent prototype space. This approach elegantly resolves semantic discrepancies across different datasets, creating a universal coordinate system that naturally facilitates zero-shot knowledge transfer. • Architectural Innovation. We propose Falcon-X, a tailored encoder-only foundation model. It features a Differential Prototype Attention which comprehensively captures both synergistic and antagonistic system dynamics, together with a gated Variate Reassembly Router that adaptively regulates global context fusion for cross-dataset robustness. • Empirical Excellence. Through extensive evaluations on comprehensive widely used benchmarks (GIFT-Eval and fev-bench), Falcon-X consistently achieves state-of-the-art forecasting performance. The results validate its superior structural adaptability and broad generalization capabilities in complex multivariate environments.
2
Related Works
In recent years, the emergence of pre-training techniques has shifted the paradigm of time series forecasting from domain-specific models to the era of foundation models. Early explorations in this domain, such as Time-LLM (Jin et al., 2023), primarily focused on tuning third-party large language models to adapt time series tasks. Subsequently, with the accumulation of massive time series data, the trend has transitioned toward learning from scratch, where models are pre-trained directly on large-scale time series data to capture inherent temporal dynamics. Chronos (Ansari et al., 2024) frames time series forecasting as a language modeling task by tokenizing real-valued observations into a discrete vocabulary. MOMENT (Goswami et al., 2024) introduces a family of foundation models using a masked prediction task, which allows for generalization across time series tasks through fine-tuning. Inspired by the success of large language models, the generative paradigm built upon decoder-only architectures has become a prominent choice for many mainstream works. These methods typically partition time series into non-overlapping patches, treating each patch as a single token (Nie et al., 2023), and leverage decoder-only structures to generate one patch at each step (Liu et al., 2024; Das et al., 2024; Liu et al., 2025a,b; Auer et al., 2025). Despite achieving competitive zero-shot performance, these models primarily rely on the channel independence strategy, which limits them to univariate forecasting and ignores the rich context of multivariate dependencies found in real-world time series data.
3
Table 1 Comparison of capabilities of TSFMs. Falcon-X Moirai 2.0 TimesFM-2.5 Chronos-2 Timer-S1
Toto
(Ours)
(2025a)
(2024)
(2025)
(2026)
(2025) (2025b)
Univariate
✓
✓
✓
✓
✓
✓
Multivariate
✓
✗
✗
✓
✗
Cross Learning1
✓
✗
✗
✓
Signed Dependence2
✓
✗
✗
Heterogeneous Unification3
Prototype Routing
✗
TSFMs
✗
Sundial Moirai 1.0 TabPFN-TS (2024)
(2025)
✓
✓
✓
✓
✗
✓
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
Group Mixing
✗
Fixed
✗
Concat.
✗
1 Transferring of universal cross-variate interactive patterns across distinct datasets. 2 Computing both positive and negative affinities to capture synergistic and antagonistic dynamics. 3 Projecting physical variates with varying dimensionalities into a dimension-agnostic latent space.
Despite these advances, a key challenge of multivariate forecasting remains the unified modeling of heterogeneous time series. Moirai 1.0 (Woo et al., 2024) flattens variates into a single sequence to capture joint interactions. Toto (Cohen et al., 2025) introduces proportional factorized space-time attention to efficiently model cross-variate dependency. Furthermore, Chronos-2 (Ansari et al., 2025) further adopts group attention, facilitating in-context learning by sharing information across related series within flexible groups. However, these approaches still operate directly within the raw variate space, fundamentally limiting their semantic alignment and relational expressivity. Specifically, they primarily capture aggregative effects and struggle to represent the antagonistic dynamics ubiquitous in real-world physical systems. To bridge this gap, our Falcon-X introduces a differential mechanism within an explicitly aligned latent prototype space, enabling the foundation model to systematically capture complex dual dynamics and facilitate robust cross-domain transfer.
3
Falcon-X
In this section, we present the architecture of Falcon-X. As shown in Figure 2, Falcon-X consists of four parts: pre-process of Normalization and Tokenization, Time Attention, Variate Attention and Forecasting Head. To maintain a concise exposition of the mathematical formulations, we present extended discussions of our underlying design philosophies in Appendix A. 3.1
Problem Formulation
Let E = {ei ∈ Rmi × L }iN=1 be a collection of N entities, where each entity ei represents a multivariate time series with a specific variate dimensionality mi and a historical look-back window L. A key challenge in TSFM is the heterogeneity of these dimensions, with mi varying significantly across diverse entities and datasets. We define the total aggregated dimensionality M across all entities as ∑iN=1 mi . The objective of Falcon-X is to learn a dimension-agnostic mapping Fθ , parameterized by θ, that transforms the heterogeneous input space into the target predictive space: Ŷ = Fθ (X),
where X ∈ R M× L , Ŷ ∈ R M×T .
4
(1)
G
ℰ = {ℯR ∈ ℝS" ×@ }G RT!, 𝑀 ≜ :
𝑚R
RT!
Entity ID
History
𝐿
Eg. 𝑁 = 3
Future
𝑇
C : Unified latent dim. D : Hidden dim. @A3 P : Patch num. 𝑃 =
Attention across time
@!
Unified Cross-Variate Modeling in Latent Prototype Space
𝐻3
ℝ=×D×E
Timestamp 𝒯 Mask 𝑀 Timestamp 𝒯 𝒕 𝒕 𝒕 𝒕 𝒕 𝒕 𝒕 $𝟒 $𝟑 $𝟐 $𝟏 𝟎 𝟏 𝟐
𝐻F
ℝ(G×F)×D×E
𝐻F’
ℝG×F×D×E
𝒙𝟑
(a) Unified Prototype Diff-Attention ℒ .789 = Sim(𝑲𝒑𝒐𝒔, 𝑲𝒏𝒆𝒈)
𝐻'
(b) Latent Entity Attention
Fixed dim. in latent prototype space.
Entity 1
3 softmax(𝑄𝐾-./ )
𝑲𝒑𝒐𝒔 Learnable tensor as prototype 𝐾-./, 𝐾012
𝑸
Proj.Q,V
𝐻&
𝑨𝒏𝒆𝒈 3 softmax(𝑄𝐾012 ) 𝑸
𝐻'
𝑲𝒏𝒆𝒈
ℝ=×D×E
𝐶 𝐶
Ground Truth Y
" 𝐻
𝑅71Q fetches specific prototypes to assemble heterogeneous variates.
𝐻) 𝐻'’
𝒢(4)
𝐶 Cross Entity 1&3 Dependency
Varying variate dim.
4 Y 𝜇, 𝜎 ℝ=×3
(c) Variate Reassembly Router
Entity 2 Entity 3 Mask Attention across latent variate with Entity-Aware Masking ℯ! ℯ" ℯ#
FFN & Norm
𝑽
Multi-head Self-Attention
[𝐴-./ − 𝜆 ( 𝐴012 ]3 𝑉 𝑨𝒑𝒐𝒔
4 𝐻
𝐻I
Forecast + Y
Variate Attention
Forecast
Quantile Loss ℒ'()*
De-Norm
𝐻
ℝ=×D×E
" = 𝐻& +𝒢(𝐻& ) ⊙ 𝐻) 𝐻
Quantile Head
𝑚- = 3
𝒙𝟐
ℒ .789
Variate Reassembly Router
X 𝜇, 𝜎 ℝ=×(@AB)
𝒙𝟏
Entity ℯ#
𝑙× Layers
Latent Entity Attention
𝒙𝟐
𝑛× Layers
Time Attention
Norm
𝑚, = 2
𝒙𝟏
Tokenization
𝑚+ = 1
Entity ℯ"
Unified Prototype Diff-Attention
Entity ℯ! 𝒙𝟏
Gated Residual
Soft Routing 𝑷𝒊𝒅𝒙 𝑹𝒓𝒆𝒒 Proj. Proj.
𝐻&
𝑺𝒄𝒕𝒙 Proj.
𝐻'’
Figure 2 The overall architecture of Falcon-X. The raw inputs are normalized, tokenized, and processed by Time Attention to extract independent temporal features. The Unified Prototype Diff-Attention (UPDA) then projects these features into a shared prototype space, enabling Latent Entity Attention (LEA) to capture global cross-variate dependencies explicitly. The Variate Reassembly Router (VRR) then dynamically reconstructs variate-specific representations. These are fused with the temporal context and fed into a quantile head to generate the final probabilistic forecasts.
3.2
Normalization and Tokenization
Normalization. We formulate the forecasting task as a unified masked reconstruction paradigm. Given the raw input X ∈ R M×( L+T ) spanning the historical window L and future horizon T, we first replace the target future steps with placeholder tokens. To ensure scale-invariance across diverse domains, we apply an arcsine transformation: X−µ X̂ = arcsin . (2) σ Crucially, instead of zero-filling or truncating sequences at missing positions (Xiaoming et al., 2025), the instance-wise mean µ and standard deviation σ are computed exclusively from observed values, preserving missing entries for explicit downstream modeling. Tokenization. To obtain robust representations, we augment X̂ by concatenating it with a relative timestamps T and a binary observation mask M. To generalize across heterogeneous sampling frequencies, T injects a normalized sequential ordering anchored at 0 for the first forecasting step: L T−1 , . . . , 0, . . . , }. (3) L+T L+T Meanwhile, the explicit inclusion of M enables the model to dynamically distinguish genuine observations from missing entries or masked future targets. Subsequently, the augmented sequence is partitioned into P = LL+pT non-overlapping patches of length L p and projected into the hidden dimension D via a residual patch embedding mechanism:
T = {−
H = ResPatchEmbed(Concat(X̂, T , M)) ∈ R M× P× D .
(4)
Unlike standard linear projections, the residual patch embedding is designed to seamlessly integrate linear properties with the complex non-linear semantics extracted from the augmented inputs. 5
3.3
Time Attention
To capture the intrinsic evolutionary patterns of each individual variate, Falcon-X utilizes an encoder-only Transformer architecture (see Appendix A.2). Specifically, the Time Attention module consists of n identical encoder layers, each applied independently along the time dimension D for all M variates. Let H(0) = H be the input from the Tokenization layer. For each layer i ∈ {1, . . . , n}, the hidden state H(i) is computed as follows: H(i) = LayerNorm H(i−1) + MHA H(i−1) , (5) where MHA denotes the multi-head attention mechanism (Vaswani et al., 2017). We apply LayerNorm (Ba et al., 2016) to every layer to stabilize the internal activations and facilitate the training of deep foundation architectures. After n successive transformations, the final output is denoted as H(n) = H T ∈ R M× P× D , encapsulating the comprehensive temporal dynamics of each variate. 3.4
Variate Attention
A core requirement of time-series foundation models is the ability to transcend rigid, datasetspecific dimensional constraints and learn a unified representation of dependencies across multivariat series. However, treating different physical variates homogeneously or simply concatenating them leads to severe semantic misalignment. To overcome this, Falcon-X introduces a unified latent space paradigm. As shown in Figure 2(a–c), these modules progressively aligns the heterogeneous patch embeddings H T into a shared prototype space, models both intra- and cross-dataset dependencies, and dynamically reassembles the global context back to the original variate dimensions. 3.4.1
Unified Prototype Diff-Attention
To overcome the semantic misalignment inherent in raw-space interactions (detailed discussion in Appendix A.3.1), Falcon-X projects the heterogeneous temporal embeddings HT into a fixed-Cdimensional latent prototype space. We employ a Differential Attention mechanism to explicitly capture both positive and negative semantic affinities. Concretely, we define two globally shared learnable parameter matrices, Kpos , Kneg ∈ RC× D , representing the synergistic and antagonistic temporal prototypes, respectively. For each entity ei , given its temporal embeddings hiT ∈ Rmi × P× D , we generate the query Qi and value Vi via linear projections. The dual-dependency attention maps, quantifying the affinity between the mi heterogeneous variates of the i-th entity and the C unified prototypes, are computed as: ! ! i K⊤ i K⊤ Q Q neg pos i i √ √ Apos = softmax , Aneg = softmax , (6) D D i , Ai mi × P×C . The unified representation hi ∈ RC × P× D for entity e is then derived where Apos i neg ∈ R C by aggregating the value features based on the differential attention score:
h i⊤ i i hiC = Apos − λ · Aneg Vi ,
(7)
where λ is a learnable scaling factor and Vi = Linear(hiT ). Notably, while this projection is logically defined at the individual entity level, we implement it concurrently across all N entities to maximize computational efficiency. By employing an entity-aware masking strategy, we completely bypass inefficient explicit loops, seamlessly and parallelly transforming the heterogeneous inputs into 6
a perfectly unified latent space HC ∈ R N ×C× P× D . To ensure semantic distinctiveness, we apply an orthogonality loss Lorth = Sim(Kpos , Kneg ) to constrain the relationship between positive and negative prototypes, where Sim(·, ·) denotes the normalized cosine similarity. 3.4.2
Latent Entity Attention
With all representations HC ∈ R( N ×C)× P× D aligned into a shared, dimension-agnostic semantic space, Latent Entity Attention naturally facilitates cross-learning, leveraging shared structural patterns across entirely different domains (see Appendix A.3.2). To be specific, we treat the combined entity and prototype dimensions as the spatial sequence for interaction, and then apply l layers of the standard MHA mechanism to capture the global cross-variate dependencies: ′ HC = LayerNorm (HC + MHA(HC )) ,
(8)
′ ∈ R N ×C × P× D denotes the refined global context matrix. Similar to the Time Attention, where HC this straightforward yet highly effective operation utilizes residual connections and layer normalization to ensure stable representations. By allowing all aligned entities to interact fully within this ′ successfully captures the holistic dynamics necessary for accurate forecasting. latent space, HC
3.4.3
Variate Reassembly Router
To accurately reconstruct variate-specific trajectories, Falcon-X orchestrates a targeted retrieval from the unified prototype space back to individual physical dimensions (mi ) via a request-and-dispatch mechanism (see Appendix A.3.3). Formally, for each entity ei , we generate the routing components through independent linear projections based on their respective source tensors: i Rreq = Linear(hiT ),
i Pidx = Linear(hC′i ),
i Sctx = Linear(hC′i ).
(9)
i ) acts as an entity identity tag conveying the unique temporal Here, the Routing Request (Rreq trajectory of the original variate. It is matched against the Prototype Index (Piidx ), an addressable map of the global prototype library, to selectively retrieve refined semantic payloads from the Source Context (Sictx ). Rather than performing dense token-level interaction, the reconstruction is then executed via a scaled dot-product soft-routing operation, where each variate dynamically allocates its representation across a compact set of latent prototypes: ! i ( Pi )⊤ Rreq idx i i i i i √ hV = Route(Rreq , Pidx )Sctx = softmax Sctx . (10) D i ∈ Rmi × P× D with high fidelity, This retrieval paradigm successfully reconstructs variate-specific hV smoothly restoring the physical dimensionality to yield HV ∈ R M× P× D . Similar to the initial prototype projection, we apply an entity-aware masking strategy during routing, enabling concurrent soft routing across all N entities without explicit loops.
Finally, to maintain robust cross-dataset performance regardless of varying dependency strengths, we introduce an explicit gated residual connection to dynamically fuse the temporal embeddings H T with the cross-variate representations HV . The final output Ĥ ∈ R M× P× D is computed as: Ĥ = H T + G(H T ) ⊙ HV ,
(11)
where G(·) is a gating mechanism with a linear projection followed by a sigmoid activation, and ⊙ is element-wise multiplication. Thus, Falcon-X effectively prevents semantic interference in weakly correlated systems while making full use of cross-variate dependencies in strongly correlated ones. 7
3.5
Forecasting Head
Following Chronos-2 (Ansari et al., 2025) and Timer-S1 (Liu et al., 2026), Falcon-X adopts a probabilistic forecasting paradigm, predicting future distributions instead of deterministic point estimates. Given the reconstructed representation Ĥ, we extract the embeddings corresponding to the masked future horizon and apply a linear projection to generate forecasts across a predefined set of quantiles Q. The model is end-to-end optimized using the standard Quantile Loss:
Lpred =
M T 1 (q) (q) max q ( Yi, t − Ŷi, t ) , ( q − 1 )( Yi, t − Ŷi, t ) , ∑∑ |Q| · M · T q∑ ∈Q i =1 t=1
(12)
(q)
where Yi,t represents the ground truth and Ŷi,t denotes the model’s prediction at the q-th quantile. The overall training objective combines this with the prototype orthogonality loss: L = Lpred + αLorth , where α is a hyper-parameter balancing the forecasting and orthogonality objectives. During inference, the predictions are mapped back to their original physical scale Ỹ(q) via a straightforward de-normalization process, applying the sine transformation followed by De-Norm using the preserved instance-wise statistics. Ỹ(q) = σ · sin(Ŷ(q) ) + µ.
4
Training Details
4.1
Pre-Training Corpus
(13)
Our pre-training corpus combines large-scale real-world and synthetic time series data spanning several domains. In addition to public datasets from GIFT-E VAL (Aksu et al., 2024) , C HRONOS (Ansari et al., 2024) and Q UITO B ENCH (Xue et al., 2026), it includes synthetic univariate and multivariate series designed to increase diversity in temporal patterns and dependency structures. Specifically, synthetic univariate data are generated through data mixing and stochastic process sampling, while multivariate series are constructed by grouping related univariate signals and injecting explicit cross-variate dependencies, including both instantaneous and temporal interactions. The details can be found in Appendix B.1. 4.2
Training Infrastructure and Config
To enable scalable pretraining on massive, heterogeneous time-series corpora, we build FalconX upon Megatron-LM (Shoeybi et al., 2019) and design a custom sampling pipeline to balance data distribution across diverse domains. Furthermore, to address the heterogeneity in variate dimensionality, we implement a runtime multivariate sampling strategy that dynamically balances the number of consuming variates per batch, thereby improving GPU utilization and training stability. Further implementation details are provided in Appendix C. Falcon-X features a hidden dimension of D = 1024, a patch length of L p = 16, and utilizes n = 16 Time Attention layers alongside l = 16 Entity Attention layers (16 heads per layer). By accommodating up to 512 input tokens and 30 output tokens, it achieves a maximum context length of L = 8192 and a prediction length of T = 480 in a single inference pass. The model is pre-trained on a cluster of NVIDIA B200-180GB GPUs for one million iterations with a global batch size of 384 using bf16 precision, optimized by a joint quantile and orthogonality loss. We adopt the AdamW optimizer (Loshchilov and Hutter, 2019) with β 1 = 0.9, β 2 = 0.95, and a weight decay 8
Figure 3 Performance of Falcon-X on the GIFT-Eval leaderboard. DeOS denotes DeOSAlphaTimeGPTPredictor-2025 and STRIDE denotes STRIDE+Chronos-2.
Figure 4 Performance (MASE) of Falcon-X on the GIFT-Eval leaderboard, grouped by the term length. Falcon-X exhibits remarkable stability across all horizons.
of 0.1. The learning rate warms up linearly to 6 × 10−5 over the first 0.1% of steps, followed by a cosine decay to 6 × 10−6 . As shown in Figure 6, the training process is highly stable, with the loss curve converging smoothly and robustly.
5
Experiments
5.1
Main Results
We evaluate Falcon-X on two comprehensive benchmarks, GIFT-Eval (Aksu et al., 2024) and fevbench (Shchur et al., 2025), using MASE and CRPS to measure point forecasting accuracy and probabilistic calibration, respectively. As shown in Figure 3, Falcon-X achieves the best overall performance on GIFT-Eval, reaching 0.666 MASE and 0.453 CRPS. Compared with the strongest competing time-series foundation models, Falcon-X consistently delivers lower errors: it improves over STRIDE by 1.2% in MASE, over Toto-2.0-FT (Khwaja et al., 2026) by 1.9% in MASE and 2.2% in CRPS, and over Timer-S1 (Liu et al., 2026) by 3.9% in MASE and 6.6% in CRPS. It also surpasses representative multivariate TSFMs such as Chronos-2 (Ansari et al., 2025) and Toto-1.0 (Cohen et al., 2025), demonstrating that explicit latent prototype alignment is more effective than raw-space 9
Figure 5 Performance of Falcon-X on the fev-bench leaderboard.
variate mixing for heterogeneous multivariate forecasting. We further analyze the robustness of Falcon-X across different prediction horizons on GIFT-Eval. Figure 4 reports the MASE results grouped into short-, medium-, and long-term forecasting settings. Falcon-X obtains 0.65 MASE in the short-term setting, tying for the best result with Toto-2.0FT (Khwaja et al., 2026). More importantly, its advantage becomes clearer as the forecasting horizon increases: Falcon-X achieves the best medium-term and long-term results, with 0.68 and 0.70 MASE, respectively. In contrast, Chronos-2 (Ansari et al., 2025) increases from 0.67 in the short-term setting to 0.76 in the long-term setting, while Toto-1.0 (Cohen et al., 2025) and TabPFN-TS (Hoo et al., 2025) degrade more substantially. These results indicate that Falcon-X not only performs well on immediate extrapolation, but also maintains stable predictive accuracy under extended horizons, suggesting that the latent prototype routing mechanism can capture durable cross-variate dynamics and mitigate horizon-wise error accumulation. On fev-bench, Falcon-X also exhibits highly competitive generalization performance, as shown in Figure 5. Falcon-X ranks closely behind Chronos-2 (Ansari et al., 2025), achieving 0.652 MASE and 0.490 CRPS, compared with Chronos-2’s 0.645 MASE and 0.485 CRPS. The gap is only about 1.1% on MASE and 1.0% on CRPS, while Falcon-X relies strictly on endogenous target series rather than additional past-only or future-known covariates. Beyond Chronos-2 (Ansari et al., 2025), Falcon-X substantially outperforms other recent foundation models, including TiRex (Auer et al., 2025), Toto-1.0 (Cohen et al., 2025), Moirai 2.0 (Liu et al., 2025a), and so on. These results confirm that the proposed dual-dependency architecture provides strong relational expressivity and structural adaptability across diverse real-world forecasting tasks. 5.2
Ablation Studies
As shown in Figure 7(a), to isolate the contributions of the architecture and optimization pipeline, we conduct comprehensive ablation studies on both model components and training strategies. Module Ablation. We first evaluate the structural designs of Falcon-X. (i) Only Kpos : Removing the negative prototype key (Kneg ) causes the most severe performance drop, verifying that modeling negative affinities is essential for heterogeneous series interactions. (ii) w/o gated residual: Removing the gated residual connection degrades performance, showing its importance in context-aware filtering cross-variate information in various datasets. (iii) w/o timestamp & mask: Excluding the relative timestamp T and mask M reduces robustness to irregular sampling and missing values. Strategy Ablation. We further analyze the impact of our data processing and training pipeline. (i) w/o sampling shuffle: Disabling variate shuffling significantly hurts performance, indicating
10
Figure 6 Stable training dynamics and performance scaling over one million iterations.
Figure 7 (a) Ablation studies validating the necessity of key architectural components and training strategies. (b) Inference paradigm comparison, highlighting our robust multivariate modeling against Chronos-2.
that random permutation is crucial for learning content-driven rather than index-dependent relationships. (ii) w/o flexible horizon: Replacing flexible horizon sampling with fixed-length prediction weakens generalization across unseen forecasting horizons. (iii) Two-stage vs. Joint training: We compare direct joint training with a two-stage curriculum consisting of: (Stage 1) univariate pre-training for temporal modeling initialization, and (Stage 2) multivariate fine-tuning with univariate replay. Joint training consistently performs better, indicating that Falcon-X can naturally unify temporal dynamics and cross-variate interactions within a single optimization process, without relying on carefully staged training curricula. 5.3
Inference Setting Analysis
We compare inference paradigms against Chronos-2 in Figure 7(b). On GIFT-Eval, Chronos2 (Ansari et al., 2025) exhibits nearly identical performance regardless of whether its group attention is enabled, indicating that raw-space variate mixing contributes little effective relational information. This reveals a severe semantic collapse, where cross-variate interaction degenerates into univariatelike behavior. In contrast, enabling multivariate inference in Falcon-X consistently improves accuracy on multivariate tasks over its univariate inference mode, demonstrating that FalconX can effectively capture transferable cross-variate dependencies. Crucially, this cross-variate enhancement fully preserves the accuracy on univariate forecasting, proving that our latent routing successfully extracts global synergistic context without corrupting individual temporal signals.
11
0.729
0.477
0.471
0.693
0.453 0.666
0.666
0.451
0.453
0.666
Figure 8 Sensitivity and scaling analysis. (a) Performance impact of layer allocation between Time Attention n and Entity Attention l. (b) Sensitivity to the latent prototype dimension C. (c) Consistent performance scaling across increasing model parameter sizes (from 59M to 591M).
5.4
Influence of Key Parameter
We analyze the sensitivity of Falcon-X to two key architectural parameters. Depth Distribution (n vs. l). We investigate the layer allocation between Time Attention (n) and Latent Entity Attention (l) under a fixed depth budget. As shown in Figure 8(a), allocating sufficient capacity to temporal modeling is essential, while excessive cross-variate routing significantly degrades performance. This indicates that robust temporal modeling is the foundation of forecasting, while cross-variate interaction provides complementary gains. Falcon-X achieves the best trade-off with a balanced 16/16 configuration. Latent Prototype Dimension (C). We evaluate the representational capacity of the unified semantic space by varying C. Figure 8(b) demonstrates that a severely restricted dimension (C ≤ 2) induces an information bottleneck, leading to semantic over-compression. Conversely, expanding the dimension enhances relational expressivity. The model achieves peak accuracy at C = 6 and C = 8, striking a perfect balance between capturing diverse dynamics and avoiding redundant noise. 5.5
Scaling Analysis
We evaluate the scaling behaviors of Falcon-X across training iterations and parameter sizes. As shown in Figure 6, the training dynamics exhibit a smooth, stable loss descent alongside continuous forecasting performance gains throughout the entire 1 × 106 steps. Furthermore, scaling the model capacity from 59M (l = n = 8, D = 512) to 253M (l = n = 12, D = 768) and up to 591M (l = n = 16, D = 1024) yields strictly predictable improvements in both MASE and CRPS (Figure 8(c)). These consistent trajectories demonstrate that our decoupled architecture strictly adheres to neural scaling laws, confirming its robust scalability and vast capacity to absorb massive heterogeneous time series without saturation. 5.6
Case Study
To qualitatively analyze Falcon-X, we present representative forecasting cases from GIFT-Eval. Multivariate vs. Univariate Inference. We compare multivariate inference with channel-independent inference on highly correlated sequences from ETT1/15T (see Figure 9). Without cross-variate interactions, the channel-independent setting gradually drifts from the ground truth under complex temporal shifts. In contrast, Falcon-X leverages its unified latent prototype space to aggregate complementary signals across variates, producing substantially more accurate trajectories.
12
Compare
multi 😊
uni ☹
multi 😊
uni ☹
Figure 9 Case study on ETT1/15T dataset. In univariate inference mode, forecasts gradually deviate from the ground truth due to the absence of global context. In contrast, Falcon-X’s multivariate inference effectively leverages cross-variate signals to calibrate trajectories.
Positive and Negative Dependency Modeling. We further examine the ability of Falcon-X to capture dual dependencies. As shown in Figure 13, Falcon-X accurately models positive correlations with synchronized trends. More importantly, in Figure 14, it successfully captures negative correlations with opposing dynamics. These results validate the effectiveness of our Unified Prototype Diff-Attention for modeling both synergistic and antagonistic relationships.
6
Conclusion
In this paper, we identify two fundamental challenges of existing multivariate time series foundation models: semantic alignment and relational expressivity. To address these issues, we propose Falcon-X, a novel modeling paradigm that decouples physical variables into a unified latent prototype space. By introducing the Unified Prototype Diff-Attention, our architecture effectively captures both synergistic and antagonistic correlations. Additionally, a Variate Reassembly Router ensures robust global context fusion across diverse domains. Extensive evaluations on GIFT-Eval and fev-bench demonstrate that Falcon-X achieves state-of-the-art performance, showcasing exceptional scalability and zero-shot transferability. We hope that this work will contribute to the development of more unified and expressive foundation models for time series.
13
References Walmart Competition Admin and Will Cukierski. Walmart recruiting - store sales forecasting. https://kaggle.com/ competitions/walmart-recruiting-store-sales-forecasting, 2014. Kaggle. Taha Aksu, Gerald Woo, Juncheng Liu, Xu Liu, Chenghao Liu, Silvio Savarese, Caiming Xiong, and Doyen Sahoo. Gift-EVAL: A benchmark for general time series forecasting model evaluation. arXiv preprint arXiv:2410.10393, 2024. Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, et al. Chronos: Learning the language of time series. Transactions on Machine Learning Research, 2024. Abdul Fatir Ansari, Oleksandr Shchur, Jaris Küken, Andreas Auer, Boran Han, Pedro Mercado, Syama Sundar Rangapuram, Huibin Shen, Lorenzo Stella, Xiyuan Zhang, Mononito Goswami, Shubham Kapoor, Danielle C. Maddix, Pablo Guerron, Tony Hu, Junming Yin, Nick Erickson, Prateek Mutalik Desai, Hao Wang, Huzefa Rangwala, George Karypis, Yuyang Wang, and Michael Bohlke-Schneider. Chronos-2: From univariate to universal forecasting. arXiv preprint arXiv:2510.15821, 2025. George Athanasopoulos, Roman A. Ahmed, and Rob J. Hyndman. Hierarchical forecasts for Australian domestic tourism. International Journal of Forecasting, 25(1):146–166, January 2009. ISSN 0169-2070. doi: 10.1016/j.ijforecast.2008.07.004. http://dx.doi.org/10.1016/j.ijforecast.2008.07.004. Andreas Auer, Patrick Podest, Daniel Klotz, Sebastian Böck, Günter Klambauer, and Sepp Hochreiter. Tirex: Zero-shot forecasting across long and short horizons with enhanced in-context learning. In Neural Information Processing Systems, 2025. Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. In arXiv preprint arXiv:1607.06450, 2016. Lawrence J. Christiano, Martin Eichenbaum, and Charles L. Evans. Monetary policy shocks: What have we learned and to what end? In Handbook of Macroeconomics, volume 1 of Handbook of Macroeconomics, pages 65–148. Elsevier, 1999. doi: https://doi.org/10.1016/S1574-0048(99)01005-8. https://www.sciencedirect.com/science/article/pii/ S1574004899010058. Ben Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi, Chris Lettieri, Charles Masson, Hugo Miccinilli, Elise Ramé, Qiqi Ren, Afshin Rostamizadeh, et al. This time is different: An observability perspective on time series foundation models. In Neural Information Processing Systems, 2025. Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. A decoder-only foundation model for time-series forecasting. In International Conference on Machine Learning, pages 10148–10167, 2024. Open Power System Data. Data package time series. version 2020-10-06, 2020. https://doi.org/10.25832/time_series/ 2020-10-06. UK COVID-19 data from official UK government sources. UK COVID-19 dashboard data. https://www.kaggle.com/ datasets/happyadam73/uk-covid19-dashboard-data-sqlite-compressed, 2022. Kaggle. Etienne David, Jean Bellot, and Sylvain Le Corff. HERMES: Hybrid error-corrector model with inclusion of external signals for nonstationary fashion time series. arXiv preprint arXiv:2202.03224, 2022. S. De Vito, E. Massera, M. Piga, L. Martinotto, and G. Di Francia. On field calibration of an electronic nose for benzene estimation in an urban pollution monitoring scenario. Sensors and Actuators B: Chemical, 129(2):750–757, 2008. ISSN 0925-4005. doi: https://doi.org/10.1016/j.snb.2007.09.060. https://www.sciencedirect.com/science/article/pii/ S0925400507007691. ECDC. Respiratory viruses weekly data. https://github.com/EU-ECDC/Respiratory_viruses_weekly_data/tree/main, 2025. Open data repository; weekly respiratory virus surveillance in the EU/EEA. Philip J Fleming and John J Wallace. How not to lie with statistics: the correct way to summarize benchmark results. Communications of the ACM, 29(3):218–221, 1986. FlorianKnauer and Will Cukierski. Rossmann store sales. https://kaggle.com/competitions/rossmann-store-sales, 2015. Kaggle. Rakshitha Wathsadini Godahewa, Christoph Bergmeir, Geoffrey I. Webb, Rob Hyndman, and Pablo Montero-Manso. Monash time series forecasting archive. In The Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021. https://openreview.net/forum?id=wEc1mgAjU-.
14
Mononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai, Shuo Li, and Artur Dubrawski. MOMENT: A family of open time-series foundation models. In International Conference on Machine Learning, 2024. Tao Hong, Pierre Pinson, and Shu Fan. Global energy forecasting competition 2012. International Journal of Forecasting, 30 (2):357–363, 2014. Shi Bin Hoo, Samuel Müller, David Salinas, and Frank Hutter. From tables to time: Extending tabpfn-v2 to time series forecasting. arXiv preprint arXiv:2501.02945, 2025. Addison Howard, Haruka Yui, Mark McDonald, and Will Cukierski. Recruit restaurant visitor forecasting. https: //www.kaggle.com/c/recruit-restaurant-visitor-forecasting, 2017a. Kaggle. Addison Howard, Haruka Yui, Mark McDonald, and Will Cukierski. Recruit restaurant visitor forecasting. https: //kaggle.com/competitions/recruit-restaurant-visitor-forecasting, 2017b. Kaggle. Jiawei Jiang, Chengkai Han, Wenjun Jiang, Wayne Xin Zhao, and Jingyuan Wang. Libcity: A unified library towards efficient and comprehensive urban spatial-temporal prediction. arXiv preprint arXiv:2304.14343, 2023. Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, et al. Time-LLM: Time series forecasting by reprogramming large language models. In International Conference on Learning Representations, 2023. Emaad Khwaja, Chris Lettieri, Gerald Woo, Eden Belouadah, Marc Cenac, Guillaume Jarry, Enguerrand Paquin, Xunyi Zhao, Viktoriya Zhukov, Othmane Abou-Amal, et al. Toto 2.0: Time series forecasting enters the scaling era. arXiv preprint arXiv:2605.20119, 2026. Xiangjie Kong, Zhenghao Chen, Weiyao Liu, Kaili Ning, Lechao Zhang, Syauqie Muhammad Marier, Yichen Liu, Yuhao Chen, and Feng Xia. Deep learning for time series forecasting: a survey. International Journal of Machine Learning and Cybernetics, 16(7):5079–5112, 2025. Siva Rama Krishna Kottapalli, Karthik Hubli, Sandeep Chandrashekhara, Garima Jain, Sunayana Hubli, Gayathri Botla, and Ramesh Doddaiah. Foundation models for time series: A survey. arXiv preprint arXiv:2504.04011, 2025. Guokun Lai, Wei-Cheng Chang, Yiming Yang, and Hanxiao Liu. Modeling long- and short-term temporal patterns with deep neural networks. In The International ACM SIGIR Conference on Research & Development in Information Retrieval, 2017. https://api.semanticscholar.org/CorpusID:4922476. lexis Cook, DanB, inversion, and Ryan Holbrook. Store sales – time series forecasting. https://www.kaggle.com/ competitions/store-sales-time-series-forecasting, 2020. Kaggle. Chenghao Liu, Taha Aksu, Juncheng Liu, Xu Liu, Hanshu Yan, Quang Pham, Silvio Savarese, Doyen Sahoo, Caiming Xiong, and Junnan Li. Moirai 2.0: When less is more for time series forecasting. arXiv preprint arXiv:2511.11698, 2025a. Yong Liu, Haoran Zhang, Chenyu Li, Xiangdong Huang, Jianmin Wang, and Mingsheng Long. Timer: generative pretrained transformers are large time series models. In International Conference on Machine Learning, pages 32369–32399, 2024. Yong Liu, Guo Qin, Zhiyuan Shi, Zhi Chen, Caiyin Yang, Xiangdong Huang, Jianmin Wang, and Mingsheng Long. Sundial: A family of highly capable time series foundation models. In International Conference on Machine Learning, pages 39295–39317. PMLR, 2025b. Yong Liu, Xingjian Su, Shiyu Wang, Haoran Zhang, Haixuan Liu, Yuxuan Wang, Zhou Ye, Yang Xiang, Jianmin Wang, and Mingsheng Long. Timer-S1: A billion-scale time series foundation model with serial scaling. arXiv preprint arXiv:2603.04791, 2026. Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. https://openreview.net/forum?id=Bkg6RiCqY7. Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos. The M4 competition: Results, findings, conclusion and way forward. International Journal of Forecasting, 2018. Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos. M5 accuracy competition: Results, findings, and conclusions. International Journal of Forecasting, 38(4):1346–1364, 2022. ISSN 0169-2070. doi: https://doi.org/10.1016/j. ijforecast.2021.11.013. https://www.sciencedirect.com/science/article/pii/S0169207021001874. Special Issue: M5 competition. Paolo Mancuso, Veronica Piccialli, and Antonio M Sudoso. A machine learning approach for forecasting hierarchical time series. Expert Systems with Applications, 182:115102, 2021.
15
AI Maverick. Renewable energy and weather conditions. renewable-energy-and-weather-conditions, 2025. Kaggle.
https://www.kaggle.com/datasets/samanemami/
Michael W. McCracken and Serena Ng. FRED-MD: A monthly database for macroeconomic research. Journal of Business & Economic Statistics, 34(4):574–589, 2016. doi: 10.1080/07350015.2015.1086655. https://doi.org/10.1080/07350015. 2015.1086655. Michael W. McCracken and Serena Ng. FRED-QD: A quarterly database for macroeconomic research. Review, 103(1): 1–44, January 2021. doi: 10.20955/r.103.1-44. https://ideas.repec.org/a/fip/fedlrv/90588.html. MichalKecera. Rohlik sales forecasting challenge. rohlik-sales-forecasting-challenge-v2, 2024. Kaggle.
https://kaggle.com/competitions/
Kamiar Mohaddes and Mehdi Raissi. Compilation, revision and updating of the global var (gvar) database. Mendeley Data, Version 1, 2024. https://doi.org/10.17632/kfp5fhgkvf.1. Yuqi Nie, Nam H Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. A time series is worth 64 words: Long-term forecasting with transformers. In International Conference on Learning Representations, 2023. Nafay Un Noor. Global life expectancy data (1950–2023). global-life-expectancy-data-1950-2023, 2025. Kaggle.
https://www.kaggle.com/datasets/nafayunnoor/
General Directorate of Health Affairs and Saudi Arabia Ministry of Health. Riyadh hospital admissions dataset (2020–2024). https://www.kaggle.com/dsv/9992619, 2024. Santosh Palaskar, Vijay Ekambaram, Arindam Jati, Neelamadhav Gantayat, Avirup Saha, Seema Nagar, Nam Nguyen, Pankaj Dayama, Renuka Sindhgatta, Prateeti Mohapatra, Harshit Kumar, Jayant Kalagnanam, Nandyala Hemachandra, and Narayan Rangaraj. Automixer for improved multivariate time-series forecasting on business and it observability data. Proceedings of the AAAI Conference on Artificial Intelligence, 38:22962–22968, 2024. Ulrik Thyge Pedersen. CO2 emissions by country. co2-emissions-by-country, 2025. Kaggle. Bushra Qurban. Tourism and economic tourism-and-economic-impact, 2025. Kaggle.
https://www.kaggle.com/datasets/ulrikthygepedersen/
impact.
https://www.kaggle.com/datasets/bushraqurban/
Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen, Lorenzo Stella, Nick Erickson, Pablo Guerron, Michael Bohlke-Schneider, and Yuyang Wang. fev-bench: A realistic benchmark for time series forecasting. arXiv preprint arXiv:2509.26468, 2025. Siqi Shen, Vincent Van Beek, and Alexandru Iosup. Statistical characterization of business-critical workloads hosted in cloud datacenters. In IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, pages 465–474. IEEE, 2015. Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019. Iain Staffell, Stefan Pfenninger, and Nathan Johnson. A global model of hourly space heating and cooling demand at multiple spatial scales. Nature Energy, 8(12):1328–1344, 2023. doi: 10.1038/s41560-023-01341-5. https://doi.org/10. 1038/s41560-023-01341-5. Artur Trindade. ElectricityLoadDiagrams20112014. https://doi.org/10.24432/C58C86.
UCI Machine Learning Repository, 2015.
DOI:
Alexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya, Wenjian Dong, Murali Narayanaswamy, Zhengchun Liu, Gaurav Saxena, Andreas Kipf, and Tim Kraska. Why TPC is not enough: An analysis of the amazon redshift fleet. Proc. VLDB Endow., 17(11):3694–3706, July 2024. ISSN 2150-8097. doi: 10.14778/3681954.3682031. https: //doi.org/10.14778/3681954.3682031. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. Jingyuan Wang, Jiawei Jiang, Wenjun Jiang, Chengkai Han, and Wayne Xin Zhao. Towards efficient and comprehensive urban spatial-temporal prediction: A unified library and performance benchmark. arXiv preprint arXiv:2304.14343, 2023.
16
Ines Wilms and Christophe Croux. Forecasting using sparse cointegration. International Journal of Forecasting, 32(4): 1256–1267, 2016. ISSN 0169-2070. doi: https://doi.org/10.1016/j.ijforecast.2016.04.005. https://www.sciencedirect. com/science/article/pii/S0169207016300589. Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, and Doyen Sahoo. Unified training of universal time series forecasting transformers. In International Conference on Machine Learning, 2024. Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In Neural Information Processing Systems, 2021. https://api. semanticscholar.org/CorpusID:235623791. Shi Xiaoming, Wang Shiyu, Nie Yuqi, Li Dianqi, Ye Zhou, Wen Qingsong, and Ming Jin. Time-MoE: Billion-scale time series foundation models with mixture of experts. In International Conference on Learning Representations, 2025. Jiyuan Xu, Wenyu Zhang, Xin Jing, Jiahao Nie, Shuai Chen, and Shuai Zhang. CPiRi: Channel permutation-invariant relational interaction for multivariate time series forecasting. In The International Conference on Learning Representations, 2026. https://openreview.net/forum?id=tgnXCCjKE3. Siqiao Xue, Zhaoyang Zhu, Wei Zhang, Rongyao Cai, Rui Wang, Yixiang Mu, Fan Zhou, Jianguo Li, Peng Di, and Hang Yu. QuitoBench: A high-quality open time series forecasting benchmark. arXiv preprint arXiv:2603.26017, 2026. Tianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun, Yi Zhu, Gao Huang, and Furu Wei. Differential transformer. In International Conference on Learning Representations, 2025. Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pages 11106–11115, 2021. Jingbo Zhou, Xinjiang Lu, Yixiong Xiao, Jian Tang, Jiantao Su, Yu Li, Ji Liu, Junfu Lyu, Yanjun Ma, and Dejing Dou. SDWPF: A dataset for spatial dynamic wind power forecasting over a large turbine array. Scientific Data, 11(1):649, 2024. doi: 10.1038/s41597-024-03427-5. https://doi.org/10.1038/s41597-024-03427-5.
17
A
Methodology Design Philosophies
A.1
Overall Architecture
As shown in Figure 2, the architecture of Falcon-X follows a hierarchical transformation pipeline that progressively aligns heterogeneous variates into a unified latent space, and ultimately reconstructs them for accurate forecasting. We formulate the forecasting task as a unified masked reconstruction paradigm. Formally, the input heterogeneous series X ∈ R M×( L+T ) is first subjected to instance normalization and tokenization to generate the time tokens H ∈ R M× P× D , where P = LL+pT is the number of patches and L p is the length of patches. These tokens are then processed through time attention layers, yielding temporal representations H T ∈ R M× P× D . To bridge the dimensionality gap, Falcon-X employs the Unified Prototype Diff-Attention (UPDA) to project the disparate N entities into a compact, unified prototype space HC ∈ R( N × P)×C× D . Following cross-variate interactions via Latent Entity Attention (LEA), which yields the refined ′ , the Variate Reassembly Router (VRR) performs a soft-routing operation to retrieve context HC and reassemble the latent representations back into the entity-specific space HV ∈ R M× P× D by ′ . matching the routing request in H T with the prototype index in HC Finally, the reassembled entities HV are fused with the temporal representations H T and mapped through a quantile forecasting head to produce the final predictive output Ŷ ∈ R M×T , completing the end-to-end flow from raw heterogeneous inputs to unified latent representations and back to structured forecasts. A.2
Time Attention
To capture the intrinsic evolutionary patterns of each individual variate, Falcon-X utilizes an encoder-only Transformer architecture. A critical design choice in Falcon-X is the deliberate decoupling of temporal and cross-variate modeling. Unlike many existing multivariate models, which interleave temporal and spatial mixing, leading to semantic entanglement whereby a variate’s subtle temporal signal is prematurely confounded by the noisy dependencies of its heterogeneous neighbors, Falcon-X prioritizes establishing a robust temporal module. Stacking n layers of Time Attention before any cross-variate interaction ensures that the temporal evolutive state of each variate is fully distilled and stabilised. This provides a clean, time-aware foundation for the subsequent Variate Attention process. A.3
Variate Attention
A core requirement of time-series foundation models is the ability to transcend rigid, datasetspecific dimensional constraints and learn a unified representation of dependencies across multivariat series. However, treating different physical variates homogeneously or simply concatenating them leads to severe semantic misalignment. To overcome this, Falcon-X introduces a unified latent space paradigm. As shown in Figure 2(a–c), these modules progressively aligns the heterogeneous patch embeddings H T into a shared prototype space, models both intra- and cross-dataset dependencies, and dynamically reassembles the global context back to the original variate dimensions. A.3.1
Unified Prototype Diff-Attention
The primary challenge in modeling the foundations of time series lies in reconciling diverse physical variates within a unified semantic space. Dataset-level group mixing (Ansari et al., 2025) relies on
18
dense intra-dataset attention, lacking structural abstraction and incurring quadratic complexity O( M2 ). To address this, Falcon-X introduces prototype alignment, projecting heterogeneous variates into a fixed set of learnable latent temporal prototypes of dimension C. Instead of relying on arbitrary physical indexing, this paradigm dynamically allocates the full representation of each independent variate across the C universal prototypes based on their intrinsic semantic affinity, thereby achieving explicit semantic unification. This explicit mapping resolves the issue of semantic misalignment by aligning variates with similar temporal dynamics to the same semantic anchors, regardless of their original dataset or spatial proximity. Furthermore, by forcing the representations through this fixed-dimensional prototype space, the model inherently performs structural denoising, filtering out localized noise and isolating the most salient temporal patterns. Importantly, this projection replaces dense intra-dataset attention with cross-attention, decoupling the computational bottleneck from the physical dimensionality. The complexity is thus reduced to a strictly linear O( M · C ), effortlessly accommodating extreme-dimensional modeling since C ≪ M. Furthermore, we observe that negative correlations among heterogeneous variates, which are critical for capturing counteracting interactions across diverse time series, are difficult to exploit in standard Transformer architectures. This limitation stems from the non-negative nature of the softmax attention function, which restricts attention scores to the range [0, 1] and thus prevents the explicit modeling of opposing trends. To address this, we draw inspiration from differential attention (Ye et al., 2025). While originally proposed for attention noise suppression, FalconX repurposes this mechanism to capture dual-dependency dynamics. By introducing positive and negative learnable keys, the model explicitly represents both synergistic and antagonistic relationships, yielding improved expressiveness of the cross-variate latent space. A.3.2
Latent Entity Attention
Following the alignment of heterogeneous variates into the unified prototype space, this module models the comprehensive interactions among different variates. As all representations now reside in a shared, dimension-agnostic semantic space rather than their original disparate physical dimensions, Latent Entity Attention naturally facilitates cross-learning. This enables Falcon-X to leverage and transfer shared structural patterns across entirely different domains, thereby significantly enhancing zero-shot cross-dataset generalization. A.3.3
Variate Reassembly Router
After capturing comprehensive dependencies in the unified prototype space, the model must reassemble this global context back into the original heterogeneous dimensions (mi ). Falcon-X formulates this reassembly as a targeted retrieval from the abstract prototype space to individual variate trajectories. The aim is to reconstruct the heterogeneous variates, each with distinct temporal patterns, by retrieving relevant information from the unified latent prototypes. This is orchestrated via a request-and-dispatch mechanism: the Routing Request (Rreq ), derived from hiT , acts as a structural query conveying the specific physical dimensionality and unique temporal trajectory of the original variate, effectively serving as entity identity tag. The request is then matched against the Prototype Index (Pidx ), which is an addressable map of the global prototype library. Meanwhile, the Source Context (Sctx ) delivers the refined semantic payloads. Using local, specific trajectories to selectively retrieve unified global prototypes enables this routing paradigm to reconstruct variate-specific patterns with high fidelity. 19
Finally, considering the significant variance in cross-variate dependencies across diverse datasets, forcing a uniform integration of the global context could introduce detrimental noise to datasets with inherently weak variate correlations. To maintain strict cross-dataset robustness, we introduce an explicit gated residual connection to dynamically fuse the temporal embeddings H T with the cross-variate representations HV . Consequently, it effectively prevents semantic interference in weakly correlated systems while making full use of cross-variate dependencies in strongly correlated ones.
B
Dataset Statistics
This section summarizes the datasets used in our experiments. Specifically, Appendix B.1 describes the corpus used for model pre-training, while Appendix B.2 and Appendix B.3 present the benchmarks used for downstream evaluation. B.1
Pre-training Corpus
Our pre-training corpus comprises both real-world and synthetic time series datasets, covering a broad range of domains and data-generation characteristics. Real-world datasets. We aggregate several large-scale time series collections, including the GIFTE VAL (Aksu et al., 2024) pre-training dataset1 , the C HRONOS (Ansari et al., 2024) training corpus2 , and the Q UITO B ENCH (Xue et al., 2026) training dataset3 , as detailed in Table 2. Collectively, these resources span seven major domains: nature, energy, transport, finance, healthcare, web, and sales. The resulting corpus contains a large number of univariate time series datasets, as well as a small set of multivariate time series datasets. Synthetic univariate datasets. We also incorporate the synthetic univariate datasets introduced by C HRONOS (Ansari et al., 2024): TSM IXUP and K ERNEL S YNTH. TSM IXUP synthesizes new time series by taking random convex combinations of samples drawn from different real-world datasets, thereby increasing diversity while preserving realistic temporal characteristics. K ERNEL S YNTH, in contrast, generates synthetic series by randomly composing Gaussian Process (GP) kernels and sampling from the resulting GP priors, producing time series with diverse trends, periodicities, and stochastic patterns. Synthetic multivariate datasets. High-quality multivariate time series datasets remain relatively scarce in existing public resources. To address this limitation, we construct a large amount of synthetic multivariate data through two complementary strategies: 1. Similarity-based multivariate construction from real univariate series. Drawing from the real univariate time series presented in Table 2, we compute pairwise similarities to group related sequences into cohesive multivariate datasets. This process enables us to derive multivariate structures from naturally occurring signals while preserving semantic coherence across dimensions. 2. Dependency injection over synthetic univariate generators. Inspired by Chronos-2 (Ansari et al., 2025), we transform multiple independently sampled univariate series from base generators (e.g., K ERNEL S YNTH) into multivariate synthetic time series by imposing explicit dependency structures. These multivariatization procedures include: 1 https://huggingface.co/datasets/Salesforce/GiftEvalPretrain 2 https://huggingface.co/datasets/autogluon/chronos_datasets 3 https://huggingface.co/datasets/hq-bench/quito-corpus
20
Table 2 Summary statistics of univariate pre-training datasets. Dataset Name Frequency Time Series Variates Time Points Domain Source BDG-2 H 611 1 9,454,968 Energy GIFT-Eval BEIJING_SUBWAY_30MIN 30T 276 2 433,872 Transport GIFT-Eval CIF 2016 M 72 1 6,334 Finance GIFT-Eval CMIP6 6H 270,336 53 1,973,452,800 Nature GIFT-Eval ERA5 H 245,760 45 2,146,959,360 Nature GIFT-Eval Electricity H, W 642 1 8,493,660 Energy Chronos HZMETRO 15T 80 2 190,160 Transport GIFT-Eval LOS_LOOP 5T 207 1 7,094,304 Transport GIFT-Eval LargeST 5T 42,333 1 4,452,510,528 Transport GIFT-Eval M1 A, M, Q 921 1 57,882 Finance GIFT-Eval M3 A, M, Q 3,003 1 209,114 Finance GIFT-Eval NN5 D, W 222 1 93,240 Finance GIFT-Eval PEMS03 5T 358 1 9,382,464 Transport GIFT-Eval PEMS04 5T 307 3 5,216,544 Transport GIFT-Eval PEMS07 5T 883 1 24,921,792 Transport GIFT-Eval PEMS08 5T 170 3 3,035,520 Transport GIFT-Eval PEMS_BAY 5T 325 1 16,941,600 Transport GIFT-Eval Q-TRAFFIC 15T 45,148 1 264,386,688 Transport GIFT-Eval Quito 10T, H 33,806 5 313,269,828 Various QuitoBench Residential Power T 504 3 271,333,509 Energy GIFT-Eval SHMETRO 15T 288 2 2,536,992 Transport GIFT-Eval Solar 5T, H 10,332 1 588,304,080 Energy Chronos Taxi 30T, H 70,412 1 56,793,348 Transport Chronos Tourism A, M, Q 1,212 1 150,822 Finance GIFT-Eval Traffic H, W 1,724 1 15,060,864 Transport GIFT-Eval Uber TLC D, H 524 1 1,176,531 Transport GIFT-Eval Weatherbench D, H, W 675,840 1 82,753,646,592 Nature Chronos Wind Farms D, H, T 1,011 1 175,154,333 Energy Chronos alibaba_cluster_trace_2018 5T 58,409 2 95,192,530 Web GIFT-Eval australian_electricity_demand 30T 5 1 1,153,584 Energy GIFT-Eval azure_vm_traces_2017 5T 159,472 1 885,522,908 Web GIFT-Eval beijing_air_quality H 12 11 420,768 Nature GIFT-Eval bitcoin_with_missing D 18 1 81,918 Finance GIFT-Eval borealis H 15 1 83,269 Energy GIFT-Eval borg_cluster_data_2011 5T 143,386 2 537,552,854 Web GIFT-Eval buildings_900k H 1,792,328 1 15,702,585,608 Energy GIFT-Eval bull H 41 1 719,304 Energy GIFT-Eval cdc_fluview_ilinet W 75 5 63,903 Healthcare GIFT-Eval cdc_fluview_who_nrevss W 74 4 41,760 Healthcare GIFT-Eval china_air_quality H 437 6 5,739,234 Nature GIFT-Eval cockatoo H 1 1 17,544 Energy GIFT-Eval covid19_energy H 1 1 31,912 Energy GIFT-Eval covid_mobility D 362 1 148,602 Transport GIFT-Eval dominick W 100,014 1 29,652,492 Sales Chronos elecdemand 30T 1 1 17,520 Energy GIFT-Eval elf H 1 1 21,792 Energy GIFT-Eval exchange_rate D 8 1 84,976 Finance Chronos extended_web_traffic_with_missing D 145,063 1 370,926,091 Web GIFT-Eval godaddy M 3,135 2 128,535 Finance GIFT-Eval hog H 24 1 421,056 Energy GIFT-Eval ideal H 217 1 1,255,253 Energy GIFT-Eval kaggle_web_traffic_weekly W 145,063 1 16,537,182 Web GIFT-Eval lcl H 713 1 9,543,553 Energy GIFT-Eval london_smart_meters_with_missing 30T 5,520 1 166,238,880 Energy GIFT-Eval mexico_city_bikes H 494 1 38,687,004 Transport Chronos oikolab_weather H 8 1 800,456 Nature GIFT-Eval pdb H 1 1 17,520 Energy GIFT-Eval pedestrian_counts H 66 1 3,130,762 Transport GIFT-Eval project_tycho W 1,258 1 1,377,707 Healthcare GIFT-Eval rideshare_with_missing H 2,304 1 859,392 Transport GIFT-Eval sceaux H 1 1 34,223 Energy GIFT-Eval smart H 5 1 95,709 Energy GIFT-Eval solar_power 4S 1 1 7,397,222 Energy GIFT-Eval spain H 1 1 35,064 Energy GIFT-Eval subseasonal D 862 4 14,197,140 Nature GIFT-Eval subseasonal_precip D 862 1 9,760,426 Nature GIFT-Eval sunspot_with_missing D 1 1 73,894 Nature GIFT-Eval ushcn_daily D 1,218 5 47,080,115 Nature Chronos vehicle_trips_with_missing D 329 1 32,512 Transport GIFT-Eval weather D 3,010 1 42,941,700 Nature GIFT-Eval wiki-rolling_nips D 47,675 1 40,619,100 Web GIFT-Eval wiki_daily_100k D 100,000 1 274,100,000 Web Chronos wind_power 4S 1 1 7,397,147 Energy GIFT-Eval Note. Frequency aliases follow common time-series conventions: S = second, T = minute, H = hourly, D = daily, W = weekly, M = monthly, Q = quarterly, and A = annual.
21
• Cotemporaneous multivariatizers, which introduce instantaneous cross-variate dependencies through linear or nonlinear transformations at the same time step; • Sequential multivariatizers, which impose temporal cross-series relations across time, such as lead–lag dependencies and cointegration. Through this combination of real-world corpora, univariate synthetic generators, and large-scale multivariate synthesis, our dataset collection supports training and evaluation across diverse domains and temporal dependency structures. B.2
GIFT-Eval Benchmark
GIFT-Eval is constructed from 15 univariate and 8 multivariate datasets, spanning 7 domains and 10 frequencies. In total, the benchmark contains 144,000 time series and 177 million observations. To support evaluation across forecasting horizons, prediction lengths are determined in two ways. For widely used benchmarks such as M4 (Makridakis et al., 2018), established prediction lengths are retained. For the remaining datasets, the short-horizon prediction length is set to 48 time steps, and the medium- and long-horizon settings are defined according to dataset frequency and domain as 10× and 15× the short-horizon length, respectively. This results in 97 unique combinations of dataset, frequency, and prediction length, with model performance reported as the geometric mean across these configurations. The benchmark is curated from 10 publicly available sources covering a diverse set of application domains. Below, the included datasets are grouped by domain and described together with their original sources. • Nature. The benchmark includes the Jena Weather dataset4 , following the preprocessing protocol used in Autoformer (Wu et al., 2021). • Web/CloudOps. This domain contains the BizITObs Application, Service, and L2C datasets5 , processed according to the pipeline introduced in AutoMixer (Palaskar et al., 2024). These datasets combine business KPIs with IT event channels, forming multivariate time series for observability-related forecasting tasks. In addition, Bitbrains datasets from the Grid Workloads Archive (Shen et al., 2015) are included in the same domain. • Sales. For the sales domain, the Restaurant dataset is adopted from the Recruit Restaurant Forecasting Competition (Howard et al., 2017b), where the objective is to predict future customer visits using reservation and visitation records. Another sales dataset is included from Mancuso et al. (2021). • Energy. The energy domain includes ETT1 and ETT2 from Informer (Zhou et al., 2021), which represent electricity transformer temperature and are widely used in long-horizon forecasting. It also includes the Electricity dataset from the UCI ML Archive (Trindade, 2015), containing electricity consumption records for 370 clients, and the Solar dataset from LSTNet (Lai et al., 2017), which focuses on forecasting solar plant power output. • Transport. Transport datasets are drawn from LibCity (Wang et al., 2023), a benchmark collection of urban spatio-temporal and time series datasets. • Econ/Fin & Healthcare. A subset of datasets is selected from the Monash repository (Godahewa et al., 2021), which provides a broad collection of time series from multiple domains. 4 https://www.bgc-jena.mpg.de/wetter/ 5 https://github.com/BizITObs/BizITObservabilityData/tree/main
22
Table 3 Statistics of the GIFT-Eval benchmark across seven domains. Entries under Short-term/Med-term/Long-term are reported as Pred/Win, denoting the prediction length and the number of rolling windows, respectively.
Nature
Domain
Dataset Name
Freq.
#Series
Avg. length
#Vars
Short-term
Med-term
Long-term
Jena Weather
10T H D D W-THU M D H D
1 1 1 1 1 1 32,072 270 270
52,704 8,784 366 23,741 3,391 780 725 10,898 455
21 21 21 1 1 1 1 1 1
48 / 20 48 / 19 30 / 2 30 / 20 8 / 20 12 / 7 30 / 3 48 / 20 30 / 2
480 / 11 480 / 2 – – – – – 480 / 2 –
720 / 8 720 / 2 – – – – – 720 / 2 –
10S 10S 5T H 5T H 5T H
1 21 1 1 1,250 1,250 500 500
8,834 8,835 31,968 2,664 8,640 721 8,640 720
2 2 7 7 2 2 2 2
60 / 15 60 / 15 48 / 20 48 / 6 48 / 18 48 / 2 48 / 18 48 / 2
600 / 2 600 / 2 480 / 7 480 / 1 480 / 2 – 480 / 2 –
900 / 1 900 / 1 720 / 5 720 / 1 720 / 2 – 720 / 2 –
15T H D W-THU 15T H D W-THU 10T H D W-FRI 15T H D W-FRI
1 1 1 1 1 1 1 1 137 137 137 137 370 370 370 370
69,680 17,420 725 103 69,680 17,420 725 103 52,560 8,760 365 52 140,256 35,064 1,461 208
7 7 7 7 7 7 7 7 1 1 1 1 1 1 1 1
48 / 20 48 / 20 30 / 3 8/2 48 / 20 48 / 20 30 / 3 8/2 48 / 20 48 / 19 30 / 2 8/1 48 / 20 48 / 20 30 / 5 8/3
480 / 15 480 / 4 – – 480 / 15 480 / 4 – – 480 / 11 480 / 2 – – 480 / 20 480 / 8 – –
720 / 10 720 / 3 – – 720 / 10 720 / 3 – – 720 / 8 720 / 2 – – 720 / 20 720 / 5 – –
5T H D 15T H H D
323 323 323 156 156 30 30
105,120 8,760 365 2,976 744 17,520 730
1 1 1 1 1 1 1
48 / 20 48 / 19 30 / 2 48 / 7 48 / 2 48 / 20 30 / 3
480 / 20 480 / 2 – 480 / 1 – 480 / 4 –
720 / 15 720 / 2 – 720 / 1 – 720 / 3 –
D D W-WED M
807 118 118 2,674
358 1,825 260 51
1 1 1 1
30 / 1 30 / 7 8/4 12 / 1
– – – –
– – – –
M4 Yearly M4 Quarterly M4 Monthly M4 Weekly M4 Daily M4 Hourly
A Q M W D H
22,974 24,000 48,000 359 4,227 414
37 100 234 1,035 2,371 902
1 1 1 1 1 1
6/1 8/1 18 / 1 13 / 1 14 / 1 48 / 2
– – – – – –
– – – – – –
Hospital COVID Deaths US Births
M D D W-TUE M
767 266 1 1 1
84 212 7,305 1,043 240
1 1 1 1 1
12 / 1 30 / 1 30 / 20 8 / 14 12 / 2
– – – – –
– – – – –
Saugeen
Web/CloudOps
Temperature Rain KDD Cup 2018 BizITObs - Application BizITObs - Service BizITObs - L2C Bitbrains - Fast Storage Bitbrains - rnd ETT1
Energy
ETT2
Solar
Electricity
Sales
Transport
Loop Seattle SZ-Taxi M_DENSE Restaurant Hierarchical Sales
Healthcare
Econ/Fin
Car Parts
23
The selected datasets are chosen to avoid any leakage between pretraining and test data. Detailed dataset statistics are provided in Table 3, including frequency, prediction length, variate setting, number of series, series length, and total number of observations. For each time series, the final 10% of observations is reserved as the test split. B.3
fev-bench Benchmark
The fev-bench benchmark comprises a total of 100 time series forecasting tasks. Detailed dataset statistics are provided in Table 4. This section summarizes the main characteristics of these tasks and provides citations for the corresponding data sources. For datasets originating from forecasting competitions, the benchmark adopts the fixed forecast horizon T specified by the original competition setup. For all other datasets, the forecast horizon is determined according to a frequency–horizon mapping. An exception is made for a subset of hourly datasets, for which T = 168 is used in order to support long-range forecasting over a one-week period. The number of evaluation windows W is then selected so as to split each series as evenly as possible while ensuring that sufficient historical context remains available for every forecast of length H. Dataset frequencies are reported using pandas frequency aliases, namely minuTely, Hourly, Daily, Weekly, Monthly, Quarterly, and Yearly. The benchmark is constructed from a diverse collection of domains, including macroeconomics, energy systems, retail and sales forecasting, epidemiology, public health, environmental monitoring, and database operations. The included datasets can be grouped into the following source categories. • GIFT-Eval. The benchmark includes datasets from the GIFT-Eval corpus (Aksu et al., 2024), which contains a mixture of univariate and multivariate forecasting tasks. The original GIFT-Eval collection draws on data sources compiled from prior benchmark and application papers (Godahewa et al., 2021; Jiang et al., 2023; Mancuso et al., 2021; Wu et al., 2021; Palaskar et al., 2024). • Macroeconomic datasets. A broad set of macroeconomic and socioeconomic datasets is included, such as GVAR (Mohaddes and Raissi, 2024), US Consumption (Wilms and Croux, 2016), Australian Tourism (Athanasopoulos et al., 2009), FRED-MD (McCracken and Ng, 2016), FRED-QD (McCracken and Ng, 2021), world CO2 emissions (Pedersen, 2025), life expectancy (Noor, 2025), and global tourism (Qurban, 2025). For both FRED-MD and FRED-QD, two separate forecasting tasks are defined. The first task follows the CEE model (Christiano et al., 1999) and focuses on forecasting employment, inflation, and federal funds rate indicators. The second task considers the joint forecasting of 51 core macroeconomic indicators. It should be noted that the benchmark uses the August 2025 snapshot of FRED-MD, which differs from the snapshot used in Monash repository (Godahewa et al., 2021). • Energy datasets. The energy-related portion of the benchmark includes several forecasting settings of practical relevance. These datasets cover the electricity price forecasting (EPF) benchmark (Fleming and Wallace, 1986), ERCOT generation data (Ansari et al., 2024), ENTSOe load data (Data, 2020) paired with weather variates obtained from Renewables.ninja (Staffell et al., 2023), and solar generation data (Maverick, 2025). Together, these datasets provide a mix of load, price, and renewable generation forecasting tasks. • BOOMLET. The benchmark also includes multivariate observability datasets from BOOMLET (Cohen et al., 2025), which is itself a subset of the larger BOOM benchmark curated by the original authors. To maintain diversity across data sources and prevent overre presentation 24
from a single benchmark family, only BOOMLET datasets with a sampling frequency of at least one minute are retained. • Forecasting competitions. A substantial portion of the benchmark is drawn from forecasting competitions, many of which were hosted on kaggle.com. These include Favorita store sales and transactions (lexis Cook et al., 2020), the M5 competition (Makridakis et al., 2022), restaurant visitor and reservation forecasting (Howard et al., 2017a), Rossmann store sales (FlorianKnauer and Cukierski, 2015), Walmart sales forecasting (Admin and Cukierski, 2014), and Rohlik sales forecasting (MichalKecera, 2024). In addition, the benchmark includes the KDD Cup 2022 dataset for wind power forecasting (Zhou et al., 2024), as well as datasets from the Global Energy Forecasting Competitions held in 2012, 2014, and 2017 (Hong et al., 2014). These competition datasets typically come with standardized train–test setups and fixed forecast horizons, making them especially useful for controlled model comparison. • Other sources. To further broaden domain coverage, the benchmark incorporates datasets from several additional sources: – Influenza-like illness case counts collected by the European Centre for Disease Prevention and Control (ECDC, 2025). – Fashion trend data from Hermes (David et al., 2022). – Hospital admissions data from Riyadh (of Health Affairs and Ministry of Health, 2024). – Query count data for Amazon Redshift database servers (van Renen et al., 2024). – Solar energy generation data with associated weather covariates (Maverick, 2025). – Air quality measurements from an Italian city together with weather variates (De Vito et al., 2008). – COVID-19 cases, hospital admissions, and deaths in the United Kingdom across multiple administrative levels (data from official UK government sources, 2022). These additional datasets complement the benchmark by introducing forecasting tasks from healthcare, epidemiology, fashion, environmental sensing, and cloud/database system monitoring, thereby increasing the breadth of real-world scenarios represented in fev-bench. Table 4 Individual statistics of the fev-bench benchmark across all datasets. Task
Domain
Freq.
T
W
Median length
cloud cloud energy energy energy energy retail retail healthcare nature nature nature mobility mobility mobility mobility mobility mobility mobility
5T H 15T H D W D W M 10T D H D 5T H D H 15T H
288 24 96 168 28 13 28 13 12 144 28 24 28 288 168 28 168 96 168
20 20 20 20 20 5 10 10 4 20 11 20 10 10 10 10 10 10 2
31,968 2,664 69,680 17,420 724 103 1,825 260 84 52,704 366 8,784 365 105,120 8,760 730 17,520 2,976 744
# series
# targets
1 1 2 2 2 2 118 118 767 1 1 1 323 323 323 30 30 156 156
7 7 7 7 7 7 1 1 1 21 21 21 1 1 1 1 1 1 1
GIFT-Eval BizITObs-L2C BizITObs-L2C ETT ETT ETT ETT Hierarchical Sales Hierarchical Sales Hospital Jena Weather Jena Weather Jena Weather Loop Seattle Loop Seattle Loop Seattle M-DENSE M-DENSE SZ Taxi SZ Taxi
Continued on next page
25
Table 4 Individual statistics of the fev-bench benchmark across all datasets. (continued) Task
Domain
Freq.
H
W
Median length
# series
# targets
Solar Solar
energy energy
W D
13 28
1 10
52 365
137 137
1 1
econ econ econ econ econ econ econ econ econ econ econ econ
Q M M Q Q Q M Q Y Y Y Y
8 12 12 8 8 8 12 8 5 5 5 5
2 20 20 20 20 10 10 10 10 9 10 2
36 798 798 266 266 178 792 262 64 60 74 21
89 1 1 1 1 33 31 31 31 191 237 178
1 3 51 3 51 6 1 1 1 1 1 1
energy energy energy energy energy energy energy energy energy energy energy energy energy energy energy energy energy
15T 30T H H H H H H D H M W H H H 15T H
96 96 168 24 24 24 24 24 28 168 12 13 168 168 168 96 24
20 20 20 20 20 20 20 20 20 20 15 20 10 20 20 20 20
175,292 87,645 43,822 52,416 52,416 52,416 52,416 52,416 6,452 154,872 211 921 39,414 17,520 17,544 198,600 49,648
6 6 6 1 1 1 1 1 8 8 8 8 11 1 8 1 1
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
cloud cloud cloud cloud cloud cloud cloud cloud cloud cloud cloud cloud cloud cloud cloud
5T 5T T 5T T 5T 30T 30T H H H T T T T
288 288 60 288 60 288 96 96 24 24 24 60 60 60 60
20 20 20 20 20 20 20 20 20 20 20 20 20 20 20
16,384 16,384 16,384 16,384 16,384 16,384 10,463 10,463 5,231 5,231 5,231 16,384 16,384 16,384 16,384
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
21 53 49 23 35 54 40 100 52 75 100 75 52 67 28
retail retail retail retail retail retail energy energy energy retail retail retail retail retail retail retail retail retail retail retail
M W D M W D D 10T 30T M W D D W D W D W D W
12 13 28 12 13 28 14 288 96 12 13 28 28 8 61 8 14 13 48 39
2 10 10 2 10 10 10 10 10 1 1 1 8 5 5 1 1 8 10 1
54 240 1,688 54 240 1,688 243 35,279 11,758 58 257 1,810 296 170 1,197 150 1,046 133 942 143
1,579 1,579 1,579 51 51 51 134 134 134 30,490 30,490 30,490 817 7 7 5,243 5,390 1,115 1,115 2,936
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
healthcare
W
13
10
201
25
1
Macroeconomic datasets Australian Tourism FRED-MD-CEE FRED-MD-Macro FRED-QD-CEE FRED-QD-Macro GVAR US Consumption US Consumption US Consumption World CO2 Emissions World Life Expectancy World Tourism Energy datasets ENTSO-e Load ENTSO-e Load ENTSO-e Load EPF-BE EPF-DE EPF-FR EPF-NP EPF-PJM ERCOT ERCOT ERCOT ERCOT GFC12 GFC14 GFC17 Solar with Weather Solar with Weather BOOMLET BOOMLET-1062 BOOMLET-1209 BOOMLET-1225 BOOMLET-1230 BOOMLET-1282 BOOMLET-1487 BOOMLET-1631 BOOMLET-1676 BOOMLET-1855 BOOMLET-1975 BOOMLET-2187 BOOMLET-285 BOOMLET-619 BOOMLET-772 BOOMLET-963 Forecasting competitions Favorita Store Sales Favorita Store Sales Favorita Store Sales Favorita Transactions Favorita Transactions Favorita Transactions KDD Cup 2022 KDD Cup 2022 KDD Cup 2022 M5 M5 M5 Restaurant Rohlik Orders Rohlik Orders Rohlik Sales Rohlik Sales Rossmann Rossmann Walmart Other datasets ECDC ILI
Continued on next page
26
Table 4 Individual statistics of the fev-bench benchmark across all datasets. (continued) Task
Domain
Freq.
H
W
Median length
# series
# targets
Hermes Hospital Admissions Hospital Admissions Redset Redset Redset UCI Air Quality UCI Air Quality UK COVID-Nation-Cumulative UK COVID-Nation-Cumulative UK COVID-Nation-New UK COVID-Nation-New UK COVID-UTLA-Cumulative UK COVID-UTLA-New
retail healthcare healthcare cloud cloud cloud nature nature healthcare healthcare healthcare healthcare healthcare healthcare
W D W 5T 15T H H D D W D W W D
52 28 13 288 96 24 168 28 28 8 28 8 13 28
1 20 16 10 10 10 20 11 20 4 20 4 5 10
261 1,731 246 25,920 8,640 2,160 9,357 389 729 105 729 105 104 721
10,000 8 8 118 126 138 1 1 4 4 4 4 214 214
1 1 1 1 1 1 4 4 3 3 3 3 1 1
C
Sampling Details
C.1
Flexible Context Length and Horizon Sampling
Unlike decoder-only transformers that inherently support variable context and prediction lengths during pre-training, encoder-only architectures face significant challenges in achieving such flexible length settings, which is quite crucial for model generalization. Theoretically, while models pre-trained with fixed target horizon can perform arbitrary-length inference via auto-regression, they are susceptible to the cumulative error propagation common in decoder-only structures. To address this, we implement flexible context length and horizon during the sampling stage. Specifically, for each sampled entity ei with context length Li and target horizon Ti , we construct each batch by left-padding input contexts with NaN values to the batch-wise maximum input length Lmax = max{ Li }iN=1 . Meanwhile, target outputs are right-padded to a pre-defined maximum horizon Tmax = 480, corresponding to 30 output tokens with patch length L p = 16. The padded context positions, together with invalid observations, constitute the binary observation mask M, while the padded target segments are excluded from the final loss calculation. Such flexible sampling strategy effectively enhances the predictive generalization of Falcon-X. C.2
Runtime Multivariate Sampling
To mitigate potential GPU memory bottlenecks arising from multivariate sampling, we adopt a dynamic variate sampling strategy at runtime. After flexible context length and horizon sampling, for each entity ei with vi variates, we first randomly permute the variate order to reduce order bias and encourage permutation-robust, content-driven cross-variate modeling, consistent with prior findings that channel shuffling improves robustness to channel ordering in multivariate forecasting (Xu et al., 2026). Then we traverse variates in the permuted order. Subsequently, we iteratively append valid variates to the training sample until all candidates are processed or a pre-defined per-sample limit, Mmax , is reached. Finally, entity ei contributes mi = min(vi , Mmax ) variates. This runtime design naturally supports heterogeneous variate dimensionalities and prevents dimensionality-related computational bottlenecks while facing excessively large vi . After variate selection, multivariate samples may still vary in variate count. To batch them efficiently, we enforce a batch-level variate budget and accumulate samples until the total number of retained variates reaches a preset threshold, stabilizing memory usage across training steps. This variatewise batching strategy substantially reduces channel-padding waste and enables efficient training on multivariate data with highly variate dimensionality.
27
D
Additional Visualization
We provide visualizations of Falcon-X’s quantile forecasts across representative datasets at different frequencies {5T, 15T, 10S, H} and prediction horizons T = {48, 60, 480, 720}. Specifically, Figure 10 shows medium-horizon (T = 480) forecasts for Loop Seattle/H, Figure11 for KDD Cup 2018/H, and Figure 12 for Electricity/H at T = 720. Figure 13 presents forecasts on Bitbrains Fast Storage/5T at T = 48, where the two channels exhibit strong positive correlation, which Falcon-X accurately captures. In contrast, Figure 14 shows forecasts on Bizitobs Application/10s at T = 60, where two channels are negatively correlated, and Falcon-X successfully models the inverse relationship. These results demonstrate that the variable-to-prototype design effectively captures both positive and negative inter-variable dependencies, and preserves temporal consistency across diverse datasets. Moreover, the visualizations highlight Falcon-X’s robustness in handling different sampling frequencies and prediction horizons without manual adjustment. Moreover, we visualize Falcon-X ’s medium-horizon (T = 480) quantile forecasts on ETTh1/15T and ETTh1/H, comparing results with and without multivariate inference across all seven channels. As shown in Figure 15 and Figure 16, enabling multivariate inference allows Falcon-X to more accurately capture complex inter-variable relationships, including both strong and subtle dependencies, which results in tighter predictive intervals and improved alignment with observed dynamics across channels.
Figure 10 Medium-horizon (T = 480) quantile forecasts on Loop Seattle.
Figure 11 Medium-horizon (T = 480) quantile forecasts on kdd cup 2018.
28
Figure 12 Long-horizon (T = 720) quantile forecasts on Electricity.
Figure 13 Short-horizon (T = 48) quantile forecasts on bitbrains fast storage.
Figure 14 Short-horizon (T = 60) quantile forecasts on bizitobs application.
29
Figure 15 Medium-horizon (T = 480) quantile forecasts on ETT1/15T with 7 channels, comparing Falcon-X with and without multivariate inference.
30
Figure 16 Medium-horizon (T = 480) quantile forecasts on ETT1/H with 7 channels, comparing Falcon-X with and without multivariate inference.
31