ConceptioArchivearXiv CS
arXiv CSopen access

The Spectrum Is Not Enough: When Context Helps Time-Series Forecasting

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

KURBAN

I N T E L L I G E N C E

L A B

PREPRINT

The Spectrum Is Not Enough When Context Helps Time-Series Forecasting Mert Onur Cakiroglu1 1 2

Mehmet Dalkilic1

Hasan Kurban

2,

Luddy School of Informatics, Computing, and Engineering, Indiana University Bloomington, Bloomington, IN, USA College of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar

arXiv:2607.13006v1 [cs.LG] 14 Jul 2026

Correspondence: [email protected]

ABSTRACT

A growing family of indices scores how predictable a series is from its spectrum. Practitioners increasingly read these scores as answering a different question: whether adding context, a longer lookback, a retrieval plug-in, or a pretrained model, will help. These are not the same question. The value of context is a property of the operating point, not of the series. Any index built from the power spectrum is invariant under phase randomization, whereas the beyond-second-order value that retrieval and foundation models supply is not, because a phase-randomized series is asymptotically Gaussian. We state this as an impossibility result and isolate it with surrogate pairs that fix the spectrum and the marginal by construction. We then give a label-free, configuration-level diagnostic, the coverage deficit, whose principal term measures beyond-spectrum structure as the gain of analog over linear prediction. On seven benchmarks the prediction holds: window-keyed retrieval’s value collapses across surrogate pairs (ECL median +33% → −35%, 𝑝<10−40 ) while every spectral index stays frozen; a foundation model’s value splits into a surviving second-order part and a small beyondlinear margin that collapses; a longer linear window’s value survives. Leave-one-dataset-out, the structure term predicts the sign of beyond-spectrum value where the spectral indices trail it, and the reverse holds for the second-order mechanism. We introduce no new forecaster; the contribution is the distinction, a controlled comparison, and a diagnostic for the deployment decision. Code: https://github.com/KurbanIntelligenceLab/SINE. time-series forecasting; predictability; phase randomization; surrogate data; retrievalaugmented forecasting; foundation models KEYWORDS

1 Introduction Two series can be equally predictable yet differ in whether a longer history improves forecasts. This distinction is at odds with how a growing family of predictability measures is being applied. Recent work scores the predictability of a series with a single inexpensive scalar: spectral predictability, which reports that large pretrained models outperform light baselines when it is high (Wang et al., 2025a); a minimum achievable error from second-order structure with a computable

Preprint • Kurban Intelligence Lab, Hamad Bin Khalifa University • https://github.com/KurbanIntelligenceLab/SINE

1

THE SPECTRUM IS NOT ENOUGH

b 𝑥 : phase-coherent

a

Kurban Intelligence Lab

d value of context 𝑉𝑀 +33%

c

phase randomization 𝑥̃ : phase-randomized

window

𝑥 𝑥̃ 0 56.7% → 55.1%

retrieval

+33.0% → −35.0%

e measured on ECL

0

helps

survives

opposite

collapses

−35%

shared power spectrum

hurts

FM margin

+9.0% → −8.8%

𝑃(𝑥) = 𝑃(𝑥) ̃

Figure 1 Identical predictability, opposite value of context. (a–c) A series 𝑥 and its phase-randomized

surrogate 𝑥̃ share the power spectrum, and after amplitude adjustment the marginal, so every powerspectrum index scores both identically. (d) The beyond-second-order value of context 𝑉𝑀 is nonetheless large on 𝑥 and negative on 𝑥. ̃ (e) The measured dissociation on ECL (channel medians; Tables 1 and 3): the longer window’s purely second-order gain survives phase randomization, while retrieval and the foundation model’s beyond-linear margin collapse through zero, the structure the spectrum cannot see.

spectral surrogate (Feng et al., 2026); an accuracy law relating window-wise complexity to the smallest error deep models reach (Wang et al., 2026); and conditional-entropy forecastability profiles over horizons (Catt, 2026). These address a well-posed question: how predictable a series is in principle. Practitioners face a different question. Given a series and a deployment configuration, will adding context pay off: a longer lookback, a retrieval plug-in that fetches from the training record (Han et al., 2025), or a foundation model that carries broad temporal priors (Ansari et al., 2024; Woo et al., 2024)? A high predictability score is easily read as license for the heavier option and a low one as a reason to stay simple, and the stakes are real: large-scale re-evaluations report that supervised long-term forecasting rankings flip under small changes of setup or metric (Brigato et al., 2026). Determining when additional context will improve forecasting remains an active deployment problem. Recent work addresses different aspects of this decision, including selecting the appropriate lookback for each task (Abdelmalak et al., 2026), explaining retrospectively when foundation models perform well (Widener et al., 2025), and redesigning retrieval methods to capture phase-dependent structure that conventional spectral representations overlook (Nguyen et al., 2026). We show that this reading is unsound: the benefit of context is not a property of the series, and we identify the property that governs it. Predictability and the benefit of added context diverge when predictability indices cannot distinguish cases in which added context yields different benefits. This limitation follows from what these indices measure. The power spectrum records how much energy sits at each frequency, but it does not preserve the phase structure needed to identify the series’ current position within those cycles or determine whether the same position leads to a repeatable future pattern. Phase is precisely what a short window may fail to carry and additional context can supply. A window spanning a full dominant period contains the complete cycle, whereas a shorter one may not, and no within-window model can recover information that is absent (Butera et al., 2026). Whether supplying that missing phase is worthwhile depends on whether it recurs across cycles, beyond the spectrum’s reach. We formalize this using phase-randomized surrogate series. Phase randomization (Theiler et al., 1992; Schreiber and Schmitz, 2000) transforms any series into a surrogate with the same power spectrum but randomized Fourier phases, while an amplitude-adjusted variant also preserves the 2

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

marginal. Any power-spectrum index is therefore identical on a series and its surrogate. Yet the surrogate is asymptotically Gaussian, so the beyond-second-order value exploited by retrieval and foundation models collapses on it. The spectrum therefore identifies what these indices preserve, but not the structure that determines whether additional context is useful. Figure 1 summarizes the consequence. We therefore introduce the coverage deficit, a configuration-level diagnostic computed before deployment without test labels. Its principal term combines a measure of beyond-spectrum structure, the gain of analog over linear prediction, with the fraction of the state-identifying motif that the window fails to observe. A second term flags distributional novelty, the regime in which any memory is stale. Across three context-extending mechanisms on seven standard benchmarks, these two terms separate the deployment question along exactly the theoretical line (Table 4). We make three contributions. Separating predictability from context value. We distinguish series-level predictability from configuration-level context value and prove an impossibility result. Any index built from the power spectrum, which covers spectral predictability (Ω) (Wang et al., 2025a; Goerg, 2013) and spectral-coherence predictability (SCP) (Feng et al., 2026), is invariant under phase randomization. Because beyond-spectrum context value is not invariant, no such index can predict it. Amplitude-adjusted surrogates extend the control to indices with a distributional term such as accuracy-law complexity (Wang et al., 2026) (Section 3). A spectrum-controlled comparison. Surrogate pairs hold the spectrum and the marginal fixed by construction while a longer window, a retrieval plug-in, and a foundation model are switched on and off. The construction fixes exactly what the competing indices read, so the comparison isolates their blind spot rather than relying on correlation (Sections 3, 5). A configuration-level diagnostic. The coverage deficit is label-free and computed before deployment, and its principal term repurposes the nonlinear-prediction statistic of Sugihara and May (1990) to measure the beyond-spectrum structure the power spectrum cannot represent. Leave-one-dataset-out, it predicts the sign of beyond-spectrum context value where Ω is at or below chance; SCP can exceed chance but trails it by 12–19 points (Section 4, Table 4). A forecaster maps a lookback window 𝑥𝑡−𝑆+1∶𝑡 ∈ ℝ𝑆×𝐷 of a 𝐷channel time series to the next 𝐻 steps. We write 𝑆 for the lookback length, 𝐻 for the prediction horizon, and 𝐿 for the dominant period, estimated per channel as the peak of the training-split periodogram. A context-extending mechanism 𝑀 enlarges the information available to a base forecaster without changing the prediction target, for example by increasing the lookback (𝑆 → 𝑆 ′ > 𝑆), retrieving similar windows from the training record, or supplying pretrained temporal knowledge. Let 𝑓 denote the base forecaster and 𝑓 ⊕ 𝑀 the same forecaster augmented with 𝑀. We define the context value of 𝑀 at operating point (𝑆, 𝐻 ) on series 𝑥 as the paired relative reduction in test mean squared error (MSE), Problem setup and notation..

𝑉𝑀 (𝑥; 𝑆, 𝐻 ) =

MSE(𝑓 ) − MSE(𝑓 ⊕ 𝑀) , MSE(𝑓 )

(1)

which is positive when 𝑀 improves prediction. Finally, a series-level predictability index 𝑃(𝑥) is any statistic intended to characterize the intrinsic predictability of 𝑥 independent of a particular 3

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

operating point. Throughout, we distinguish these two quantities: 𝑃(𝑥) characterizes the series itself, whereas 𝑉𝑀 depends on both the series and the deployment configuration.

2 Related Work A growing line scores how predictable a series is. Spectral predictability traces to forecastable-component analysis (Goerg, 2013); Wang et al. (2025a) revive it as Ω and show, across 51 models and 28 datasets, that foundation models beat light baselines when Ω is high. Feng et al. (2026) derive a per-instance linear MSE lower bound from spectral coherence. Wang et al. (2026) relate a window-wise complexity to the smallest error deep models attain. The information-theoretic and dynamical route runs from model-free quantification with weighted permutation entropy (Garland et al., 2014) and largest-Lyapunov measures (Wang et al., 2025b) to horizon-resolved forecastability profiles conditioned on a declared information set (Catt, 2026). That profile bounds the total improvement over the unconditional predictor, whereas our gap Δ isolates the component beyond the best linear predictor on the same access, which no power spectrum represents. We do not dispute these limits, but prove the power-spectrum ones cannot answer the deployment question they are increasingly used for. Entropy and higher-order scores fall outside the impossibility yet stay series- or access-level, and E2 tests them head to head. Our Δnl is the configuration-level analogue. Series-level predictability indices..

The deployment question is now studied directly. Abdelmalak et al. (2026) show a mis-specified lookback inverts rankings and tune it by search. Butera et al. (2026) attribute long-context benefit to generative-process identification and prove a window must strictly exceed a process memory to reach the minimum error. Our spectral/beyond-spectral split refines that benefit: its second-order part is spectrum-visible and survives phase randomization (our longer-linear-window mechanism), the remainder is not. Widener et al. (2025) rate foundation models post hoc. Symbolic memories make the operating point concrete: a de Bruijn graph over the discretized training record recovers cross-window structure at windows as short as 𝑆=12, handling at test time exactly the out-of-vocabulary event our novelty term measures (Cakiroglu et al., 2025). Dynamical-systems forecasters revive delay-coordinate embedding (Majeedi et al., 2025; Hu et al., 2024), exploiting the beyond-spectrum structure our result concerns. None provides a label-free, pre-deployment statistic paired with a statement of what no spectral index can do. When does a heavier option help?.

Retrieval plug-ins inject cross-window structure by frequency statistics (Ye et al., 2024), learned cycle embeddings (Lin et al., 2024), corpus lookup (Han et al., 2025; Tire et al., 2026), diffusion guidance (Liu et al., 2024a), or per-channel retrieval (Kang et al., 2026). Stationarity-aware variants adapt retrieval under non-stationarity (Zhou et al., 2026), and long-context comparisons place retrieval against very long windows (Ahuja et al., 2026). A recent redesign carries amplitude and phase in the retrieval similarity metric (Nguyen et al., 2026), independently pointing to phase as the relevant axis. Foundation models are the pretraining route to the same end (Ansari et al., 2024; Woo et al., 2024). Context parroting shows copying from a long context can beat them (Zhang and Gilpin, 2026), and their failures track spectral shift (Wang et al., 2025c). We treat all of these as context-extending mechanisms and ask a single question across them. Retrieval and pretraining as context..

4

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

Phase-randomized and amplitude-adjusted surrogates are the classical instrument for separating linear from nonlinear structure, including assessing the significance of a nonlinear prediction gain (Theiler et al., 1992; Schreiber and Schmitz, 2000). We repurpose them, with the nonlinear-prediction statistic of Sugihara and May (1990) as Δnl , to control the exact quantities the predictability indices read; E2 adds generic catch22 features (Lubba et al., 2019) as a selection baseline. Surrogate data and selection..

3 Predictability Does Not Determine Context Value We now prove that no power-spectrum index can predict the value of beyond-spectrum context. The argument turns on the gap between what a linear predictor and the best possible predictor achieve, which the spectrum cannot see and a surrogate erases. Throughout, 𝑥 is a real, secondorder-stationary, finite-variance series; the classical steps and all regularity conditions are deferred to the appendix. Fix a horizon and an information set  available to a forecaster at forecast origin 𝑡: a length-𝑆 window, optionally augmented by retrieved context or information supplied by a pretrained model. Throughout this section, let ℎ denote a predictor based on the information set . The minimum mean-squared error achievable by any linear predictor is Two error floors..

2 𝜎lin () = min 𝔼 ‖𝑥𝑡+1∶𝑡+𝐻 − ℎ()‖2 , ℎ∈lin

(2)

where lin denotes the class of all linear predictors based on . Let 𝜎∗2 () denote the Bayes error, i.e., the minimum mean-squared error over all measurable predictors based on . The quantity 2 () depends on 𝑥 only through its autocovariance (App. A.2). Their difference, 𝜎lin 2 Δ() = 𝜎lin () − 𝜎∗2 () ≥ 0,

(3)

is the component of predictability beyond second order. Call a mechanism 𝑀 beyond-spectrum if the predictability it exploits lies past second order, as analog and similarity retrieval and foundation models do and a longer linear window does not. The following bound is the theoretical core of the paper: it ties the value of any such mechanism to the gap Δ, the one quantity a power spectrum cannot see. 2 () > 0. Let ℎ be measurable with respect to the informaTheorem 1 (Value ceiling). Assume 𝜎lin

tion set  and have finite MSE. Then the relative error reduction of ℎ over the best linear predictor on  satisfies 2 () − MSE(ℎ) 𝜎lin Δ() ≤ 2 =∶ 𝑉̄ (), (4) 2 𝜎lin () 𝜎lin () with equality iff ℎ attains the Bayes error 𝜎∗2 (). Consequently, for any context-extending mechanism 𝑀 with access 𝑀 , the relative error reduction of 𝑓 ⊕𝑀 over the best linear predictor on 𝑀 is at most 𝑉̄ (𝑀 ). 2 () and normalProof sketch. Every ℎ measurable in  has MSE(ℎ) ≥ 𝜎∗2 (); subtracting from 𝜎lin izing gives (4), with equality exactly at the Bayes error. Randomized predictors are covered by

5

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

Jensen. Full proof: App. A.4. Because the bound is over the linear predictor with the same access, 𝑉̄ (𝑀 ) measures the beyond-second-order component of context value. For retrieval with memory 2 ( ) = 𝜎 2 ( conditioned upon and the window as key, 𝜎lin 𝑀 lin base ), where base is the base 𝑆-window. Thus, the mechanism’s entire value is beyond second order (App. A.8). A mechanism that also enlarges the linear information set (a longer window, or a foundation model reading a long context) keeps a spectrum-visible second-order component. The theorem then governs its margin over the best linear predictor on that access, which E1 records. To isolate the beyond-second-order gap, we next introduce surrogate time series. The phaserandomized surrogate 𝑥̃ preserves the Fourier amplitudes while randomizing the phases. The iterative amplitude-adjusted Fourier transform (IAAFT) surrogate additionally preserves the marginal distribution (App. A.1). Proposition 2 (Spectral invariance). The periodogram, the full autocovariance, and therefore 2 () for every  drawn from the series are identical for 𝑥 and its phase-randomized surrogate 𝜎lin 𝑥. ̃ Hence any index 𝑃 that is a functional of the power spectrum or the autocovariance satisfies 𝑃(𝑥) ̃ = 𝑃(𝑥); this covers spectral predictability Ω exactly, and spectral-coherence predictability insofar as it reads the preserved per-channel spectra (empirically frozen to |ΔSCP| ≤ 0.015 in E1). The amplitude-adjusted variant additionally fixes the marginal, up to the reported residual.

The amplitudes |𝑋𝑘 | are untouched, so the periodogram and its inverse transform, the auto2 , which solves the linear normal equations in the covariance, are preserved at every lag, and 𝜎lin autocovariance, follows; App. A.3 gives the computation. Lemma 3 (The phase-randomized surrogate erases the gap). Assume the normalized spec-

tral mass is not concentrated on finitely many frequencies (the Lindeberg condition max𝑘 𝑎2𝑘 /𝑠𝑇2 → 0). Then the finite-dimensional laws of 𝑥̃ converge to those of the stationary Gaussian process 𝑥𝐺 with the autocovariance of 𝑥, second moments are preserved exactly along the sequence, and for every fixed degree 𝐷 the best degree-≤ 𝐷 polynomial predictor asymptotically gains nothing over the linear one (Δ𝐷 () → 0) for every finite . Under Condition (M) of App. A.5 (convergence of conditional means in 𝐿2 ), the full gap closes as well: Δ() → 0. The surrogate is a sum of independent-phase sinusoids with the covariance of 𝑥 at every length; the Lindeberg condition kills every standardized joint cumulant of order three and up, so all joint moments converge to Gaussian ones and each fixed-degree least-squares problem converges to its Gaussian counterpart, where the linear predictor is already optimal (App. A.5). Closing the gap over all measurable predictors needs more than moments; Condition (M) (App. A.5) supplies it, and no downstream claim uses it: the impossibility is anchored at the exact endpoint 𝑥𝐺 . Proposition 4 (Beyond-spectrum context value is not spectral). For the stationary Gaussian

process 𝑥𝐺 with the autocovariance of 𝑥, 𝑉̄ () = 0 exactly for every , so by Theorem 1 the beyondsecond-order value of every mechanism is zero on 𝑥𝐺 : retrieval keyed on the operating window has no value, while a mechanism that also enlarges the linear information set keeps its spectrum-visible second-order gain and loses exactly its margin. For any 𝑥 with structure beyond second order, 𝑉̄ (𝑀 ) = 2 ( ) > 0, and such 𝑥 exist. Beyond-spectrum context value therefore separates the pair Δ(𝑀 )/𝜎lin 𝑀 (𝑥, 𝑥𝐺 ). 6

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

Proof sketch. Under 𝑥𝐺 the coordinates of any window and target are jointly Gaussian, so conditional 2 on every ; window-measurable augmentations do not enlarge expectations are affine and 𝜎∗2 = 𝜎lin the conditioning 𝜎-algebra (App. A.8). Existence: for 𝑥𝑡+1 = 𝑓 (𝑥𝑡 ) + 𝜀𝑡 with i.i.d. noise and 𝑓 non2 strictly exceeds it; affine on the support of the stationary law, 𝜎∗2 is the noise variance while 𝜎lin E3’s generator instantiates this. Full proof: App. A.6. Corollary 5 (Impossibility). No predictability index that is a functional of the power spectrum or

the autocovariance can determine beyond-spectrum context value: any such 𝑃 is constant across the pair (𝑥, 𝑥𝐺 ) (Proposition 2) while 𝑉̄ differs across it whenever Δ > 0 (Proposition 4). This covers Ω exactly and SCP up to the reported per-channel residual; the surrogate realizes the comparison at finite length (Lemma 3; proof: App. A.7). Corollary 5 is the central result. It does not say the indices of Section 2 are wrong about predictability; it says the deployment question requires a statistic sensitive to the gap Δ and the window, not the spectrum alone. This refines rather than contradicts prior work: an index reported to predict when foundation models beat baselines (Wang et al., 2025a) is, by Proposition 2, blind to the gap those models exploit; Section 4 estimates it directly. The phase-randomized surrogate alters the marginal, so an index with a distributional term, such as accuracy-law complexity, is not constant across that pair and not covered exactly by Corollary 5. The IAAFT surrogate fixes the marginal too, holding everything such an index reads; being a static monotonic transform of a Gaussian process rather than Gaussian, its gap is small but nonzero. We compare Δnl (𝑥) against the IAAFT ensemble as a standard surrogate test (E1) and cross-check against phase-randomized (FT) surrogates (E6). A marginal term yields no reliable handle on the gap either. Indices that also read the marginal..

Remark 6. The result applies where context value arises from the gap Δ, the recurring nonlinear

motifs and deterministic dynamics that similarity retrieval and in-context completion exploit. It is vacuous where value is purely second-order or Δ = 0 leaves nothing to separate. The diagnostic below carries one term for Δ and one for novelty.

4 The Coverage-Deficit Diagnostic A useful diagnostic must be computable before deployment, without test labels, and must read the configuration, not only the series. We define the coverage deficit Γ(𝑆, 𝐻 ) from a coverage term Γcov and a novelty term Γoov , each matched to a way context value goes to zero (Figure 2). The key quantity is the gap Δ of Section 3: the structure a similarity-retrieval or foundation-model context can exploit and the spectrum cannot represent. 2 = 𝑉̄ of Eq. (4), label-free on the training split, as the analog We estimate the normalized gap Δ/𝜎lin prediction gain MSEanalog Δnl = 1 − , (5) MSElinear 2 with a least-squares predictor and MSE 2 where MSElinear estimates 𝜎lin analog upper-bounds 𝜎∗ with a fixed-𝑘 nearest-neighbour predictor, so Δnl is a conservative (lower-bound) estimate of 𝑉̄ , crossvalidated within the training split with no test labels (estimator details: App. A.9). By Proposition 2 Beyond-spectrum structure term..

7

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

and Lemma 3, Δnl is the term the power spectrum cannot see. A series and its phase-randomized surrogate share Ω, yet Δnl is large for a series with deterministic motifs and → 0 for the surrogate. The surrogate is the Gaussian process with that spectrum, on which analog matching offers no improvement. The analog predictor is the simplex/nearest-neighbour method of empirical dynamic modeling (Sugihara and May, 1990; Takens, 1981), so Δnl is the classical nonlinear-versus-linear prediction gain. This term is what lets Γ separate cases the indices treat alike. Beyond-spectrum structure is worth supplying only when the window is too short to capture it directly. Let 𝑚 be the motif length that identifies the local state, with the dominant period 𝐿 from the training periodogram as the default proxy, and let Coverage term..

𝑢(𝑆) = max(0, 𝑚−𝑆 𝑚 )

(6)

be the fraction of that motif a length-𝑆 window does not observe; a window shorter than the motif cannot form the delay embedding the analog predictor needs (Takens, 1981), and an input strictly longer than the process memory is necessary even in principle (Butera et al., 2026). Then Γcov (𝑆) = Δnl ⋅ 𝑢(𝑆).

(7)

Γcov is large only when there is beyond-spectrum structure to exploit and the window is too short to reach it on its own. The operating point enters through 𝑢(𝑆); the part the indices miss enters through Δnl . Even with exploitable structure, a memory is useless if deployment inputs are unlike the training record. Following the symbolic route, discretize each channel into 𝑏 quantile bins, index training tuples, and let Novelty term..

Γoov (𝑆) = Pr [window tuple ∉ index]

(8)

be the out-of-vocabulary rate over the symbolic index, with quantile-bin discretization in the SAX tradition (Lin et al., 2003); it counts the same event a symbolic training-set memory must handle when a test tuple is absent from its graph (Cakiroglu et al., 2025). Γoov is high for memory-hostile, non-recurring distributions, where context value is near zero regardless of structure. The predicted sign of context value is a threshold (or logistic) rule on (Γcov , Γoov ), fit on a set of datasets and evaluated leave-one-dataset-out (LODO): high Γcov and low Γoov predict that context helps. Augmented Dickey–Fuller (ADF) (Dickey and Fuller, 1979) on the training split supplies a trend-domination check where no motif length is well defined, in which case 𝑢(𝑆) → 0 and Γcov → 0 by convention. Decision rule..

5 Experimental Protocol The protocol tests, in order, that the spectrum-controlled gap is real (Proposition 4), that the diagnostic predicts context-value sign where the indices cannot (Corollary 5), and that both hold across mechanisms, each the simplest standard instance of its class under one protocol. The protocol uses seven benchmarks (D7: ETTh1/h2, ETTm1/m2, Weather, ECL, Traffic), operating windows 𝑆 ∈ 8

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

series level

1

series 𝑥 (training split)

configuration level

2

power-spectrum indices

Ω SCP accuracy-law

achievable accuracy of 𝑥

∅ Corollary 5 indices constant across the pair, so 𝑉𝑀 cannot be recovered

series 𝑥 and

coverage deficit Γ

window (𝑆, 𝐻 )

Δnl ⋅ 𝑢(𝑆) Γoov

value of context 𝑉𝑀

Figure 2 Two levels, two questions. Power-spectrum indices read the series and predict its achievable

accuracy (top band). The coverage deficit reads the series and the window and predicts the value of context (bottom band, highlighted). The dashed link is the impossibility result: the top-band indices are constant across a phase-randomization pair whose context value differs (Corollary 5), so the bottom-band quantity cannot be recovered from them.

{12, 24, 48, 96}, and direct multi-step prediction at the four standard horizons 𝐻 ∈ {96, 192, 336, 720} (detailed tables at 𝐻 =96; the collapse is verified at all four). It uses three seeds for the surrogate draws, paired MSE on 𝑧-normalized channels with 𝑉𝑀 = (MSE(𝑓 ) − MSE(𝑓 ⊕𝑀))/MSE(𝑓 ) per cell, and the last 20,000 points per channel. The base forecaster 𝑓 is the direct-𝐻 least-squares predictor on the 𝑆-window. The mechanisms are (a) a longer linear window (4𝑆 lags; purely second-order), (b) analog retrieval keyed on the 𝑆-window over the training record (the simplex predictor of Sugihara and May, 1990; adds no linear information), and (c) a zero-shot foundation model (Chronos-Bolt, Ansari et al., 2024) reading a long context of 512 points. For mechanism (c), Theorem 1 bounds the population margin over the best linear predictor on the same access; we report its empirical counterpart, the margin over the train-fit linear predictor on the same context window. For each benchmark channel and each 𝑆, generate 𝐾 =20 amplitude-adjusted surrogates (IAAFT, 1000 iterations) per series, preserving the periodogram and marginal; a cell whose mean periodogram residual exceeds 0.02 is excluded as not spectrumcontrolled and counted. Measure 𝑉𝑀 on the original and on every surrogate for the three mechanisms, the foundation model on an eight-channel-per-dataset subsample with 𝐾 =10. Report Ω, SCP (identical across arms by Proposition 2, up to the residual), and Γcov for both arms. Statistics are medians over channels, a bootstrap 95% CI on the median paired gap, and a one-sided Wilcoxon signed-rank over paired cells. E1 is the positive control for the whole argument: absent a spectrum-controlled gap, none of the downstream claims can hold. E1. The spectrum-controlled gap..

Leave-one-dataset-out prediction of the sign of 𝑉𝑀 across all (dataset, channel, 𝑆) cells. Each rule is a one-dimensional threshold on its statistic, with the threshold and direction fit on the six training datasets by balanced accuracy: the structure term Δnl at the operating window (a motif-embedding variant is compared qualitatively in Limitations), Γcov , Ω (Wang et al., 2025a), SCP (Feng et al., 2026), bicoherence (Nikias and Raghuveer, 1987), permutation entropy (Bandt and Pompe, 2002), catch22 with gradient boosting (Lubba et al., 2019), and a per-fold majority baseline. Metric: balanced sign accuracy; significance by a label-permutation test against chance (𝐵=2000). It also tests whether a phase-sensitive higher-order index or generic E2. Sign prediction, head to head..

9

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

features would suffice. Spearman correlation of each index with measured 𝑉𝑀 across cells, on the original arm and on the matched surrogate arm. Supports Corollary 5. E3. Correlation with measured value..

Repeat E2 separately for 𝑀 = longer lookback, 𝑀 = retrieval, and 𝑀 = foundation model, the last for its beyond-linear margin (Theorem 1). The theory predicts the boundary: one Δnl rule should predict the sign of every beyond-spectrum component, while spectral rules should predict the purely second-order mechanism. E4. Generality across mechanisms..

Freeze the Γ rule and the index rules fit on the seven benchmarks; evaluate on withheld Exchange, ILI, and M5 retail series, none seen in development. E5. Out-of-distribution transfer..

Γcov with Δnl forced to one; Γcov and Γoov alone; sensitivity to the surrogate count 𝐾 and the neighbor count; a cross-check of FT against IAAFT surrogates, which lack the remapping artifact (Räth and Monetti, 2009); and MAE–MSE agreement on in-regime cells. E6. Ablations..

6 Results Table 1 is the paper’s spine: on the two benchmarks where window-keyed retrieval has material value, the value collapses to near or below zero. The median 𝑉𝑀 falls from +33.0% to −35.0% on ECL and from +33.8% to +0.7% on Traffic, passing through zero to the analog estimator’s negative finite-sample floor as the Gaussian endpoint of Lemma 3 predicts, and each paired gap is large and overwhelmingly significant. The effect is broad (86% of ECL and 93% of Traffic channels carry positive value; the other five benchmarks have honestly negative medians), seed-stable, and horizon-stable (Table 2). The invariance is not approximate: across every ECL and Traffic surrogate pair the spectral indices are frozen to two decimal places (|ΔΩ| ≤ 0.014, |ΔSCP| ≤ 0.015), while the context value they are meant to predict swings by up to seventy points, the empirical face of Corollary 5. Nor is the collapse a remapping artifact: plain phase-randomized (FT) surrogates reproduce it (App. A.10), and forcing Δnl =1 so that coverage acts alone drops sign agreement with retrieval value from 0.90 to 0.09, confirming that the structure term, not the operating-point fraction, carries the signal. That term also tracks the collapse per dataset, agreeing in sign with the measured change in 𝑉𝑀 on all seven benchmarks (four at ≥ 0.90 cell-level agreement, Traffic at 0.90, the lowest 0.67). The longer linear window, by contrast, survives phase randomization intact (ECL 56.7% → 55.1%), as Proposition 4 requires, and the foundation model splits into a dominant second-order part that survives (ECL: +66% of its +77% total) and a small beyond-linear margin that collapses wherever the model carries one (Table 3). For E3, the structure term correlates with measured retrieval value across all 4,622 gated cells, with Spearman 𝜌 rising from 0.75 on the original arm to 0.90 on the surrogate arm, while Ω carries no signal (𝜌=0.03) and SCP is negatively correlated throughout. The gap is real and spectral indices are blind to it (E1, E3)..

Table 4 reports leave-one-datasetout sign accuracy, and the boundary matches the split Corollary 5 predicts. One structure-term rule predicts both beyond-spectrum components (78.1% retrieval, 73.2% foundation margin; both permutation-significant, caption), where Ω is below chance on retrieval (39.5%) and at chance on the foundation margin (52.4%). The mirror image holds for the purely second-order mechanism, The diagnostic predicts the sign where indices do not (E2, E4)..

10

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

𝑉𝑀 (%) Dataset ECL Traffic

𝑥

𝑥̃

paired gap pts

95% CI

𝑝

+33.0 −35.0 𝟔𝟗.𝟖 [64.5, 73.1] 3.5×10−47 +33.8 +0.7 𝟑𝟐.𝟒 [31.2, 33.2] 3.4×10−125

Table 1 The spectrum-controlled gap on the two benchmarks where retrieval has material value (𝑆=12,

𝐻 =96, 𝑀 = window-keyed retrieval; channel medians, 𝐾 =20 IAAFT; 𝑛=301 ECL and 813 Traffic channels). 𝑉𝑀 collapses through zero under phase randomization, whereas the spectral index Ω is unchanged by construction (Prop. 2). Gaps are paired medians with bootstrap 95% confidence intervals. retrieval gap (pts) at horizon 𝐻 Dataset

96

192

336

ECL Traffic D7 (pooled)

69.8 32.4 35.0

66.6 31.3 33.7

64.5 30.4 32.9

720 62.4 29.7 31.8

Table 2 The collapse holds at every standard horizon (𝑆=12; paired median, points). It declines only

gently with 𝐻 and stays overwhelmingly significant (𝑝 < 10−46 on ECL, 𝑝 < 10−124 on Traffic, at all four).

where SCP reaches 79.2% and the structure term is rightly silent. The comparison also answers a natural objection: since our impossibility covers only power-spectrum functionals, would a phase-sensitive higher-order index suffice? Bicoherence, a third-order phase-coupling measure, does beat Ω on retrieval (66.8 vs 39.5%), confirming it sees structure the power spectrum cannot. But it, permutation entropy, a catch22 gradient-boosting stack, and the two newest indices we cite (accuracy-law complexity and Catt’s profile) all remain series-level and trail the configuration-level structure term on both beyond-spectrum mechanisms (caption). Rules frozen on the seven development benchmarks and applied to withheld Exchange, ILI, and M5 (untouched in development) meet a finding of their own. Retrieval carries beyond-spectrum value in almost none of their cells (positive in 0% of Exchange cells, 4% of ILI, 22% of M5), so the correct call is overwhelmingly “context will not help.” The in-distribution prior, carrying the development “helps” majority, is therefore wrong almost everywhere (0–22% accuracy), the exact failure we warn against. The frozen structure-term rule beats it on raw sign accuracy on all three sets and correctly calls Exchange’s near-total absence of value. It is not a clean sweep: on balanced accuracy the structure term clears chance only on M5 and falls below it on ILI’s 28 cells (Exchange has no positive cells, so balanced accuracy there is not a chance comparison), and we read E5 as an honest specificity test the diagnostic mostly passes, not a second positive control (full table in App. A.12). Transfer (E5)..

7 Discussion The mechanism split reconciles our impossibility result with the strongest empirical finding in this literature. Wang et al. (2025a) report that Ω predicts when foundation models beat light baselines; our decomposition explains why both facts hold at once. The foundation model’s value on these benchmarks is predominantly second-order, content the spectrum represents and Ω can rank. Its beyond-linear margin, the part Corollary 5 says no spectral index can see, is several times smaller 11

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

margin (%) Dataset ECL Traffic

𝑥

paired gap pts

𝑥̃

95% CI

𝑝

+9.0 −8.8 𝟏𝟖.𝟓 [13.2, 37.2] 0.016 +10.9 −10.8 𝟏𝟓.𝟑 [3.5, 30.6] 0.012

Table 3 The foundation model’s beyond-linear margin collapses under phase randomization on

both benchmarks where the model carries one (𝑆=12, 𝐻 =96, 𝑀 = Chronos-Bolt on a 512-point context; channel medians, 𝐾 =10 IAAFT). The margin is the value over the train-fit linear predictor on the same context window, the empirical counterpart of the bound in Thm. 1; the total value (+77% median on ECL) is dominated by a second-order part that survives (+66%). The paired gap is also significant on ETTm2 (𝑝=0.031). balanced sign accuracy (%) Predictor Δnl (ours) Ω SCP bicoherence majority

lookback

retrieval

foundation

46.9 24.0 79.2 46.7 50.0

78.1 39.5 66.1 66.8 50.0

73.2 52.4 54.5 67.3 26.4

Table 4 Leave-one-dataset-out balanced sign accuracy for predicting the sign of context value across

mechanisms (𝑆=12, 𝐻 =96). Δnl performs best on the two beyond-spectrum mechanisms (retrieval and foundation-model margin), while SCP performs best on the second-order lookback. The foundation column reports the beyond-linear margin (Thm. 1); majority is the per-fold majority-class baseline.

(5.7× on Traffic, 7.3× on ECL) and collapses under phase randomization wherever the model carries it. The bound constrains what a spectral index can see, not any one model, so it holds for current foundation-model successors (Liu et al., 2026; Ansari et al., 2025) as well, which the frontier Chronos-2 confirms: its beyond-linear margin likewise collapses on ECL and Traffic (both 𝑝 < 0.02). Retrieval sits at the opposite pole: for mechanisms whose promise is beyond-spectrum structure, a spectral index is provably, and now measurably, silent.

8 Limitations The coverage term relies on a motif length (the trend-robust first-difference period), degrading on multi-scale states; Δnl inherits nearest-neighbor sensitivity to embedding and length. The novelty term is inert in-distribution and earns its keep only under shift. IAAFT preserves the periodogram up to a gated residual (gate 0.02), cross-checked against FT surrogates (Räth and Monetti, 2009); the Gaussianization is asymptotic and degree-wise (App. A.5), excluding finitely supported spectra. Our probes use a linear direct-𝐻 base and, per class, the simplest clean mechanism (the analog retriever adds no linear predictability, so Theorem 1 attributes its whole value to the gap). A learned SOTA retriever (Han et al., 2025; Nguyen et al., 2026) may add linear predictability and is left as a consistent extension; the foundation-model margins rest on eight-channel Chronos-Bolt and frontier Chronos-2 subsamples. The collapse survives a nonlinear MLP base (𝑝 < 0.01 on ECL and Traffic), not just the linear one; deep SOTA backbones (Zeng et al., 2023; Nie et al., 2023; Liu et al., 2024b) and frontier models (Liu et al., 2026) remain untested. Channels are treated univariately;

12

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

the multivariate surrogate (Prichard and Theiler, 1994) covers the cross-channel case. The theory assumes second-order stationarity, and channels are treated as exchangeable: nonstationarity is absorbed only partially (ADF gate, 𝑧-normalization), and cross-channel dependence makes the CIs and 𝑝-values nominal rather than conservative.

9 Conclusion Series-level predictability and the value of added context are different quantities: one belongs to the series, the other to the operating point. Because every power-spectrum index is invariant under phase randomization while beyond-second-order context value is not, none can decide the deployment question. We isolate the gap with surrogate pairs that fix spectrum and marginal by construction, and fill it with the coverage-deficit diagnostic: computed pre-deployment without labels, its structure term predicts the sign of beyond-spectrum value leave-one-dataset-out where spectral indices trail it. The boundary is honest: second-order value is already spectrum-visible, and the diagnostic reports whether context helps, not how much. DATA AND CODE AVAILABILITY

The surrogate generator, the coverage-deficit diagnostic, every predictability index, and the experiment runners are released at https://github.com/KurbanIntelligenceLab/SINE, with a smoke test and seeded surrogate draws. All datasets (ETT, Weather, ECL, Traffic, Exchange, ILI, and M5) are public and cited in Sec. 5; no new human data were collected. COMPETING INTERESTS

The authors declare no competing interests.

References Ibram Abdelmalak, Kiran Madhusudhanan, Jungmin Choi, Christian Klötergens, Vijaya Krishna Yalavarthi, Maximilian Stubbemann, and Lars Schmidt-Thieme. Channel dependence, limited lookback windows, and the simplicity of datasets: How biased is time series forecasting? In Raymond Chi-Wing Wong, Hanghang Tong, Hua Lu, James Kwok, Flora Salim, Yuanfeng Song, and Man Lung Yiu, editors, Advances in Knowledge Discovery and Data Mining, pages 585–597, Singapore, 2026. Springer Nature Singapore. ISBN 978-981-92-1462-4. Rishi Ahuja, Kumar Prateek, Simranjit Singh, and Vijay Kumar. Retrieval mechanisms surpass long-context scaling in time series forecasting, 2026. URL https://arxiv.org/abs/2605.08217. Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C. Maddix, Hao Wang, Michael W. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke-Schneider, and Yuyang Wang. Chronos: Learning the language of time series, 2024. URL https://arxiv.org/abs/2403.07815. Abdul Fatir Ansari, Oleksandr Shchur, Jaris Küken, Andreas Auer, Boran Han, Pedro Mercado, Syama Sundar Rangapuram, Huibin Shen, Lorenzo Stella, Xiyuan Zhang, Mononito Goswami, Shubham Kapoor, Danielle C. Maddix, Pablo Guerron, Tony Hu, Junming Yin, Nick Erickson, Prateek Mutalik Desai, Hao Wang, Huzefa Rangwala, George Karypis, Yuyang Wang, and Michael Bohlke-Schneider. Chronos-2: From univariate to universal forecasting, 2025. URL https://arxiv.org/abs/2510.15821. Christoph Bandt and Bernd Pompe. Permutation entropy: A natural complexity measure for time series.

13

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

Phys. Rev. Lett., 88:174102, Apr 2002. doi: 10.1103/PhysRevLett.88.174102. URL https://link.aps.org /doi/10.1103/PhysRevLett.88.174102.

Lorenzo Brigato, Rafael Morand, Knut Joar Strømmen, Maria Panagiotou, Markus Schmidt, and Stavroula Mougiakakou. There are no champions in supervised long-term time series forecasting. Transactions on Machine Learning Research, 2026. ISSN 2835-8856. URL https://openreview.net/forum?id=yO1JuBpTBB. Luca Butera, Giovanni De Felice, Andrea Cini, and Cesare Alippi. Why do time series models need long context windows?, 2026. URL https://arxiv.org/abs/2606.01999. Mert Onur Cakiroglu, Idil Bilge Altun, Mehmet Dalkilic, Elham Buxton, and Hasan Kurban. Multivariate de bruijn graphs: A symbolic graph framework for time series forecasting, 2025. URL https://arxiv.org/ abs/2505.22768. ICML 2025 Workshop on Foundation Models for Structured Data. Peter Maurice Catt. Forecastability as an information-theoretic limit on prediction, 2026. URL https: //arxiv.org/abs/2603.27074. David A. Dickey and Wayne A. Fuller. Distribution of the estimators for autoregressive time series with a unit root. Journal of the American Statistical Association, 74(366a):427–431, 1979. doi: 10.1080/01621459.1 979.10482531. URL https://doi.org/10.1080/01621459.1979.10482531. Wanjin Feng, Yuan Yuan, Jingtao Ding, and Yong Li. Beyond model ranking: Predictability-aligned evaluation for time series forecasting, 2026. URL https://arxiv.org/abs/2509.23074. Joshua Garland, Ryan James, and Elizabeth Bradley. Model-free quantification of time-series predictability. Phys. Rev. E, 90:052910, Nov 2014. doi: 10.1103/PhysRevE.90.052910. URL https://link.aps.org/doi/1 0.1103/PhysRevE.90.052910. Georg Goerg. Forecastable Component Analysis. In Sanjoy Dasgupta and David McAllester, editors, Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Research, pages 64–72, Atlanta, Georgia, USA, June 2013. PMLR. URL https://proceedings. mlr.press/v28/goerg13.html. Part 2. Sungwon Han, Seungeon Lee, Meeyoung Cha, Sercan O Arik, and Jinsung Yoon. Retrieval augmented time series forecasting. In Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu, editors, Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 21774–21797. PMLR, 13–19 Jul 2025. URL https://proceedings.mlr.press/v267/han25d.html. Jiaxi Hu, Yuehong Hu, Wei Chen, Ming Jin, Shirui Pan, Qingsong Wen, and Yuxuan Liang. Attractor memory for long-term time series forecasting: A chaos perspective. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 20786–20818. Curran Associates, Inc., 2024. doi: 10.52202/079017- 0655. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/24ef004f733548db6b3197d9f68dc b85-Paper-Conference.pdf. Junhyeok Kang, Jun Seo, Soyeon Park, Sangjun Han, Seohui Bae, Hyeokjun Choe, and Soonyoung Lee. Channel-wise retrieval for multivariate time series forecasting. In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1336–1340, 2026. doi: 10.1109/ICAS SP55912.2026.11463178. Jessica Lin, Eamonn Keogh, Stefano Lonardi, and Bill Chiu. A symbolic representation of time series, with implications for streaming algorithms. In Proceedings of the 8th ACM SIGMOD Workshop on Research Issues in Data Mining and Knowledge Discovery, DMKD ’03, page 2–11, New York, NY, USA, 2003. Association for Computing Machinery. ISBN 9781450374224. doi: 10.1145/882082.882086. URL https://doi.org/10 .1145/882082.882086. Shengsheng Lin, Weiwei Lin, Xinyi Hu, Wentai Wu, Ruichao Mo, and Haocheng Zhong. Cyclenet: Enhancing time series forecasting through modeling periodic patterns. In A. Globerson, L. Mackey, D. Belgrave,

14

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 106315–106345. Curran Associates, Inc., 2024. doi: 10.52202/079017-3373. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/bfe7998398779dde03cad7a73b1f8 1b6-Paper-Conference.pdf. Chenghao Liu, Taha Aksu, Juncheng Liu, Xu Liu, Hanshu Yan, Quang Pham, Silvio Savarese, Doyen Sahoo, Caiming Xiong, and Junnan Li. Moirai 2.0: When less is more for time series forecasting, 2026. URL https://arxiv.org/abs/2511.11698. Jingwei Liu, Ling Yang, Hongyan Li, and Shenda Hong. Retrieval-augmented diffusion models for time series forecasting. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 2766–2786. Curran Associates, Inc., 2024a. doi: 10.52202/079017-0091. URL https://proceedings.neurips.cc/paper_files/paper/2024/ file/053ee34c0971568bfa5c773015c10502-Paper-Conference.pdf. Yong Liu, Tengge Hu, Haoran Zhang, Haixu Wu, Shiyu Wang, Lintao Ma, and Mingsheng Long. itransformer: Inverted transformers are effective for time series forecasting. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun, editors, International Conference on Learning Representations, volume 2024, pages 11116–11140, 2024b. URL https://proceedings.iclr.cc/paper_files/paper/2024/file/ 2ea18fdc667e0ef2ad82b2b4d65147ad-Paper-Conference.pdf. Carl H Lubba, Sarab S Sethi, Philip Knaute, Simon R Schultz, Ben D Fulcher, and Nick S Jones. catch22: Canonical time-series characteristics selected through highly comparative time-series analysis. bioRxiv, 2019. doi: 10.1101/532259. URL https://www.biorxiv.org/content/early/2019/01/28/532259. Abrar Majeedi, Viswanatha Reddy Gajjala, Satya Sai Srinath Namburi GNVV, Nada Magdi Elkordi, and Yin Li. Lets forecast: Learning embedology for time series forecasting, 2025. URL https://arxiv.org/abs/ 2506.06454. Huu Hiep Nguyen, Minh Hoang Nguyen, Dung Nguyen, and Hung Le. Spectral retrieval-augmented time-series forecasting, 2026. URL https://arxiv.org/abs/2606.19412. Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. A time series is worth 64 words: Long-term forecasting with transformers, 2023. URL https://arxiv.org/abs/2211.14730. C.L. Nikias and M.R. Raghuveer. Bispectrum estimation: A digital signal processing framework. Proceedings of the IEEE, 75(7):869–891, 1987. doi: 10.1109/PROC.1987.13824. Dean Prichard and James Theiler. Generating surrogate data for time series with several simultaneously measured variables. Phys. Rev. Lett., 73:951–954, Aug 1994. doi: 10.1103/PhysRevLett.73.951. URL https://link.aps.org/doi/10.1103/PhysRevLett.73.951. C. Räth and R. Monetti. Surrogates with Random Fourier Phases, pages 274–285. World Scientific, 2009. doi: 10.1142/9789814271349_0031. URL https://www.worldscientific.com/doi/abs/10.1142/9789814271 349_0031. Thomas Schreiber and Andreas Schmitz. Surrogate time series. Physica D: Nonlinear Phenomena, 142 (3):346–382, 2000. ISSN 0167-2789. doi: https://doi.org/10.1016/S0167-2789(00)00043-9. URL https://www.sciencedirect.com/science/article/pii/S0167278900000439. George Sugihara and Robert M. May. Nonlinear forecasting as a way of distinguishing chaos from measurement error in time series. Nature, 344(6268):734–741, Apr 1990. ISSN 1476-4687. doi: 10.1038/344734a0. URL https://doi.org/10.1038/344734a0. Floris Takens. Detecting strange attractors in turbulence. In Dynamical Systems and Turbulence, Warwick 1980, volume 898 of Lecture Notes in Mathematics, pages 366–381. Springer, 1981. James Theiler, Stephen Eubank, André Longtin, Bryan Galdrikian, and J. Doyne Farmer. Testing for nonlinearity in time series: the method of surrogate data. Physica D: Nonlinear Phenomena, 58(1):

15

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

77–94, 1992. ISSN 0167-2789. doi: https://doi.org/10.1016/0167-2789(92)90102-S. URL https: //www.sciencedirect.com/science/article/pii/016727899290102S.

Kutay Tire, Ege Onur Taga, Muhammed Emrullah Ildiz, and Samet Oymak. Retrieval augmented time series forecasting, 2026. URL https://arxiv.org/abs/2411.08249. Oliver Wang, Pengrui Quan, Kang Yang, and Mani Srivastava. Spectral predictability as a fast reliability indicator for time series forecasting model selection, 2025a. URL https://arxiv.org/abs/2511.08884. Rui Wang, Steven Klee, and Alexis Roos. Time series forecastability measures, 2025b. URL https://arxiv. org/abs/2507.13556. Tianze Wang, Sofiane Ennadir, John Pertoft, Gabriela Zarzar Gandler, Lele Cao, Zineb Senane, Styliani Katsarou, Sahar Asadi, Axel Karlsson, and Oleg Smirnov. Frequency matters: When time series foundation models fail under spectral shift, 2025c. URL https://arxiv.org/abs/2511.05619. Yuxuan Wang, Haixu Wu, Yuezhou Ma, Yuchen Fang, Ziyi Zhang, Yong Liu, Shiyu Wang, Zhou Ye, Yang Xiang, Jianmin Wang, and Mingsheng Long. Exploring accuracy law for deep time series forecasters: An empirical study, 2026. URL https://arxiv.org/abs/2510.02729. Michael Widener, Kausik Lakkaraju, John Aydin, and Biplav Srivastava. On identifying why and when foundation models perform well on time-series forecasting using automated explanations and rating, 2025. URL https://arxiv.org/abs/2508.20437. Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, and Doyen Sahoo. Unified training of universal time series forecasting transformers. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=Yd8eHMY1wz. Weiwei Ye, Songgaojun Deng, Qiaosha Zou, and Ning Gui. Frequency adaptive normalization for nonstationary time series forecasting. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 31350–31379. Curran Associates, Inc., 2024. doi: 10.52202/079017-0985. URL https://proceedings.neurips.cc/pap er_files/paper/2024/file/37c6d0bc4d2917dcbea693b18504bd87-Paper-Conference.pdf. Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. Are transformers effective for time series forecasting? Proceedings of the AAAI Conference on Artificial Intelligence, 37(9):11121–11128, Jun. 2023. doi: 10.1609/aa ai.v37i9.26317. URL https://ojs.aaai.org/index.php/AAAI/article/view/26317. Yuanzhao Zhang and William Gilpin. Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning, 2026. URL https://arxiv.org/abs/2505.11349. Shiqiao Zhou, Holger Schöner, Zipeng Wu, Edouard Fouché, IAG Wilson, and Shuo Wang. Stationarityaware retrieval-augmented time series forecasting, 2026. URL https://arxiv.org/abs/2606.04135.

16

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

APPENDIX

A Surrogates, Proofs, and Estimator This appendix supplies the surrogate construction (A.1); the autocovariance representation of the linear floor (A.2); complete proofs of spectral invariance (A.3), the value ceiling (A.4), the gap-erasure lemma with the explicit Condition (M) for full Bayes-risk convergence (A.5), beyondspectrum non-invariance (A.6), and the impossibility corollary (A.7); the formal retrieval information set (A.8); the gap estimator (A.9); and the ablation and full-table material (A.10–A.12). Every main-text result states its assumptions in place and is sketched there; every proof below is complete. A.1 Phase randomization and IAAFT

Definition 7 (Phase randomization). For a finite real series 𝑥 of length 𝑇 with discrete Fourier

transform 𝑋𝑘 = |𝑋𝑘 |𝑒 𝑖𝜙𝑘 , a phase-randomized surrogate 𝑥̃ has transform |𝑋𝑘 |𝑒 𝑖𝜓𝑘 with the 𝜓𝑘 drawn i.i.d. uniform on [0, 2𝜋), subject to conjugate symmetry so that 𝑥̃ is real. The amplitude-adjusted (IAAFT) variant additionally rank-maps the surrogate onto the empirical values of 𝑥. The IAAFT iteration (Schreiber and Schmitz, 2000): given sorted values 𝑐 of 𝑥 and target amplitudes 𝐴 = |DFT(𝑥)|, initialize 𝑠 as a random permutation of 𝑥, then alternate (1) 𝑠 ← IDFT(𝐴 𝑒 𝑖∠DFT(𝑠) ) to impose the spectrum and (2) a rank-map of 𝑠 onto 𝑐 to impose the marginal, until the relative amplitude change falls below tolerance. We report the residual 𝜀 = ‖‖ |DFT(𝑠)|−𝐴 ‖‖/‖𝐴‖. As a cross-check we also use plain phase-randomized (FT) surrogates, which omit the rank-map and therefore carry no remapping-induced phase correlations (Räth and Monetti, 2009). A.2 The linear floor is an autocovariance functional

Let 𝛾 (ℎ) = cov(𝑥𝑡 , 𝑥𝑡+ℎ ). For one-step prediction (𝐻 =1) from an information set  that is a finite collection of coordinates of 𝑥, the best linear predictor solves the normal equations Σ 𝑏 = 𝑐, where Σ stacks the 𝛾 (⋅) among the coordinates of  and 𝑐 stacks the 𝛾 (⋅) between  and the target. The 2 () = 𝛾 (0) − 𝑐 ⊤ Σ−1 𝑐 depends on 𝑥 only through {𝛾 (ℎ)}, hence only through the resulting error 𝜎lin  power spectrum; the 𝐻 -step case replaces this scalar expression with the trace of the analogous block form and is identical in its dependence on {𝛾 (ℎ)}. A.3 Proof of spectral invariance 2 Since |||𝑋𝑘 |𝑒 𝑖𝜓𝑘 || = |𝑋𝑘 |2 , the periodogram 𝐼 (𝑘) = |𝑋𝑘 |2 /𝑇 is unchanged for every 𝑘, and the circular sample autocovariance, its inverse transform, is preserved at every lag (the ordinary sample autocovariance differs only in edge terms of order 1/𝑇 ), including the past–future lags that link a window 2 () is then identical for 𝑥 and 𝑥̃ for every . Any 𝑃 that to its horizon. By Appendix A.2, 𝜎lin is a measurable functional of {𝐼 (𝑘)} or {𝛾 (ℎ)} inherits the invariance: spectral concentration and entropy (Ω) exactly; per-channel spectral-coherence quantities exactly; cross-channel magnitudes |𝑋𝑘 𝑌̄𝑘 | = |𝑋𝑘 ||𝑌𝑘 | per frequency exactly, while smoothed cross-coherence, which averages random relative phases, is preserved only up to the reported residual; and the covariance component of window-wise complexity. The IAAFT rank-map fixes the empirical marginal exactly, at the cost

17

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

of the periodogram residual 𝜀; hence any functional of the power spectrum and the marginal is invariant up to 𝜀. □ A.4 Proof of the value ceiling

Every predictor ℎ measurable with respect to  has MSE(ℎ) ≥ 𝜎∗2 (), the infimum over all measur2 () and using 𝜎 2 () = 𝜎 2 ()−Δ() gives 𝜎 2 ()−MSE(ℎ) ≤ able predictors. Subtracting from 𝜎lin ∗ lin lin 2 () > 0 yields the bound, with equality iff MSE(ℎ) = 𝜎 2 (). A randomized Δ(); dividing by 𝜎lin ∗ predictor ℎ(, 𝑈 ) with auxiliary noise 𝑈 independent of (, target) satisfies MSE(ℎ) ≥ MSE(𝔼[ℎ ∣ ]) ≥ 𝜎∗2 () by Jensen, so the bound covers stochastic forecasters. Taking  = 𝑀 and ℎ = 𝑓 ⊕𝑀 gives the mechanism statement. □ A.5 Proof of the gap-erasure lemma

Write the surrogate as 𝑥̃𝑡 = ∑𝑚 𝑘=1 𝑎𝑘 cos(2𝜋𝑘𝑡/𝑇 + 𝜓𝑘 ) with 𝑎𝑘 = |𝑋𝑘 |, the 𝜓𝑘 i.i.d. uniform on [0, 2𝜋), and the mean removed (𝑎0 =0). Second moments are exact. For any lag ℎ, taking the expectation over the independent uniform phases kills every cross-frequency term and leaves 𝔼[𝑥̃𝑡 𝑥̃𝑡+ℎ ] = 21 ∑𝑘 𝑎2𝑘 cos(2𝜋𝑘ℎ/𝑇 ) = 𝛾 (ℎ), the target autocovariance, at every 𝑇 . Hence the covariance of any finite coordinate vector of 𝑥̃ equals 2 () is the same for 𝑥 and 𝑥. that of 𝑥 exactly, and by Appendix A.2 𝜎lin ̃ Higher cumulants vanish. Fix a finite index set {𝑡1 , … , 𝑡𝑛 } and set 𝑠𝑇2 = 12 ∑𝑘 𝑎2𝑘 . Because the contributions from distinct frequencies are independent and cumulants of a sum of independent terms add, the joint cumulant of order 𝑟 of the coordinates is ∑𝑘 of the order-𝑟 cumulant of the frequency-𝑘 term, each bounded in modulus by 𝐶𝑟 𝑎𝑘𝑟 with 𝐶𝑟 absorbing the bounded trigonometric factor. Standardizing by 𝑠𝑇𝑟 and using ∑𝑘 𝑎𝑘𝑟 ≤ (max𝑘 𝑎𝑘 )𝑟−2 ∑𝑘 𝑎2𝑘 = 2𝑠𝑇2 (max𝑘 𝑎𝑘 )𝑟−2 , 𝑎2𝑘 (𝑟−2)/2 𝑇 →∞ |cum𝑟 | −−−−−→ 0 ≤ 2 𝐶 max 𝑟 ( 𝑘 𝑠2 ) 𝑠𝑇𝑟 𝑇 for every 𝑟 ≥ 3 under the no-dominant-frequency (Lindeberg) condition max𝑘 𝑎2𝑘 /𝑠𝑇2 → 0, while the order-two cumulants stay fixed at 𝛾 (⋅). By the multivariate Lindeberg–Feller theorem the finite-dimensional laws of 𝑥̃ converge to those of the Gaussian process with covariance 𝛾 (⋅). From law to risk. For the Gaussian endpoint 𝑥𝐺 a conditional expectation is affine, so 𝜎∗2 () = 2 () and the gap is exactly zero; Proposition 4 and Corollary 5 use only this endpoint, so the 𝜎lin impossibility needs no asymptotics. Along the sequence, fix a degree 𝐷 and the polynomial feature map of degree at most 𝐷 on the standardized (window, target) coordinates. The degree-𝐷 least-squares risk is a fixed polynomial in the joint moments of order at most 2𝐷+2; cumulant convergence gives moment convergence, and the Gaussian Gram matrix of the feature map is nonsingular for a nondegenerate covariance, so the optimal degree-𝐷 risk converges to its Gaussian 2 () − inf value, which the affine predictor already attains. Hence Δ𝐷 () = 𝜎lin deg≤𝐷 MSE → 0 for every fixed 𝐷. Condition (M). Termwise decay does not by itself close the gap over all measurable predictors: Δ() → 0 additionally requires the conditional means to converge, 𝔼[𝑌𝑇 ∣ 𝑋𝑇 ] → 𝔼𝐺 [𝑌 ∣ 𝑋 ] in 𝐿2 (Condition (M)); a local limit theorem for the standardized vector together with uniform integrability of 𝑌𝑇2 suffices. We state (M) explicitly rather than assume it silently; no main-text claim uses it, and E1 measures the finite-length gap directly. A finitely supported spectrum (𝑚 18

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

lines) is excluded by the Lindeberg condition, and rightly so: such a series obeys an order-2𝑚 linear recurrence, so for 𝑆 ≥ 2𝑚 it is perfectly linearly predictable and carries no gap, while for 𝑆 < 2𝑚 it carries a genuine gap that phase randomization does not erase. □ A.6 Proof of beyond-spectrum non-invariance

Under 𝑥𝐺 , any finite collection of series coordinates is jointly Gaussian with the target, so 𝔼[𝑌 ∣ ] 2 (), i.e. 𝑉̄ () = 0, for every such . A window-keyed retrieval reads a is affine and 𝜎∗2 () = 𝜎lin 𝜎(𝜅)-measurable augmentation (Appendix A.8), so it does not enlarge the conditioning 𝜎-algebra and, by Theorem 1 at 𝑉̄ = 0, has value 0 on 𝑥𝐺 ; a mechanism that widens the linear span retains its second-order gain and, again by Theorem 1, loses exactly its beyond-linear margin. Existence: let 𝑥𝑡+1 = 𝑓 (𝑥𝑡 ) + 𝜀𝑡 with 𝜀𝑡 i.i.d., mean zero, variance 𝜎𝜀2 , independent of the past, and 𝑓 bounded with 2 = 𝜎 2 + min a stationary solution. Then 𝔼[𝑥𝑡+1 ∣ 𝑥𝑡 ] = 𝑓 (𝑥𝑡 ), so 𝜎∗2 = 𝜎𝜀2 , while 𝜎lin 𝑎,𝑏 𝔼 (𝑓 (𝑥𝑡 ) − 𝜀 2 2 𝑎 − 𝑏𝑥𝑡 ) > 𝜎𝜀 whenever 𝑓 is not almost surely affine on the support of the stationary law; hence Δ > 0 and 𝑉̄ > 0 on the one-step window. E3’s generator instantiates this construction. □ A.7 Proof of the impossibility corollary

Let 𝑃 be any functional of the power spectrum or the autocovariance. By Proposition 2, 𝑃 depends on the process only through {𝛾 (ℎ)}, so 𝑃(𝑥) = 𝑃(𝑥𝐺 ) for the Gaussian process 𝑥𝐺 sharing that autocovariance. By Proposition 4, 𝑉̄ = 0 on 𝑥𝐺 while 𝑉̄ (𝑀 ) > 0 for any 𝑥 with Δ(𝑀 ) > 0, and such 𝑥 exist. A single value 𝑃(𝑥) = 𝑃(𝑥𝐺 ) cannot determine a quantity that differs across the pair, so no such 𝑃 determines beyond-spectrum context value. The statement covers Ω exactly and SCP up to the reported per-channel residual. □ A.8 The retrieval information set

A retrieval mechanism carries a fixed training memory  (built once, not re-estimated at test time) and reads a key 𝜅 at prediction time; its information set is 𝑀 = 𝜎(𝜅) with  conditioned upon, 2 ( ) is the residual of the best affine function of 𝜅. When the key is the operating window, and 𝜎lin 𝑀 2 ( ) = 𝜎 2 ( 𝜅 = 𝑥𝑡−𝑆+1∶𝑡 and 𝜎lin 𝑀 lin base ): the memory contributes no linear predictability beyond the window’s own, so by Theorem 1 the mechanism’s entire relative value is the beyond-second-order margin 𝑉̄ (base ), which Lemma 3 sends to zero on the surrogate. A mechanism that also widens the linear span (a longer window, or a long-context model reading more than 𝑆 points) instead keeps a spectrum-visible second-order component that phase randomization preserves; Theorem 1 then bounds only its margin over the best linear predictor on that wider access, and E1 reports that margin. A.9 Analog-gain estimator Δnl

Delay-embed 𝑥 at dimension 𝑑 (default 𝑑 = clip(𝐿, 4, 32) for dominant period 𝐿) and split the embedded pairs 60/40 in time. MSElinear is the one-step error of the least-squares AR(𝑑) predictor 2 . MSE fit on the first split and tested on the second, an estimator of 𝜎lin analog is the error of a fixed-𝑘 nearest-neighbour predictor (𝑘 = 4) over the same library; being one specific measurable 2 ,Δ = predictor it satisfies MSEanalog ≥ 𝜎∗2 . Hence, at the population level where MSElinear → 𝜎lin nl 2 2 clip(1 − MSEanalog /MSElinear , 0, 1) ≤ 1 − 𝜎∗ /𝜎lin = 𝑉̄ is a conservative (lower-bound) estimate of the normalized gap of Theorem 1; we do not claim consistency at fixed 𝑘. It is the standard 19

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

nonlinear-prediction statistic of surrogate-data analysis (Theiler et al., 1992; Schreiber and Schmitz, 2000); under the linear-Gaussian null of Lemma 3 (the phase-randomized surrogate) it is ≈ 0 in expectation, and against the IAAFT ensemble it is calibrated as a surrogate test (Appendix above), which is why it separates a structured series from its surrogates. The primary E2/E4 rule (Table 4) is the operating-window structure term: Δnl with delayembedding dimension 𝑑=𝑆, so 𝑑=12 at the reported operating point (column dnl_S_train in the released shards). The motif-embedding variant instead sets 𝑑 from the first-difference motif length 𝑚 (𝑑=clip(𝑚, 4, 96), column delta_nl) and is the form compared qualitatively in the main-text Limitations. The released analysis_phasepower.py exposes both through the embedding argument (pass the window length 𝑆 for the operating-window primary). A.10 Ablation numbers (E6)

All at 𝑆=12, 𝐻 =96, retrieval mechanism unless noted. Structure vs. coverage. Sign agreement of a single rule with sign(𝑉 >0) over all seven-benchmark (D7) cells: the operating-window Δnl alone 0.90, Γcov alone 0.73, and the coverage fraction 𝑢(𝑆) alone (i.e. Δnl forced to 1) only 0.09. The structure term, not coverage, carries the signal. FT vs. IAAFT surrogates. Median retrieval 𝑉 (original → surrogate): FT gives ECL +0.43 → −0.32 and Traffic +0.37 → −0.03; IAAFT gives +0.43 → −0.26 and +0.37 → +0.02 on the same 52-channel subsample. The collapse is present under both, so it is not an IAAFT remapping artifact (Räth and Monetti, 2009). Surrogate count. On the gated Table 1 population, the median paired retrieval gap on ECL is +0.667, +0.686, +0.698 for 𝐾 = 5, 10, 20 (𝑛=301) and on Traffic +0.322, +0.323, +0.324 (𝑛=813); the estimate is stable in 𝐾 and reaches the main-text Table 1 value at 𝐾 =20. Neighbours and metric. Median 𝑉 rises monotonically with the neighbour count (𝑘=2, 4, 8) on every dataset without changing sign, and the sign of 𝑉 under MAE agrees with that under MSE in every cell on ECL, ETTh2, ETTm2 and Traffic, 0.86 on ETTh1, and 0.71/0.75 on ETTm1/Weather. A.11 Full spectrum-controlled gap table (E1)

Table A1 gives the per-dataset E1 numbers for all seven benchmarks (the main text shows the two with material retrieval value). 𝑉𝑀 is the window-keyed retrieval value at 𝑆=12, 𝐻 =96; medians over surrogate-valid channels, 𝐾 =20 IAAFT. Ω(𝑥)=Ω(𝑥) ̃ by construction. 𝑉𝑀 (%)

Ω

Γcov

Dataset

𝑥

𝑥̃

𝑥=𝑥̃

𝑥

𝑥̃

ECL Traffic ETTm1 ETTh1 Weather ETTh2 ETTm2

+33.0 +33.8 −14.5 −18.2 −35.0 −40.9 −47.4

−35.0 +0.7 −25.3 −29.2 −43.7 −34.8 −38.7

0.639 0.571 0.541 0.522 0.603 0.671 0.646

0.000 0.129 0.000 0.000 0.000 0.000 0.000

0.000 0.017 0.000 0.000 0.000 0.000 0.000

Table A1 Full E1 spectrum-controlled table (all seven benchmarks). Only ECL and Traffic carry positive

retrieval value on 𝑥; Ω is frozen throughout and Γcov is zero everywhere except Traffic.

20

THE SPECTRUM IS NOT ENOUGH

Kurban Intelligence Lab

A.12 Out-of-distribution transfer table (E5)

Rules frozen on the seven development benchmarks, evaluated on three withheld datasets (Table A2, all 𝑆 pooled); each cell is the fraction of channels whose context-value sign the rule calls correctly. Subheads give cell count and the fraction with positive retrieval value. These sets carry almost no beyond-spectrum value, so “predict no help” scores high by default (SCP’s 96.4% on ILI is exactly this); the discriminating comparison is against the in-distribution prior. Frozen rule Δnl (ours) Γcov (ours) Ω SCP prior

Exchange

ILI

M5

𝑛=32, 0%+

𝑛=28, 4%+

𝑛=400, 22%+

100 100 100 100 0.0

28.6 21.4 3.6 96.4 3.6

50.7 48.2 22.0 22.0 22.0

Table A2 Full E5 out-of-distribution transfer. Frozen rules on three withheld sets; the prior row carries

the development majority.

21

Record · ID 366240 · SHA-256 59a722011522fdb0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.