Time series generation with spectrally aligned latent flow matching Camilo Carvajal Reyesa,∗ , Felipe Tobarb
ARTICLE INFO
ABSTRACT
Keywords: time series flow matching generative modelling latent space Fourier transform wavelet signatures
Latent flow models have proven to be a reliable and cost-effective method for time series generation. However, the latent compression induces unwanted artefacts, such as a spectral mismatch with respect to the underlying dataset, thus hindering their use as training surrogates. In this article, we propose a spectrally-aligned latent-flow time series generator, where the latent space for flow matching is trained to preserve dynamical properties that are relevant for the suitability of synthetic samples. We find that incorporating fine-tuning losses based on canonical signal representations such as the Fourier, wavelet and signature transforms helps overcome these issues. The interpretability of these transformations allows us to ensure that the synthetic signals are aligned with the true ones in terms of relevant features, such as smoothness or targeted spectral content, as opposed to relying on pointwise reconstruction losses only. We compare the proposed aligned models against a base latent-flow model and the state of the art over real-world long-range univariate and multivariate benchmark datasets. Our quantitative results validate the superiority of the proposed method in terms of its performance on metrics reflecting signal realness and computational efficiency, while being aligned to the training set with respect to its local structure.
1. Introduction Probabilistic generative modelling for time series (TS) aims to synthesise signals following a reference distribution which is only available through observed samples. Under the unconditional generation setting, the synthetic series can be used for data augmentation or as training surrogates for largescale downstream tasks. Upon conditioning, the generating distribution can be leveraged for canonical time series tasks such as imputation, forecasting or denoising. Recent probabilistic models for multidimensional time series build on two related perspectives (Lipman, Chen, BenHamu, Nickel and Le, 2022). First, diffusion models (SohlDickstein, Weiss, Maheswaranathan and Ganguli, 2015; Song and Ermon, 2019; Ho, Jain and Abbeel, 2020), which generate series via denoising a Gaussian noise source, incurring a high sampling cost. Second, flow-matching models (Lipman et al., 2022), which train a velocity field that transports from any suitable source distribution (e.g., Gaussian) towards a desired target distribution. In real-world applications involving largedimensional data such as time series, both diffusion- and flow-based models may operate on latent spaces (Rombach, Blattmann, Lorenz, Esser and Ommer, 2022; Dao, Phung, Nguyen and Tran, 2023) as a means to bound computational complexity. Naturally, this latent representation can be seen as a form of lossy compression, where achieving the desired computational improvement comes at the cost of a degradation in the quality of the synthetic samples. See Fig. 1 for an illustration of these artefacts. Hence, an effective latent flow method to sample realistic time series implies addressing these artefacts. These are a ∗ Corresponding author
[email protected] (C. Carvajal Reyes); [email protected] (F. Tobar)
(C. Carvajal Reyes); (F. Tobar) ORCID (s): 0009-0000-8334-1668 (C. Carvajal Reyes);
0000-0003-2486-3583 (F. Tobar)
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
104 Power
arXiv:2609.21989v1 [cs.LG] 18 Sep 2026
Department of Mathematics, Imperial College London, 180 Queen’s Gate, London SW7 2HR, , United Kingdom
Ground truth
Base rec.
Aligned rec.
101 0.0
0.2 0.4 Frequency
0.0
0.2 0.4 Frequency
0.0
0.2 0.4 Frequency
Figure 1: Fourier power spectrum of the Exchange rate dataset (Lai et al., 2018) (mean ± 1 st. dev. , logscale): true samples (left), example of a non-aligned base flow model (centre), and our proposed aligned flow model (right). Observe how the non-aligned model exhibits a spectral artefact in the form of a peak at around frequency 0.25, which is largely minimised by the proposed alignment strategy.
consequence of the spectral bias, which postulates that lower frequencies are prioritised when training neural networks (Rahaman, Baratin, Arpit, Draxler, Lin, Hamprecht, Bengio and Courville, 2019). We hypothesise that the latent space can be designed to reduce the detrimental effect that the latent compression has on the relevant dynamical features of the synthetic signals, such as the signal regularity, e.g., the smoothness of the underlying continuous path of the time series. The key element for this is using signal representations allowing us to identify those key features in order to shift the optimisation weights accordingly. In other words, given that the compressed representation can be interpreted as a budget constraint, we seek to allocate the computational budget to features that matter most. We restrict ourselves to flow-based models with autoencoderinduced latent spaces, and implement the above rationale by training the encoder-decoder pair through a novel set of transform-consistency losses that promote the preservation of high-frequency features, directly in the latent space. This is implemented as a fine-tuning step that we refer to as “alignment”. Our losses are based on different spectral Page 1 of 18
Spectral alignment for time series latent flows
representations of the series, leveraging the complementary properties of the Fourier, wavelet and signature transforms. Intrinsically, this methodology aims to guide the compression of the series to maintain the spectral content that determines regularity, rather than only minimising the Euclidean discrepancy between the true and reconstructed samples as customary in the design of the latent space. We provide an illustration of our pipeline in Figure 2.
a) Learning the latent representation
E
reconstructions latent codes +
The rest of the article is organised as follows: Sec. 2 presents flow models, latent-space models and spectral representations; Sec. 3 establishes the general latent flow matching procedure; Sec. 4 presents our proposed alignment method addressing the drawbacks of latent models, while Sec. 5 shows the theoretical connections of the proposed losses with other discrepancies. The proposed approach is Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Reconstruction loss (Euclidean)
-
b) Aligning the latent representation
The main contributions of this work are: 1. The implementation of a base latent-flow model using an encoder-decoder pair trained with the Euclidean loss that performs well in temporal metrics at a remarkable reduction in sampling cost. This benchmarking of latent flow matching for TS has been largely unexplored and serves as motivation to tackle the issues arising from this compression. 2. A performance analysis of the base latent space, showing that the share of the reconstruction error concentrates in high-detail components. This indicates that the standard training objective under-penalises them, as a consequence of the well-known spectral bias (Rahaman et al., 2019), which undermines the samples realism. 3. A family of loss functions, namely the transformconsistency losses, that promote the preservation of relevant spectral features in the latent representation. The flexibility of the losses allows us to combine them with a weighting scheme that prioritises higherorder features. The losses are then used to improve the latent space via fine-tuning. To the best of our knowledge, the use of general time series transforms to induce geometric structure in the context of generative modelling has not been studied. 4. Theoretical guarantees that our weighted transformconsistency losses can be regarded as matching the roughness of the input signal in the Fourier case, as assessed by the discrete Sobolev norm; and that the Signature-consistency loss can be regarded as an approximation of the norm corresponding to a signature kernel. 5. Empirical quantitative evidence that fine-tuning the encoder-decoder pair with the transform-consistency losses improves the quality of the generated signals in terms of the discriminative score (which measures how difficult it is to distinguish real signals from synthetic ones), while maintaining the computational gain of the base flow-based model, as well as their performance in waveform-based metrics.
D
input series
E
D
input series
reconstructions latent codes +
Consistency loss (Fourier, wave, etc)
-
c) Learning the latent distribution
Gaussian samples
encoded dataset
Rectified flow matching
d) Generating new time series neural ODE
D Gaussian samples
synthetic codes
synthetic series
Figure 2: Latent flow matching for time-series generation. We a) train an encoder-decoder pair (, ) for reconstruction, then b) fine-tune it with a transform-consistency loss. We then c) train a flow-matching model in the latent space to transport a Gaussian prior to the distribution of encoded latents, and d) sample by drawing from the prior, applying the flow, and decoding to obtain new time-series samples.
empirically validated in Section 6. We highlight recent work on autoencoders and flow-based TS generation in Section 7, and an analysis of the nature of the improvements in Section 8, along with the conclusions and future work in Sec 9.
2. Background 2.1. Flow models Given datapoints {𝑥(𝑖) }𝑁 ⊂ ℝ𝑑 , flow matching seeks 𝑖=1 to sample from the data distribution 𝑝data by defining a transport from a source (e.g., Gaussian) distribution 𝑝0 to the target distribution 𝑝1 ≈ 𝑝data . In this context, a collection of intermediate probability densities {𝑝𝑡 }𝑡∈[0,1] is defined by applying the pushforward operator1 𝑝𝑡 = [𝜙𝑡 ]# 𝑝0 . The map 𝜙𝑡 (⋅) ∶ [0, 1] × ℝ𝑑 → ℝ𝑑 is defined as the solution to the ordinary differential equation (ODE) (Lipman et al., 2022): 𝑑 𝜙 (𝑥) = 𝑢𝑡 (𝜙𝑡 (𝑥)); 𝑑𝑡 𝑡
𝜙0 (𝑥) = 𝑥,
(1)
1 equivalently, 𝜙 (𝑥) ∼ 𝑝 for 𝑥 ∼ 𝑝 . 𝑡 𝑡 0
Page 2 of 18
Spectral alignment for time series latent flows
where 𝑢𝑡 ∶ ℝ𝑑 → ℝ𝑑 , 𝑡 ∈ [0, 1] is a time-dependent vector field. The corresponding continuous normalising flow, or simply “flow”, 𝜙𝑡 returns the position of 𝑥 at time 𝑡 when following the vector field 𝑢𝑡 (⋅). Furthermore, 𝜙𝑡 is said to be constructed by the velocity field 𝑢𝑡 when eq. (1) is satisfied. In practice, a conditional version of 𝑢𝑡 , expressed in terms of the samples from source and target distributions, is approximated for general 𝑥𝑡 with a neural network 𝑢𝜃 (𝑥𝑡 , 𝑡) (Lipman et al., 2022; Tong, Fatras, Malkin, Huguet, Zhang, Rector-Brooks, Wolf and Bengio, 2023) via regression. The model 𝑢𝜃 is required to be continuously differentiable with bounded derivatives, a condition that can be easily satisfied when using neural networks (Holderrieth and Erives, 2025) with activation functions such as ReLU. This ensures the existence and uniqueness of a flow 𝜙𝑡 solving eq. (1) (Perko, 2013; Coddington, Levinson and Teichmann, 1956). After training, sampling is achieved by drawing 𝑥0 ∼ 𝑝0 and then solving the ODE in eq. (1) with 𝑥0 as the initial value. In practice, the flow model is implemented such that the intermediate probability density follows a Gaussian transition, i.e., 𝑝𝑡 (𝑥 ∣ 𝑥0 , 𝑥1 ) = (𝑥; 𝜇(𝑥0 , 𝑥1 , 𝑡), 𝜎(𝑥0 , 𝑥1 , 𝑡)2 𝐼) , where the mean is the linear interpolation between source and target (data) samples 𝑥0 and 𝑥1 , respectively (Tong et al., 2023). The resulting conditional velocity field reduces to 𝑥1 − 𝑥0 , as a consequence of this choice of mean. Subsequently, the velocity field model 𝑢𝜃 is trained to match this straight line in the ambient space. Hence, we obtain the following conditional flow-matching objective: = 𝔼𝑡∼ ([0,1]),𝑥1 ∼𝑝1 ,𝑥0 ∼𝑝0 ,𝑥𝑡 ∼𝑝𝑡 (⋅|𝑥1 ,𝑥0 ) ‖𝑢𝜃 (𝑥𝑡 , 𝑡)−(𝑥1 −𝑥0 )‖2 . The limit case where 𝜎 ≡ 0, namely rectified flow (Liu, Gong and Liu, 2022), has been widely adopted in practice. Remark 1. A flow given by a straight trajectory in the Euclidean space would not necessarily be a straight transport with respect to the signal in other representations. Therefore, we propose embedding a latent space with dynamics modelled by suitable TS transforms, allowing for a flow that would be trained by linearly transporting the source distribution to a (latent) target one that does reflect these signal representations.
2.2. Latent spaces for scaling flow models Flow models have benefited from training and sampling in a lower-dimensional latent space, notably leading to faster sampling than models defined directly in the ambient space (Rombach et al., 2022). The learning strategy is separated into two stages: (1) learning the latent space by training an encoder-decoder structure (also referred to as the autoencoder, AE), and (2) defining the flow model in this lower-dimensional space. The compression model introduces latent variables 𝑧 ∈ ℝ𝑙 , 𝑙 ≪ 𝑑, via an encoder ∶ ℝ𝑑 → ℝ𝑙 and a decoder ∶ ℝ𝑙 → ℝ𝑑 . These are trained by minimising a reconstruction error recon = −𝔼𝑧∼𝑞𝜙 (𝑧|𝑥) [log 𝑝𝜃 (𝑥|𝑧)], as well as a weighted adversarial term. The latter corresponds to the loss of a discriminator Carvajal Reyes and Tobar: Preprint submitted to Elsevier
that is trained to distinguish between synthetic and real inputs (Goodfellow, Pouget-Abadie, Mirza, Xu, Warde-Farley, Ozair, Courville and Bengio, 2014), and it has been adopted in the context of image synthesis (Esser, Rombach and Ommer, 2021; Yu, Li, Koh, Zhang, Pang, Qin, Ku, Xu, Baldridge and Wu, 2022). In this setting, which is similar to adversarial autoencoders (Makhzani, Shlens, Jaitly, Goodfellow and Frey, 2015), the adversarial loss effectively acts as regularisation (Lamb, Dumoulin and Courville, 2016). This has proven to be successful in regularising latent spaces for transformer-based probabilistic forecasting (Zhang and Dai, 2022). This way, denoting Adv = log 𝐷(𝑥) + log(1 − 𝐷(𝑥)), ̄ where 𝐷 is the discriminator network, and a weighting factor 0 < 𝜆 < 1, the loss can be written as follows: AE = recon + 𝜆Adv .
(2)
2.3. Time series representations A single time series will be represented by 𝑥 ∈ ℝ𝑚×𝑑 , that is, 𝑚 (uniform in time) realisations of an ℝ𝑑 -valued stochastic process. The (𝑚 × 𝑑)-dimensional vector will be referred to as the time representation of the time series (TS), also called signal or path.
Fourier transform. For a 1-dimensional series, the discrete Fourier transform is defined as ∶ ℝ𝑑×𝑚 → ℂ𝑑×𝑚 , 1 ∑ 𝑥 𝑒−𝜉2𝜋𝑖𝑡 , 𝑚 𝑡=0 𝑡 𝑚−1
[𝑥](𝜉) =
(3)
where 𝑖 is the imaginary unit. Key properties of the Fourier transform are linearity and symmetry. The latter is exploited by the Fast Fourier Transform algorithm (FFT) (Cooley and Tukey, 1965), which is the most widely used method to compute discrete frequency samples from the time representation of a TS.
The wavelet transform. Decompositions of TS based on Fourier coefficients are able to capture precise frequency values, but present low sensitivity in the time domain. This is called frequency localisation. The converse is called time localisation, and the trade-off between both is known as the uncertainty principle (Gabor, 1946). The wavelet decomposition (Meyer, 1992) is a method that presents moderately accurate localisation in both time and frequency domains. In this setting, a reconstruction associated with the discrete wavelet transform (DWT) at decomposition level 𝐾 represents a time series as: 𝐾 ∑ ∑ ∑ 𝑥(𝑡) = 𝐴𝐾 + 𝐷𝑗 = 𝑐𝐴𝐾 [𝑘]𝜙𝐾,𝑘 (𝑡)+ 𝑐𝐷𝑗 [𝑘]𝜓𝑗,𝑘 (𝑡) , 𝑗=1
𝑘
𝑗,𝑘
where 𝜓𝑗,𝑘 = 2𝑗∕2 𝜓(2𝑗 𝑡 − 𝑘) are derived from a mother wavelet 𝜓(𝑡) by shifting and scaling (indices 𝑘 and 𝑗 respectively), and likewise for 𝜙𝐾,𝑘 . The term 𝐴𝐾 corresponds to an approximation coefficient, i.e., modelling coarse aspects of the TS, equivalent to applying a low-pass filter. On the ∑ other hand, 𝐾 𝑗=1 𝐷𝑗 is a sum of detail coefficients, computed Page 3 of 18
Spectral alignment for time series latent flows
recursively by applying a band-pass filter over many levels with downsampling (Daubechies, 1988). The coefficients 𝐴𝐾 and {𝐷𝑗 }𝐾−1 are calculated via the dot product between 1 𝑥 and the wavelets 𝜓 and 𝜙𝑗 , respectively. The resulting orthonormal series is called the wavelet transform, and it fully determines the time series 𝑥. Crucially, the magnitude of wavelet coefficients determines the regularity of the underlying path, similarly to Fourier coefficients (Mallat, 1999). We refer to Appendix A for a brief overview of wavelet families.
Signature Transform (ST). Signatures are representations of paths, whose motivation comes from the method of Picard iterations to obtain solutions to controlled differential equations (Cass and Salvi, 2024). For paths defined over [0, 1] with bounded 𝑝-variation (see Appendix A for a more formal presentation), the 𝑘-fold iterated integral 𝑆(𝑥)(𝑘) is denoted by 𝑆(𝑥)(𝑘) =
∫0<𝑡1 <𝑡2 <…<𝑡𝑘 <1
𝑑𝑥𝑡1 ⊗ 𝑑𝑥𝑡2 ⊗ … ⊗ 𝑑𝑥𝑡𝑘 .
The signature transform )is then the collection 𝑆(𝑥) = ( 1, 𝑆(𝑥)(1) , … , 𝑆(𝑥)(𝑘) , … . The signature can be seen as a basis for functions over curve spaces, rather than being a linear decomposition in a pre-existing basis, which is the case for the Fourier and wavelet transforms (Kidger, Bonnier, Perez Arribas, Salvi and Lyons, 2019).
3. Latent flow matching for TS generation Recall that our goal is to generate high-quality time series that avoid the spectral bias, while maintaining the computational advantages of latent generative models. To this end, we propose a four-stage procedure, outlined in Figure 2 and presented as follows. I. Encoder-decoder training: An encoder-decoder pair (, ) is defined according to Section 2.2. This compression model without fine-tuning will be referred to as the base model. II. Alignment of the latent space: The autoencoder is then fine-tuned by incorporating an additional loss, namely the transform-consistency loss, thereby improving the smoothness by allocating more computational resources to higher-order features. We refer to this process as the alignment. The strategy we adopt for this step is presented in Section 4. III. Latent flow matching: A (rectified) flow model is trained as outlined in Section 2.1, approximating the constant velocity 𝑧1 − 𝑧0 from the interpolation 𝑧𝑡 = 𝑡𝑧1 + (1 − 𝑡)𝑧0 , where 𝑧1 and 𝑧0 denote an encoded data point and a latent prior sample. That is, 𝑢𝜃 (𝑧𝑡 , 𝑡) ≈ 𝑧1 − 𝑧0 , where 𝑧0 is the realisation of an isotropic Gaussian. The target 𝑧1 will correspond to signals represented via our aligned latent space. IV. Sampling: Samples are generated by the ODE in Equation 1 by replacing 𝑢𝑡 with the trained network while using an isotropic Gaussian realisation 𝑧0 as the initial value. This yields latent samples that are close to the encoded data. Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Synthetically generated time series are the result of decoding these back to the ambient space.
4. Transform-consistency losses for spectral alignment In addition to their computational advantage, the flexibility of latent spaces presents the opportunity to incorporate geometries beyond the Euclidean distance in its definition. Therefore, by employing time series transforms that enable us to prioritise high frequencies in the loss computation, we propose a means of counteracting the spectral bias exhibited by Euclidean-only latent spaces. To this end, we define a transform-consistency loss in its general form in Subsection 4.1, followed by specific variants based on the Fourier, Wavelet and Signature transforms (Subsections 4.2, 4.3 and 4.4 respectively). These losses will be used to align (i.e., fine-tune) the base latent space for flow matching. Crucially, the definition of the transform-consistency loss will allow us to prioritise higher-order features, which is key to compensating for the spectral bias Rahaman et al. (2019), which postulates that low frequencies are learnt first in neural network training. This priorisation of higher levels will be done via a Sobolev-inspired weighting scheme 𝑤𝑘 , defined in Section 5.
4.1. General transform-consistency loss Let 𝑇 ∶ ℝ𝑑×𝑚 → ℝ𝐿 be a time series transform, i.e., a mapping as those introduced in Section 2.3. We will consider transforms allowing us to model various levels of detail from the TS. Consequently, we assume that 𝑇 can be decomposed as 𝑇1 ∶ ℝ𝑑×𝑚 → ℝ𝑙1 ⋮ 𝑇𝑘 ∶ ℝ𝑑×𝑚 → ℝ𝑙𝑘 with 𝐿 = 𝑙1 + ⋯ + 𝑙𝑘 . We can then consider a weighted loss of the form: =
𝐾 ∑
𝑤2𝑘 ‖𝑇𝑘 (𝑥) − 𝑇𝑘 (𝑥)‖ ̄ 2,
𝑘=1
for some appropriate norm ‖ ⋅ ‖ and weighting scheme 𝑤𝑘 , 𝑘 = 1, ..., 𝐾. Here 𝑥̄ = ((𝑥)) is the reconstruction of 𝑥 through the autoencoder. This general, transform-dependent evaluation will be referred to as the transform-consistency loss, since it assesses how close a reconstruction is to the original signal according to features computed from time series representations beyond the waveform, such as higher frequencies or decomposition levels. The structure of the transforms considered in this paper is such that higher levels 𝑘 represent finer levels of detail. Consequently, we adopt a weighting scheme that will increase the weight as 𝑘 increases, hence prioritising local structure over global structure when adjusting the latent space. We refer to Section 5.2 to present the scheme and its related theoretical notions, and to Section 6.3 for an empirical view on the effect of upweighting higher levels of detail. Page 4 of 18
Spectral alignment for time series latent flows
4.2. Fourier-consistency loss Let be the discrete Fourier transform, as defined in (3), we can define the Fourier-consistency loss as =
𝐾 ∑ ∑ 𝑘=1 𝜉∈𝐵𝑘
| |2 𝑤2𝑘 |𝐹𝜉 (𝑥) − 𝐹𝜉 (𝑥) ̄ | | |
(4)
with | ⋅ |2 the square complex modulus and 𝐵𝑘 , 𝑘 = 1, … , 𝐾 denotes a partition of the frequency grid.
4.3. Wavelet-consistency loss Let 𝐴(𝑥) denote the linear projection arising from projecting 𝑥 into the basis function 𝜙, as defined in Section 2.3. Similarly, let 𝐷𝑘 (𝑥) be given by projecting 𝑥 onto 𝜓𝑘 , for 𝑘 = 1, … , 𝐾. Then weighted wavelet-consistency loss may be expressed as: 𝑊 (𝑥, 𝑥) ̄ =𝑤2𝐾+1 ‖𝐴(𝑥) − 𝐴(𝑥)‖ ̄ 2 +
𝐾 ∑ 𝑘=1
2 𝑤2𝑘 ‖ ̄ ‖ ‖𝐷𝑘 (𝑥) − 𝐷𝑘 (𝑥) ‖ .
We considered the Daubechies orthogonal family with 4 vanishing moments and five levels of decomposition as a wavelet basis (see Appendix A), as implemented by PyWavelets (Lee, Gommers, Waselewski, Wohlfahrt and O’Leary, 2019).
4.4. Signature-consistency loss Let 𝑆 (≤𝑘) (𝑥) be the signature transform of 𝑥 up to level 𝑘. This representation of a path is composed of a collection of iterated integrals on several levels, as defined in Section 2.3. We define the weighted signature-consistency loss as:
𝑆 =
𝐾 ∑ 𝑘=1
‖ ‖2 𝑤2𝑘 ‖𝑆 (≤𝑘) (𝑥) − 𝑆 (≤𝑘) (𝑥) ̄ ‖ . ‖ ‖
Notice that the injectivity of the full (untruncated) signature only holds when one of the axes is a monotone path shared between paths (Cass and Salvi, 2024), and when all paths share a common departing point. Hence, a common practice is to augment the time series with the time coordinate and to append a starting point to every sequence. To apply this, we consider series to be of dimension 𝑑 + 1 by simply adding the timestamps 0 ≤ 𝑡𝑖 ≤ 1, 𝑖 = 1, … , 𝑚 as the first coordinate, but only to one-dimensional series (for which the signature would not be defined otherwise). We do not append a starting point, meaning that 𝑆 (≤𝑘) will be invariant to path translation (Kidger and Lyons, 2020).
5. Weighting scheme and theoretical justification of the consistency losses 5.1. Weighting scheme definition The structure of the transforms presented in Section 4 allows us to focus the computational resources on features Carvajal Reyes and Tobar: Preprint submitted to Elsevier
that are less emphasised by the Euclidean reconstruction error of vanilla autoencoders. To this end, we will consider weighting schemes to increase the importance of higher-level features, while allowing us to draw theoretical connections between our losses and existing norms in the literature. Indeed, we define the weighting scheme for the Fourier loss as ( )𝑠 𝑤𝑘 = 1 + |𝜉𝑘 |2 . This scheme will be referred to as the Sobolev weighting scheme, inspired by the connection with Sobolev norms, which we outline in Section 5.2. For the Signature and Wavelet transforms, where the levels consist of several components, the weighting scheme is set to: ( ( )2 ) 𝑠 𝑘 𝑤𝑘 = 1 + , 𝑘max with 𝑤𝑘 denoting the level-indexed counterpart of the frequency bin-indexed 𝑤𝑘 . This choice emulates the effect of the frequency weighting 𝑤𝑘 = (1 + |𝜉𝑘 |2 )𝑠 as it augments the weights as the level of detail increases, without fully neglecting low levels. Dividing by 𝑘max , i.e., the maximum level considered in practice, ensures that the scale remains similar with different numbers of levels. The effect of varying the exponent 𝑠 is included in Appendix G, with 𝑠 = 1 being used throughout the experiments. In the rest of this section, we will establish the aforementioned theoretical connections for the Fourier and Signature consistency losses (Sections 5.2 and 5.4 respectively).
5.2. Links to Sobolev norms Remarkably, our transform-consistency losses amount to constraining the functional class of the generated samples. Indeed, our alignment pushes the underlying path of samples to lie in the same Sobolev space as the data, hence forcing them to share the same degree of regularity. To formalise this, we present the discrete form of the Sobolev norm, which can be written in terms of the Fourier transform: ‖𝑥‖2𝐻 𝑠 ∶=
𝑀−1 ∑
(1 + |𝜉𝑘 |2 )𝑠 ‖ [𝑥](𝑘)‖2 ,
𝑘=1
where 𝜉𝑘 corresponds to the 𝑘th frequency. Proposition 1. Let 2 ,𝑠 denote the transform-consistency loss in Equation 4, with disjoint frequency bins 𝐵𝑘 . Denote 𝑚𝑘 ∶= inf 𝜉∈𝐵𝑘 (1 + |𝜉|2 )𝑠∕2 and 𝑀𝑘 ∶= sup𝜉∈𝐵𝑘 (1 + |𝜉|2 )𝑠∕2 . Then, for any weighting scheme {𝑤𝑘 }𝐾 such that 𝑚𝑘 ≤ 𝑤𝑘 ≤ 1 𝑀𝑘 , ‖𝑥 − 𝑥‖ ̄ 2𝐻 𝑠 is equivalent to 2 ,𝑠 (𝑥, 𝑥). ̄ The proof is provided in Appendix C. This holds in particular for separating the frequency bins independently, in which case 𝑤𝑘 ∶= (1+|𝜉𝑘 |2 ) would yield exactly the Sobolev norm. Using the triangle inequality, it is straightforward to verify the following corollary: Page 5 of 18
Spectral alignment for time series latent flows
Corollary 2. Minimising (𝑥, 𝑥) ̄ under the assumptions of Proposition 1 reduces the roughness increase that the reconstructions 𝑥̄ show with respect to the input 𝑥, as measured by the 𝑠-Sobolev norm.
closer with respect to 𝑑𝑇 , which arguably decreases the 𝑑𝑇 Lipschitz constant 𝐿,𝑇 in the support of 𝑝𝑧 . Consequently, the full 𝑇 -based Wasserstein distance is reduced by the proposed alignment.
In 𝐿2 , for a real-valued function 𝑓 in 𝐿2 , the Sobolev ∑ norm can be written as ‖𝑓 ‖2𝐻 𝑠 = |𝛼|≤𝑠 ‖𝐷𝛼 𝑓 ‖2 , where 𝐷𝛼 𝑓 are the partial derivatives of order 𝛼 of 𝑓 . In the context of machine learning, this norm has been used to reinforce training with derivative information of the target values with respect to the input (Czarnecki, Osindero, Jaderberg, Swirszcz and Pascanu, 2017). Here, a higher 𝑠 entails a stronger suppression of high frequencies in the context of signal processing.
5.4. The signature kernel
5.3. Wasserstein bound The fine-tuning procedure makes the latent space reconstructions closer to the dataset according to norms prioritising higher frequencies. We will now see that this effectively reduces an upper bound of the distance between the data distribution and the distribution generated by latent flow matching. Indeed, let 𝑑𝑇 (𝑥1 , 𝑥2 ) =
(𝐾 ∑
)1∕2 𝜔2𝑘 ‖𝑇𝑘 (𝑥1 ) − 𝑇𝑘 (𝑥2 )‖
𝑘=1
for a given representation 𝑇 = 𝑇1 , … , 𝑇𝐾 . Let us denote by 𝑊 (𝑇 ) the 2-Wasserstein distance that has 𝑑𝑇 as the cost function, that is, ( ) 𝑊2(𝑇 ) (𝑝1 , 𝑝2 ) = min 𝑑𝑇 (𝑥1 , 𝑥2 )2 𝑑𝜋(𝑥1 , 𝑥2 ) . 𝜋∈Π(𝑝1 ,𝑝2 ) ∫ℝ𝑑 ×ℝ𝑑 Proposition 3. Let 𝑝̃1 be the distribution induced by latent flow matching, that is, 𝑝̃1 = # 𝑝𝑧̃ with 𝑝𝑧̃ the distribution of latent samples produced by the trained flow in the latent space. Let 𝑝𝑧 = # 𝑝data be the distribution of encoded data, and assume that both the decoder and the flow models are Lipschitz with constants 𝐿 and 𝐿𝜃 , respectively. Then the 2-Wasserstein distance with a 𝑇 -based metric between 𝑝data and 𝑝̃1 can be bounded by: 1+2𝐿𝜃 [ ]1∕2 𝑊2(𝑇 ) (𝑝data , 𝑝̃1 ) ≤ 𝔼𝑝data 𝑇 (𝑥, 𝑥) ̃ +𝐿,𝑇 𝑒 2 𝐻(𝑢𝜃 )1∕2 ,
𝑑 ((𝑧),(𝑧′ ))
where 𝐿,𝑇 ∶= sup𝑧≠𝑧′ 𝑇 ‖𝑧−𝑧′ ‖ is the Lipschitz con2 stant of 𝑑𝑇 , 𝑥̃ = (𝐸(𝑥)) and 𝐻(𝑢𝜃 ) is the flow matching objective: 1
𝐻(𝑢𝜃 ) =
∫0 ∫ℝ𝑑
‖𝑢𝑡 (𝑧𝑡 ) − 𝑢𝜃 (𝑧, 𝑡)‖2 𝑑𝑞𝑡 (𝑧𝑡 ) 𝑑𝑡
given the marginal distribution 𝑞𝑡 of latents 𝑧𝑡 . The proof can be found in Appendix D. Remark 4. Our transform-consistency losses decrease the bound on the Wasserstein distance by directly minimising the term 𝔼𝑝data [𝑇 (𝑥, 𝑥)]. ̃ Moreover, the loss is also implicitly acting on the second term by pulling reconstructions and data Carvajal Reyes and Tobar: Preprint submitted to Elsevier
We next show that the Signature-consistency loss approximates the action of a kernel, that is, a similarity function for a pair of data points. Let us consider the signature kernel given by (Cass and Salvi, 2024) 𝑘𝑤 ([𝑥], [𝑦]) = ⟨𝑆(𝑥), 𝑆(𝑦)⟩𝑤 ,
(5)
and refer to Appendix E.1 for further details of this kernel. A notable property is that 𝑘𝑤 allows us to define a reproducible kernel Hilbert space (RKHS, Aronszajn (1950)), which we briefly introduce in Appendix E.2. We formalise this connection in Proposition 5 below. Proposition 5. Let 𝑠 ∈ ℝ+ and set 𝑤𝑘 = (1 + ( 𝑘 𝑘 )2 )𝑠 for max 𝑘max ∈ ℕ+ . Then, the 𝑤-signature kernel given by 𝑘𝑤 (𝑥, 𝑦) =
∞ ∑
𝑤𝑘 ⟨𝑆(𝑥)(𝑘) , 𝑆(𝑦)(𝑘) ⟩
𝑘=0
defines a unique RKHS 𝑤 satisfying the reproducing property via 𝑘𝑤 . In particular, minimising the infinite transformconsistency loss with weighting function 𝑤𝑘 defined as ∞ 𝑆 =
∞ ∑
𝑤2𝑘 ‖𝑆 (𝑘) (𝑥) − 𝑆 (𝑘) (𝑥)‖ ̄ 2
𝑘=1
amounts to minimising ‖𝑥−𝑥‖ ̄ 2 , where ‖𝑥‖𝑤 = 𝑤
√ 𝑘𝑤 (𝑥, 𝑥).
The proof of Proposition 5 is provided in Appendix E.3. The kernel associated with the signature loss has the property of uniquely characterising the underlying probability measure of paths. This characterisation holds via the kernel mean embedding 𝑀(𝜇) ∶= 𝔼𝑥∼𝜇 [𝑘𝜃 (𝑥, ⋅)], for 𝜇 in the Borel set (𝐾) of probability measures on a compact set 𝐾. In the setup of Proposition 5, 𝑀 is injective on (𝐾), hence distinguishing between measures on (𝐾). We refer to Cass and Salvi (2024) and Simon-Gabriel and Schölkopf (2018) for a formal statement and proof. On the other hand, Lemma 2.3.1 from Cass and Salvi (2024) ensures that the truncated signature transform-consistency loss that we use in practice converges to ∞ as the number of levels 𝐾 grows to infinity. 𝑆
6. Experiments 6.1. Quantitative evaluation Architecture. We begin by training a base model (unaligned) as specified in Section 3. The AE consists of an encoder-decoder pair, where each consists of two convolutional layers with a stride equal to 2, after which a linear layer projects the output to the final latent representation. The flow model regressing the velocity vector field is composed of a multilayer perceptron-based U-NET (Ronneberger, Fischer and Brox, 2015) with depth 2. Specific dimensionalities of the network are specified in Table 4. Page 6 of 18
Spectral alignment for time series latent flows
AE alignment. We then align the encoder-decoder pair by fine-tuning the base model for 25 epochs. This separate optimisation stage corresponds to minimising the transformconsistency losses as an additional term to the reconstruction and adversarial losses introduced in Section 2.2. This joint optimisation yielded more stable training than only minimising the transform-consistency loss. The weighting schemes are set to Sobolev for all three variants, as specified in Section 5.2, with 𝑠 = 1. Other details, such as training times and dimensionality of the transforms, are included in Appendix F.
Flow model training and sampling. The flow model was trained for 500 epochs for multivariate series and 250 for single-channel ones, after which the models showed no improvement. A new flow is trained for each new aligned latent space, although the fact that the geometry of the latent space does not change drastically suggests that fine-tuning an existing flow is probably a positive prospect. We emphasise that our alignment framework is intended to be carried out before training the flow; the fine-tuning of the autoencoder is the only extra computational overhead. Sampling was carried out using a 10-step Euler method (Hairer, Wanner and Nørsett, 1993). Datasets. This study focuses on long-range signals. To this end, we considered the datasets Weather (Kolle, 2024), Exchange rates (Lai et al., 2018) and HEPC (Repository, 2024). These datasets are comprised of multiple channels (1, 8 and 14 for HEPC, exchange rates and weather, respectively), and they reflect different real-world scenarios. Consequently, their frequency structure is different, which highlights the advantages of a methodology that aligns representations to the underlying dataset.
Metrics and baselines. We followed the evaluation protocol used in Yoon, Jarrett and van der Schaar (2019) and Barancikova, Huang and Salvi (2024), adapted for unconditional time series generation. The discriminative score reflects how well synthetic data can fool a classifier, measured via the out-of-sample accuracy of a recurrent neural network (RNN). The predictive score is the loss of an RNN trained on real data to forecast the next point, evaluated on generated samples. We also applied a Kolmogorov–Smirnov (KS) test to compare the empirical distributions of real test series with those of generated data. We report the KS statistic and the percentage of points where the null hypothesis (equal distributions) is rejected at a 5% significance level, applied to the distribution of values of the time series evaluated at a given time stamp (𝑡 = 300, 500, 700 and 900). For readability, the discriminative score is shifted by subtracting 0.5, so all metrics improve as they move closer to 0. Our models are compared with four diffusion-based baselines: SigDiffusion (Barancikova et al., 2024), diffusion-TS (Yuan and Qiao, 2023), CSPD-GP (Biloš, Rasul, Schneider, Nevmyvaka and Günnemann, 2023), and DDO (Lim, Kim, Park and Park, 2023).
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
0.8 0.6 0.4 0.2 0
Real signals Base reconstructions Spectrally aligned reconstructions 200 400 600 800 1000
Figure 3: Signals from the weather dataset (Kolle, 2024) and their reconstructions from the base and Fourier-aligned autoencoders.
Results. Table 1 reports the performance of our proposed models and baselines. The base model and the transformaligned variants show clear gains over these state-of-the-art methods, achieving stronger discriminative and predictive results. The regularised versions usually improve on the base model, especially on the discriminative score. We include a discussion of why this is the case in Section 6.3. Furthermore, our proposed method generates 1000 samples in under one second on an RTX3090 GPU, which represents a significant improvement over the considered benchmarks: between 22 and 97 seconds for SigDiffusion and more than one minute for the rest using the same hardware. Training times are also included in Appendix F. Fine-tuning with the signature-consistency loss incurs longer training times, which is discussed in Appendix H.
6.2. Qualitative evaluation Waveform visualisation. Figure 4 shows real and generated samples for the Weather dataset (Kolle, 2024) using both base and aligned models (via the Fourier transform). The aligned version shows a significant improvement over the base model in terms of smoothness. This is further noticeable when observing the reconstructions, shown in Figure 3, where it is evident that the compression induces a mismatch in local structure.
Density and t-SNE visualisations. We include visualisations of the generated samples compared to SigDiffusions Barancikova et al. (2024) in Figures 5, 6 and 7 for the HEPC, exchange rates and weather datasets, respectively. These consist of probability density estimates and t-SNE dimensionality reduction. Our samples show a better match to the structure of the underlying datasets. We have shown the Fourier-aligned variants for the plots.
6.3. Ablation: The role of high frequencies in improving generation In this subsection, we ablate the effect of assigning stronger weights to higher-order features of the considered TS transforms. For this purpose, we define alternative weighting schemes for the Fourier and wavelet consistency losses Page 7 of 18
Spectral alignment for time series latent flows Table 1 Discriminative, predictive and KS test metrics (KS score and rejection percentage in brackets). Less is better for all metrics. The results are computed for 1000 synthetically generated time series of length 1000, compared to 1000 unseen series from the original datasets. The KS tests are computed over the temporal indices 𝑡 = 300, 500, 700, 900. dataset
HEPC
Exchange
Weather
Model
Discriminative score
Predictive score
KS t=300
KS t=500
KS t=700
KS t=900
Base + Fourier cons. (ours) Base + wavelet cons. (ours) Base + signature cons. (ours)
0.018±.012 0.022±.015 0.208±.100
0.043±.001 0.043±.000 0.042±.000
0.15 (3%) 0.15 (3%) 0.17 (12%)
0.15 (4%) 0.17 (12%) 0.15 (3%)
0.14 (2%) 0.15 (4%) 0.16 (5%)
0.15 (3%) 0.15 (2%) 0.15 (4%)
Base model (ours)
0.211±.103
0.043±.000
0.16 (7%)
0.15 (5%)
0.14 (2%)
0.16 (7%)
SigDiffusion DDO (𝛾 = 1) Diffusion-TS CSPD-GP (RNN) CSPD-GP (Transformer)
0.070±.032 0.081±.019 0.438±.057 0.415±.045 0.500±.000
0.050±.012 0.044±.001 0.066±.022 0.108±.002 0.551±.028
0.20 (16%) 0.25 (46%) 0.85 (100%) 0.52 (100%) 1.00 (100%)
0.18 (9%) 0.23 (38%) 0.87 (100%) 0.53 (100%) 1.00 (100%)
0.19 (12%) 0.25 (46%) 0.82 (100%) 0.55 (100%) 1.00 (100%)
0.21 (22%) 0.25 (51%) 0.88 (100%) 0.56 (100%) 1.00 (100%)
Base + Fourier cons. (ours) Base + wavelet cons. (ours) Base + signature cons. (ours)
0.072±.015 0.069±.023 0.114±.026
0.030±.001 0.029±.001 0.029±.001
0.18 (18%) 0.18 (17%) 0.19 (21%)
0.18 (17%) 0.18 (18%) 0.19 (18%)
0.19 (18%) 0.18 (15%) 0.19 (18%)
0.19 (19%) 0.19 (19%) 0.19 (21%)
Base model (ours)
0.074±.026
0.029±.001
0.19 (18%)
0.18 (18%)
0.18 (16%)
0.18 (17%)
SigDiffusion DDO (𝛾 = 1) Diffusion-TS CSPD-GP (RNN) CSPD-GP (Transformer)
0.278±.062 0.326±.102 0.401±.196 0.500±.001 0.500±.000
0.057±.001 0.094±.004 0.120±.016 0.273±.100 0.432±.074
0.31 (80%) 0.24 (41%) 0.72 (100%) 0.59 (100%) 1.00 (100%)
0.28 (67%) 0.24 (42%) 0.71 (100%) 0.56 (100%) 0.99 (100%)
0.28 (65%) 0.25 (45%) 0.70 (100%) 0.55 (100%) 0.98 (100%)
0.31 (74%) 0.25 (45%) 0.69 (100%) 0.56 (100%) 0.99 (100%)
Base + Fourier cons. (ours) Base + wavelet cons. (ours) Base + signature cons. (ours)
0.350±.153 0.282±.097 0.227±.190
0.158±.002 0.163±.005 0.158±.003
0.22 (20%) 0.22 (18%) 0.23 (22%)
0.22 (16%) 0.21 (14%) 0.21 (16%)
0.21 (16%) 0.22 (18%) 0.23 (17%)
0.20 (15%) 0.21 (20%) 0.21 (18%)
Base model (ours)
0.379±.063
0.176±.015
0.23 (24%)
0.23 (18%)
0.24 (22%)
0.23 (20%)
SigDiffusion DDO (𝛾 = 1) Diffusion-TS CSPD-GP (RNN) CSPD-GP (Transformer)
0.350±.080 0.356±.196 0.498±.003 0.500±.000 0.500±.000
0.168±.001 0.307±.007 0.438±.035 0.505±.007 0.490±.000
0.35 (82%) 0.26 (45%) 0.49 (100%) 0.57 (100%) 0.91 (100%)
0.34 (80%) 0.27 (52%) 0.50 (100%) 0.56 (100%) 0.92 (100%)
0.33 (76%) 0.26 (46%) 0.50 (100%) 0.56 (100%) 0.91 (100%)
0.34 (78%) 0.26 (47%) 0.49 (100%) 0.56 (100%) 0.91 (100%)
of Section 4. Recall first that we consider decomposable transforms in 𝐾 levels. Consider the approximation-detail structure of the Wavelet transform (Section 4.3). Indeed, a common approach to data compression would be to remove the detail coefficients corresponding to higher levels of detail (removing half of the wavelet coefficients). Akin to this compression, we define a truncated version of the transformconsistency loss by simply removing the lower level of details (i.e., setting their weight to zero). A version of this is implemented for the Fourier representation by partitioning the frequency bins into “dyadic levels”, i.e., such that each level of detail (frequencies with lower absolute value) corresponds approximately to 1∕2𝑘 of the frequency grid. The result is a structure that resembles the wavelet transform, to which we can apply the truncation by removing the top high-frequency bins. We show the results in Table 2 for the discriminative and predictive scores. We can observe, for both variants, that removing the last level of detail/frequency is detrimental to sample quality, particularly for the Fourier transform. This shows the relevance of high frequencies for improving samples, as these generally control the functional class to which the time series belongs. Moreover, we observe the benefits of stronger weights in higher frequencies or detail levels by comparing the Sobolev weighting to a vanilla Euclidean fine-tuning of the encoderdecoder pair (referred to as the uniform version). The result of not emphasising any level of detail lies between the truncated and the Sobolev variants in terms of the discriminative score. Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Table 2 Performance of the truncated version of the Fourier and wavelet consistency losses, using the Weather dataset. The abbreviations are disc. for the discriminative score, and pred. for the predictive score, respectively. transform
weights
disc. score ↓
pred. score ↓
wavelet
Sobolev truncated
0.28±.10 0.47±.03
0.16±.00 0.16±.00
Fourier
Sobolev truncated
0.35±.15 0.5±.00
0.16±.00 0.47±.01
None
uniform
0.37±.11
0.16±.00
Hence, we can conclude that emphasising these over coarse levels is beneficial for new samples to resemble the data points, thus validating our choice of weighting scheme.
6.4. The location of the reconstruction error To reflect on how the models are benefiting from the finetuning stage, we analyse what proportion of the reconstruction error is reflected in each detail/decomposition level. By analysing the percentage of the reconstruction error in each frequency bin (see Figure 8), we can see that for the HEPC and Weather datasets, the share of the reconstruction loss in high frequencies is disproportionate, even though they carry limited energy. We claim this is a consequence of neural networks’ tendency to fit lower frequencies first, Page 8 of 18
Spectral alignment for time series latent flows
7. Related work
0.8
0.6
0.4 Real signal Base flow samples
0.2 0
200
400 600 time step
800
1000
800
1000
(a) Base model 0.8
0.6
0.4 Real signal Fourier-aligned samples
0.2 0
200
400 600 time step
(b) Fourier-aligned model Figure 4: First dimension of FM-generated samples vs testing data for the weather dataset (Kolle, 2024). The subcaption shows the underlying models used. The five samples from both models were drawn by solving the FM ODE from the same Gaussian source sample.
known as spectral bias Rahaman et al. (2019). The transformconsistency losses with Sobolev weighting emphasise gradients towards components that contribute to perceptual fidelity, but are less prioritised in the standard objective. Overall, some of the spectral artefacts are attenuated by our alignment procedure. For instance, the aligned reconstructions show a less pronounced peak at the higher frequencies, for all three variants. Moreover, some peaks at intermediate frequencies are diminished, and some improvement is even observed at low frequencies. The effect of higher levels of detail having a high share of the reconstruction error is arguably even higher for the HEPC and weather datasets, when using the signature and wavelet transforms as underlying representations. This indicates the representation invariance of the spectral bias phenomenon. Figure 9 shows the per-level share of the reconstruction error for the wavelet transform, while Figure 10 shows it for the Signature transform. In both cases, we include the error shares for the aligned models relative to the original error.
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Several works have tackled signal generation with flowbased models. We provide an overview of the main approaches in this section. Firstly, the denoising diffusion framework (Ho et al., 2020) was applied to predict each future hidden state sequentially (Rasul, Seward, Schuster and Vollgraf, 2021). Latent diffusion has been used for unconditional generation of short time series (Lim et al., 2023) and for forecasting (Feng, Miao, Zhang and Zhao, 2024). Flow matching has likewise been explored for TS tasks, such as autoregressive forecasting (El-Gazzar and Gerven, 2025) and conditional generation (Hu, Wang, Ding, Wu, Zhang, Li, Wang, Zhang, Li and Chen, 2024) via rectified flows (Liu et al., 2022). The flexibility of FM to allow for arbitrary source distributions has led to Gaussian processes being used as priors (Kollovieh, Lienen, Lüdke, Schwinn and Günnemann, 2025). Attempts to use other TS representations include the use of seasonal and trend components (Cleveland, Cleveland, McRae and Terpenning, 1990; Shen, Chen and Kwok, 2023; Yuan and Qiao, 2023), the use of spectral filters as normalising flows that act on the frequency domain (Alaa, Chan and Schaar, 2021), and how the use of frequency features compares to temporal ones for diffusion-based generation (Crabbé, Huynh, Stanczuk and Schaar, 2024). In this line, Fourier components have been used to encode time series as images to perform diffusion-based generation in that domain (Naiman, Berman, Pemper, Arbiv, Fadlon and Azencot, 2024). Some works do use both time and other TS features, such as periodograms, to condition diffusion models for TS generation (Fons, Sztrajman, El-Laham, Ferrer, Vyetrenko and Veloso, 2025). Recently, Khodakarami, Oommen, Bora and Karniadakis (2026) have proposed decomposing the latent space into mean and high-frequency components, albeit to counteract the smoothing effect that spectral bias has in neural operators. Moreover, augmenting high-frequency components has been used to counteract the effect that noise addition has on these features (Galib, Tan and Luo, 2024), and the Fourier transform has been used to complement the time-based loss of diffusion models (Shen et al., 2023). Our work takes a step further in this direction. Directly aligning representations for unconditional TS generation for more realistic local geometry in cost-efficient generation has not been addressed in the literature to the best of our knowledge. Works about the use of other losses for training autoencoders for time series and other data modalities are presented in Appendix B.
8. Discussion The convenience of spectral alignment. The origin of the problem of noise in latent flow matching samples can primarily be attributed to (1) a distributional discrepancy between generated data and the ground truth signals as a consequence of using flow matching with a Gaussian prior, or (2) as a result of artefacts introduced by the compression stage. The latter happens since the decoder needs to reconstruct Page 9 of 18
Spectral alignment for time series latent flows Density plot
2 1 0
3
2
1 0 Data Value
1
2
(a) Ours (density)
t-SNE plot Original Synthetic
4 2 0 2
0.0
0.2
0.4 0.6 Data Value
0.8
4
1.0
Original Synthetic
2 y_tsne
3
t-SNE plot
Original Synthetic samples
6 5 4 3 2 1 0
y_tsne
Data Density Estimate
Data Density Estimate
Density plot Original Synthetic samples
4
0 2
10
(b) SigDiffusions (density)
5
0 x-tsne
5
10
10
(c) Ours (t-SNE)
5
0 x-tsne
5
(d) SigDiffusions (t-SNE)
Figure 5: Comparison of density and t-SNE plots on the HEPC dataset. Density plot
2 1 0
0.0
0.2
0.4 0.6 Data Value
0.8
4 2 1
(a) Ours (density)
0 5
5
0.4 0.6 Data Value
0.8
10
(b) SigDiffusions (density)
0 5
10
0.2
t-SNE plot
10
Original Synthetic
5
3
0
1.0
10
y_tsne
3
t-SNE plot
Original Synthetic samples
5
y_tsne
Original Synthetic samples
Data Density Estimate
Data Density Estimate
Density plot
5
0 x-tsne
5
10
10
(c) Ours (t-SNE)
10
5
0 x-tsne
Original Synthetic 5 10
(d) SigDiffusions (t-SNE)
Figure 6: Comparison of density and t-SNE plots on the Exchange dataset.
Align vs regularise. Despite the fact that Sobolev norms
the full signal from a lower-dimensional representation, leading to artificial elements in the final TS sample. Our alignment tackles this, which improves the final samples as demonstrated in Section 5.3. Our work is not the first to integrate notions of signal processing and paths with diffusion-based generation. For instance, SigDiffusions (Barancikova et al., 2024) carry out the diffusion training and sampling entirely in the space of log-signatures. While this leads to improved performance and sampling times over long time series compared to other time-based approaches, the fact that the (log)signature needs to be inverted usually leads to signals that are smoother than the dataset signals. Similarly, Diffusion-TS (Yuan and Qiao, 2023) decomposes the time series into seasonal and trend components, focusing on global rather than local structure. We believe that this may be beneficial in certain settings, but that work aiming for improved fine-grained sampling is lacking in the literature.
3 2 1 0
0.3
0.4
0.5 0.6 Data Value
0.7
(a) Ours (density)
0.8
12 10 8 6 4 2 0
t-SNE plot 4 2
Original Synthetic samples
y_tsne
4
spectral alignment will be more evident in datasets where the local structure plays a predominant role, and this is not well captured by the time-based training. Moreover, the criteria used to evaluate success are key. For instance, fewer benefits are observed when evaluating samples with the predictive score, as opposed to the discriminative score. Indeed, the
Density plot Data Density Estimate
Data Density Estimate
Original Synthetic samples
5
Domain adaptability. We can infer that the success of our
Original Synthetic y_tsne
Density plot
6
have been used as regularisers (Czarnecki et al., 2017), we emphasise that our transform-consistency losses are not equivalent to a Sobolev regularisation of the latent space. Instead, in the event that reconstructions are smoother, these losses will encourage reconstructions to follow the regularity of the input. Regularisations only decrease the norm without following a data-based reference. Hence, our alignment procedure can be regarded as adapting the latent space so that the regularity of the data is preserved, which induces a regularisation effect as a particular case.
0 2
0.3
0.4 0.5 Data Value
0.6
(b) SigDiffusions (density)
0.7
4
10
5
0 x-tsne
5
(c) Ours (t-SNE)
10
3 2 1 0 1 2 3
t-SNE plot
Original Synthetic 10 5
0 x-tsne
5
10
(d) SigDiffusions (t-SNE)
Figure 7: Comparison of density and t-SNE plots on the Weather dataset. Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Page 10 of 18
10 2
Base reconstructions
Aligned reconstructions
10 3 0.0
0.2 0.4 frequency
0.0
0.2 0.4 frequency
Percentage
share of loss (log)
Spectral alignment for time series latent flows
0.2 0.1
A1 D1 D2 D3 D4 D5 Wavelet level
Aligned reconstructions
10 2
(a) HPEC
0.0
0.2 0.4 frequency
0.0
0.2 0.4 frequency
share of loss (log)
(b) Exchange 10 1
Base reconstructions
Percentage
10 3
Aligned reconstructions
Original Aligned
0.3 0.2 0.1
10 2
A1 D1 D2 D3 D4 D5 Wavelet level
10 3 0.0
0.2 0.4 frequency
0.0
(b) Exchange
0.2 0.4 frequency
(c) Weather Figure 8: Sum of the reconstruction error for each frequency bin. The subcaption indicates the underlying dataset. The sections where the alignment helps decrease the reconstruction error are highlighted in red.
use of neural networks may easily detect discrepancies in local structure. This suggests that spectrally-aligned samples may be better suited for data augmentation for deep learning training. The reconstruction error location plots included in Section 6.4 highlight where the alignment is most beneficial. For instance, the aligned reconstructions for the HEPC dataset show a decrease in both coarse and detailed features. By contrast, the high-level features for the exchange dataset inherently carry less of the loss share. Interestingly, the signature-alignment is able to decrease the error share at all three levels, which may explain the improvements observed for this variant. Regardless of the transform, some unwanted artefacts are still observed. This reflects on the limitations of our method, but also on the fact that minimising all these peaks would imply that the reconstruction is perfect, which is hard to achieve in the context of a compression model. Nevertheless, the fact that we can characterise the error in seemingly different ways for the same dataset shows the flexibility of our approach, and is at the core of the contribution of this paper.
9. Conclusion This work focuses on improving the fidelity of synthetic time series generation through latent-space alignment. We Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Percentage
share of loss (log)
(a) HPEC Base reconstructions
Original Aligned
0.2
Original Aligned
0.1 A1 D1 D2 D3 D4 D5 Wavelet level (c) Weather
Figure 9: Sum of the reconstruction error averaged over each Wavelet level. The subcaption indicates the underlying dataset.
have presented the effectiveness of tuning latent spaces to focus on fine-detail elements, as modelled by time representations beyond the waveform. Indeed, through the use of our transform-consistency losses, aligned models have shown suitability for data augmentation by improving in terms of the discriminative score, which measures indistinguishability with respect to a neural network. Moreover, the generated samples have shown an out-of-sample error lower than that of existing diffusion models for TS generation, whilst showing a significant decrease in sampling times. Potential positive impacts include improved data augmentation and benchmarking in settings where real time series data is limited or costly to collect. We believe that this study can be extended by analysing the effect of changing the prior to other sources, such as Gaussian processes. This may further decrease the mismatch between flow samples and real signals. Furthermore, the regularity of the autoencoder (e.g., as assessed by its Lipschitz
Page 11 of 18
Percentage
Spectral alignment for time series latent flows
0.6 0.4 0.2 0.0
A. Signal representations
Original Aligned
1
Wavelet families. The form of the wavelets 𝜙 and 𝜓
2 3 Signature level
4
Percentage
(a) HPEC
Original Aligned
0.4 0.2
introduced in Section 2.3 will depend on the type of the particular wavelet family chosen. Common families include Haar (Haar, 1909), Daubechies, and Symlets, a more symmetric version of Daubechies (Daubechies, 1988). These are orthogonal wavelet families, which implies that the signal can be reconstructed from its coefficients. The Daubechies family, in particular, is used as a general-purpose family and allows for focusing on increasing degrees of regularity in the signal, allowing for the detection of a high degree of regularity without excessive loss of time localisation (Mallat, 1999). Typically, only the approximation and coarse levels of 𝑖 detail coefficients are kept in order to have a more compact and less noisy representation of the signal.
Signatures. This section is based on Cass and Salvi (2024).
1
2 3 Signature level
4
Let 𝑥 ∶ [𝑎, 𝑏] → 𝑉 be a path in 𝐶𝑝 ([𝑎, 𝑏], 𝑉 ), i.e., a continuous path with finite 𝑝-variation: (
Percentage
(b) Exchange
sup
∑
𝐷⊆[𝑠,𝑡] 𝐷,𝑡
0.4
Original Aligned
0.3 0.2 1
2 Signature level
3
(c) Weather Figure 10: Sum of the reconstruction error averaged over the Signature level. The subcaption indicates the underlying dataset.
constant) may play a role in the spectral profile of the underlying samples.
Appendix This appendix comprises complementary material to the main body. In Section A we include additional background related to the wavelet and signature transforms. In Section B we provide more related work, particularly regarding latent spaces that are defined with additional loss functions. Sections C and D include the proofs of Proposition 1 and 3, while Section E has further definitions and background that complement the connection between our Signatureconsistency loss and signature kernels. Experimental details are reported in Section F, and the variation of results for different Sobolev exponents is shown in Section G. Finally, Section H complements the discussion for the specific case of the signature-consistency loss and its limitations.
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
)1∕𝑝 ‖𝑥𝑡𝑖+1 − 𝑥𝑡𝑖 ‖
𝑝
< ∞,
𝑖
where 𝐷 ⊆ [𝑠, 𝑡] denotes a partition of [𝑠, 𝑡]. Given any subinterval [𝑠, 𝑡] ⊆ [𝑎, 𝑏], the signature transform 𝑆(𝑥)[𝑠,𝑡] will be given by ( ) 𝑆(𝑥)[𝑠,𝑡] = 1, 𝑆(𝑥)(1) , … , 𝑆(𝑥)(𝑘) ,… , [𝑠,𝑡] [𝑠,𝑡] where 𝑆(𝑥)(𝑘) = [𝑠,𝑡]
∫𝑠<𝑡1 <𝑡2 <…<𝑡𝑘 <𝑡
𝑑𝑥𝑡1 ⊗ 𝑑𝑥𝑡2 ⊗ … ⊗ 𝑑𝑥𝑡𝑘 .
𝑆(𝑥)(𝑘) is called the 𝑘-fold iterated integral. [𝑠,𝑡] Since the magnitude of the coefficients shows a factorial decay, it is appropriate to truncate the signature transform at a certain “depth” 𝑛. Indeed, 𝑛-step ST will be given by ( ) (1) (𝑛) 𝑆(𝑥)≤𝑛 = 1, 𝑆(𝑥) , … , 𝑆(𝑥) , [𝑠,𝑡] [𝑠,𝑡] [𝑠,𝑡] Note that 𝑆 (𝑘) (𝑥) lies in the 𝑘-fold tensor product of ℝ𝑑 (Cass and Salvi, 2024), i.e., it has dimensionality 𝑑 𝑘 . Signatures, by definition, do not depend on the specific set of coordinates employed. Other practical properties of STs include invariance under reparameterisations and injectivity for certain classes of paths. We refer to Cass and Salvi (2024) for a formal presentation of these results. Chen’s relation is an essential result for the implementation of signatures in practice. It states that for 1 ≤ 𝑝 < 2, 𝑥 ∈ 𝐶𝑝 ([𝑎, 𝑏]), 𝑦 ∈ 𝐶𝑝 ([𝑏, 𝑐]), then 𝑆(𝑥 ⊗ 𝑦)[𝑎,𝑐] = 𝑆(𝑥)[𝑎,𝑏] ⊗ 𝑆(𝑦)[𝑏,𝑐] , where 𝑥 ⊗ 𝑦 = 𝟏𝑡∈[𝑎,𝑏] 𝑥𝑡 + 𝟏𝑡∈[𝑐,𝑏] (𝑦𝑡−𝑏+𝑎 ) denotes the concatenation of 𝑥 and 𝑦. Page 12 of 18
Spectral alignment for time series latent flows
A consequence of Chen’s relation is that for piecewise linear 𝑥: 𝑆(𝑥)[𝑡0 ,𝑡𝑛 ] = exp(𝑥𝑡1 −𝑥𝑡0 )⋅exp(𝑥𝑡2 −𝑥𝑡1 )⋅…⋅exp(𝑥𝑡𝑛 −𝑥𝑡𝑛−1 ) which is used to compute signatures. In practice, one approximates the signals with piecewise linear paths. Indeed, given the output continuous piecewise linear function 𝑓 ∶ [0, 1] → ℝ𝑑 satisfying 𝑓 (𝑡𝑖 ) = 𝑥𝑡𝑖 ∈ ℝ𝑑 , Chen’s identity (Lyons, 1998) allows us to write: ( 𝑆 (𝑘) (𝑥) =
𝑘 𝑑𝑓 (𝑡 ) ∏ 𝑖𝑗 𝑗
∫0<𝑡1 <⋯<𝑡𝑘 <1
𝑗=1
𝑑𝑡
) 𝑑𝑡1 … 𝑑𝑡𝑘
,
as features in multi-scale TS representation learning (Liang, Zhang, Liang, Wang, Liang and Pan, 2023). These works have, however, focused on universal representations that can lead to performance in downstream tasks, rather than tackling generative modelling.
C. Proof of Proposition 1 By definition of the Sobolev norm and using the fact that 𝐵𝑘 , 𝑘 = 1, … , 𝐾 partitions the frequency domain, ‖𝑥 − 𝑥‖ ̄ 𝐻𝑠 =
B. Additional loss functions for AEs A wide range of variations on the training strategy have been proposed to improve representations. In the context of image generation, latent spaces are often trained using perceptual losses (Dieleman, 2025). Indeed, the availability of pre-trained networks having an outstanding correlation with human perception (Zhang, Isola, Efros, Shechtman and Wang, 2018) has provided the community with a useful toolkit to model high-level features. These perceptual losses have proved effective to use in the context of image autoencoders (Pihlgren, Sandin and Liwicki, 2020). Moreover, perceptual features have been used to ensure the consistency between input and output images along the intermediate hidden representations (Hou, Shen, Sun and Qiu, 2017). In the context of video generation, additional losses have been implemented to train the autoencoder (HaCohen, Chiprut, Brazowski, Shalem, Moshe, Richardson, Levin, Shiran, Zabari, Gordon, Panet, Weissbuch, Kulikov, Bitterman, Melumian and Bibi, 2024), enforcing the consistency of the L1 distance between the discrete wavelet transforms of the input and the reconstruction. In order to improve time series representations, other features have been leveraged in the context of contrastive learning (Trirat, Shin, Kang, Nam, Na, Bae, Kim, Kim and Lee, 2024). For instance, frequency invariance is improved by minimising the distance between frequency embeddings, as well as between time embeddings (Zhang, Zhao, Tsiligkaridis and Zitnik, 2022). This is done for samples and between samples and augmentations (positive pairs for contrastive learning), which aims at obtaining representations that are generalizable to unseen datasets via fine-tuning. Time-frequency invariance has also been used for TS representations that allow for anomaly detection (Li, Chen, Wen, Xiao and Zhu, 2026). On the other hand, TS2Vec learns representations of each timestep contrastively, by sampling overlapping segments of a TS and matching the vectors within the common subsegment (Yue, Wang, Duan, Yang, Huang, Tong and Xu, 2022). Crucially, their temporal and instance-wise contrastive losses are computed at multiple granularity levels via iterative max pooling. In addition, learnt shapelets have also been used Carvajal Reyes and Tobar: Preprint submitted to Elsevier
(1 + |𝜉|2 )𝑠 | [𝑥 − 𝑥](𝜉 ̄ 𝑘 )|2 .
𝑘=1 𝜉𝑘 ∈𝐵𝑘
𝑖1 ,…,𝑖𝑘 ≤𝑑
which follows the notation of Kidger et al. (2019).
𝐾 ∑ ∑
Since, by construction, we have that ∑ ∑ 𝑚2𝑘 | [𝑥− 𝑥](𝜉 ̄ 𝑘 )|2 ≤ (1+|𝜉|2 )𝑠 | [𝑥− 𝑥](𝜉 ̄ 𝑘 )| , 𝜉𝑘 ∈𝐵𝑘
𝜉𝑘 ∈𝐵𝑘
using the linearity of , we obtain that 𝐾 ∑ ∑
𝑚2𝑘 | [𝑥](𝜉𝑘 ) − [𝑥](𝜉 ̄ 𝑘 )|2 ≤ ‖𝑥 − 𝑥‖ ̄ 𝐻𝑠 .
𝑘=1 𝜉𝑘 ∈𝐵𝑘
∑ ∑ 2 The inequality ‖𝑥 − 𝑥‖ ̄ 𝐻𝑠 ≤ 𝐾 𝜉𝑘 ∈𝐵𝑘 𝑀𝑘 | [𝑥](𝜉𝑘 ) − 𝑘=1 2 [𝑥](𝜉 ̄ 𝑘 )| follows analogously. We conclude that ,𝑠 is comparable to ‖ ⋅ ‖𝐻 𝑠 .
D. Proof of Proposition 3 Proof. By the triangle inequality, we have: 𝑊2(𝑇 ) (𝑝data , 𝑝̃1 ) ≤ 𝑊2(𝑇 ) (𝑝data , 𝑝̄1 ) + 𝑊2(𝑇 ) (𝑝̄1 , 𝑝̃1 ) for 𝑝̄1 = (◦)# 𝑝data . On the one hand, since the Wasserstein distance takes the minimum value of ∫ 𝑑𝑇 (𝑥1 , 𝑥2 )2 𝜋(𝑥1 , 𝑥2 ), in particular it holds that for the coupling given by taking 𝑥 ∼ 𝑝data and 𝑥̃ = ((𝑥)): 𝑊2(𝑇 ) (𝑝1 , 𝑝̄1 ) ≤ 𝔼𝑥∼𝑝data [𝑑𝑇 (𝑥1 , 𝑥)] ̃ 1∕2 , which corresponds to the expected value of the transformconsistency loss. For the second term of the right-hand side, we have ( )1∕2 𝑊2(𝑇 ) (𝑝̄1 , 𝑝̃1 ) ≤ 𝐿,𝑇 𝑊2 (𝑝𝑧̃ , 𝑝𝑧 ) ≤ 𝐿,𝑇 𝑒1+2𝐿𝜃 𝐻(𝑢𝜃 ) where we have used Albergo and Vanden-Eijnden (2022) for the second inequality, by assuming the flow and decoder to be Lipschitz.
E. The signature-consistency loss as a kernel In this section, we highlight the connection between the signature-consistency loss and signature kernels.
Page 13 of 18
Spectral alignment for time series latent flows
E.1. Unparameterised paths
E.3. Proof of proposition 5
This signature kernel, as introduced in Eq. 5 will be welldefined if
Proof. Let 𝑤̃ 𝑘 = (1 + 𝑎2 𝑘2 )𝑠 for 𝑎 > 0 and 𝑠 > 0, and let ∑ 𝐶 𝑘 𝑤̃ 𝑘 𝐶 > 0. The series ∞ 𝑘=0 (𝑘!)2 is known to converge. Indeed, the ratio test implies that if
∞ ∑
‖𝑆(𝑥)‖2𝑤 =
𝑤𝑘 ‖𝑆(𝑥)(𝑘) ‖2𝑉 ⊗𝑘 < ∞,
𝑘=0
∑ 𝐶 𝑘 𝑤𝑘 which holds when ∞ 𝑘=0 (𝑘!)2 is a convergent series for any 𝐶 > 0. Denote by ∼𝑡 the tree-like equivalence relation for absolutely continuous paths, i.e., 𝑥 and 𝑦 will be tree-like equivalent if there exists a time reparametrisation 𝜏(⋅) such that 𝑥(𝑡) = 𝑦(𝜏(𝑡)). Since 𝑆(𝑥) = 𝑆(𝑦) if and only if 𝑥 ∼𝑡 𝑦 (Boedihardjo, Geng, Lyons and Yang, 2016), the signature kernel will be defined for elements of the quotient space 𝐶0,𝑝 ∕ ∼𝑡 , where 𝐶0,𝑝 is the subspace of 𝐶𝑝 (𝑉 ) such that paths start at 0. This space is often referred to as the set of unparameterised paths, and its elements are denoted by [𝑥]. Note that, even though 𝑘𝑤 is defined for unparameterised paths, it can be extended to continuous paths with bounded 1variation (Cass and Salvi, 2024). Consequently, the notation 𝑘𝑤 (𝑥, 𝑦) can be adopted.
E.2. RKHS and the reproducing property This section is based on Cass and Salvi (2024). Recall that a Hilbert space of functions defined over is said to be a reproducing kernel Hilbert space (RKHS) if 𝑓 ↦ 𝑓 (𝑥) is a continuous linear functional, that is, ∃ 𝐶𝑥 ≥ 0 such that |𝑓 (𝑥)| ≤ 𝐶𝑥 ‖𝑓 ‖ ,
∀𝑓 ∈ .
The Riesz representation theorem guarantees the existence of a unique functional 𝑘(𝑥, ⋅) in the RKHS satisfying ⟨𝑘(𝑥, ⋅), 𝑓 ⟩ = 𝑓 (𝑥),
∀𝑓 ∈ , ∀𝑥 ∈ .
(6)
Equation (6) is referred to as the reproducing property. This then defines a kernel 𝑘(𝑥, 𝑦) = ⟨𝑘(𝑥, ⋅), 𝑘(𝑦, ⋅)⟩ . Conversely, thanks to the Moore-Aronszajn theorem (Aronszajn, 1950), given a positive semi-definite kernel, i.e., such that 𝑘 ∶= (𝑘(𝑥𝑖 , 𝑥𝑗 ))𝑖,𝑗 is a positive semi-definite matrix for any 𝑛 ∈ ℕ and any 𝑥1 , … , 𝑥𝑛 ∈ , there exists a unique RKHS such that the reproducing property in Equation (6) holds. Signatures can be used to define such spaces. Given 𝑣 = (𝑣1 , … , 𝑣𝑘 ) and 𝑣̃ = (𝑣̃1 , … , 𝑣̃𝑘 ) in 𝑉 ⊗𝑘 and a weight function 𝑤 ∶ ℕ ∪ {0} → ℝ+ , the 𝑤-inner product ⟨⋅, ⋅⟩𝑤 over the tensor algebra 𝑇 (𝑉 ) is given by ⟨𝑣, 𝑣⟩ ̃ 𝑤 ∶=
∞ ∑
𝑤𝑘 ⟨𝑣𝑘 , 𝑣̃𝑘 ⟩𝑉 ⊗𝑘 ,
| 𝐶 𝑘+1 𝑤(𝑘 ̃ + 1) 𝑘! || | lim | ⋅ 𝑘 | < 1, 𝑘→∞ | 𝐶 𝑤̃ 𝑘 || | (𝑘 + 1)! then the series converges absolutely. Since ( )𝑠 ̃ + 1) 𝑘! 1 + 𝑎2 (𝑘 + 1)2 𝐶 𝑘+1 𝑤(𝑘 𝐶 = (𝑘 + 1)! 𝐶 𝑘 𝑤̃ 𝑘 (𝑘 + 1)2 1 + 𝑎2 𝑘2 ( )𝑠 2𝑘 + 1 𝐶 1+ = (𝑘 + 1)2 𝑘2 ⟶𝑘→∞ 0, ∑∞
𝐶 𝑘 𝑤̃ 𝑘 𝑘=0 (𝑘!)2 converges absolutely. This, in particular, holds for 𝑎 = 2𝑘1 . Consemax
this is indeed the case. Then, for any
quently, by Lemma 2.1.10 in Cass and Salvi (2024), 𝑘𝑤 is a positive semi-definite kernel. Then the Moore-Aronszajn theorem (Aronszajn, 1950) implies that there is a unique RKHS such that 𝑘𝑤 has the reproducing property.
F. Datasets and experimental details Following the experimental setup described in Section 6, we provide details on the datasets, running times and parameters in Table 3. Times are reported using an NVIDIA GeForce RTX3090 GPU. The hidden dimensionality of the latent space was set to 64 for high-dimensional sets, and 32 for HEPC. The last linear layer of the encoder projects the output of the convolutions to a latent space of 128 dimensions (64 in the case of HEPC). The base autoencoders were trained for 500 epochs. While training the AE for considerably longer yielded cleaner samples, the spectral alignment achieves this sample quality with a very quick fine-tuning stage as opposed to relying on longer training times. Model details are provided in Appendix F, Table 4. The encoder-decoder pair architecture corresponds to mirrored convolutional neural networks with a stride of 2, which halves the temporal resolution after each of the two layers. The hidden dimensionality is set to half the latent space dimensionality, and the models were trained for 500 epochs with a learning rate of 0.0001. The velocity field of the flow matching model is parameterised via a UNET whose blocks are multilayer perceptrons. The learning rate used is 0.001, and the time is included using a learnable embedding of the same dimensionality as the latent codes. The batch size was set to 128 for both latent space and flow training.
𝑘=0
where ⟨𝑣𝑘 , 𝑣̃𝑘 ⟩𝑉 ⊗𝑘 is the canonical Hilbert-Schmidt inner product ⟨𝑣, 𝑣⟩ ̃ 𝑉 ⊗𝑘 =
𝑘 ∏ ⟨𝑣𝑖 , 𝑣̃𝑖 ⟩𝑉 . 𝑖=1
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
G. The exponent of the Sobolev weighting Our weighting scheme is inspired by the theoretical connections with the Sobolev norms outlined in Section 5.2. Throughout our experiments, we set the Sobolev exponent to 1. As can be seen in Figure 11, only a slight increase in the Page 14 of 18
Spectral alignment for time series latent flows Table 3 Experimental details for the empirical evaluation of the transform-consistency losses. We present each dataset and its dimensionality. This, in turns, modifies the dimensionality of their transformed inputs, which are detailed along with the depths/decomposition levels used. The training times for each stage are also provided. Dataset HEPC
Exchange rates
Weather
Dim. 1
8
14
N. train
Variant
Levels (dim.)
8242
Base Fourier Wavelet Signature
- (1000) 5 (1040) 5 (62)
5088
Base Fourier Wavelet Signature
- (1000 × 8) 5 (1040 × 8) 3 (584)
8340
Base Fourier Wavelet Signature
- (1000 × 14) 5 (1040 × 14) 3 (2954)
Train time 4m 5s
3m 2s
5m 12s
Fine-tuning 14s 34s 6m 35s 9s 29s 3m 54s 17s 37s 3m 6s
flow training
Total time
5m 34s
9m 39s 9m 53s 10m 13s 16m 14s
3m 24s
6m 26s 6m 35s 6m 55s 10m 20s
2m 52s
8m 4s 8m 21s 8m 21s 11m 10s
Table 4 Specifications of epochs and latent and hidden dimensionality for each dataset. The abbreviation “dim.” stands for dimensionality and “params.” for parameters. The number of epochs used to train the laten flow models is shown as “FM epochs”. Dataset
Dim.
Latent dim.
Inner dim. flow
Total params.
FM epochs
HEPC
1
64
258
2.021 M
500
Exchange rates
8
128
512
8.037 M
500
Weather
14
128
512
8.037 M
250
marginal score, which corresponds to the absolute difference between the real and generated marginal distributions, can be seen as 𝑠 increases. By contrast, while the discriminative score does not vary significantly for the HEPC and exchange rates datasets, less improvement is observed for 𝑠 = 2.5 for the linear transforms applied to the Weather dataset. Overall, we propose 𝑠 = 1 as the standard choice. The performance of the signature variants for HEPC and exchange datasets is discussed in Section H.
H. Remarks on the use of signatures The signature transform is unique with respect to the Fourier and Wavelet transforms, in the sense that it is an infinite series of coefficients. High-dimensional series may prove too costly for the use of the signature transform as a transform-consistency loss when intended to be used with more depth in high-dimensional datasets. For instance, for the Weather dataset (of dimensionality 14), we are only able to use up to a signature depth of 3 before this representation becomes too large for a consumer-level GPU. While the dimensionality of log-signatures is lower than that of signatures, their computation in practice relies on computing full signatures first, thus not alleviating the GPU memory problem. As reported in Appendix F, fine-tuning with the signature loss entails a significantly higher cost when compared to using linear transforms such as Fourier Carvajal Reyes and Tobar: Preprint submitted to Elsevier
or wavelet, linked to their computation (and their gradients) more than to the dimensionality of the vectors themselves. The resolution-free nature of the signature transform has the advantage that its dimension does not increase as the number of measurements of a signal increases, but it also means that the signature may end up being an underrepresentation of the signal and not capture the high-level behaviour of it if the depth is not enough. While the signature alignment presents a similar improvement to that of the linear transforms for the weather dataset (14 dimensions), it presents less improvement for the HEPC and exchange rates datasets. For the former, as shown in Table 5, for which the original signals are one-dimensional, applying the signature transform to the time-augmented path leads to dimensionalities that can be insufficient to characterise the local behaviour. The performance does improve for depth 7, but the long alignment times that this entails make it impractical. While this is a limitation of this particular variant, the signature alignment might still prove useful for signals with high dimensionality, such as the weather dataset, where even a depth of 3 encodes sufficiently useful information for alignment.
CRediT authorship contribution statement Camilo Carvajal Reyes: Writing - original draft, Writing: review & editing, Methodology, Conceptualization, Formal
Page 15 of 18
Spectral alignment for time series latent flows
exchange rates
0.25 0.00 0.5 1 1.5 2 2.5 s
0.5 1 1.5 2 2.5 s
Fourier
HEPC
exchange rates
weather
0.5 1 1.5 2 2.5 s
0.5 1 1.5 2 2.5 s
0.5 1 1.5 2 2.5 s
weather value
value
HEPC
0.5 1 1.5 2 2.5 s
wavelet
signature
(a) Discriminative score
0.3 0.2 0.1
Fourier
wavelet
signature
(b) Marginal score
Figure 11: Effect of varying the Sobolev exponent 𝑠.
Table 5 Change in fine-tuning times (25 epochs), dimensionality, and discriminative score for varying signature transform depth as observed for the HEPC dataset. The abbreviation disc. stands for the discriminative score, and the 0 before the decimal is omitted from the discriminative score results. depth
3
4
5
6
7
time dim disc.
2m 33s 14 .21±.13
4m 15s 30 .17±.12
6m 35s 62 .18±.14
9m 29s 126 .17±.12
12m 50s 254 .12±.11
analysis, Visualisation. Felipe Tobar: Writing: review & editing, Writing, Validation, Formal analysis, Supervision.
References Alaa, A., Chan, A.J., Schaar, M.v.d., 2021. Generative time-series modeling with Fourier flows, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id=PpshD0AXfA. Albergo, M.S., Vanden-Eijnden, E., 2022. Building normalizing flows with stochastic interpolants, in: The Eleventh International Conference on Learning Representations. Aronszajn, N., 1950. Theory of reproducing kernels. Transactions of the American mathematical society 68, 337–404. Barancikova, B., Huang, Z., Salvi, C., 2024. SigDiffusions: Score-based diffusion models for time series via log-signature embeddings, in: International Conference on Learning Representations. URL: https: //openreview.net/forum?id=Y8KK9kjgIK. Biloš, M., Rasul, K., Schneider, A., Nevmyvaka, Y., Günnemann, S., 2023. Modeling temporal data as continuous functions with stochastic process diffusion, in: International Conference on Machine Learning, PMLR. pp. 2452–2470. URL: https://proceedings.mlr.press/v202/bilos23a.html. ISSN: 2640-3498. Boedihardjo, H., Geng, X., Lyons, T., Yang, D., 2016. The signature of a rough path: uniqueness. Advances in Mathematics 293, 720–737. Cass, T., Salvi, C., 2024. Lecture notes on rough paths and applications to machine learning. arXiv preprint arXiv.2404.06583 URL: http: //arxiv.org/abs/2404.06583. Cleveland, R.B., Cleveland, W.S., McRae, J.E., Terpenning, I., 1990. STL: A seasonal-trend decomposition. Journal of Official Statistics 6, 3–73. Coddington, E.A., Levinson, N., Teichmann, T., 1956. Theory of Ordinary Differential Equations. volume 9. American Institute of Physics. URL: https://doi.org/10.1063/1.3059875. Cooley, J.W., Tukey, J.W., 1965. An algorithm for the machine calculation of complex Fourier series. Mathematics of computation 19, 297–301. Crabbé, J., Huynh, N., Stanczuk, J.P., Schaar, M.V.D., 2024. Time series diffusion in the frequency domain, in: International Conference on Machine Learning, PMLR. pp. 9407–9438. URL: https://proceedings. mlr.press/v235/crabbe24a.html.
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Czarnecki, W.M., Osindero, S., Jaderberg, M., Swirszcz, G., Pascanu, R., 2017. Sobolev training for neural networks. Advances in neural information processing systems 30. Dao, Q., Phung, H., Nguyen, B., Tran, A., 2023. Flow matching in latent space. arXiv preprint arXiv:2307.08698 URL: http://arxiv.org/abs/ 2307.08698. Daubechies, I., 1988. Orthonormal bases of compactly supported wavelets. Communications on pure and applied mathematics 41, 909–996. Dieleman, S., 2025. Generative modelling in latent space. URL: https: //sander.ai/2025/04/15/latents.html. online. Accessed on July 07, 2025. El-Gazzar, A., Gerven, M.v., 2025. Probabilistic forecasting via autoregressive flow matching. arXiv preprint arXiv:2503.10375 URL: http: //arxiv.org/abs/2503.10375. Esser, P., Rombach, R., Ommer, B., 2021. Taming transformers for high-resolution image synthesis, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12873–12883. URL: https://openaccess.thecvf.com/content/CVPR2021/html/Esser_Taming_ Transformers_for_High-Resolution_Image_Synthesis_CVPR_2021_paper. html?ref=.
Feng, S., Miao, C., Zhang, Z., Zhao, P., 2024. Latent diffusion transformer for probabilistic time series forecasting, in: AAAI Conference on Artificial Intelligence, pp. 11979–11987. URL: https://ojs.aaai.org/index.php/ AAAI/article/view/29085, doi:10.1609/aaai.v38i11.29085. number: 11. Fons, E., Sztrajman, A., El-Laham, Y., Ferrer, L., Vyetrenko, S., Veloso, M., 2025. LSCD: Lomb–scargle conditioned diffusion for time series imputation, in: Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J. (Eds.), Proceedings of the 42nd International Conference on Machine Learning, PMLR. pp. 17411–17436. URL: https://proceedings.mlr.press/v267/fons25a.html. Gabor, D., 1946. Theory of communication. Journal of the Institution of Electrical Engineers 93, 429–457. Galib, A.H., Tan, P.N., Luo, L., 2024. FIDE: Frequency-inflated conditional diffusion model for extreme-aware time series generation, in: Advances in Neural Information Processing Systems, pp. 114434–114457. Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y., 2014. Generative adversarial nets, in: Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., Weinberger, K. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. p. 2672–2680. URL: https://proceedings.neurips.cc/paper_ files/paper/2014/file/f033ed80deb0234979a61f95710dbe25-Paper.pdf. Haar, A., 1909. Zur theorie der orthogonalen funktionensysteme. GeorgAugust-Universitat, Gottingen. HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., Panet, P., Weissbuch, S., Kulikov, V., Bitterman, Y., Melumian, Z., Bibi, O., 2024. LTX-video: Realtime video latent diffusion. arXiv preprint arXiv.2501.00103 URL: http://arxiv.org/abs/2501.00103. Hairer, E., Wanner, G., Nørsett, S.P., 1993. Solving ordinary differential equations I: Nonstiff problems. Springer. Ho, J., Jain, A., Abbeel, P., 2020. Denoising diffusion probabilistic models, in: Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 6840–6851. URL: https://proceedings.neurips.cc/paper/2020/ hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html.
Page 16 of 18
Spectral alignment for time series latent flows Holderrieth, P., Erives, E., 2025. An introduction to flow matching and diffusion models. arXiv preprint arXiv:2506.02070 URL: http://arxiv. org/abs/2506.02070. Hou, X., Shen, L., Sun, K., Qiu, G., 2017. Deep feature consistent variational autoencoder, in: IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1133–1141. URL: https://ieeexplore.ieee.org/ abstract/document/7926714, doi:10.1109/WACV.2017.131. Hu, Y., Wang, X., Ding, Z., Wu, L., Zhang, H., Li, S.Z., Wang, S., Zhang, J., Li, Z., Chen, T., 2024. FlowTS: Time series generation via rectified flow. arXiv preprint arXiv.2411.07506 URL: http://arxiv.org/abs/2411. 07506. Khodakarami, S., Oommen, V., Bora, A., Karniadakis, G.E., 2026. Mitigating spectral bias in neural operators via high-frequency scaling for physical systems. Neural Networks 193, 108027. URL: https: //www.sciencedirect.com/science/article/pii/S0893608025009074, doi:https://doi.org/10.1016/j.neunet.2025.108027. Kidger, P., Bonnier, P., Perez Arribas, I., Salvi, C., Lyons, T., 2019. Deep signature transforms, in: Advances in Neural Information Processing Systems, Curran Associates, Inc. URL: https://proceedings.neurips.cc/paper_files/paper/2019/hash/ d2cdf047a6674cef251d56544a3cf029-Abstract.html.
Kidger, P., Lyons, T., 2020. Signatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU, in: International Conference on Learning Representations. URL: https: //openreview.net/forum?id=lqU2cs3Zca. Kolle, O., 2024. Documentation of the weather station on top of the roof of the institute building of the Max-Planck-Institute for biogeochemistry. URL: https://www.bgc-jena.mpg.de/wetter/. Kollovieh, M., Lienen, M., Lüdke, D., Schwinn, L., Günnemann, S., 2025. Flow matching with Gaussian process priors for probabilistic time series forecasting, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id=uxVBbSlKQ4&. Lai, G., Chang, W.C., Yang, Y., Liu, H., 2018. Modeling long- and shortterm temporal patterns with deep neural networks, in: International ACM SIGIR Conference on Research & Development in Information Retrieval, pp. 95–104. URL: https://doi.org/10.1145/3209978.3210006, doi:10.1145/3209978.3210006. Lamb, A., Dumoulin, V., Courville, A., 2016. Discriminative regularization for generative models. arXiv preprint arXiv:1602.03220 . Lee, G., Gommers, R., Waselewski, F., Wohlfahrt, K., O’Leary, A., 2019. Pywavelets: A Python package for wavelet analysis. Journal of Open Source Software 4, 1237. Li, Y., Chen, Z., Wen, Z., Xiao, X., Zhu, M., 2026. Time-frequency contrastive learning with context modeling for time series anomaly prediction. Neural Networks 197, 108494. URL: https://www.sciencedirect.com/ science/article/pii/S0893608025013759, doi:https://doi.org/10.1016/j. neunet.2025.108494. Liang, Z., Zhang, J., Liang, C., Wang, H., Liang, Z., Pan, L., 2023. A shapeletbased framework for unsupervised multivariate time series representation learning. Proceedings of the VLDB Endowment 17, 386–399. URL: https://doi.org/10.14778/3632093.3632103. Lim, H., Kim, M., Park, S., Park, N., 2023. Regular time-series generation using SGM. arXiv preprint arXiv.2301.08518 URL: http://arxiv.org/ abs/2301.08518. Lipman, Y., Chen, R.T.Q., Ben-Hamu, H., Nickel, M., Le, M., 2022. Flow matching for generative modeling, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id= PqvMRDCJT9t. Liu, X., Gong, C., Liu, Q., 2022. Flow straight and fast: Learning to generate and transfer data with rectified flow, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id= XVjTT1nw5z. Lyons, T.J., 1998. Differential equations driven by rough signals. Revista Matemática Iberoamericana 14.2 . Makhzani, A., Shlens, J., Jaitly, N., Goodfellow, I., Frey, B., 2015. Adversarial autoencoders. arXiv preprint arXiv:1511.05644 . Mallat, S., 1999. A wavelet tour of signal processing. Academic Press. Meyer, Y., 1992. Wavelets and Operators. Cambridge University Press.
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Naiman, I., Berman, N., Pemper, I., Arbiv, I., Fadlon, G., Azencot, O., 2024. Utilizing image transforms and diffusion models for generative modeling of short and long time series. Advances in Neural Information Processing Systems 37, 121699–121730. Perko, L., 2013. Differential Equations and Dynamical Systems. Springer Science & Business Media. Pihlgren, G.G., Sandin, F., Liwicki, M., 2020. Improving image autoencoder embeddings with perceptual loss, in: International Joint Conference on Neural Networks, pp. 1–7. URL: https://ieeexplore.ieee.org/ abstract/document/9207431, doi:10.1109/IJCNN48605.2020.9207431. ISSN: 2161-4407. Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., Courville, A., 2019. On the spectral bias of neural networks, in: International conference on machine learning, PMLR. pp. 5301–5310. Rasul, K., Seward, C., Schuster, I., Vollgraf, R., 2021. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting, in: International Conference on Machine Learning, PMLR. pp. 8857–8868. URL: https://proceedings.mlr.press/v139/rasul21a. html. ISSN: 2640-3498. Repository, U.M.L., 2024. Household electric power consumption. URL: https://www.kaggle.com/datasets/uciml/ electric-power-consumption-data-set. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B., 2022. High-resolution image synthesis with latent diffusion models, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10674–10685. doi:10.1109/CVPR52688.2022.01042. ISSN: 2575-7075. Ronneberger, O., Fischer, P., Brox, T., 2015. U-net: Convolutional networks for biomedical image segmentation, in: International Conference on Medical image computing and computer-assisted intervention, Springer. pp. 234–241. Shen, L., Chen, W., Kwok, J., 2023. Multi-resolution diffusion models for time series forecasting, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id=mmjnr0G8ZY. Simon-Gabriel, C.J., Schölkopf, B., 2018. Kernel distribution embeddings: Universal kernels, characteristic kernels and kernel metrics on distributions. Journal of Machine Learning Research 19, 1–29. Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., Ganguli, S., 2015. Deep unsupervised learning using nonequilibrium thermodynamics, in: 32nd International Conference on Machine Learning, PMLR. pp. 2256–2265. URL: https://proceedings.mlr.press/v37/sohl-dickstein15.html. ISSN: 1938-7228. Song, Y., Ermon, S., 2019. Generative modeling by estimating gradients of the data distribution, in: Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 11918–11930. URL: https://proceedings.neurips.cc/paper/2019/hash/ 3001ef257407d5a371a96dcd947c7d93-Abstract.html. Tong, A., Fatras, K., Malkin, N., Huguet, G., Zhang, Y., Rector-Brooks, J., Wolf, G., Bengio, Y., 2023. Improving and generalizing flow-based generative models with minibatch optimal transport. Transactions on Machine Learning Research URL: https://openreview.net/forum?id= CD9Snc73AW. Trirat, P., Shin, Y., Kang, J., Nam, Y., Na, J., Bae, M., Kim, J., Kim, B., Lee, J.G., 2024. Universal time-series representation learning: A survey. arXiv preprint arXiv.2401.03717 URL: http://arxiv.org/abs/2401.03717. Yoon, J., Jarrett, D., van der Schaar, M., 2019. Time-series generative adversarial networks, in: Advances in Neural Information Processing Systems, pp. 5508–5518. URL: https://proceedings.neurips.cc/paper/ 2019/hash/c9efe5f26cd17ba6216bbe2a7d26d490-Abstract.html?ref=https: //githubhelp.com.
Yu, J., Li, X., Koh, J.Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., Wu, Y., 2022. Vector-quantized image modeling with improved VQGAN, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id=pfNyExj7z2. Yuan, X., Qiao, Y., 2023. Diffusion-TS: Interpretable diffusion for general time series generation, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id=4h1apFjO99.
Page 17 of 18
Spectral alignment for time series latent flows Yue, Z., Wang, Y., Duan, J., Yang, T., Huang, C., Tong, Y., Xu, B., 2022. TS2vec: Towards universal representation of time series, in: AAAI Conference on Artificial Intelligence, pp. 8980–8987. URL: https://ojs.aaai.org/index.php/AAAI/article/view/20881, doi:10.1609/ aaai.v36i8.20881. number: 8. Zhang, J., Dai, Q., 2022. Latent adversarial regularized autoencoder for highdimensional probabilistic time series prediction. Neural Networks 155, 383–397. URL: https://www.sciencedirect.com/science/article/pii/ S0893608022003252, doi:https://doi.org/10.1016/j.neunet.2022.08.025. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O., 2018. The unreasonable effectiveness of deep features as a perceptual metric, in: IEEE Conference on Computer Vision and Pattern Recognition, pp. 586–595. URL: https://openaccess.thecvf.com/content_cvpr_2018/ html/Zhang_The_Unreasonable_Effectiveness_CVPR_2018_paper.html. Zhang, X., Zhao, Z., Tsiligkaridis, T., Zitnik, M., 2022. Self-supervised contrastive pre-training for time series via time-frequency consistency, in: Advances in Neural Information Processing Systems, pp. 3988–4003. URL: https://papers.nips.cc/paper_files/paper/2022/ file/194b8dac525581c346e30a2cebe9a369-Paper-Conference.pdf.
Carvajal Reyes and Tobar: Preprint submitted to Elsevier
Page 18 of 18