Heavy-Tailed Flow Matching via Random Clocks Zhouhao Yang*
Yezhen Wang†
Kenji Kawaguchi†
‡
Vladimir Braverman
Haoyang Cao*
arXiv:2607.13841v1 [cs.LG] 15 Jul 2026
Abstract Heavy-tailed data arise in many domains where rare events carry disproportionate importance, such as imbalanced image datasets, financial returns, and weather extremes. Standard diffusion and flow-matching models typically begin from Gaussian noise or Gaussian source distributions, which yield tractable training targets but provide a poor inductive match for heavy-tailed data. We propose Heavy-Tailed Flow Matching via Random Clocks (HTFM), a framework that portrays heavy-tailed sources as mixtures of clock-conditioned Gaussian sources. Conditioning on a given clock path, the source distribution and flow are Gaussian; marginalizing over the clock gives a Gaussian scale mixture covering Gaussian, α-stable, and Student-t families. To make the clock-conditioned vector field practical, we encode the path-valued clock using truncated logsignature features, allowing the velocity field to adapt to the realized conditional space with negligible overhead. Empirically, on 2D imbalanced α-stable mixtures, CIFAR10-LT, and HRRR weather fields, HTFM improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and competitive heavy-tailed baselines, while retaining the low-NFE sampling advantage of flow matching. Moreover, the random-clock formulation further provides a practical tail-control interface: by varying only the clock law or tail parameter, the same architecture can calibrate the “heaviness” of generated tails across different distribution families.
1
Introduction
In many real-world applications, rare or atypical observations carry disproportionately large impact relative to typical ones [8, 40]: hurricanes and heatwaves in operational weather and climate models [41, 12, 37, 15], the tail of asset returns in financial risk modeling [9, 13, 10], and minority classes in long-tailed image datasets [6, 32, 46]. A generative model is useful in such regimes only insofar as it reproduces these rare events with the right frequency and magnitude. This requirement is difficult for two related reasons. First, extreme observations are rare by construction, so empirical training data contains few samples from the relevant tail regions [15]. Second, many deep generative modeling pipelines begin from simple light-tailed randomness: variational autoencoders and generative adversarial networks commonly use Gaussian or uniform latent priors [22, 14], while state-of-the-art diffusion and flow-matching models typically use Gaussian noising processes or Gaussian source distributions [18, 43, 29]. When the data distribution has substantially heavier tails, this light-tailed source is a poor inductive match: the model must learn tail behavior through the learned transport, which can make rare events difficult to reproduce accurately. A natural response is to replace the Gaussian source or noise with a heavy-tailed alternative. In diffusion and score-based modeling, Lévy-Itô models [45, 39] use α-stable noise but require ∗
Department of Applied Mathematics and Statistics, Johns Hopkins University. School of Computing, National University of Singapore. ‡ Department of Computer Science, Johns Hopkins University. †
1
nonlocal fractional-score machinery and are tied to a particular stable family. Denoising Lévy probabilistic models (DLPM) [42] adapt the DDPM [18] construction to α-stable noise and use a denoising objective with robust aggregation over heavy-tailed randomness. Student-t-based diffusion and flow models [36] use Student-t kernels with a γ-divergence training loss and are restricted to the finite-variance regime ν > 2. Other constructions specialize the source or noise to heavy-tailed families such as Cauchy noise [28] or finite-activity jump diffusions [4]. Related tail limitations have also been studied for normalizing flows [19, 17]. These approaches demonstrate the value of heavy-tailed generative mechanisms, but they are typically built around a specific tail family and the corresponding score, kernel, sampler, or loss aggregation. In practice, the appropriate tail mechanism is rarely known in advance; a method built for one heavy-tailed family may be misspecified for another. This leaves a gap for flow matching [29, 31, 1]. Flow matching learns an ordinary differential equation (ODE) that transports a source distribution to the data distribution by regressing tractable velocity targets along prescribed probability paths, rather than simulating a long reverse diffusion chain. This often enables high-quality sampling with a substantially lower number of function evaluations (NFE), and rectified-flow and optimal-transport variants further sharpen this sample-efficiency advantage [30, 44, 38, 24]. A heavy-tailed analogue of flow matching should therefore preserve the low-NFE advantages of flow matching while allowing source laws with heavier tails than Gaussians. Existing heavy-tailed diffusion and flow constructions make important progress, but their targets or divergences are usually derived separately for each chosen heavy-tailed family. In this paper, we propose Heavy-Tailed Flow Matching via Random Clock (HTFM). A random clock is a nondecreasing path-valued latent variable that controls the conditional covariance of the source distribution. Conditioning on a given clock path, the source distribution and the corresponding affine flow between source and data are Gaussian, so the endpoint-conditioned velocity target has the same algebraic form as in Gaussian conditional flow matching. After marginalizing over the clock, the source and flow can be heavy-tailed. The random-clock viewpoint gives a unified construction for several representative tail mechanisms. A deterministic clock recovers vanilla Gaussian flow matching; an α-stable subordinator clock yields α-stable source marginals; and an inverse-gamma clock recovers Student-t marginals for any ν > 0, including the infinite-variance regime ν ≤ 2 not covered by finite-variance Student-t constructions [36]. Across these choices, the endpoint-conditioned target is derived in the same clock-conditioned Gaussian space, avoiding family-specific score formulas, direct use of generally unavailable α-stable densities, or Student-t-specific kernels. The remaining difficulty is that the pathwise optimal vector field is determined by the realized clock path, an infinite-dimensional object. Passing the full path into the velocity field would amount to an operator-learning problem [25] and would undermine the practical simplicity of flow sampling. We instead characterize the clock path by a finite-dimensional truncated logsignature feature [33, 7]. Path signatures enable universal finite-dimensional representations of stream-valued data and have been used as kernels, feature maps, and differentiable neural components [27, 23, 20, 21, 35]. This lets the velocity field adapt to the realized clock at negligible training and sampling cost. Empirically, our experiments support three findings about the random clock: (i) heavy-tailed clocks improve over the Gaussian-clock variant across benchmarks; (ii) different clock families behave differently even at matched polynomial tail-decay rates, indicating that the clock family is a meaningful design knob beyond tail heaviness; and (iii) explicit clock conditioning of the velocity field provides gains beyond simply sampling from a heavy-tailed source. The contributions of our paper can be summarized as follows: 1. A random-clock framework for heavy-tailed flow matching. We introduce HTFM, a flow-matching framework in which a path-valued random clock controls the conditional source covariance. Conditioning on the realized clock gives a Gaussian conditional source and a tractable endpoint-conditioned
2
affine target; marginalizing over clocks recovers Gaussian scale-mixture families, including Gaussian, α-stable, and Student-t sources. 2. Signature-based clock conditioning. We encode the path-valued clock using truncated logsignature features, allowing the neural velocity field to adapt to the realized conditional Gaussian space with negligible additional cost. This provides a finite-dimensional interface for clock-conditioned vector fields. 3. Empirical results. Experiments on 2D imbalanced α-stable mixtures, CIFAR10-LT, and HRRR weather fields show that heavy-tailed clock laws improve over the Gaussian-clock variant and that the clock family remains important even at matched polynomial tail-decay rates. On CIFAR10-LT, HTFM compares favorably with existing heavy-tailed diffusion and ODE-based baselines, especially at low NFE; ablations further show that signature conditioning improves over using a heavy-tailed source alone.
2
Preliminaries
2.1
Flow matching
Flow matching learns a time-dependent vector field whose ODE transports a source distribution p0 to the data distribution pdata . Let Xt ∈ Rd denote the state of the flow at time t ∈ [0, 1], a stochastic process with X0 ∼ p0 and marginal law pt = Law(Xt ). Concretely, consider the ODE flow dXt = ut (Xt ), t ∈ [0, 1], X0 ∼ p0 , (2.1) dt whose marginal density evolves according to the continuity equation. The goal is to choose ut so that p1 = pdata . Given the probability flow (pt )t∈[0,1] and its marginal velocity field ut , the ideal marginal flow-matching (FM) objective is i h 2 (2.2) LFM (θ) = Et∼Unif(0,1) EXt ∼pt uθ (Xt , t) − ut (Xt ) 2 . However, the marginal field ut is usually intractable. Conditional flow matching (CFM) therefore conditions on a data endpoint X1 ∼ pdata and specifies an endpoint-conditioned path pt (· | X1 ) with tractable conditional velocity vt (· | X1 ). The corresponding objective is h i 2 LCFM (θ) = Et∼Unif(0,1) EX1 ∼pdata EXt ∼pt (·|X1 ) uθ (Xt , t) − vt (Xt | X1 ) 2 . (2.3) Under standard L2 -integrability and differentiability assumptions, the conditional and marginal objectives differ only by a θ-independent constant, so their gradients agree [29, 44]. This L2 projection principle is the basic mechanism we preserve in our framework proposed later in Section 3.
2.2
Heavy-tailed noises as Gaussian scale mixtures
Gaussian and heavy-tailed laws are primarily distinguished by their tail decay. Gaussian tails are exponentially small, whereas many heavy-tailed laws exhibit only polynomial decay. Such laws are often better suited to modeling rare but non-negligible large deviations. A convenient latent-Gaussian framework for symmetric heavy-tailed modeling is provided by Gaussian scale mixtures. Let V be a latent scale variable and let G ∼ N (0, Id ) be independent of V . A random vector X ∈ Rd is called a Gaussian scale mixture if it admits a representation of the form X = L(V )G, where L(V ) ∈ Rd×d is a measurable random matrix. Equivalently, X | V ∼ N 0, Σ(V ) , Σ(V ) := L(V )L(V )⊤ . (2.4) Thus, conditioning on the latent variable V , the law remains Gaussian, whereas marginalizing over specific choices of V can recover a heavy-tailed marginal law. The mechanism is simple: 3
rare large values of the latent scale generate broad conditional Gaussians, and the mixture assigns non-negligible probability to large deviations even though every conditional component is light-tailed. Representative examples include symmetric α-stable laws and Student-t laws; their concrete scale-mixture forms are recalled in Appendix A.3. This latent-Gaussian viewpoint suggests a natural strategy for flow matching. Rather than trying to regress directly against a heavy-tailed marginal flow, we condition on the latent mixing object so that the interpolation becomes conditionally Gaussian. We may then derive the conditional flow-matching target in this Gaussian regime and exploit the L2 projection principle from Section 2.1. This conditional-Gaussian perspective is also closely related to [42], which uses a conditioned representation of α-stable noise to build tractable denoising targets for diffusion models. The Gaussian scale-mixture viewpoint is already broad enough to cover many symmetric heavy-tailed laws and will be the focus of the main development in this paper. A broader latentGaussian framework, obtained by allowing a nonzero conditional mean in (2.4) and hence passing from Gaussian scale mixtures to Gaussian mean-variance mixtures, can also accommodate more general families of asymmetric heavy-tailed laws; we defer this extension to Appendix A.4.
3
Heavy-Tailed Flow Matching via Random Clocks
This section develops the random-clock framework for heavy-tailed flow matching; a related diffusion-model extension is given in Appendix B. The latent-Gaussian representation in Section 2.2 suggests replacing the scalar mixing variable V by a path-valued latent object, if we want to work on flow matching. We call this object a random clock. Formally, let D([0, 1]; R+ ) := {τ : [0, 1] → R+ τ is càdlàg and nondecreasing}, equipped with its Borel σ-field. A random clock is a random element on a probability space (ΩT , FT , PT ), T : ΩT → D([0, 1]; R+ ), and we write T[0,1] = (Tt )t∈[0,1] for the full realized clock path. Throughout, T0 = 0 almost surely. The clock is sampled independently of the base randomness used to construct the source and target variables, such as a standard Gaussian G and X1 ∼ pdata . Clock-dependent sources are nevertheless allowed: for example, X0 = L(T[0,1] )G is conditionally Gaussian given the realized clock path but is not independent of the clock after it is formed.
3.1
Clock-conditioned dynamics and objectives
Motivated by Section 2.2, we formulate heavy-tailed flow matching pathwise, that is, conditioning on a realized clock path, and reserve the outer averaging over the random clock for the final objective. Let T = (Tt )t∈[0,1] be a random clock1 , and fix a realization of the full clock path T[0,1] . We write XtT , pTt , and uTt for the clock-conditioned process, its law pTt := Law(XtT ), and its velocity field under this realized clock path. Due to the independence of the data law and the random clock, we write X1T := X1 for the terminal data variable, and pdata (· | T[0,1] ) = pdata . We consider the clock-conditioned ODE dXtT = uTt (XtT ), dt
t ∈ [0, 1],
X0T ∼ pT0 := p0 (· | T[0,1] ),
(3.1)
whose conditional density satisfies the continuity equation ∂t pTt + ∇ · (pTt uTt ) = 0. Equivalently, one may view uTt (·) as the evaluation at T[0,1] of a measurable map ut : Rd × D([0, 1]; R+ ) → Rd . Our goal is to construct, for each realized T[0,1] , a flow from pT0 at time t = 0 to the terminal law pdata at time t = 1. 1
When appearing as a superscript, T indicates the clock path and is not an exponent or transpose.
4
Marginal Noise (Heavy-Tailed) Random Clock Tt
Data
Small Random Clock Large Random Clock
Clock-Conditioned Noise (Light-Tailed)
Clock-Conditioned Time
T0
T1
Original Time
1
0
Figure 1: Overview of random-clock conditioning for heavy-tailed flow matching. The right column shows two realized clock paths: a large clock realization (top) and a small clock realization (bottom). For each fixed clock path T , the source distribution is conditionally Gaussian, and the model learns a clock-conditioned flow from this light-tailed conditional source to the data distribution. Larger clock realizations induce broader conditional source distributions, while smaller clock realizations induce more concentrated ones. After marginalizing over clock paths, the mixture of these conditional Gaussian sources yields a heavy-tailed marginal source distribution. Intuitively, conditioning on a realized clock path locally stretches or compresses physical time. This conditioning moves the heavy-tailedness into the outer randomness of the clock, while the inner pathwise problem becomes conditionally Gaussian and therefore amenable to the usual L2 -projection argument of flow matching. An illustration is provided in Figure 1. Training objective. For brevity, after fixing T[0,1] we write EXtT for expectation under pTt , and EXtT |X T for expectation under Law(XtT | X1T , T[0,1] ). The pathwise marginal quadratic risk 1 and the corresponding heavy-tailed flow-matching objective are h i LHTFM (θ) := Et∼Unif(0,1) ET EXtT ∥uθ (XtT , t, T[0,1] ) − uTt (XtT )∥22 . (3.2) | {z } =: Rtmarg (θ;T )
The risk Rtmarg (θ; T ) is typically intractable because the pathwise marginal velocity uTt is unknown. As in [29, 31], we therefore introduce a tractable endpoint-conditioned formulation: for each fixed realized clock path T[0,1] , denote a conditional path pTt (· | x1 ) := Law(XtT | X1T = x1 ) and a corresponding pathwise conditional velocity field vtT (· | x1 ) : Rd → Rd such that, for a.e. t ∈ [0, 1], pTt (x) =
Z Rd
pTt (x | x1 ) pdata (dx1 ),
pTt (x) uTt (x) =
Z Rd
pTt (x | x1 ) vtT (x | x1 ) pdata (dx1 ).
The corresponding endpoint-conditioned heavy-tailed flow-matching objective writes h i LCHTFM (θ) := Et∼Unif(0,1) ET,X T EXtT |X T ∥uθ (XtT , t, T[0,1] ) − vtT (XtT | X1T )∥22 . 1 1 {z } | =: Rtcond (θ;X1T ,T )
5
(3.3)
(3.4)
The following proposition is the clock-conditioned analogue of the standard conditional– marginal equivalence [29, 31]. The standard identity requires a finite-second-moment marginal velocity and breaks down for heavy-tailed sources; recasting it pathwise in T is what makes the equivalence available in our setting, as discussed after the proposition. Assumptions and proof are in Appendix C.3. Proposition 3.1 (Pathwise conditional–marginal equivalence). Under the integrability and differentiability conditions in Appendix C.3, for a.e. t there exists a nonnegative Ct (T ), independent of θ, such that E Rtcond (θ; X1T , T ) T[0,1] = Rtmarg (θ; T ) + Ct (T ). Consequently, whenever the risks and Ct (T ) are integrable over (t, T ) with t ∼ Unif(0, 1), LCHTFM (θ) = LHTFM (θ) + Et∼Unif(0,1) ET [Ct (T )].
(3.5)
Thus, whenever the additive constant is finite, the two objectives have the same minimizers; when differentiation may be exchanged with the expectations, their gradients in θ are also identical.
Why a clock-conditioned formulation? Without this conditioning, one would attempt to learn the unconditional marginal velocity u∗t through an L2 -risk of the form EXt [∥uθ (Xt , t) − u∗t (Xt )∥22 ], which may fail to be finite under heavy tails, since heavy-tailed laws need only admit moments of order p < 2. In practice, this infinite-variance effect can appear as exploding or highly unstable quadratic training losses, as observed in Lévy-based models such as LIM and addressed explicitly by DLPM through robust aggregation [45, 42]. Replacing the L2 -loss by a generic Lp -loss is not a satisfactory theoretical remedy either, because it destroys the L2 projection structure underlying the conditional–marginal equivalence. HTFM instead fixes the realized clock path before applying flow matching. In this clock-conditioned space, the source is Gaussian and the standard conditional-flow-matching projection identity applies pathwise under Proposition 3.1. The heavy-tailedness is then pushed to the outer clock average, whose finiteness is handled by separate integrability assumptions.
3.2
Signature feature for path conditioning
The objective (3.4) fixes the inner conditional regression space through its dependence on the path-valued variable T[0,1] . Because T[0,1] is infinite-dimensional, it cannot be passed directly to a neural field; we therefore summarize the realized clock path by a finite-dimensional pathsignature feature, taking the truncated logsignature ℓm (T ) defined below. Path signatures originate in rough path theory [33, 7] and have been widely used as universal finite-dimensional encodings of stream-valued data in machine learning, both as differentiable neural-network layers [20, 21] and as kernels and feature maps for sequential data [27, 23]; they are computationally cheap to evaluate yet expressive, separating paths up to tree-like equivalence on suitable compact families (Proposition D.1 in Appendix D.1). Time-augmented signature. Let Ts := (s, Ts ) ∈ R2 , s ∈ [0, 1], be the time-augmented clock path. Its truncated signature of order m is the iterated-integral tensor Z m (≤m) ϕm (T ) := S (T) = 1, dTu1 ⊗ · · · ⊗ dTuk ∈ RDm . (3.6) 0<u1 <···<uk <1
k=1
Time augmentation captures how clock mass is allocated in physical time. For càdlàg clocks observed on a discretization 0 = t0 < t1 < · · · < tN = 1, we linearly interpolate the sampled N time-augmented points (ti , Tti ) i=0 and take the truncated signature of the resulting piecewise linear path; this is the operational meaning of ϕm (T ) used throughout the paper. For a more 6
compact representation we additionally use the truncated logsignature ℓm (T ) := log⊗ ϕm (T ) , computed in the truncated tensor algebra. The dimension reduction is substantial: for q = 2 and m = 3, the raw signature has 15 coordinates whereas the nonconstant logsignature has only 5. Precise jump conventions, dimension formulas, and the lead–lag and piecewise-constant variants are deferred to Appendix D. Network conditioning. The neural velocity field uses ℓm (T ) as an auxiliary conditioning input alongside the scalar time t and the state XtT . In the experiments we use the concat-std variant: ℓm (T ) is standardized using running statistics collected during training, projected to a fixed embedding size by a small learnable linear map, and concatenated with the time embedding before being injected into every residual block of the velocity-field backbone. Replacing the full path argument in (3.4) by ℓm (T ) gives the practical signature-conditioned endpoint objective h h ii 2 T T T T Lsig (θ) := E E E u (X , t, ℓ (T )) − v (X | X ) . (3.7) T T T m θ t∼Unif(0,1) T,X t t t 1 Xt |X CHTFM 2 1
1
Computational overhead. The signature pipeline adds a one-shot logsignature computation of complexity O(N q m ) (with N clock observation points and q = 2, m ≤ 3 in our experiments), plus a small linear projection of ℓm (T ) concatenated with the time embedding. This cost is independent of the data dimension and negligible compared to a single forward pass of the velocity-field backbone on XtT ; no extra denoising or solver step is introduced.
3.3
Flow design and transport interpretation
We now specify the affine flow family used in this paper. Let Σ : D([0, 1]; R+ ) → Sd+ be a measurable map into the cone of symmetric positive semidefinite d × d matrices. Suppose that, conditioned on the realized clock path, X0T ∼ N 0, Σ(T[0,1] ) , and that X0T and X1T are conditionally independent given T[0,1] . Let αtT and βtT be clock-dependent schedules such that, for almost every realized clock path, t 7→ αtT and t 7→ βtT are absolutely continuous and satisfy α0T = 0, β0T = 1, α1T = 1, β1T = 0, βtT > 0 for t ∈ [0, 1). The clock-conditioned affine path is then XtT = αtT X1T + βtT X0T , t ∈ [0, 1]. (3.8) Under (3.8), for each x1 ∈ Rd and each realized clock path T[0,1] , the conditional law of XtT is Gaussian (derivation in Appendix C.2): pTt (· | x1 ) = Law(XtT | X1T = x1 ) = N αtT x1 , (βtT )2 Σ(T[0,1] ) . (3.9) The random clock thus influences the interpolation through the conditional source covariance Σ(T[0,1] ) and through the schedule pair (αtT , βtT ); in both cases the path remains conditionally Gaussian once the clock is fixed. For a.e. t ∈ [0, 1), differentiating (3.8) at fixed clock realization and substituting X0T = (XtT − αtT X1T )/βtT yields the explicit endpoint-conditioned target ! dβtT dβtT dXtT dαtT T T T T dt dt vt (x | x1 ) := E Xt = x, X1 = x1 = T x + (3.10) − T αt x1 . dt dt βt βt This is the same algebraic target as in Gaussian flow matching, with the deterministic schedules replaced by their clock-dependent versions. Training samples are generated by drawing T and then XtT from (3.8); the model regresses uθ (XtT , t, ℓm (T )) toward (3.10) in (3.7). For generation, we sample T , compute ℓm (T ), draw X0T ∼ N (0, Σ(T[0,1] )), and use a numerical ODE solver on dXtT = uθ∗ (XtT , t, ℓm (T )) dt. We now specialize to the straight-line flow used throughout our experiments; a more general signature-linear schedule family is deferred to Appendix D.1. 7
1 −x Straight-line flow. Set αtT = t, βtT = 1 − t. Then (3.10) reduces to vtT (x | x1 ) = x1−t , t∈ [0, 1), and the endpoint-conditioned target has no explicit clock dependence. The clock still enters through the conditional law (3.9) via the source covariance, and hence through the projection defining the pathwise marginal optimum uTt (x) = E[vtT (XtT | X1T ) | XtT = x]. A network that does not observe the clock can therefore only learn a clock-averaged field, whereas the clock-conditioned formulation learns the family of optima indexed by the realized clock path. This straight-line schedule cleanly recovers the heavy-tailed marginal source families used in our experiments, and additionally admits a pathwise Wp -geodesic interpretation under a poptimal endpoint coupling that remains valid in the infinite-second-moment regime p < 2. See Appendix C.4 for the additional theorem.
Heavy-tailed marginals from clock laws. With the straight-line schedule, choosing the clock law and source covariance recovers the named heavy-tailed marginal families used in our experiments. Throughout this paragraph, Xt denotes the clock-marginal variable obtained by T first sampling T and then sampling Xt Tin theT corresponding clock-conditioned space. Equivalently, Law(Xt | X1 = x1 ) = ET Law(Xt | X1 = x1 , T[0,1] ) .2 For the representative examples, R1 we use the area-clock covariance ΣA (T[0,1] ) := 2A(T )Id , where A(T ) := 0 Ts ds. The following theorem states that the same clock-conditioned Gaussian construction can recover three commonly used source families: Gaussian, α-stable, and Student-t. The proof is deferred to Appendix C.5. Theorem 3.2 (Recovered clock marginals). Assume the affine path (3.8), the clock-independent straight-line schedule, and the covariance map ΣA (T[0,1] ). For notational brevity, write pt (· | x1 ) := Law(Xt | X1 = x1 ). For every t ∈ [0, 1) and x1 ∈ Rd , the following identities hold. (i) If Tt = t, then pt (· | x1 ) = N (αt x1 , βt2 Id ), with sub-Gaussian tail decay. (ii) Let α ∈ (0, 2), ρ = α/2, and let T be a ρ-stable subordinator normalized by E[e−sTt ] = exp(−(ρ+1)tsρ /2). Then pt (· | x1 ) is a location-shifted isotropic α-stable law with location αt x1 , scale βt . In particular, its radial tail satisfies P (∥Xt − αt x1 ∥2 > r | X1 = x1 ) = Θ(r−α ), r → ∞. (iii) Let ν > 0. If Vν ∼ InvGamma(ν/2, ν/2) and Tt = tVν , then pt (· | x1 ) is a multivariate Student-t law with location αt x1 , scale βt2 Id , ν degrees of freedom. In particular, its radial tail satisfies P (∥Xt − αt x1 ∥2 > r | X1 = x1 ) = Θ(r−ν ), r → ∞. Beyond clock-independent straight-line schedules. The schedules αtT , βtT may themselves depend on the realized clock path through its truncated signature; concrete signaturelinear and VE-type parameterizations and a near-linear boundary-corrected variant are collected in Appendices D.1, D.2, and A.5. Such clock-dependent schedules generally produce broader Gaussian-mixture sources rather than exact α-stable or Student-t laws, so the named recoveries of Theorem 3.2 are specific to the clock-independent case. A detailed comparison of the random-clock formulation with DLPM, LIM, and Student-t-based heavy-tailed diffusion families, together with a discussion of data-adaptive clock laws, is given in Appendix A.2.
4
Experiments
To assess the effectiveness and generality of the proposed heavy-tailed flow matching model via random clock, we conduct extensive experiments across three benchmarks of increasing dimension and distinct noise structure. Full experimental details are provided in Appendix E. Ablation studies on flow type, number of function evaluations (NFE), and ODE solver are deferred to Appendix F. 2
When ET is applied to laws, it denotes the clock mixture: (ET µT )(B) := set B ⊆ Rd . This does not require finite moments of the clock-marginal law.
8
R
µT (B) PT (dT ) for every Borel
4.1
Experimental Setup
Datasets We use three benchmarks of increasing dimension and distinct heavy-tail structure. (i) 2D imbalanced α-stable mixture [42] : a 9-component isotropic α-stable mixture on a 3×3 grid with imbalanced weights, swept over per-component stability index αD ∈ {1.5, 1.8} (smaller αD gives heavier tails). (ii) CIFAR10-LT [6, 26] : a long-tailed subsample of CIFAR-10 at 32×32 with imbalance ratio ρ=0.01. CIFAR10-LT also serves as the testbed for baseline comparison and ablations. (iii) HRRR VIL [5, 12, 36] : hourly Vertically Integrated Liquid (VIL) analyses from the High-Resolution Rapid Refresh archive, cropped to 128×128 over the Central US; VIL is a strongly right-tailed quantity tied to convective weather extremes, following [36]. Sample sizes, splits, augmentations, backbones, optimizers, and samplers are detailed in Appendix E. Tasks and Metrics On 2D α-stable mixtures, we evaluate mode coverage ability across models using the precision–recall f1pr score [42], computed between real and generated samples. On CIFAR10-LT, we report Fréchet Inception Distance [16] (FID), computed against the full longtailed training split. On HRRR VIL, we follow [37, 36] and evaluate unconditional generation of VIL fields. For quantitative tail analysis, we compare flattened generated and test-set intensity samples using three one-dimensional tail-aware statistics: the kurtosis ratio (KR), skewness ratio (SR), and a tail-restricted two-sample Kolmogorov–Smirnov statistic (KS). All results are averaged over 3 random seeds. Formal definitions and implementation details are provided in Appendix E.3. Methods Following Theorem 3.2, different specifications of clock lead to different underlying source distribution for HTFM. In the following, we refer to the deterministic-clock variant Tt =t as HTFM-G, essentially recovering the standard Gaussian flow matching; the α/2-stable subordinator clock as HTFM-α; and the inverse-Gamma clock Tt =tVν for student-t, where Vν ∼ InvGamma(ν/2, ν/2), as HTFM-t. Each variant is indexed by the polynomial tail-decay rate of its source: pα for α-stable, pν for Student-t, and p=∞ for Gaussian (since Gaussians essentially have exponential tail decay). Thus, the polynomial tail-decay rate serves as the tail-calibration parameter across clock families.
4.2
Cross-Domain Tail Calibration with Random Clocks
We first study whether HTFM can Method 2D synthetic (f1pr ↑) CIFAR10-LT HRRR VIL act as a single heavy-tailed flowVariant p αD =1.5 αD =1.8 FID ↓ KR ↓ SR ↓ KS ↓ matching framework across differ0.893 19.29 0.72 0.68 0.42 ent datasets. For each benchmark, HTFM-G ∞ 0.888 1.6 0.943 0.934 16.19 0.04 0.11 0.11 we keep the training protocol fixed HTFM-α 1.7 0.941 0.923 16.10 0.47 0.06 0.04 and vary only the random-clock law 1.8 0.931 0.921 15.89 0.06 0.31 0.03 and its tail parameter, thereby com1.7 0.962 0.959 16.31 0.56 0.14 0.14 paring the Gaussian clock HTFM- HTFM-t 2.0 0.954 0.950 16.76 0.43 0.52 0.21 3.0 0.943 0.939 16.14 0.16 0.41 0.16 G against the two heavy-tailed clock families HTFM-α and HTFM-t. We Table 1: HTFM main results across benchmarks. p is report one sampling budget per the polynomial tail-decay exponent (pα for HTFM-α and benchmark: f1pr at NFE=10 on the pν for HTFM-t). Bold indicates the best result in each 2D mixture, FID at NFE=100 on column. CIFAR10-LT, and HRRR VIL tail statistics at NFE=50. See Appendix E for detailed information. Table 1 reveals two patterns. (i) HTFM with a heavy-tailed clock consistently outperforms HTFM-G across three datasets, which indicates that the random clock provides a useful inductive bias for heavy-tailed generation. (ii) The tail-decay exponent alone does not determine the best source family. At the matched decay rate p=1.7, HTFM-t performs 9
Noise Support
Method
Polynomial Tail Decay Exponent
NFE
α-stable Student-t Gaussian
pα = 1.6 pα = 1.7 pα = 1.8 pα = 1.9 pν = 1.7 pν = 2.0 pν = 3.0
Baselines
p=∞
LIM
✓
✗
✗
20 100 500
135.9 34.70 20.22
130.2 34.11 19.18
Heavy Tailed 126.8 132.5 32.40 34.10 18.23 20.60
DLPM
✓
✗
✓
20 100 500
29.16 18.18 15.26
28.55 18.23 14.91
32.49 18.29 15.35
30.96 18.74 15.43
– – –
– – –
– – –
35.49 22.28 19.47
DLIM
✓
✗
✓
20 100 500
21.30 19.69 17.22
23.90 20.38 17.73
22.75 18.46 17.90
24.50 19.60 18.14
– – –
– – –
– – –
25.45 21.63 19.92
t-Flow
✗
✓
✗
20 100 500
– – –
– – –
– – –
– – –
– – –
– – –
22.49 19.05 17.66
– – –
✓
20 100 500
17.52 16.19 15.27
17.13 16.10 15.04
16.91 15.89 14.99
20.48 16.76 16.51
19.35 16.14 15.38
Gaussian 22.67 19.29 19.04
Our Framework HTFM
✓
✓
– – –
– – –
– – –
Gaussian – – –
Heavy Tailed 16.76 19.91 16.21 16.31 15.67 15.97
Table 2: Main CIFAR10-LT results. We report FID at each NFE while sweeping the polynomial tail-decay parameter p. A dash “–” marks method–tail pairs that are not applicable. best on the 2D mixture, whereas HTFM-α performs better on CIFAR10-LT and HRRR VIL. Thus, the polynomial tail-decay rate captures only part of the modeling choice: the shape of the source beyond its asymptotic tail rate, such as the jump structure of α-stable clocks versus the scale-mixture geometry of Student-t clocks, also interacts with the data distribution. This suggests that the clock family is a meaningful design knob even after the nominal tail heaviness is fixed.
4.3
Comparison with Baselines
In Table 2, we compare HTFM against four heavy-tailed generative baselines on CIFAR10-LT under matched backbone and training budget: LIM [45, 39], DLPM/DLIM [42], and t-Flow [36]. Columns sweep the polynomial tail-decay rate across the union of the α-stable range pα ∈ {1.6, 1.7, 1.8, 1.9}, the Student-t range pν ∈ {1.7, 2.0, 3.0}, and the Gaussian limit p = ∞. Three observations can be made. (i) HTFM provides a versatile interface for tail calibration. LIM and DLPM/DLIM are tied to α-stable laws, while t-Flow is tied to the finitevariance Student-t regime. In contrast, HTFM recovers the Gaussian, α-stable, and Student-t regimes by changing only the random-clock law, while keeping the same backbone and training objective. (ii) HTFM inherits the low-NFE efficiency of flow matching. At low NFE, HTFM substantially improves over the α-stable diffusion baselines LIM and DLPM; at larger NFE, HTFM and DLPM become more comparable. (iii) HTFM outperforms the ODE-based heavy-tailed samplers. Compared with DLIM and t-Flow, HTFM achieves consistently lower FID across the reported NFE budgets.
4.4
Ablation: Signature Order for Clock Conditioning
This ablation tests whether HTFM’s gains come solely from using a heavy-tailed source, or whether the velocity field also benefits from observing the realized clock. On CIFAR10-LT, we report FID at NFE= 100 while varying the order m of the logsignature feature ℓm (T ) supplied to the network (cf. (3.7)). The case m = 0 denotes no clock feature: the model is still trained with samples from the random-clock source, but only observes (Xt , t). For m ≥ 1, the model additionally receives logsignature features of the realized clock path, allowing the vector field to adapt to the corresponding conditional Gaussian space. For HTFM-t, the affine clock Tt = tVν 10
is already determined by first-order path information, so higher-order features are omitted. Table 3 demonstrates the effectiveness of using Variant Tail exp. m = 0 m = 1 m = 2 signature features for clock-conditioning in HTFM. pα = 1.6 16.78 16.19 16.21 Across all reported HTFM-α and HTFM-t settings, HTFM-α pα = 1.7 16.65 16.10 16.46 using m=1 improves over the clock-independent m=0 pα = 1.8 16.49 15.89 15.94 variant. Meanwhile, increasing the order from m=1 pν = 1.7 18.72 16.31 – to m=2 does not provide a consistent additional gain, HTFM-t pν = 2.0 17.30 16.75 – suggesting that first-order clock information already pν = 3.0 16.58 16.14 – captures the dominant conditional-scale variation for Table 3: Signature-order ablation on this benchmark. CIFAR10-LT. Bold indicates the best result in each row.
5
Conclusion
We introduced Heavy-Tailed Flow Matching via Random Clocks (HTFM), which represents different families of heavy-tailed sources as mixtures of clock-conditioned Gaussian sources within a single flow-matching objective. Using truncated logsignature features, HTFM conditions the velocity field on the realized clock with negligible overhead. Across 2D synthetic data, CIFAR10-LT, and HRRR weather fields, HTFM improves over the Gaussian-clock variant, enables controllable tail calibration, and achieves favorable low-NFE performance against heavy-tailed baselines.
11
References [1] Michael S Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. arXiv preprint arXiv:2209.15571, 2022. [2] D. F. Andrews and C. L. Mallows. Scale mixtures of normal distributions. Journal of the Royal Statistical Society: Series B (Methodological), 36(1):99–102, 1974. [3] O. Barndorff-Nielsen, J. Kent, and M. Sørensen. Normal variance-mean mixtures and z-distributions. International Statistical Review, 50(2):145–159, 1982. [4] Adrian Baule. Generative modelling with jump-diffusions. arXiv:2503.06558, 2025.
arXiv preprint
[5] Stanley G Benjamin, Stephen S Weygandt, John M Brown, and ... A north american hourly assimilation and model forecast cycle: The rapid refresh. Monthly Weather Review, 144(4):1669–1694, 2016. [6] Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma. Learning imbalanced datasets with label-distribution-aware margin loss. In Advances in Neural Information Processing Systems, 2019. [7] Ilya Chevyrev and Andrey Kormilitzin. A primer on the signature method in machine learning. In Signature Methods in Finance: An Introduction with Computational Applications, pages 3–64. Springer, 2025. [8] Stuart Coles, Joanna Bawa, Lesley Trenner, and Pat Dorazio. An introduction to statistical modeling of extreme values, volume 208. Springer, 2001. [9] Rama Cont. Empirical properties of asset returns: stylized facts and statistical issues. Quantitative Finance, 1(2):223–236, 2001. [10] Rama Cont, Mihai Cucuringu, Renyuan Xu, and Chao Zhang. Tail-gan: Learning to simulate tail risk scenarios. Management Science, 72(4):2917–2936, 2026. [11] Christa Cuchiero, Francesca Primavera, and Sara Svaluto-Ferro. Universal approximation theorems for continuous functions of càdlàg paths and lévy-type signature models. Finance and Stochastics, 29(2):289–342, 2025. [12] David C Dowell, Curtis R Alexander, Eric P James, Stephen S Weygandt, Stanley G Benjamin, Geoffrey S Manikin, Benjamin T Blake, John M Brown, Joseph B Olson, Ming Hu, et al. The high-resolution rapid refresh (hrrr): An hourly updating convection-allowing forecast model. part i: Motivation and system description. Weather and Forecasting, 37(8):1371–1395, 2022. [13] Paul Embrechts, Claudia Klüppelberg, and Thomas Mikosch. Modelling extremal events: for insurance and finance, volume 33. Springer Science & Business Media, 2013. [14] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, volume 27, 2014. [15] Gaby J Gründemann, Nick van de Giesen, Lukas Brunner, and Ruud van der Ent. Rarest rainfall events will see the greatest relative increase in magnitude under future climate change. Communications Earth & Environment, 3(1):235, 2022.
12
[16] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, 2017. [17] Tennessee Hickling and Dennis Prangle. Flexible tails for normalizing flows. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 23155–23178. PMLR, 2025. [18] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. [19] Priyank Jaini, Ivan Kobyzev, Yaoliang Yu, and Marcus Brubaker. Tails of Lipschitz triangular flows. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 4673–4681. PMLR, 2020. [20] Patrick Kidger, Patric Bonnier, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons. Deep signature transforms. Advances in neural information processing systems, 32, 2019. [21] Patrick Kidger and Terry J. Lyons. Signatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU. In International Conference on Learning Representations, 2021. [22] Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. In International Conference on Learning Representations, 2014. [23] Franz J. Kiraly and Harald Oberhauser. Kernels for sequentially ordered data. Journal of Machine Learning Research, 20(31):1–45, 2019. [24] Nikita Kornilov, Petr Mokrov, Alexander Gasnikov, and Alexander Korotin. Optimal flow matching: Learning straight trajectories in just one step. Advances in Neural Information Processing Systems, 37:104180–104204, 2024. [25] Nikola Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar. Neural operator: Learning maps between function spaces with applications to pdes. Journal of Machine Learning Research, 24(89):1–97, 2023. [26] Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009. [27] Daniel Levin, Terry Lyons, and Hao Ni. Learning from the past, predicting the statistics for the future, learning an evolving system. arXiv preprint arXiv:1309.0260, 2013. [28] Qi Lian, Yu Qi, and Yueming Wang. Cauchy diffusion: A heavy-tailed denoising diffusion probabilistic model for speech synthesis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 24549–24557, 2025. [29] Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022. [30] Qiang Liu. Rectified flow: A marginal preserving approach to optimal transport. arXiv preprint arXiv:2209.14577, 2022. [31] Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022.
13
[32] Ziwei Liu, Zhongqi Miao, Xiaohang Zhan, Jiayun Wang, Boqing Gong, and Stella X Yu. Large-scale long-tailed recognition in an open world. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2537–2546, 2019. [33] Terry Lyons. Rough paths, signatures and the modelling of functions on streams. arXiv preprint arXiv:1405.4537, 2014. [34] Frank J Massey Jr. The Kolmogorov–Smirnov test for goodness of fit. Journal of the American Statistical Association, 46(253):68–78, 1951. [35] James Morrill, Cristopher Salvi, Patrick Kidger, and James Foster. Neural rough differential equations for long time series. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 7829–7838. PMLR, 2021. [36] Kushagra Pandey, Jaideep Pathak, Yilun Xu, Stephan Mandt, Michael Pritchard, Arash Vahdat, and Morteza Mardani. Heavy-tailed diffusion models. arXiv preprint arXiv:2410.14171, 2024. [37] Jaideep Pathak, Shashank Subramanian, Peter Harrington, Sanjeev Raja, Ashesh Chattopadhyay, Morteza Mardani, Thorsten Kurth, David Hall, Zongyi Li, Kamyar Azizzadenesheli, Pedram Hassanzadeh, Karthik Kashinath, and Animashree Anandkumar. FourCastNet: A global data-driven high-resolution weather model using adaptive Fourier neural operators. arXiv preprint arXiv:2202.11214, 2022. [38] Aram-Alexandre Pooladian, Heli Ben-Hamu, Carles Domingo-Enrich, Brandon Amos, Yaron Lipman, and Ricky TQ Chen. Multisample flow matching: Straightening flows with minibatch couplings. arXiv preprint arXiv:2304.14772, 2023. [39] Vadim Popov, Assel Yermekova, Tasnima Sadekova, Artem Khrapov, and Mikhail Sergeevich Kudinov. Improved sampling algorithms for lévy-itô diffusion models. In The Thirteenth International Conference on Learning Representations. [40] Sidney I Resnick. Heavy-tail phenomena: probabilistic and statistical modeling. Springer, 2007. [41] Sonia I Seneviratne, Xuebin Zhang, Muhammad Adnan, Wafae Badi, Claudine Dereczynski, A Di Luca, Subimal Ghosh, Iskhaq Iskandar, James Kossin, Sophie Lewis, et al. Weather and climate extreme events in a changing climate. Climate change 2021: The physical science basis: Working group I contribution to the sixth assessment report of the intergovernmental panel on climate change, pages 1513–1766, 2021. [42] Dario Shariatian, Umut Simsekli, and Alain Oliviero Durmus. Heavy-tailed diffusion with denoising levy probabilistic models. In The Thirteenth International Conference on Learning Representations. [43] Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. [44] Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, and Yoshua Bengio. Improving and generalizing flow-based generative models with minibatch optimal transport. arXiv preprint arXiv:2302.00482, 2023.
14
[45] Eun Bi Yoon, Keehun Park, Sungwoong Kim, and Sungbin Lim. Score-based generative models with lévy processes. Advances in Neural Information Processing Systems, 36:40694– 40707, 2023. [46] Tianjiao Zhang, Huangjie Zheng, Jiangchao Yao, Xiangfeng Wang, Mingyuan Zhou, Ya Zhang, and Yanfeng Wang. Long-tailed diffusion models with oriented calibration. In The twelfth international conference on learning representations, 2024.
15
Appendix Contents A
Additional Discussions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 A.1 Related work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 A.2 Detailed comparison with existing heavy-tailed models. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .17 A.3 Representative Gaussian scale mixtures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 A.4 Gaussian mean-variance mixtures and asymmetric extensions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 A.5 Other admissible clock-conditioned paths . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 B All-clock extension to diffusion models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 C Additional results and proofs of main theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 C.1 A note on conditioning on the clock history. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .25 C.2 Conditional Gaussian law of the affine path . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 C.3 Detailed pathwise equivalence result and proof. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .26 C.4 Clock-conditioned straight-line geodesic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 C.5 Proof of Theorem 3.2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 D Additional Information on Log-Signature . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 D.1 Signature-linear schedules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 D.2 VE-type signature-linear schedule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 E Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 E.1 Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 E.2 Baselines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 E.3 Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 E.4 Denoiser Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 E.5 Training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 E.6 Sampling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 F Ablation Studies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 F.1 Effect of Flow Type . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 F.2 Effect of Number of Function Evaluations (NFE) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 F.3 Effect of ODE Solver . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 G Visualizations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
A
Additional Discussions
A.1
Related work
Heavy-tailed diffusion and flow models. Most diffusion and flow-matching generative models use Gaussian reference distributions or Gaussian perturbation kernels. Several recent works replace this Gaussian mechanism by heavy-tailed alternatives. Lévy score models use α-stable Lévy processes and replace the ordinary score by a fractional score, leading to nonlocal reverse-time dynamics and deterministic fractional probability-flow samplers [45, 39]. DLPM and DLIM instead build a discrete-time α-stable analogue of DDPM/DDIM; their key practical step is a Gaussian scale-mixture augmentation that makes the conditional reverse kernels Gaussian while the marginal noising law remains heavy-tailed [42]. Student-t-based diffusion and flow models replace Gaussian perturbation kernels with Student-t kernels and derive family-specific denoising posteriors and training divergences, including t-Flow as a heavy-tailed flow construction [36]. Cauchy Diffusion studies the Cauchy case in speech synthesis, using Cauchy noise and a ratio-of-Gaussians view to support stochastic and deterministic sampling in a DDPMstyle vocoder [28]. Related jump-diffusion generative models consider finite-activity jumps and marginal-matching reverse dynamics [4]. Our random-clock formulation follows the same broad motivation of using heavy-tailed perturbation families, but differs in where the tractability is imposed: after conditioning on a realized clock path, the flow-matching problem is Gaussian and uses the ordinary conditional velocity target, while marginalizing the clock produces the heavy-tailed source or path family.
16
Latent Gaussian mixtures and random time changes. Gaussian scale mixtures provide a classical route from conditionally Gaussian models to heavy-tailed marginals [2]. More general normal variance-mean mixtures can also be interpreted as Brownian motions with drift stopped at independent random times, and include many skewed and heavy-tailed families [3]. This latent-Gaussian perspective is exactly the mechanism exploited by DLPM for α-stable noising [42]. Our construction extends this idea from a scalar mixing variable to a path-valued clock. The clock can determine the conditional source covariance, the interpolation schedule, or both. Thus, the model keeps a Gaussian conditional regression problem but allows the outer clock law to encode family choice, tail heaviness, and path-dependent allocation of noise. Signature and logsignature transforms. Path signatures originate in rough path theory and give an order-sensitive sequence of iterated-integral features for streams [33, 7]. In statistics and machine learning, signature features have been used as universal, truncated representations for stream-valued inputs [27], as kernels for sequential data [23], and as differentiable neuralnetwork components [20, 21]. Logsignature features are especially relevant when one wants a compact nonredundant representation of path increments; for example, neural rough differential equations summarize long driving signals over subintervals by logsignatures [35]. Recent approximation results also justify signature-linear representations of path functionals on suitable compact path sets, including càdlàg settings [11]. Our use of signatures is narrower: the clock path is not the generated sample path, but a latent conditioning environment. We use low-order logsignature features to give the neural velocity field finite-dimensional access to the realized random clock, and, optionally, to parameterize clock-dependent affine schedules.
A.2
Detailed comparison with existing heavy-tailed models
The related-work overview in Appendix A.1 surveys heavy-tailed diffusion and flow constructions. Here we elaborate on the technical contrast between the random-clock formulation and the closest prior families. DLPM and Lévy-based diffusion. DLPM [42] obtains heavy-tailed diffusion transitions via α-stable innovations and a robust outer power applied to empirical losses. As shown in Theorem 3.2(ii), our stable-subordinator clock with the area-clock covariance reproduces the same α-stable marginal source family. The two constructions differ at two levels. At the flowmatching target level, DLPM averages over noise realizations and learns a clock-marginal vector field, whereas conditioning on the realized clock path turns the regression into a clock-indexed family of pathwise optima uTt ; a network with access to the clock can learn the appropriate optimum for each noise realization, and the random-clock view can therefore be read as a refinement of the DLPM stable-noise construction with the same marginal family but a betterspecified flow-matching target. We expect empirical gains when different clock realizations induce meaningfully different optima, and the formulation also avoids direct use of the generally unavailable α-stable density. At the loss level, the outer robust power z 7→ z r with 0 < r < 1 used by DLPM destroys the additive projection decomposition in Proposition 3.1; in contrast, the random-clock formulation keeps the inner problem in a conditionally Gaussian L2 -projection space and treats any robustification over sampled clocks as an estimator-level stabilization rather than a change of the population target. Compared with Lévy score and Lévy–Itô diffusion models [45, 39], the flow-matching construction further avoids a nonlocal fractional-score target: after conditioning on the clock, the endpoint-conditioned target remains the explicit affine expression (3.10). Student-t diffusion and t-Flow. Student-t diffusion and t-Flow models replace Gaussian perturbations by Student-t heavy-tailed perturbations [36]. Their construction is tailored to the
17
Student-t family: it derives Student-t-specific perturbation kernels and uses a γ-divergence objective, with the practical formulation restricted to the finite-variance regime ν > 2. Extending that derivation formally to ν ≤ 2 is mathematically problematic, rather than merely an implementation detail, because the covariance-based normalizations and finite-variance quantities used by the Student-t perturbation model cease to be well-defined. In contrast, the inversegamma clock in Theorem 3.2(iii) is well-defined for every ν > 0, so our pathwise construction covers the extremely heavy-tailed infinite-variance regime ν ≤ 2 without changing the conditional Gaussian target. Toward data-adaptive clocks. The most general use of random clocks would not prescribe the clock law in advance. Instead, one could estimate a clock distribution from data, or infer a data-dependent posterior over clock paths with an amortized model, and then use the inferred clock as the conditioning variable in (3.7). This would turn the clock into a learned latent description of local tail behavior rather than a hand-chosen family such as a stable subordinator or inverse-gamma slope. Developing such data-adaptive clock inference requires additional identifiability and estimation assumptions, so we leave it as a future extension.
A.3
Representative Gaussian scale mixtures
We recall two standard Gaussian scale-mixture examples that are only mentioned briefly in Section 2.2. For α ∈ (0, 2], a centered symmetric isotropic α-stable random vector Sα ∈ Rd is characterized by i h ξ ∈ Rd . (A.1) E ei⟨ξ,Sα ⟩ = exp −∥ξ∥α2 /2 , The case α = 2 reduces to the Gaussian law, whereas for α < 2 the law is heavy-tailed, with radial tail probability P (∥Sα ∥2 > r) ≍ r−α , r → ∞. With this normalization, Sα admits the Gaussian scale-mixture representation d p
Sα =
Vα G,
where G ∼ N (0, Id ) and Vα is a positive α/2-stable mixing variable independent of G, with Laplace transform E[e−sVα ] = exp −2α/2−1 sα/2 , s ≥ 0. Likewise, a centered d-dimensional Student-t random vector Zν with ν > 0 degrees of freedom and identity scale has density −(ν+d)/2 Γ (ν + d)/2 ∥x∥22 fν,d (x) = 1 + , ν Γ(ν/2) (νπ)d/2
x ∈ Rd ,
(A.2)
and satisfies P (∥Zν ∥2 > r) ≍ r−ν , It can be written as
d p
Zν =
Vν G,
r → ∞.
Vν ∼ InvGamma(ν/2, ν/2),
with Vν independent of G. These examples make explicit why conditioning on the latent scale leaves a Gaussian problem, while marginalizing the latent scale produces polynomial tails.
18
A.4
Gaussian mean-variance mixtures and asymmetric extensions
The main text focuses on centered Gaussian scale mixtures, which already cover a broad class of symmetric heavy-tailed laws. A natural extension is to allow the latent variable to affect both the conditional mean and the conditional covariance, leading to the class of Gaussian mean-variance mixtures. Let V be a latent random variable taking values in a measurable space V, and let G ∼ N (0, Id ) be independent of V . A random vector X ∈ Rd is called a Gaussian mean-variance mixture if it admits a representation of the form X = m(V ) + L(V )G,
(A.3)
where m : V → Rd and L : V → Rd×d are measurable. Equivalently, X | V ∼ N m(V ), Σ(V ) , Σ(V ) := L(V )L(V )⊤ .
(A.4)
The heavy-tailed or asymmetric nature of the marginal law is produced by marginalizing over the latent variable V , while conditional on V the distribution remains Gaussian. The centered Gaussian scale-mixture subclass from the main text is recovered by setting m(V ) ≡ 0. When m(V ) ̸= 0, the latent variable affects not only the covariance but also the location, and the resulting marginal law may be asymmetric. Thus Gaussian mean-variance mixtures provide a natural route to skewed heavy-tailed extensions. In the random-clock setting, the corresponding extension would be to replace the centered conditional source law X0T ∼ N 0, Σ(T[0,1] ) by the more general form X0T ∼ N m(T[0,1] ), Σ(T[0,1] ) , where m : D([0, 1]; R+
(A.5)
) → Rd is a measurable mean functional. Under the affine interpolation XtT = αt X1T + βt X0T ,
we then have XtT | X1T = x1 ∼ N αt x1 + βt m(T[0,1] ), βt2 Σ(T[0,1] ) .
(A.6)
Hence the path remains conditionally Gaussian after conditioning on the clock path, so the latent-Gaussian principle underlying the flow-matching construction continues to hold. Moreover, the explicit affine target formula is unaffected by this extension. Indeed, for a.e. t ∈ [0, 1), the identity X T − αt X1T X0T = t βt still holds pathwise, so substituting into dαt T dβt T dXtT = X + X dt dt 1 dt 0 yields dβt dXtT T T E Xt = x, X1 = x1 = dt x + dt βt
dβ
t dαt − dt αt dt βt
! x1 .
(A.7)
Thus, even in the more general mean-variance mixture setting, the affine clock-conditioned target retains the same closed-form expression as in the centered case. What changes is the induced marginal family after integrating out the clock, which may now be skewed as well as heavy-tailed. This extension suggests that the random-clock framework is not limited to symmetric Gaussian scale mixtures. By allowing a nonzero conditional mean functional m(T[0,1] ), one can in principle accommodate a broader class of asymmetric heavy-tailed laws while preserving conditional Gaussianity. We do not pursue this direction further in the main text. 19
A.5
Other admissible clock-conditioned paths
The main text focuses on the straight-line interpolation, which is the most common choice in conditional flow matching and already suffices for the heavy-tailed framework developed here. More generally, however, the same clock-conditioned construction applies to a broader class of affine paths. Let α, β : [0, 1] → R be absolutely continuous functions such that α0 = 0,
β0 = 1,
α1 = 1,
β1 = 0,
and assume βt > 0 for all t ∈ [0, 1). Define XtT = αt X1T + βt X0T ,
t ∈ [0, 1].
(A.8)
As in Section 3.3, if X0T ∼ N (0, Σ(T[0,1] )) and X0T and X1T are conditionally independent given T[0,1] , then pTt (· | x1 ) = Law(XtT | X1T = x1 ) = N αt x1 , βt2 Σ(T[0,1] ) . (A.9) Moreover, for a.e. t ∈ [0, 1), dXtT dαt T dβt T = X + X , dt dt 1 dt 0
X0T =
XtT − αt X1T , βt
so the corresponding conditional flow-matching target is vtT (x | x1 ) =
dβt dt
βt
dβ
t dαt − dt αt dt βt
x+
! x1 .
(A.10)
Thus any admissible pair (α, β) yields an explicit clock-conditioned path/target pair. VP-type path. A natural diffusion-inspired choice is the variance-preserving schedule. Let ᾱ : [0, 1] → [0, 1] be absolutely continuous with ᾱ0 = 0,
ᾱ1 = 1,
and assume ᾱt < 1 for t ∈ [0, 1). Define αt = ᾱt ,
q βt = 1 − ᾱt2 .
Then XtT = ᾱt X1T +
q 1 − ᾱt2 X0T ,
(A.11)
which is the analogue of a variance-preserving interpolation. In this case, (A.10) becomes vtT (x | x1 ) = −
dᾱt ᾱt ᾱt ddt dt x + x1 . 1 − ᾱt2 1 − ᾱt2
(A.12)
VE-type path. A variance-exploding analogue can also be accommodated within the affine framework. Let σ : [0, 1] → [0, 1] be an absolutely continuous schedule such that σ0 = 1,
σ1 = 0,
σt > 0
for t ∈ [0, 1).
Set αt = 1 − σt ,
20
βt = σt .
Then the corresponding affine interpolation is XtT = (1 − σt )X1T + σt X0T ,
t ∈ [0, 1].
(A.13)
Conditioning on X1T = x1 and the realized clock path T[0,1] , we have pTt (· | x1 ) = Law(XtT | X1T = x1 ) = N (1 − σt )x1 , σt2 Σ(T[0,1] ) .
(A.14)
Thus the conditional variance is directly controlled by the scale schedule σt2 . Although our interpolation is written in the generative direction from source to data, this corresponds to a variance-exploding schedule when viewed backward from t = 1 to t = 0. Differentiating (A.13) gives dXtT dσt T dσt T =− X + X , dt dt 1 dt 0 Substituting into (A.10), or directly into vtT (x | x1 ) =
X0T =
XtT − (1 − σt )X1T . σt
dXtT dt , yields the conditional flow-matching target dσt dt
σt
(x − x1 ),
t ∈ [0, 1).
(A.15)
A simple concrete choice is the polynomial schedule σt = (1 − t)γ ,
γ > 0,
for which XtT = 1 − (1 − t)γ X1T + (1 − t)γ X0T , and
γ (x − x1 ). (A.16) 1−t More generally, any absolutely continuous decreasing schedule σt with the above endpoint conditions yields an admissible VE-type path within the same clock-conditioned affine framework. vtT (x | x1 ) = −
B
All-clock extension to diffusion models
In this appendix we record how the random-clock viewpoint extends to score-based diffusion models. The point of this subsection is slightly different from the flow-matching construction in the main text. Flow matching only needs a clock-conditioned interpolation and an explicit velocity target; it does not require reversing a stochastic process. Diffusion models do require a reverse-time dynamics, and for general càdlàg clocks this dynamics is not, in general, a purely local score SDE. The correct all-clock statement is a reverse-kernel formulation. The familiar local score SDE is recovered on the continuous part of the clock, and in particular for absolutely continuous clocks. Let W be a standard d-dimensional Brownian motion independent of the clock T . Given drift f : [0, 1] × Rd → Rd and deterministic diffusion matrix G : [0, 1] → Rd×d , consider the clock-conditioned forward dynamics dXt = ft (Xt ) dt + Gt dWTt ,
X0 ∼ p0 (· | T[0,1] ).
(B.1)
For a fixed realized clock path, WTt is a semimartingale with quadratic variation Tt . If T is continuous, this is a continuous martingale; if T has jumps, then WTt has Gaussian jumps with T covariance proportional to the clock jump sizes. Write at := Gt G⊤ t and pt := Law(Xt | T[0,1] ). T Whenever pt has a density, we define the clock-conditioned score by sTt (x) := ∇x log pTt (x). 21
(B.2)
As in the main text, the density is conditioned on the full clock path. This is the right object for clock-dependent initial laws such as X0 = L(T[0,1] )G; in general it cannot be replaced by conditioning only on T[0,t] . We first state the reverse dynamics in a form that is valid for arbitrary realized càdlàg clock T (x, dy) denote the conditional forward transition kernel of paths. For 0 ≤ s < t ≤ 1, let Ps,t T = pT . The reverse transition (B.1) from time s to time t, given the full clock path. Thus pTs Ps,t t from t back to s is the Bayes kernel associated with this joint law. Theorem B.1 (All-clock reverse dynamics in kernel form). Fix a realized clock path T ∈ D([0, 1]; R+ ). Assume that the forward equation (B.1) is well posed, that regular conditional T exist, and that pT P T = pT for 0 ≤ s < t ≤ 1. For s < t, define a reverse transition kernels Ps,t s s,t t T (y, dx) by the measure identity kernel Rt,s T T pTt (dy) Rt,s (y, dx) = pTs (dx) Ps,t (x, dy)
(B.3)
on Rd × Rd . Equivalently, when transition densities exist, T Rt,s (y, dx) =
pTs (x)pTs,t (y | x) pTt (y)
dx.
(B.4)
Let (Yr )r∈[0,1] be the reverse-time Markov process initialized by Y0 ∼ pT1 and, for 0 ≤ r < r′ ≤ 1, using the transition T Law(Yr′ ∈ dx | Yr = y, T[0,1] ) = R1−r, 1−r′ (y, dx).
(B.5)
Then the reverse process is conditional marginal-matching: Law(Yr | T[0,1] ) = pT1−r ,
r ∈ [0, 1].
(B.6)
Proof. Fix 0 ≤ s < t ≤ 1. Integrating (B.3) over y ∈ Rd gives, for every measurable set A ⊆ Rd , Z Z T T Rt,s (y, A) pTt (dy) = Ps,t (x, Rd ) pTs (dx) = pTs (A). Rd
A
Thus a single reverse step transports pTt back to pTs . Taking t = 1 − r and s = 1 − r′ , this says that if Yr ∼ pT1−r , then the transition (B.5) gives Yr′ ∼ pT1−r′ . Since Y0 ∼ pT1 , iterating this identity over any finite partition of [0, 1] yields Law(Yr | T[0,1] ) = pT1−r for every r ∈ [0, 1]. The Markov kernels are consistent because they are the regular conditional kernels of the same forward Markov law read backward in time, so the corresponding reverse process exists by the usual extension construction. The theorem is deliberately stated without a score. A score is a local object, whereas a jump of the clock produces a finite, nonlocal update of the law. This can be seen explicitly. Let t be a jump time of the realized clock and set h := ∆Tt = Tt − Tt− > 0. Since the drift term has no atom in physical time, the forward jump is Zt ∼ N (0, hId ),
Xt = Xt− + Gt Zt ,
so the jump kernel is QTt (x, dy) = N (y; x, hat )(dy). Therefore Z T pt (dy) = QTt (x, dy) pTt− (dx), and the reverse jump is the Bayes kernel characterized by T pTt (dy)Rt,t− (y, dx) = pTt− (dx)QTt (x, dy).
22
(B.7)
When at is nonsingular, this becomes the density formula T Rt,t− (y, dx) = R
pTt− (x)φhat (y − x) dx, pTt− (z)φhat (y − z) dz
(B.8)
where φhat is the Gaussian density with covariance hat . Thus reverse jumps are posterior sampling steps. They cannot be represented exactly by the local drift term at sTt dTt . We next recover the usual score-based formula from the kernel theorem on the continuous part of the clock. Write X ∆Ts Ttc := Tt − 0<s≤t
for the continuous part of the realized clock, and define the reversed continuous clock Tbrc := c . At reverse times corresponding to jumps of T , the process uses the jump kernel (B.7); T1c −T1−r between these jumps, or if the clock is continuous, the reverse dynamics have the following local representation. Corollary B.2 (Local score form on the continuous clock part). Assume the hypotheses of Theorem B.1. In addition, assume that the continuous part of the forward law admits strictly positive smooth densities, that the usual decay and integrability conditions for integration by parts hold, and that the local reverse equations below are well posed. On the continuous clock part, with τ = 1 − r, the reverse process has the local score representation dYr = −fτ (Yr ) dr + aτ sTτ (Yr ) dTbrc + Gτ dW̄Tbc ,
(B.9)
r
where W̄ is a standard Brownian motion. If T is continuous, this local equation alone gives Law(Yr | T[0,1] ) = pT1−r for all r ∈ [0, 1]. If T has jumps, it must be supplemented by the reverse jump kernels (B.7) at the corresponding reverse jump times. For the continuous part, the deterministic probability-flow equation is 1 dZr = −fτ (Zr ) dr + aτ sTτ (Zr ) dTbrc . 2
(B.10)
When T has jumps, an exact deterministic flow across a jump additionally requires choosing a transport map from pTt to pTt− ; this choice is not canonical and is not determined by the score alone. Proof. We prove the continuous part; the jump updates are already identified by Theorem B.1. Let ψ ∈ Cc∞ (Rd ). Along the continuous part of the forward dynamics, the conditional law satisfies the weak equation Z Z Z 1 T T d ψ(x)pt (x) dx = ft (x) · ∇ψ(x)pt (x) dx dt + Tr(at ∇2 ψ(x))pTt (x) dx dTtc . 2 c , this forward weak equation gives Set qr = pT1−r and τ = 1 − r. Since Tbrc = T1c − T1−r Z Z Z 1 d ψ(x)qr (x) dx = − fτ (x) · ∇ψ(x)qr (x) dx dr − Tr(aτ ∇2 ψ(x))qr (x) dx dTbrc . 2
Now let ρr = Law(Yr | T[0,1] ) for the local SDE (B.9). Its weak equation is Z Z d ψ(x)ρr (x) dx = − fτ (x) · ∇ψ(x)ρr (x) dx dr Z + aτ sTτ (x) · ∇ψ(x)ρr (x) dx dTbrc Z 1 + Tr(aτ ∇2 ψ(x))ρr (x) dx dTbrc . 2 23
Substituting ρr = qr = pTτ and using pTτ sTτ = ∇pTτ , integration by parts gives Z Z Z aτ sTτ · ∇ψ pTτ dx = aτ ∇pTτ · ∇ψ dx = − Tr(aτ ∇2 ψ)pTτ dx. R Therefore the two dTbrc -terms combine to − 12 Tr(aτ ∇2 ψ)pTτ dx dTbrc , which matches the weak equation for qr . With matching initial law at r = 0, uniqueness gives the stated marginal matching. equation (B.10), the weak equation has the same drift term but only R For Tthe deterministic 1 c b aτ sτ · ∇ψ ζr dx dTr in the clock direction. Substituting ζr = qr and applying the same 2 integration by parts yields exactly the same weak equation for qr . Hence Law(Zr | T[0,1] ) = pT1−r on the continuous part. Rt If the clock is absolutely continuous, Tt = 0 ℓs ds, then dTbrc = ℓτ dr, and (B.9) reduces to p dYr = −fτ (Yr ) + ℓτ aτ sTτ (Yr ) dr + ℓτ Gτ dB̄r , τ = 1 − r, while the probability flow becomes dZr = −fτ (Zr ) + 21 ℓτ aτ sTτ (Zr ) dr. This is the classical reverse-SDE/probability-flow structure with the diffusion rate modulated by the realized clock rate. The learning target associated with the local continuous dynamics is the clock-conditioned score sTt . Accordingly, for continuous-clock diffusion implementations, or for the continuous part of a jump-clock implementation, a pathwise score-matching objective is ii h h 2 (B.11) LDSM (θ) := Et∼Unif(0,1) ωt ET EXt |T sθ (Xt , t, ηT ) − sTt (Xt ) 2 , where ωt > 0 is an optional weighting schedule. For clocks with jumps, this score objective does not by itself specify the reverse jump kernels in (B.7); exact reverse simulation for such clocks requires those nonlocal kernels, or an additional approximation to them. In practice, the marginal score sTt is typically unavailable. One may therefore use a clockconditioned denoising score-matching objective whenever the forward transition kernel given (X0 , T[0,1] ) is tractable: LCDSM (θ) h h ii (B.12) 2 :=Et∼Unif(0,1) ωt EX0 ,T EXt |X0 ,T sθ (Xt , t, ηT ) − ∇xt log pt (xt | X0 , T[0,1] ) xt =Xt 2 . Under the same square-integrability and projection conditions as standard denoising score matching, the conditional score in (B.12) projects onto the marginal score sTt after conditioning on (Xt , T[0,1] ). For linear-Gaussian forward schedules, the conditional kernel is Gaussian. If Xt | (X0 , T[0,1] ) ∼ N mt (X0 , T[0,1] ), Qt (T[0,1] ) , then ∇xt log pt (xt | X0 , T[0,1] ) = − Qt (T[0,1] )−1 xt − mt (X0 , T[0,1] ) , and (B.12) becomes LCDSM (θ) 2 −1 :=Et∼Unif(0,1) ωt EX0 ,T EXt |X0 ,T sθ (Xt , t, ηT ) + Qt (T[0,1] ) Xt − mt (X0 , T[0,1] ) . 2
(B.13) 24
C
Additional results and proofs of main theorems
C.1
A note on conditioning on the clock history
The main construction conditions on the full realized clock path T[0,1] . In general this conditioning cannot be replaced by conditioning only on the past clock history T[0,t] , because the clock-conditioned source may itself depend on future clock values such as T1 . The following lemma records the additional non-anticipativity condition under which such a reduction is valid. Lemma C.1 (Clock-history reduction under non-anticipative initialization). Consider the forward construction (B.1). Assume that, for every t ∈ [0, 1], the initial variable X0 is conditionally independent of σ(T(t,1] ) given σ(T[0,t] ), and that the Brownian motion driving the SDE is independent of the clock and of the base randomness used to form X0 . Then, for every t ∈ [0, 1], Law(Xt | T[0,1] ) = Law(Xt | T[0,t] )
a.s.
Equivalently, whenever both sides admit conditional densities, pt (· | T[0,1] ) = pt (· | T[0,t] )
a.s.
Proof. Fix t ∈ [0, 1] and set Gt := σ(X0 , T[0,t] ) ∨ σ(Wu : 0 ≤ u ≤ Tt ). By non-anticipativity of the forward dynamics, Xt is measurable with respect to Gt . Since Tt is σ(T[0,t] )-measurable and the Brownian motion W is independent of the entire clock, σ(Wu : 0 ≤ u ≤ Tt ) is conditionally independent of σ(T(t,1] ) given σ(T[0,t] ). By assumption the same is true for X0 . Hence Gt , and therefore Xt , is conditionally independent of σ(T(t,1] ) given σ(T[0,t] ). For any bounded measurable test function φ, a.s., E φ(Xt ) σ(T[0,1] ) = E φ(Xt ) σ(T[0,t] ) which proves the claim. The non-anticipativity assumption in Lemma C.1 is essential. For example, if X0 | T[0,1] ∼ N (0, T1 Id ), then already at t = 0 the conditional law given T[0,1] is N (0, T1 Id ), whereas the conditional law given T[0,0] is the corresponding mixture over T1 . These laws are generally different. This is why the main flow-matching construction consistently uses pTt = Law(XtT ) = Law(Xt | T[0,1] ) rather than replacing it by a past-clock conditional law.
C.2
Conditional Gaussian law of the affine path
We verify the conditional Gaussian law (3.9) stated in the main text. Recall the assumptions: conditional on the realized clock path T[0,1] , X0T ∼ N (0, Σ(T[0,1] )); X0T and X1T are conditionally independent given T[0,1] ; and the affine path is XtT = αtT X1T +βtT X0T with deterministic schedules αtT , βtT given T[0,1] . Proof of (3.9). Fix T[0,1] and X1T = x1 . Under this conditioning, αtT and βtT are deterministic, and the conditional independence assumption gives X0T | (X1T = x1 , T[0,1] ) ∼ N (0, Σ(T[0,1] )). Hence XtT | (X1T = x1 , T[0,1] ) = αtT x1 + βtT X0T (X1T = x1 , T[0,1] ) is an affine transformation of a Gaussian, with mean αtT x1 and covariance (βtT )2 Σ(T[0,1] ). Therefore XtT (X1T = x1 , T[0,1] ) ∼ N αtT x1 , (βtT )2 Σ(T[0,1] ) , which is precisely (3.9). 25
C.3
Detailed pathwise equivalence result and proof
Proposition C.2 (Detailed pathwise conditional–marginal equivalence). Assume (3.3) and suppose that, for every θ in the parameter class and for a.e. t ∈ [0, 1], E ∥uθ (XtT , t, T[0,1] )∥22 + ∥vtT (XtT | X1T )∥22 T[0,1] < ∞ a.s. Then, for a.e. t, there exists a nonnegative σ(T[0,1] )-measurable random variable Ct (T ), independent of θ, such that h i E Rtcond (θ; X1T , T ) T[0,1] = Rtmarg (θ; T ) + Ct (T ). Consequently, whenever the risks and Ct (T ) are integrable over (t, T ) with t ∼ Unif(0, 1), LCHTFM (θ) = LHTFM (θ) + Et∼Unif(0,1) ET [Ct (T )] . In particular, if the additive constant is finite, the two population objectives have the same minimizers. If uθ is differentiable in θ and differentiation may be exchanged with the conditional expectations, the clock expectation, and the uniform-time expectation, then ∇θ LCHTFM (θ) = ∇θ LHTFM (θ). Proof of Propositions 3.1 and C.2. Fix a time t for which the assumptions hold. By (3.3), uTt (x) = E vtT (XtT | X1T ) XtT = x for pTt -a.e. x, where expectations are taken in the fixed-clock conditional probability space. The same identity is understood in the weak sense when densities are not used. Thus, within the clock-conditioned environment indexed by T , uTt (XtT ) is the L2 -projection of the pathwise conditional velocity vtT (XtT | X1T ) onto the sigma-field σ(XtT ). Applying the conditional Pythagorean identity gives h i h i 2 E Rtcond (θ; X1T , T ) T[0,1] = Rtmarg (θ; T ) + E vtT (XtT | X1T ) − uTt (XtT ) 2 T[0,1] = Rtmarg (θ; T ) + Ct (T ), where Ct (T ) := E
h
i 2 vtT (XtT | X1T ) − uTt (XtT ) 2 T[0,1] ≥ 0.
If the risks and Ct (T ) are integrable over (t, T ) with t ∼ Unif(0, 1), then Fubini’s theorem yields LCHTFM (θ) = LHTFM (θ) + Et∼Unif(0,1) ET [Ct (T )] . (C.1) The last term is independent of θ. Hence, whenever it is finite, the two population objectives have the same minimizers. If uθ is differentiable in θ and differentiation may be exchanged with the conditional expectations, the clock expectation, and the uniform-time expectation, then ∇θ LCHTFM (θ) = ∇θ LHTFM (θ).
C.4
Clock-conditioned straight-line geodesic
Theorem C.3 (Clock-conditioned straight-line geodesic). Fix p ∈ [1, ∞). For a.e. realized clock path T , assume p0 (· | T[0,1] ), pdata ∈ Pp (Rd ), and let π T be a p-optimal coupling between them. Then µTt := (1 − t)x0 + tx1 # π T , t ∈ [0, 1], is a constant-speed Wp -geodesic from p0 (· | T[0,1] ) to pdata .
26
Proof. Fix a realized clock path T for which p0 (· | T[0,1] ) ∈ Pp (Rd ), and abbreviate µ0 := p0 (· | T[0,1] ),
µr := (1 − r)x0 + rx1 # π T ,
µ1 := pdata ,
r ∈ [0, 1].
Then µTt = µt . By assumption µ0 , µ1 ∈ Pp (Rd ), so the p-optimal transport problem is well posed and Wp (µ0 , µ1 ) < ∞. For any r, r′ ∈ [0, 1], the pushforward (1 − r)x0 + rx1 , (1 − r′ )x0 + r′ x1 # π T is a coupling of µr and µr′ . Hence, using the optimality of π T at the last step, Z p Wp (µr , µr′ ) ≤ ∥(r − r′ )(x1 − x0 )∥p dπ T = |r − r′ |p Wp (µ0 , µ1 )p .
(C.2)
Now fix 0 ≤ r ≤ r′ ≤ 1. Applying (C.2) to the pairs (0, r), (r, r′ ), and (r′ , 1), we obtain Wp (µ0 , µr ) + Wp (µr , µr′ ) + Wp (µr′ , µ1 ) ≤ r Wp (µ0 , µ1 ) + (r′ − r) Wp (µ0 , µ1 ) + (1 − r′ ) Wp (µ0 , µ1 ) = Wp (µ0 , µ1 ). On the other hand, the triangle inequality yields Wp (µ0 , µ1 ) ≤ Wp (µ0 , µr ) + Wp (µr , µr′ ) + Wp (µr′ , µ1 ). Therefore equality holds throughout, and in particular, Wp (µr , µr′ ) = |r − r′ | Wp (µ0 , µ1 )
for all r, r′ ∈ [0, 1].
Thus (µTt )t∈[0,1] is a constant-speed Wp -geodesic between p0 (· | T[0,1] ) and pdata . Remark C.4. For each fixed clock path with finite covariance matrix, p0 (· | T[0,1] ) = N (0, Σ(T[0,1] )) belongs to Pq (Rd ) for every q ≥ 1. Thus Theorem C.3 is pathwise in T : after conditioning on the clock, the source side does not restrict the choice of p, and the admissible p is governed by the data moment assumption. This does not imply that the clock-marginal source law has all moments; after integrating out T , stable or Student-t clock choices can again produce heavytailed sources with limited or infinite moments. The classical W2 -statement is recovered when the data law has a finite second moment and the argument is applied pathwise.
C.5
Proof of Theorem 3.2
Proof of Theorem 3.2. By (3.9) and the source covariance ΣA (T[0,1] ) := 2A(T )Id , where A(T ) := R1 T 0 Ts ds, for every fixed clock path and endpoint X1 = x1 , XtT | X1T = x1 ∼ N
αt x1 , 2βt2 A(T )Id ,
Z 1 A(T ) =
Ts ds. 0
Thus the clock-marginal conditional law is a Gaussian scale mixture. For the deterministic clock Tt = t, A(T ) = 1/2, so the covariance above is βt2 Id . This proves the Gaussian case. For the stable case, let αR ∈ (0, 2), ρ = α/2, and let T be the ρ-stable subordinator in 1 the theorem. With A(T ) = 0 Ts ds, the integration-by-parts identity for Lebesgue–Stieltjes integrals gives Z 1
(1 − s) dTs .
A(T ) = 0
27
Using the independent increments of the subordinator, for any λ ≥ 0, ρ Z h i λ ρ+1 1 (λ(1 − s))ρ ds = exp − . E e−λA(T ) = exp − 2 2 0 Therefore, for any ξ ∈ Rd , h i E ei⟨ξ,Xt ⟩ | X1 = x1 = exp iαt ⟨ξ, x1 ⟩ ET exp −βt2 A(T )∥ξ∥22 βtα ∥ξ∥α2 . = exp iαt ⟨ξ, x1 ⟩ exp − 2
(C.3)
(C.4)
This is the characteristic function of a location-shifted isotropic α-stable law with location αt x1 and scale βt , under the normalization used in the main text. The radial tail asymptotics P(∥Xt − αt x1 ∥2 > r | X1 = x1 ) = Θ(r−α ) as r → ∞ then follow from the standard tail behavior of isotropic α-stable laws [40]. For the Student-t case, Tt = tVν gives A(T ) = Vν /2, and hence ΣA (T[0,1] ) = Vν Id . Since Vν ∼ InvGamma(ν/2, ν/2), marginalizing over Vν yields Z ∞ 1 ∥x − αt x1 ∥22 (ν/2)ν/2 −ν/2−1 −ν/(2v) pXt |X1 =x1 (x) = exp − v e dv Γ(ν/2) 2βt2 v (2πβt2 v)d/2 0 −(ν+d)/2 ∥x − αt x1 ∥22 Γ((ν + d)/2) 1+ . (C.5) = νβt2 Γ(ν/2)(νπ)d/2 βtd This is the multivariate Student-t density with location αt x1 , scale βt2 Id , and ν degrees of freedom, whose radial tail satisfies P(∥Xt − αt x1 ∥2 > r | X1 = x1 ) = Θ(r−ν ) as r → ∞ [40].
D
Additional Information on Log-Signature
This appendix fixes the signature-feature convention used in Section 3.2. The main paper specializes the clock feature to the truncated logsignature, ηT = ℓm (T ). For completeness we keep the generic notation ηT throughout this appendix and discuss alternative choices (raw signatures, lead–lag transforms, or concatenations with simple summaries). The feature ηT has two roles: it conditions the learned vector field uθ (XtT , t, ηT ), and, when a clock-dependent trajectory is used, it may also parameterize the interpolation schedules αtT and βtT . Throughout this appendix, the feature is a summary of the full realized clock path T[0,1] , sampled before drawing the clock-conditioned source and before solving the generation ODE. Let z : [0, 1] → Rq be a continuous path of bounded variation. Its signature is the tensor series Z Z S(z) = 1, dzu1 , dzu1 ⊗ dzu2 , . . . , 0<u1 <1
0<u1 <u2 <1
and its truncation to order m is denoted by S (≤m) (z). The truncated logsignature is LogSig(≤m) (z) := log⊗ S (≤m) (z) , where log⊗ is computed in the truncated tensor algebra. Unlike the raw signature, the logsignature lies in the truncated free Lie algebra and removes many algebraic redundancies among tensor coordinates. In our application, q = 2 and the path is the time-augmented clock Ts = (s, Ts ). Since the random clocks considered in the paper may be càdlàg and may have jumps, we do not rely on a canonical signature of the jump path itself. Instead, for a discretization 0 = t0 < t1 < · · · < tN = 1, we form the sampled time-augmented path Tt0 , Tt1 , . . . , TtN = (t0 , Tt0 ), (t1 , Tt1 ), . . . , (tN , TtN ) 28
and compute signatures after linearly interpolating these sampled points. With this convention, ϕm (T ) := S (≤m) (Tlin ) is the truncated signature feature used in (3.6), where Tlin denotes the piecewise linear interpolation of the sampled time-augmented clock. A logsignature version of the same feature is ℓm (T ) := LogSig(≤m) (Tlin ). Thus the model feature ηT may be chosen as ϕm (T ), as ℓm (T ),Ror as a concatenation of ei1 ther representation with simple clock summaries such as T1 or 0 Ts ds. When invoking the signature-linear approximation result in Proposition D.1, one should take the raw signature coordinates ϕm (T ), because the proposition is stated for linear functionals of signatures. In neural-network conditioning, we often use the logsignature ℓm (T ), since it is more compact; using logsignature coordinates in the learned schedules should therefore be understood as a practical finite-dimensional parameterization, not as the literal raw-signature universality statement. Under the piecewise linear convention, a jump in the sampled clock is represented by a steep linear segment between two consecutive observation times. Other conventions, such as inserting explicit vertical jump segments or using a lead–lag transform, would define different finite-dimensional features. They do not change the flow-matching identities in the main text: after conditioning on T[0,1] , the path remains Gaussian as in (3.9), and the endpoint-conditioned target remains the affine expression (3.10). The feature ηT only summarizes the realized clock environment for the network and, when used, for the schedule parameterization. The first logsignature level is the increment of the time-augmented path, so it records (1, T1 − T0 ), which is (1, T1 ) under the standing assumption T0 = 0. Higher levels encode how clock mass is allocated over physical time. For example, in two dimensions the second Lie level corresponds to the signed area between the physical-time coordinate and the clock coordinate; for monotone clocks this separates paths with the same terminal value T1 but different temporal allocations of their increments. This is why low-order signature information can represent path-dependent R1 functionals such as the area-clock quantity 0 Ts ds used in Section 3.3. The dimension reduction from signature to logsignature Pm k can be substantial. For a qdimensional path, the raw truncated signature has 1 + k=1 q coordinates including the constant coordinate, whereas the nonconstant truncated logsignature has dimension m X
ℓq (k),
ℓq (k) :=
k=1
1X µ(r)q k/r , k r|k
where µ is the Möbius function. Thus, for the time-augmented clock (q = 2) and truncation order m = 3, the raw nonconstant signature has dimension 2+4+8 = 14, while the logsignature has dimension 2 + 1 + 2 = 5. This compactness is the main reason that the experiments use truncated logsignature features for ηT when conditioning the neural velocity field.
D.1
Signature-linear schedules
This subsection gives the schedule details that are summarized in Section 3.3. The compactness assumption in Proposition D.1 below is the usual setting for uniform approximation. For heavytailed clocks, which need not live on a compact path set globally, the proposition should be read either on compact truncations of the clock space, on high-probability compact subsets, or after applying the feature normalizations used in training. The time augmentation is important: it R1 records the order in which clock mass is accumulated, so path functionals such as 0 Ts ds can be represented by low-order signature information.
29
Proposition D.1 (Signature-linear approximation of clock functionals). Let K be a compact set of time-augmented, piecewise linear clock paths T = (s, Ts )s∈[0,1] on which the signature separates paths up to tree-like equivalence. Let F : [0, 1] × K → R be continuous. Then for every ε > 0, there exist a truncation level m and a continuous coefficient curve c : [0, 1] → RDm such that sup sup |F (t, T ) − ⟨c(t), ϕm (T )⟩| < ε. (D.1) t∈[0,1] T ∈K
In particular, clock-dependent schedule functionals such as T 7→ αt (T ) and T 7→ βt (T ) can be represented, up to arbitrary uniform accuracy on compact clock families, by linear functionals of a sufficiently high-order truncated signature: αt (T ) ≈ ⟨a(t), ϕm (T )⟩,
βt (T ) ≈ ⟨b(t), ϕm (T )⟩.
(D.2)
The proposition motivates representing clock-dependent interpolation coefficients by signaturelinear functions. A direct form is αtT,m = ⟨a(t), ϕm (T )⟩,
βtT,m = ⟨b(t), ϕm (T )⟩,
(D.3)
where a(t), b(t) ∈ RDm are coefficient curves. If these curves are absolutely continuous in t, then for a fixed realized clock path the physical-time derivatives are explicit: dαtT,m da(t) dβtT,m db(t) = , ϕm (T ) , = , ϕm (T ) . (D.4) dt dt dt dt Thus the flow-matching target differentiates the scalar-time coefficient curves, not the clock path itself. The direct form (D.3) does not automatically enforce the endpoint constraints α0T = 0, β0T = 1, α1T = 1, β1T = 0, βtT > 0 for t ∈ [0, 1). A boundary-corrected near-linear schedule is obtained as follows. Let λ : [0, 1] → R be absolutely continuous with λ0 = 0, and set At (T ) := ⟨a(t), ηT ⟩,
Bt (T ) := ⟨b(t), ηT ⟩.
(D.5)
βtT = 1 − t + λt (1 − t)Bt (T ) = (1 − t) 1 + λt Bt (T ) .
(D.6)
Then define αtT = t + λt (1 − t)At (T ),
The endpoint constraints are automatic: α0T = 0,
α1T = 1,
β0T = 1,
β1T = 0.
(D.7)
The path is admissible whenever t ∈ [0, 1).
1 + λt Bt (T ) > 0,
(D.8)
Writing dAt (T ) = dt
da(t) , ηT dt
dBt (T ) = dt
,
db(t) , ηT dt
,
the required derivatives are dαtT dλt dAt (T ) =1+ (1 − t) − λt At (T ) + λt (1 − t) , dt dt dt dβtT dλt dBt (T ) = − 1 + λt Bt (T ) + (1 − t) Bt (T ) + λt , dt dt dt dβtT dt βtT
(D.9) (D.10)
dB (T )
=−
dλt t Bt (T ) + λt dt 1 + dt 1−t 1 + λt Bt (T )
30
.
(D.11)
The singularity at t = 1 is the same terminal singularity present in the straight-line flowmatching target and is irrelevant for the time-integrated objective, where t = 1 is not sampled. A symmetric special case uses a single clock-dependent time warp. Let γtT := t + λt (1 − t)Ct (T ),
(D.12)
βtT = 1 − γtT = (1 − t) 1 − λt Ct (T ) .
(D.13)
Ct (T ) := ⟨c(t), ηT ⟩, and set αtT = γtT ,
This preserves the straight-line form XtT = γtT X1T +(1−γtT )X0T , while allowing the realized clock path to accelerate or decelerate the interpolation. The endpoint-conditioned target simplifies to T vtT (x | x1 ) =
D.2
dγt dt
1 − γtT
(x1 − x),
t ∈ [0, 1).
(D.14)
VE-type signature-linear schedule
For the clock-aware trajectory ablation, we use a variance-exploding (VE)-type affine path written in the signature-linear form (D.3). This is a deterministic affine flow path inspired by VE noise schedules, not an additional stochastic SDE. Since all clocks in this experiment satisfy T1 > 0, define the normalized clock T̄s := Ts /T1 and the normalized time-augmented path T̄s = (s, T̄s ). Let ϕm (T̄ ) denote the raw truncated signature of T̄. For m ≥ 2, let ect and etc be the unit vectors selecting the level-two signature coordinates with words (c, t) and (t, c), where t denotes physical time and c denotes the normalized clock coordinate. Then Z Z 1 ect , ϕm (T̄ ) = dT̄u dv = T̄u du, (D.15) 0<u<v<1 0 Z Z 1 etc , ϕm (T̄ ) = du dT̄v = 1 − T̄v dv. (D.16) 0<u<v<1
0
Thus the signed clock-shape score Rm (T ) := rm , ϕm (T̄ ) , satisfies
rm := ect − etc ,
(D.17)
Z 1 T̄s ds − 1.
Rm (T ) = 2
(D.18)
0
It is zero for a linear normalized clock, positive when clock mass accumulates early, negative when it accumulates late, and bounded by [−1, 1] for nondecreasing normalized clocks. We use the normalized linear VE decay σVE (t) = 1 − t and introduce independent signaturelinear clock-shape corrections: αtT,m = t + λα t(1 − t)Rm (T ), βtT,m = (1 − t) 1 + λβ tRm (T ) . (D.19) Equivalently, with e∅ denoting the constant signature coordinate, aVE (t) = te∅ + λα t(1 − t)rm ,
(D.20)
bVE (t) = (1 − t)e∅ + λβ t(1 − t)rm ,
(D.21)
so that αtT,m = aVE (t), ϕm (T̄ ) ,
βtT,m = bVE (t), ϕm (T̄ ) .
(D.22)
In the experiments we use λα = λβ = 1/4. More generally, since |Rm (T )| ≤ 1, any |λβ | < 1 gives βtT,m ≥ (1 − t)(1 − |λβ |) > 0 for t < 1, so the affine path is admissible. 31
Writing Rm = Rm (T ), the derivatives are dαtT,m = 1 + λα (1 − 2t)Rm , dt and hence
dβtT,m dt LT,m := T,m t βt
=
dβtT,m = −1 + λβ (1 − 2t)Rm , dt −1 + λβ (1 − 2t)Rm . (1 − t) 1 + λβ tRm
(D.23)
(D.24)
Substituting these quantities into (3.10) gives vtT,m (x | x1 ) = LT,m x+ t
dαtT,m − LT,m αtT,m t dt
! x1
h i T,m = LT,m x + 1 + λ (1 − 2t)R − L t + λ t(1 − t)R x1 , α m α m t t
t ∈ [0, 1). (D.25)
When λα = λβ = 0, this reduces to the linear VE path αt = t, βt = 1 − t. The one-warp path is recovered only in the special symmetric case λβ = −λα ; the ablation above uses the more general affine form XtT = αtT,m X1T + βtT,m X0T .
E
Experimental Setup
This section documents the full configuration shared across all reported HTFM and baseline runs. Unless explicitly noted, any HTFM clock variant (Gaussian, α-stable, Student-t) uses the same backbone and optimizer for a given dataset; the only differences are the random-clock law and the optional path-feature conditioning. All experiments are implemented in PyTorch.
E.1
Dataset
We use three benchmarks at increasing scale. 2D imbalanced α-stable mixture. A 9-component isotropic α-stable mixture in R2 . Components sit on a uniformly spaced 3×3 grid with spacing 2.0, each component is a 2-D isotropic α-stable variable with scale 0.1, and the mixture weights are: {0.18, 0.08, 0.14, 0.06, 0.09, 0.10, 0.16, 0.07, 0.12} (range 0.06–0.18). We sweep the per-component stability index αD ∈ {1.5, 1.8} and draw 32,000 training samples per setting; an additional 24,000-sample reference set is held out for evaluation. CIFAR10-LT. The long-tailed CIFAR-10 split [6] of the original CIFAR-10 train set [26]. With imbalance ratio ρ = 0.01, head count 5000, and 10 classes, the per-class counts follow a geometric schedule and total ≈ 12,406 training images at 32×32. A deterministic subset is precomputed and reused across all runs to keep the baselines identical at the dataset level. Random horizontal flips are the only augmentation. HRRR VIL. Vertically Integrated Liquid (VIL) hourly analyses from the High-Resolution Rapid Refresh weather system [5], cropped to a 128×128 window over the Central US following the protocol of [36]. Train/test split mirrors theirs: years {2019, 2020} for training (≈ 17,400 valid hourly fields after dropping NaN slots) and year {2021} for test (≈ 8,700). Per-channel z-score statistics are computed on the training years and applied to both splits.
32
E.2
Baselines
We compare HTFM against four heavy-tailed generative baselines, each run under the same backbone and training budget as HTFM (see Sections E.4 and E.5): • LIM [45, 39]: Lévy-Itô stochastic differential equation with α-stable driving noise; trained with a fractional-score regression loss and sampled with the model’s α-stable reverse SDE solver at reverse-step budget 1000. • DLPM [42]: discrete-time α-stable analogue of DDPM; uses 4000 training-time reverse steps with the Gaussian scale-mixture augmentation, sampling via its native stochastic predictor. • DLIM [42]: deterministic α-stable sampler sharing the DLPM predictor backbone; We set η=0 for deterministic generation. • t-Flow [36]: heavy-tailed flow with a finite-variance Student-t source; reported only on the Student-t parameter range pν > 2 where its construction is well-defined.
E.3
Evaluation
Each benchmark uses the standard tail-aware metric for its scale. 2D mixture. We use the precision–recall f1pr score of [42]. Let Xref = {xi }N i=1 and Xgen = N {yj }j=1 be the reference and generated sample sets (N = 24,000 here). For a set X and a point z, write rk (z; X ) for the distance from z to its k-th nearest neighbor in X , and define the k-NN manifold [ Mk (X ) = B z, rk (z; X ) , z∈X
i.e. the union of ℓ2 -balls around each sample with radius given by its own k-th nearest-neighbor distance. The precision and recall relative to this support estimate are P =
X 1 1 y ∈ Mk (Xref ) , |Xgen |
R =
y∈Xgen
X 1 1 x ∈ Mk (Xgen ) , |Xref | x∈Xref
and the score is their harmonic mean f1pr =
2PR . P+R
Higher is better. We use k = 10 throughout. CIFAR10-LT. We report Fréchet Inception Distance (FID) [16], the standard image-generation metric. Each 32×32 image (real and generated) is passed through a pre-trained Inception-V3 network and embedded as a 2048-dimensional feature vector at the final average-pooling layer. Letting (µdata , Σdata ) and (µgen , Σgen ) be the empirical mean and covariance of the two feature sets, FID is the squared 2-Wasserstein distance between the two fitted Gaussians, FID = ∥µdata − µgen ∥22 + Tr Σdata + Σgen − 2(Σdata Σgen )1/2 . (E.1) The reference set is the full long-tail train split (12,406 images at 32×32); the generated set is 12,406 images sampled from the EMA-averaged model copy at the listed NFE budget. Lower is better.
33
HRRR VIL. Following [36], we report three one-dimensional tail statistics aligned to the operational interest in HRRR extremes. We generate 20,000 samples from each model, flatten them together with the ≈ 8,760 test fields into 1-D intensity samples, and compute the metrics below. Lower is better for all three. Kurtosis Ratio (KR). Sample kurtosis is the (normalized) fourth-order moment, characterizing how heavy the distribution’s tails are relative to the central mass. Let kdata and ksim denote the empirical kurtosis of the reference and generated samples. The kurtosis ratio is sim KR = 1 − kkdata .
(E.2)
Skewness Ratio (SR). Sample skewness is the (normalized) third-order moment, characterizing the asymmetry of a tailed distribution. With sdata , ssim the empirical skewness on reference and generated, sim SR = 1 − ssdata . (E.3) Tail Kolmogorov–Smirnov (KS). The two-sample KS statistic [34] measures the supnorm distance between two empirical CDFs. For heavy-tailed distributions we evaluate it on the tails to focus the metric on the rare-extreme regime that both KR and SR weight only indirectly. We retain generated and reference samples lying above the 99.9-th percentile (right tail) or below the 0.1-th percentile (left tail), compute the two-sample KS statistic per side, and average over the two sides for an overall score. For the VIL channel the underlying distribution has no left tail, so we report the right-tail KS statistic only.
E.4
Network Architecture
2D mixture. A small MLP with 4 residual blocks of width 64, SiLU activations, GroupNorm, and a learnable time embedding of size 32. The same backbone is shared across HTFM and the FM baselines. CIFAR10-LT. The DDPM-style U-Net of [18]: 128 base channels, channel multipliers (1, 2, 2, 2), two residual blocks per scale, 4 attention heads, self-attention at resolution 16, dropout 0. The same U-Net is used by HTFM, LIM, DLPM, DLIM, and t-Flow so that any FID difference reflects the noising/clock formulation rather than backbone capacity. For HTFM, the U-Net is additionally equipped with the path-feature conditioning module described below. HRRR VIL. The NCSN++ backbone of [43] on 128×128 single-channel inputs, with base channels 32, channel multipliers (1, 2, 2, 4, 4), four residual blocks per scale, 4 attention heads, self-attention at resolution 16, FIR (1, 3, 3, 1) up/down-sampling, swish nonlinearities, GroupNorm, and BigGAN-style residual blocks. Sigma range [0.01, 50] on 1000 noise scales when used with score-based baselines. Path-feature conditioning (HTFM-specific). HTFM’s velocity network conditions on a clock-derived path feature ϕ that summarizes the realized clock trajectory of each sample. The feature source depends on the random clock: • Student-t clock Tt =tVν , Vν ∼ InvGamma(ν/2, ν/2). The path is fully determined by the single scalar Vν , so we take ϕ= log Vν ∈ R (dϕ =1); no signature is needed. • α-stable subordinator clock. We discretize the time-augmented path {(ti , Tti )}N i=0 on a fixed N =200 grid, standardize each of the two channels along time, and take its truncated logsignature of order m (default m=1, i.e. ϕ ∈ R2 collecting the increments of the two channels). The construction matches Section 3.2.
34
• Deterministic clock Tt =t (HTFM-G, p=∞). No path feature; the conditioning module is omitted entirely and the network reduces to a standard time-conditioned velocity field. The raw feature ϕ ∈ Rdϕ is mapped through a shared two-layer MLP ϕpath : Rdϕ → R128 with SiLU activations, dropout 0, and an internal hidden width of 128. The 128-dimensional path embedding is injected into the velocity network at every residual block by concatenation with the time embedding. In the original block the time embedding temb ∈ Rdt is mapped to a per-channel scale/shift via a linear Wt ∈ RCout ×dt ; with path conditioning this is replaced by a wider linear Wt,ϕ ∈ RCout ×(dt +128) acting on the concatenation [temb ∥ ϕpath ]. The same concatenation-at-every-block scheme is used identically by the DDPM U-Net (CIFAR10-LT), the NCSN++ U-Net (HRRR VIL), and the residual-MLP backbone (2D mixture); this keeps the comparison against LIM/DLPM/DLIM/t-Flow clean, since their backbones have no analogous trajectory-summary input and therefore see only the time embedding. For completeness, our implementation also supports a zero-initialized FiLM modulation interface that produces a per-channel (scale, shift) from ϕpath and applies it at the middle block. All HTFM rows reported in the paper use only the concatenation interface; FiLM is left disabled.
E.5
Training
All runs use AdamW with default β’s. Per-dataset hyperparameters are listed below. 2D mixture. 100.
Learning rate 5×10−3 , batch size 1024, EMA decay 0.99, total number of epochs
CIFAR10-LT. Learning rate 2×10−4 with 500 warmup steps and a step-LR schedule (decay factor 0.99 every 1000 steps); batch size 100; EMA decay 0.9999 used for evaluation; gradient clipping at ℓ2 norm 1.0. Models are trained for 2000 epochs. HRRR VIL. Learning rate 10−4 with 1000 warmup steps and no LR decay; weight decay 0; batch size 128; EMA decay 0.9999; gradient clipping at 5.0. Models are trained for 1500 epochs. HTFM-specific knobs. For HTFM the random clock is one of: deterministic Tt = t (Gaussian limit, HTFM-G), an α-stable subordinator with decay degree pα ∈ {1.6, 1.7, 1.8, 1.9} on CIFAR10-LT and pα ∈ {1.6, 1.7, 1.8} on HRRR, or a Student-t random-slope clock Tt = tVν , Vν ∼ InvGamma(ν/2, ν/2), with decay degree pν ∈ {1.7, 2.0, 3.0} on CIFAR10-LT and HRRR (extended to pν = 1.5 on the 2D mixture where a degenerate variance is acceptable). The training loss is the clock-conditioned L2 flow-matching objective with velocity prediction.
E.6
Sampling
All HTFM and FM-based variants integrate the learned velocity field with a fixed-step ODE solver; LIM, DLPM, and DLIM use their respective (stochastic) reverse processes. 2D mixture. Deterministic Euler ODE solver; 24,000 samples generated per setting and compared to a held-out 24,000-sample reference set. CIFAR10-LT. Deterministic Euler ODE solver for HTFM and the baselines. DLPM uses its native predictor; DLIM uses the deterministic (η=0) variant. Higher-order solvers are reported only in the solver ablation (Appendix F.3). Unless otherwise stated, all FID results use the final EMA checkpoint after the fixed training budget.
35
HRRR VIL. Heun’s second-order ODE solver at 25 reverse steps (NFE = 50 per sample), following [36]. Tail statistics are evaluated on 20,000 generated samples.
E.7
Compute Resources
All experiments were conducted on a mix of NVIDIA L40S (48GB) and A100 (80GB) GPUs. CIFAR10-LT runs use up to 8 L40S GPUs in parallel, while HRRR VIL runs use A100 80G GPUs to accommodate the larger 128 × 128 single-channel inputs and the NCSN++ backbone. The total training budget is fixed by the per-dataset epoch counts reported in Appendix E.5 (100 epochs for the 2D mixture, 2000 epochs for CIFAR10-LT, and 1500 epochs for HRRR VIL). Sampling cost is dominated by the listed NFE budgets and the ODE/SDE solver choices in Appendix E.6.
F
Ablation Studies
Unless otherwise stated, ablations are conducted on CIFAR10-LT because it exposes both global generation quality and tail-class coverage while remaining more economical than the HRRR sweep. The same protocol can be repeated on 2D toy distributions for visualization and on HRRR for the final robustness check.
F.1
Effect of Flow Type
We next ablate the interpolation schedule itLevel 0 Level 1 Level 2 self, holding the random clock and the source Flow type covariance fixed. We compare two flow types Straight line 16.65 16.10 16.46 implemented in our codebase: (i) the stan- Signature-linear VE 16.43 16.19 16.71 dard straight-line interpolation XtT = tX1 + (1 − t)X0T , which ignores the realized clock Table 4: Flow-type ablation on CIFAR10-LT and moves linearly in physical time t; and (ii) under HTFM-α at pα = 1.7. We compare a signature-linear VE schedule whose inter- two clock-conditioned interpolation schedules in polation coefficients are convex combinations FID performance: the clock-blind straight line parameterized by truncated logsignature fea- and a signature-linear VE schedule whose mixtures of the clock path, with mixing weights ing weights depend on log-signature features λα , λβ controlling how aggressively the sched- of the realized clock path. Columns vary the ule deviates from straight line. Comparing signature order provided to the velocity netthese two schedules under a matched signa- work. Each cell reports the FID at sampling ture feature isolates the value of clock-aware NFE= 100 using the Euler ODE solver. versus clock-blind interpolation. We run this ablation on the HTFM-α family at pα = 1.7 as a representative slice of the heavy-tailed regime.
F.2
Effect of Number of Function Evaluations (NFE)
We next vary the number of function evaluations (NFE) used by the sampler, holding training fixed and changing only the sampling budget. This ablation measures each method’s quality– cost tradeoff: Figure 2 plots FID against NFE for HTFM-α, DLPM, and DLIM at three decay degrees pα ∈ {1.6, 1.7, 1.8}. Three trends are visible across all three panels. (i) HTFM-α dominates at low NFE. The red curve sits well below both baselines from the smallest sampling budget, and the gap to DLPM is largest precisely where compute is scarcest – consistent with HTFM inheriting the low-NFE sample efficiency of flow matching rather than the slow burn-in of α-stable diffusion. (ii) DLPM closes the gap only at high NFE. The blue curve falls steeply with more solver steps and crosses or matches HTFM-α near the right edge of the budget range, indicating that 36
p = 1.7
p = 1.6 26
24
24
22
22
20
20
18
18
16
16 20
50
100
NFE
200
500
HTFMDLPM DLIM
30.0 27.5
FID
26
FID
FID
28
p = 1.8
32.5
28
25.0 22.5 20.0 17.5
20
50
100
NFE
200
500
15.0
20
50
100
NFE
200
500
Figure 2: NFE-vs-FID curves on CIFAR10-LT. One panel per decay degree pα ∈ {1.6, 1.7, 1.8}; lines compare HTFM-α (ours) with the DLPM and DLIM baselines. Lower FID is better. Solver Euler (1st order) Heun (2nd order) Midpoint (2nd order) AB2 (2nd order, multi-step)
NFE = 20
NFE = 50
NFE = 100
20.80 18.55 17.82 18.02
17.92 17.44 16.51 16.47
16.31 15.87 15.81 15.82
Table 5: ODE-solver ablation on CIFAR10-LT under HTFM-α at pα = 1.7. Each cell reports FID (12,406 generated vs 12,406 reference, EMA, lower is better) at the listed total NFE budget; for the higher-order solvers the number of integration steps is ⌊NFE/NFE-per-step⌋. Bold indicates per-column best. the two methods become broadly comparable at the asymptotic-quality end of the curve. The crossover happens later for heavier tails (smaller pα ), where the sampler has more high-frequency content to integrate through. (iii) DLIM plateaus. The green curve flattens early and stays above both other methods for the rest of the range, so its deterministic sampler does not scale further with extra NFE on this benchmark. Together these curves give the same message as the main-text CIFAR10-LT comparison: HTFM-α matches the asymptotic quality of the strongest α-stable diffusion baseline while strictly dominating it under tight NFE budgets.
F.3
Effect of ODE Solver
This ablation compares the integration scheme used during sampling, holding the trained model fixed and varying only the solver and the matched NFE budget. We compare four fixed-step ODE solvers implemented in our codebase: explicit Euler (1st order), Heun and midpoint (2ndorder single-step Runge–Kutta), and Adams–Bashforth AB2 (2nd-order multi-step). Higherorder solvers reduce the per-step discretization error of the velocity integration, but each step costs more network evaluations: a step of Heun/midpoint/AB2 costs 2 NFE. At a matched total NFE budget the higher-order solvers therefore take fewer integration steps, so the ablation captures the per-NFE quality–accuracy tradeoff rather than a per-step comparison. We report HTFM-α at pα = 1.7 as a representative slice.
G
Visualizations
This section collects qualitative visualizations that complement the quantitative results in the main text.
37
G.1
2D toy distributions: clock-degree sweep
Figure 3 compares HTFM-α samples generated under clocks of different decay degrees (rows: p = ∞, pα = 1.8, pα = 1.5) against ground-truth samples (top row) across five 2D toy targets (columns: checkerboard, two moons, swiss roll, rings, olympic rings). The same backbone, training schedule, and NFE budget are used for every cell; only the random-clock law is swept.
Figure 3: HTFM-α samples on five 2D toy distributions (checkerboard, two moons, swiss roll, rings, olympic rings) under decay exponents p ∈ {∞, 1.8, 1.5} (α in the figure denotes the α-stable stability index of the random clock; the top row is ground truth). Same backbone, training, and NFE across cells.
G.2
2D SαS mixture: baseline comparison at NFE=20
Figure 4 compares HTFM against DLPM and DLIM on the 2D imbalanced α-stable mixture at NFE=20, on two stability indices αdata ∈ {1.5, 1.8} (rows). Each panel overlays generated samples (blue) on the corresponding real data (orange in the leftmost column). HTFM reproduces the imbalanced grid pattern more cleanly.
G.3
CIFAR10-LT: generated samples
Figure 5 shows generated 32×32 samples from HTFM-α on CIFAR10-LT, organised so that each column collects samples from a single CIFAR-10 class.
38
Figure 4: Generated samples on the 2D imbalanced α-stable mixture at NFE=20. Rows: αdata = 1.8 (top), αdata = 1.5 (bottom). Columns: real data, HTFM, DLPM, DLIM. W1 distances to the real data are reported below each generated panel.
39
Figure 5: CIFAR10-LT generated samples from HTFM-α. Each column corresponds to one CIFAR-10 class; from left to right: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck. This is also the CIFAR10-LT head-to-tail order, i.e. per-class training-sample counts decrease from the leftmost to the rightmost column. Sample fidelity broadly tracks the perclass training-set size: the head classes airplane and automobile (leftmost two columns) produce the sharpest, most coherent images, while progressively rarer classes (rightward) exhibit more artefacts and weaker class identity.
40