ConceptioArchivearXiv CS
arXiv CSopen access

LatentFlow: A General Framework for Conditioning Stochastic Processes

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Preprint. Work in progress.

LatentFlow: A General Framework for Conditioning Stochastic Processes Louis Sharrock University College London

Lachlan Astfalck University of New South Wales

Henry Moss Lancaster University

arXiv:2607.12922v1 [stat.ML] 14 Jul 2026

Abstract Stochastic-process models are, as a rule, far easier to simulate than to condition. Non-linear observations, non-Gaussian likelihoods, black-box information, and global constraints all induce intractable conditional laws, requiring bespoke, model-specific constructions. We introduce LatentFlow, a single framework for conditioning stochastic processes, with no learned neural approximations and no training. Our starting point is to write the stochastic process as the deterministic image of a tractable latent innovation, f0 = Tϑ (ξ0 ), with ξ0 sampled from a simple reference distribution. This reduces process-level conditioning to latent-space inference: pull the likelihood back through Tϑ , sample the resulting latent law with a tractable guided probability flow, and push the samples forward. This construction is provably exact at the level of the target law; in practice, approximation enters only through finite terminal noising, Monte Carlo guidance, and time discretisation of the continuous-time dynamics, each of which is explicit and systematically reducible. As LatentFlow is training-free, conditioning reduces to solving a single reverse-time SDE. This enables conditional sampling in seconds on a single desktop CPU across model classes that have never shared a scalable method: classical spatial priors, nonlinear stochastic dynamics, mechanistic models from the physical and life sciences, stochastic PDEs, heavy-tails and extremes, point and discrete-state processes, and neural or simulator-defined processes.

1

Introduction

Sampling from a stochastic process only requires running it forward; conditioning it, until now, has remained the preserve of bespoke, model-specific constructions. Examples include diffusion paths conditioned on fixed endpoints or partial observations (Delyon & Hu, 2006; Schauer et al., 2017); functions with shape (Riihimäki & Vehtari, 2010) or physics-informed (Chen et al., 2021; Hamelijnck et al., 2024) constraints; and simulator outputs conditioned on indirect measurements, rare events, or user-defined criteria (Botev & L’Ecuyer, 2020; Finzi et al., 2023). In each case the conditioning information cannot be absorbed into the model in closed form, demanding its own specialised, expensive, or approximate sampler; for instance, those based on MCMC (Andrieu et al., 2003), SMC (Doucet et al., 2001), variational inference (Blei et al., 2017) or Laplace approximations (Rue et al., 2009). Recently, Moss et al. (2026) introduced FlowGP, which addresses conditioning for Gaussian process (GP) priors under intractable information by recasting conditional sampling as a differential equation with closed-form Gaussian dynamics and a likelihood-dependent guidance term. The resulting sampler applies to any likelihood that can be evaluated pointwise, and requires no trained network to approximate the guidance field. Underlying the model is a simple mechanism: the GP sample is the linear, deterministic image of a Gaussian innovation, and conditioning is a re-weighting of that innovation’s density by the pulled-back likelihood. Our key observation is that Gaussianity of the stochastic process prior is not the essential ingredient in this construction; instead, it is the availability of a tractable latent innovation representation on which guided dynamics can be defined. Rather than constructing 1

Preprint. Work in progress.

conditioning dynamics directly in the original process space, this allows us to express the process as the deterministic image of a tractable innovation variable, pull the conditioning information back through this generator, and sample the resulting reweighted law in latent space. The conditioned process is then obtained by pushing the guided samples in the latent space forward through the same generator. We refer to this method as LatentFlow. Our construction does not require learning a process prior, score model, or conditional generator (e.g., Song et al., 2021; Chung et al., 2023; Zammit-Mangion et al., 2025); the sampler is built directly from the known map Tϑ and pointwise evaluations of the pulled-back likelihood. Our method also provides a natural way to learn the parameters of the process generator. We obtain formal theoretical guarantees for LatentFlow, and demonstrate its application to a wide variety of stochastic process priors. These include GPs; heavy-tailed Student-t and Cauchy-convolution processes; discrete-valued Potts processes; linear and nonlinear diffusion processes including a model for cell-differentiation, a FitzHugh–Nagumo model for excitable neurons, and a Heston stochastic volatility model; mechanistic models of SIR epidemics and Lotka–Volterra predator–prey dynamics; and Allen–Cahn and advection–diffusion SPDEs. See Figure 1 for several examples. Until now, beyond simple or contrived cases, conditioning each of these processes required either low-fidelity approximations or expensive computation. We sample in single-seconds time on a single desktop CPU, providing the first general method for conditioning stochastic processes in real time which is exact up to an explicit and systematically reducible numerical error.

Gaussian process (censored) Student-t (areal mean)

Cauchy (areal max)

Heston (Market Shock)

Lotka-Volterra (prey boom)

Spatial processes

SIR (fast epidemic)

price vol

S I R

prey pred

Dynamical systems

peak I > 0.32 final R > 0.62 min price < 87 max vol > 0.42

Allen-Cahn (phase front) τ=0

Spatiotemporal τ

= 0.66

τ = 0.31

τ=1

peak prey > 7

FitzHugh (sequential firing) τ=0

τ = 1.3

τ = 2.7

τ=4

Advection (terminal detctor hit) τ=0

τ = 0.71

τ = 1.4

τ = 2.2

Figure 1: LatentFlow across diverse stochastic-process priors and conditions. Top: spatial processes conditioned on censored point measurements for a Gaussian process; prescribed areal means for a Student-t process; and prescribed areal maxima for a Cauchy convolution field. Centre: SDEs conditioned on observations and events. A Heston model under a joint market-crash and volatility spike; an SIR model under a fast epidemic peak; and a Lotka–Volterra system given a large prey boom. Bottom: Spatio-temporal PDEs. Allen– Cahn conditioned on a prescribed phase front (negative left, positive right) at final time; FitzHugh–Nagumo conditioned on sequential source firing (bottom-left then top-right); and an advection–diffusion plume conditioned on reaching a downstream detector (box). 2

Preprint. Work in progress.

2

Conditioning via a Latent Representation

Let f ∈ E denote a stochastic process, viewed as a random element in a measurable space E, and let f0 ∈ Em denote a finite-dimensional discretisation of f. In typical examples, Em ⊆ Rm or Em ⊆ Rmd , where m is the number of input locations and d is the state dimension at each location. Let pϑ (f0 ) denote the prior density of f0 on Em , with hyperparameters ϑ. We are interested in sampling the conditional distribution of this process, given some additional information C encoded by a likelihood LC : Em → [0, ∞). The desired conditional density is thus given by Z LC (f0 ) pϑ (f0 ) πϑ (f0 | C) = , Zϑ = LC (f0 ) pϑ (f0 ) df0 . (1) Zϑ Em Our key assumption is that pϑ can be obtained by transforming a simple reference density. Assumption 2.1. There exists a latent dimension mξ , a reference density r on Rmξ , and a map Tϑ : Rmξ → Em such that, if ξ0 ∼ r, then f0 = Tϑ (ξ0 ) has density pϑ . Assumption 2.1 is very mild; after discretisation, it holds for essentially any process of interest. Consider a scalar target with continuous distribution function F . By the probability integral transform, the variable F −1 (R(ξ0 )) with ξ0 ∼ r has distribution F , where R is the distribution function of r. The Rosenblatt transform (Rosenblatt, 1952) generalises this to multivariate f0 and writes any absolutely continuous law on Em as the deterministic image of some ξ0 ∼ r. The choice of r is arbitrary; we fix r(ξ0 ) = N (ξ0 ; 0, Imξ ) throughout. It is worth noting that we use densities here only to simplify the presentation. In fact, we require only a measure-theoretic version of Assumption 2.1, which also applies when the prior is singular or discrete; see Remark B.1 for details. Once such a representation is available, conditioning the process on C can be transferred to the latent space via the latent likelihood Λϑ,C (ξ0 ) = (LC ◦ Tϑ )(ξ0 )., i.e. the standard likelihood pulled back to Rmξ . The corresponding latent target density is Z Λϑ,C (ξ0 ) r(ξ0 ) ρϑ (ξ0 | C) = , Zϑ = Λϑ,C (ξ0 ) r(ξ0 ) dξ0 , (2) Zϑ Rmξ where the posterior is recovered by pushing ρϑ through Tϑ . That is, if ξ0 ∼ ρϑ (· | C), then f0 = Tϑ (ξ0 ) has density πϑ (· | C). We formalise this equivalence in Proposition 2.2, which states that the latent and process targets agree in distribution. Proposition 2.2. Suppose that Assumption 2.1 holds and that 0 < Zϑ < ∞. Then, if ξ0 ∼ ρϑ (· | C), the random variable f0 = Tϑ (ξ0 ) has the conditioned process density πϑ (· | C). 2.1

Examples of latent innovation representations

Below we provide three examples of stochastic processes that satisfy Assumption 2.1. Further examples, including heavy-tailed and max-stable processes, state-space models and SPDEs, and neural processes and black-box simulators, are given in Appendix A. Gaussian process. Let f ∼ GP(µϑ , kϑ ), and let f0 = (f (s1 ), . . . , f (sm )) ∈ Rm denote its values at inputs s1:m . Thus f0 ∼ N (mϑ , Kϑ ), where mϑ = (µϑ (s1 ), . . . , µϑ (sm )) and (Kϑ )ij = kϑ (si , sj ). Let Lϑ be a matrix satisfying Lϑ L⊤ ϑ = Kϑ . We can then write f0 = Tϑ (ξ0 ) = mϑ + Lϑ ξ0 , ξ0 ∼ N (0, Im ). Thus, the latent innovation is the whitened Gaussian vector ξ0 , and the generator Tϑ is the usual affine GP whitening map. Student-t process. Let f ∼ TPν (µϑ , kϑ ), where kϑ denotes the Student-t scale kernel, and let f0 = (f (s1 ), . . . , f (sm )) ∈ Rm denote its values at inputs s1:m . We then have f0 = mϑ + ω −1/2 Lϑ v, where v ∼ N (0, Im ) and ω ∼ Gamma(ν/2, ν/2) independently. We can also express f0 in −1 terms of a standard Gaussian reference. Let a ∼ N (0, 1), and define Gν (a) = FΓ,ν (ΦN (a)), where FΓ,ν is the Gamma(ν/2, ν/2) CDF, and ΦN is the standard normal CDF. Then f0 = Tϑ (ξ0 ) = mϑ + Gν (a)−1/2 Lϑ v, 3

ξ0 = (v, a) ∼ N (0, Im+1 ).

Preprint. Work in progress.

Thus, m Gaussian variables generate the spatial Gaussian component, and one additional Gaussian variable generates the random scale. SDE-driven process. Consider a process defined via the stochastic differential equation dxτ = bϑ (τ, xτ ) dτ + Σϑ (τ, xτ ) dwτ , where w = (wτ )τ ≥0 is a standard Brownian motion. Define a time grid 0 = τ0 < τ1 < · · · < τN = T, with ∆n = τn+1 − τn . The Euler–Maruyama discretisation of this SDE is p xn+1 = xn + bϑ (τn , xn )∆n + Σϑ (τn , xn ) ∆n ηn , i.i.d.

where ηn ∼ N (0, I). If x0 is random, we also write x0 = χϑ (η−1 ), with η−1 ∼ N (0, I), and omit η−1 when x0 is fixed. The latent innovation is the collection of Gaussian variables ξ0 = (η−1 , η0 , . . . , ηN −1 ), with the obvious modification for deterministic x0 . Meanwhile, the generator Tϑ is the numerical solver itself: it takes ξ0 , applies the Euler–Maruyama recursion, and returns the discretised path (x0 , x1 , . . . , xN ) = Tϑ (ξ0 ).

3

LatentFlow: Inference with Guided Latent Flows

We now develop the method on which the paper rests. Whenever the latent likelihood is differentiable, or its score admits a tractable estimator, a guided diffusion draws samples ξ0 ∼ ρϑ (· | C). We call the resulting framework LatentFlow. The statements provided here are formal; see Appendix C for a more rigorous treatment. 3.1

Noising bridge, potential, and guidance field

We first define a Gaussian noising process on the latent space. Let ξ0 ∼ r and ε ∼ N (0, Imξ ) be independent of ξ0 . The variance-preserving (VP) latent noising bridge is given by ξt = α(t)ξ0 + σ(t)ε,

(3)

R 1 t

where α(t) = exp{− 2 0 β(s) ds} and σ(t)2 = 1 − α(t)2 , for a locally integrable function β : [0, 1) → (0, ∞) such that limt↑1 α(t) = 0. Thus ξ0 is the clean latent variable and ξt for t ↑ 1 is increasingly noised. Since r(ξ0 ) = N (ξ0 ; 0, Imξ ), we have ξt ∼ r for every t ∈ [0, 1). Moreover, the conditional density of ξ0 | ξt is given by  r0|t (ξ0 ) := p(ξ0 | ξt = u) = N ξ0 ; α(t)u, σ(t)2 Imξ . (4) The additional condition enters through the latent likelihood Λ. For t ∈ [0, 1), we define the latent conditioning potential as ψ(t, u) = Eξ0 ∼r [Λϑ,C (ξ0 ) | ξt = u] = Eε∼r [Λϑ,C (α(t)u + σ(t)ε)] . This can be viewed as a smoothed version of the latent likelihood. This function also has a density-ratio interpretation. Let ξt ∼ ρt be the random variable obtained by drawing ξ0 ∼ ρ0 t (u) and then applying the VP noising bridge in (3). Then ψ(t, u) = Zϑ ρr(u) . Finally, suppose ψ(t, u) > 0 and that ψ(t, ·) is differentiable at u. We define for t ∈ [0, 1) the latent guidance field t (u) g(t, u) = ∇u log ψ(t, u) = ∇u log ρr(u) . This will later be used to steer the dynamics in the direction of the increased expected future likelihood. 3.2

Exact latent dynamics

Our goal is a sampling procedure for ξ0 ∼ ρ0 with ρ0 ≡ ρϑ . First recall that the family of guided densities (ρt )0≤t<1 are precisely the one-time marginals of the VP-SDE p dξt = − 12 β(t) ξt dt + β(t) dwt , ξ0 ∼ ρ0 , t ∈ [0, 1), 4

Preprint. Work in progress.

where w is a standard Rmξ -valued Brownian motion (e.g., Song et al., 2021). By Anderson (1982), the time reversal of these dynamics is also an SDE, given by p   dξt = β(t) 12 ξt − g(t, ξt ) dt + β(t) dwt , (5) where w is a reverse-time Brownian motion, and this equation is integrated backwards in t. In particular, if (5) is initialised at ξt1 ∼ ρt1 and integrated backwards from t1 to 0, then ξt ∼ ρt for every t ∈ [0, t1 ], and hence ξ0 ∼ ρ0 . Since ρt1 ⇒ N (0, Imξ ) as t1 ↑ 1, initialising at N (0, Imξ ) in place of ρt1 is asymptotically exact, yielding a practical sampler. The same family of marginal densities can also be generated by a deterministic process known as the probability-flow ODE (see Song et al., 2021). In practice, however, the SDE is often a preferable generating mechanism; see Appendix C for further details. 3.3

Monte Carlo guidance

To generate ξ0 using the reverse-time SDE in (5) we first need g(t, u) = ∇u log ψ(t, u). Typically, neither g(t, u) nor ψ(t, u) are available in closed form. We can, however, obtain estimates of both via a Monte Carlo estimator of (4). For a given pair (t, u), draw ε(s) ∼ N (0, Imξ ), Then ψbS (t, u) = S1 then

(s)

ξ0|t = α(t)u + σ(t)ε(s) ,

s = 1, . . . , S.

(s) s=1 Λϑ,C (ξ0|t ) is an unbiased estimator of ψ(t, u). If Λϑ,C is differentiable,

PS

  Eε∼r ∇ξ Λϑ,C α(t)u + σ(t)ε   . ∇u log ψ(t, u) = α(t) Eε∼r Λϑ,C α(t)u + σ(t)ε

When Λϑ,C > 0, this motivates the self-normalised estimator gbS (t, u) = α(t)

S X

(s)

(s)

w̄s ∇ξ log Λϑ,C (ξ0|t ),

s=1

Λϑ,C (ξ0|t ) w̄s = PS . (ℓ) ℓ=1 Λϑ,C (ξ0|t )

(6)

Since Λϑ,C = LC ◦ Tϑ , the latent score can be computed by differentiating through the process generator; in practice, this is typically available via automatic differentiation. In particular, when Tϑ and LC are differentiable and LC > 0, the chain rule gives ∇ξ log Λϑ,C (ξ) = [Dξ Tϑ (ξ)]⊤ ∇f0 log LC (f0 ) f0 =T (ξ) , ϑ

where Dξ Tϑ (ξ) is the Jacobian of Tϑ (ξ). For hard constraints, discontinuous likelihoods, or non-differentiable simulators, this gradient-based form must be replaced by a smooth relaxation, a differentiable surrogate, or a derivative-free guidance estimator; see Section D.7. In practice, we replace g(t, u) by gbS (t, u) and solve (5) numerically. The resulting error decomposes into three terms: terminal initialisation, guidance approximation, and numerical discretisation; see Appendix D. Under a common exact terminal coupling, if the Monte Carlo guidance field is uniformly root-S consistent and the SDE solver has strong order qsde , then the coupled latent grid error is OP (S −1/2 + hqsde ). The error contribution of each can be made arbitrarily small, so the sampling error is limited only by available compute. These bounds pass to process space whenever Tϑ is Lipschitz.

4

Guidance-aware parameter estimation

We have, thus far, treated the parameter ϑ as fixed. Adapting it based on the information C offers further advantages of our training-free approach. A learned conditional generator or amortised sampler must be retrained as ϑ is updated, or trained in advance over a prescribed range of parameter values. In LatentFlow, Tϑ is a pre-defined generative prior, and ϑ enters through the pulled-back likelihood. Parameter estimation can thus be performed using the same latent conditional samples, without retraining a conditional generator. The normalising constant Zϑ in (1) and (2) can be interpreted as the evidence for C under the prior. This suggests the empirical Bayes objective J (ϑ) = log Zϑ . This balances satisfaction 5

Preprint. Work in progress.

of C with the relative-entropy cost of deforming pϑ into πϑ (· | C); see Proposition E.1. This relative entropy cost also admits a dynamic interpretation as the guidance energy; see Theorem E.5. Thus, parameters are compatible with the condition when they can be realised with high likelihood and low guidance energy. Given ∇ϑ J (ϑ) = ∇ϑ log Zϑ , or an approximation thereof, we can update parameters using an outer-loop scheme. Fix ϑ0 . For each k ≥ 0, run LatentFlow with ϑk fixed, and update ϑk+1 = ϑk + ηk ∇ϑ J (ϑ) ϑ=ϑ , k

where ηk > 0 is the step size. When ∇ϑ J (ϑ) is unknown, we require an approximation. If Λϑ,C is differentiable in ϑ and differentiation and integration are interchangeable, then ∇ϑ J (ϑ) = Eξ0 ∼ρϑ (·|C) [∇ϑ log Λϑ,C (ξ0 )] . Thus samples from the latent target provide a direct estimator of the conditional-evidence gradient. If the likelihood LC is not a function of ϑ, then ∇ϑ log Λϑ,C (ξ0 ) = [Dϑ Tϑ (ξ0 )]⊤ ∇f0 log LC (f0 ) f =T (ξ ) , 0

ϑ

0

where Dϑ Tϑ (ξ0 ) is the Jacobian of Tϑ (ξ0 ). The same gradient can also be represented at intermediate noising times. In particular, we have that ∇ϑ J (ϑ) = ∇ϑ log Zϑ = Eu∼ρϑ,t (·|C) [∇ϑ log ψϑ (t, u)] . This intermediate-time gradient can be estimated using the same bridge samples used for guidance; see Section 3.3. In particular, differentiating the estimator of the conditioning potential with respect to ϑ yields b aϑ,S (t, u) := ∇ϑ log ψbϑ,S (t, u) =

S X

(r)

(r) w̄r ∇ϑ log Λϑ,C (ξ0|t ),

r=1

Λϑ,C (ξ0|t )

w̄r = PS

(ℓ) ℓ=1 Λϑ,C (ξ0|t )

.

(7)

Thus, the same weighted bridge candidates provide two derivatives: (6) guides the latent state, while (7) adapts the process parameter. Given M particles and selected noising times tj , a practical estimator is M

∇ϑ J (ϑ) ≈

5

J

1 XX (i)  b aϑ,S tj , uj . M J i=1 j=1

Related work

The closest antecedent to our work is FlowGP (Moss et al., 2026). Our framework contains FlowGP as a special case, but also applies to a substantially broader class of stochastic processes; see Appendix A. More broadly, our approach relates to diffusion models, flow matching, and stochastic interpolants, which generate samples via continuous-time transport (Sohl-Dickstein et al., 2015; Ho et al., 2020; Song et al., 2021; Lipman et al., 2023; Albergo et al., 2025). Unlike pretrained diffusion-model guidance (e.g., Ho & Salimans, 2022), however, LatentFlow does not need to be pre-trained on a neural estimate of an unconditional data score or reverse process. The noising bridge is imposed in innovation space, and the guidance is the gradient of a smoothed pulled-back likelihood under the known latent bridge. A closely related line of work conditions generative models by performing inference in the latent noise space (Whang et al., 2021; Holden et al., 2022; Patel et al., 2022; Achituve et al., 2025; Askari et al., 2025; Purohit et al., 2025; Venkatraman et al., 2025; Xia et al., 2026). These methods share the idea of conditioning a generator by changing the distribution of its input noise, but focus only on learned neural generators. Our framework is significantly more general: it includes pre-trained neural generators as a special case, but also extends to a much broader class of genuine stochastic-process priors. The papers above also rely on different, often more expensive, samplers, for instance, Langevin dynamics (Purohit et al., 2025), HMC (Patel et al., 2022; Xia et al., 2026), pCN (Holden et al., 2022), or SMC (Achituve et al., 2025). LatentFlow instead targets the latent posterior using a guided, training-free diffusion process, whose guidance field is constructed from the pulled-back likelihood. Our work is also related to many other areas, including diffusion bridges, rare-event simulation, and transport-assisted MCMC; see Appendix F for a more detailed discussion. 6

Preprint. Work in progress.

6

Numerical Experiments

We now present a range of numerical experiments designed to assess the flexibility, accuracy, and scalability of LatentFlow. We first evaluate LatentFlow on three standard sampling benchmarks; see Appendix G.1 for full details. Across all three benchmarks, we achieve MMD2 scores competitive with or better than NUTS. As LatentFlow generates independent samples, it scales to problems where standard MCMC either collapses modally or explores very slowly. In this section, we consider higher-dimensional and more challenging processes. All timings are reported on a 13th Gen Intel i9-13900K CPU. Additional results are provided in Appendix G. 6.1

Spatial Processes

Spatial processes are central to many applications, including environmental modelling, climate science, and ecology (e.g., Cressie, 1993). While GPs provide a convenient prior in such settings, many applications exhibit extremal or heavy tails, discontinuities, sharp local anomalies, or discrete observations. This motivates alternatives such as Student-t processes, Lévy random fields, and Potts models (Potts, 1952; Besag, 1974; Rajput & Rosiński, 1989; Wolpert et al., 2011; Shah et al., 2014). The difficulty here has two sources. The prior alone can break conjugacy as non-Gaussian priors induce conditional laws involving scale mixtures or jump-like innovations (Walchessen et al., 2025). Additionally, non-linear conditioning statements such as areal maxima, threshold exceedances, or censored measurements do not admit closed-form conditionals, even under a Gaussian prior. In Figure 2, we use LatentFlow to condition four spatial priors on four conditions. The conditioning statements include point observations, regional summaries, and censored observations. The resulting samples qualitatively respect the imposed conditioning information, whilst also preserving the distinctive behaviour of each prior. The speed of LatentFlow is notable here: conditional samples are generated in just 3 − 4 seconds per sample.

Prior

Sparse observations

Areal max

Gaussian Process

Areal mean

Censored observations 2

Smooth, light tails

1

Student-t Process

Heavy tails, spatially coupled

0

Cauchy Convolution Heavy tails

1

2

Potts Model Discrete labels, ordered domains

3

Figure 2: Posterior samples from four spatial process priors under four observation types via LatentFlow. Rows: Gaussian process (Matérn- 32 ); Student-t process (ν = 4); Cauchy convolution field; Potts model (q = 3). Columns: unconditional; 30 point observations; areal maxima (8 regions); areal means (20 regions); 60 censored observations (+/−). Thumbnails show areal structures. Sampling time: 3 − 4s per sample.

7

Preprint. Work in progress.

6.2

Temporal processes

We next consider temporal processes. Examples include diffusion processes (e.g., Øksendal, 2003), and dynamical systems with random initial conditions or exogenous forcing (e.g., Law et al., 2015). This presents a different use case for LatentFlow: the object being conditioned is no longer a static field, but a trajectory generated by propagating latent randomness through a dynamical model. Diffusion bridges. Diffusion bridges condition diffusion processes on a fixed terminal condition. They arise in many applications, including computational chemistry (Bolhuis et al., 2002), econometrics (Elerian et al., 2001), developmental dynamics (Wang et al., 2011), and shape analysis (Arnaudon et al., 2022). Except in a few tractable cases, the bridge drift depends on an unavailable transition density, motivating bespoke guided proposals and pathspace MCMC methods (Delyon & Hu, 2006; Schauer et al., 2017), as well as recent neural approaches that approximate the bridge score or guidance field (Baker et al., 2024; Heng et al., 2025; Yang et al., 2025b). In Figure 3, we assess whether LatentFlow can generate diffusion bridges by acting on the exogenous simulation noise. In our case, exact endpoint conditioning is approximated by a narrow finite-width likelihood. Our method recovers the analytic OU bridge behaviour, and produces nonlinear cell-differentiation and FitzHugh– Nagumo paths consistent with rejection-sampled references. It also agrees qualitatively with existing methods (e.g., Yang et al., 2025b) in cases where no valid rejection samples were found after simulating the unconditioned process 100,000 times; see Figure 15 in Section G.3. Trajectory-level conditioning. We next consider more general trajectory-level conditioning, where the information may involve observations, terminal events, or pathwise threshold constraints (e.g., Bolhuis & Swenson, 2021; Finzi et al., 2023; Grafke, 2025). Such constraints act on the realised trajectory after the latent randomness has been propagated through a nonlinear dynamical simulator, and need not reduce to a standard filtering, smoothing, or bridge-simulation problem. In Figure 4, we consider Lotka–Volterra dynamics (Lotka, 1925; Volterra, 1926), an SIR epidemic model (Kermack & McKendrick, 1927), and the Heston stochastic-volatility model (Heston, 1993). The resulting conditional paths satisfy the imposed trajectory-level information while retaining the characteristic dynamics of each model. Sampling remains very fast, taking roughly 1 second per sample for the SIR and Lotka–Volterra models, and roughly 3 seconds per sample for the Heston model. 6.3

Spatio-Temporal Processes

prior posterior analytic

0.0

0.2

x1 x2 rejection ref

2.0 1.5

0.4

0.6

time

0.8

(a) Ornstein–Uhlenbeck.

1.0

1.0 0.5

1.0

state

1.2 1.0 0.8 0.6 0.4 0.2 0.0 −0.2

state

state

We finally consider spatio-temporal models, which describe the random evolution of spatial fields. Such models arise in turbulent transport, environmental dispersion, materials science, excitable media, and phase separation (Majda & Kramer, 1999; Allen & Cahn, 1979; Nagumo et al., 1962). Conditioning them on sparse observations, future events, or aggregate space– time constraints is central to data assimilation, inverse modelling, and rare-event prediction (Dashti & Stuart, 2017; Cérou & Guyader, 2007; Cérou et al., 2012; Botev & L’Ecuyer, 2020).

0.0

0.5

−0.5

0.0

−1.0 0

1

2

time

3

(b) Cell differentiation.

4

x1 x2 rejection ref

0.0

0.5

1.0

time

1.5

2.0

(c) FitzHugh–Nagumo.

Figure 3: Diffusion bridges generated by LatentFlow. (a) An Ornstein–Uhlenbeck bridge. (b) A two-dimensional nonlinear cell-differentiation diffusion conditioned to reach a prescribed terminal state. (c) A FitzHugh–Nagumo diffusion conditioned on a rare transition between metastable states. Red markers indicate terminal targets, blue & orange curves indicate conditioned state components; grey curves indicate prior or rejection-reference trajectories.

8

Lotka-Volterra predator-prey

Preprint. Work in progress.

Prior prey predator

12 10 8 6 4 2 0

Heston stochastic vol.

SIR stochastic epidemic

0

1

2

Conditioned on data

Conditioned on event peak prey

3

4

5

0

1

2

3

4

5

0

S I R

0.8 0.6

Conditioned on both

>7

1

peak final final

2

3

4

5

0

10

12

0

1

2

3

4

5

10

12

I > 0.32 I < 0.18 R > 0.62

0.4 0.2 0.0 0

2

4

6

8

10

12

0

2

4

6

8

10

12

price index volatility

0.6

0

2

4

6

min price max vol

8

2

4

6

8

120

< 87

110

> 0.42

100

0.4

90

0.2

80 70

0.0 0.0

0.2

0.4

0.6

time

0.8

1.0

0.0

0.2

0.4

0.6

time

0.8

1.0

0.0

0.2

0.4

0.6

time

0.8

1.0

0.0

0.2

0.4

0.6

time

0.8

1.0

Figure 4: LatentFlow can condition SDE priors on data and rare events. Rows: Lotka– Volterra predator–prey dynamics, a stochastic SIR epidemic, and Heston stochastic volatility. Columns: the prior, the posterior given observations (black circles), the posterior given rare events: a prey boom, a fast growing but fast terminating epidemic, and a stress event (a market crash followed by a volatility spike). Sampling time: < 1s for SIR and Lotka–Volterra; 3s Heston. We consider three examples: a stochastic Allen–Cahn field for noisy phase separation, an advection–diffusion plume model, and a stochastic FitzHugh–Nagumo excitable-media model that produces travelling activation waves (FitzHugh, 1961; Nagumo et al., 1962; Allen & Cahn, 1979; Majda & Kramer, 1999). Representative samples are shown in Figure 5. Across all three models, LatentFlow produces fields that remain consistent with the prior dynamics while satisfying the imposed conditions.

Time τ = 0.7

τ = 1.1

τ = 1.6

τ = 2.0

τ = 2.2

0.6

concentration

τ = 0.3

0.3

activator u

0.0 1.5

0.0

−1.5

0

field value

1

−1

Figure 5: Posterior samples from three spatio-temporal systems using our LatentFlow: (top) a concentration plume conditioned on triggering a downstream detector evolving under advection–diffusion dynamics (160 latent dimensions); (middle) a FitzHugh–Nagumo excitable medium conditioned on sequential pacemaker firing (40K latent dimensions); and (bottom) an Allen–Cahn phase field conditioned on a prescribed phase front (122K latent dimensions). Each row shows four independent draws, over six time snapshots. Sampling time: 1s for the advection–diffusion; 5s for FitzHugh–Nagumo, and 2s for Allen–Cahn.

9

Preprint. Work in progress.

lengthscale

fitted true

0.6 0.4 0.2 0

(c) Post-fit

prey 30 predator

30

population

(b) Pre-fit

20

40

fit step

60

(d) Optimisation trajectory

30 2.0

20

20

rate

(a) Ground truth

20

10

10

10

0

0 4 0

0 4 0 5 1

1.5

fitted α

fitted γ

fitted δ

fitted β

1.0 0.5

0

1

2

time

3

(e) Ground truth

1

2

time

(f) Pre-fit

3

2

time

(g) Post-fit

3

4

50

20

40

60

fit step

80

100

(h) Optimisation trajectory

Figure 6: Guidance-aware parameter estimation recovers reasonable process hyperparameters, modulo identifiability. Top: Cauchy-convolution field. (a) Truth, observed only via the maximum value within each of K = 32 irregular regions. (b) Guided posterior samples under a mis-specified lengthscale (ℓ = 0.7). (c) Samples after adapting the lengthscale against the regional-maximum condition. (d) Lengthscale trajectory converging to the true value (ℓ = 0.1, dashed). Bottom: stochastic Lotka–Volterra process. (e) Truth, with dense early observations (dots). (f) Guided posterior samples under mis-specified rates. (g) Samples after jointly fitting all four rates against the same early-window condition. (h) Identifiable rates converge towards true values (dashed). 6.4

Parameter Estimation

Finally, we assess the guidance-aware parameter adaptation framework introduced in Section 4. We revisit two of the processes considered in earlier sections, and compare the results obtained with fixed parameters, with the hyperparameter selected after incorporating the full conditioning information. The results are shown in Figure 6. Across both examples, our guidance-aware scheme recovers the ground truth, resulting in samples that more closely respect the target observations and constraints.

7

Discussion

LatentFlow replaces the bespoke, model-by-model samplers that have previously been required for conditioning stochastic processes with a single, unified approach. The central idea is to move conditioning from process space to a latent innovation space, with the likelihood pulled back through the process generator and the resulting conditional law sampled using guided reverse-time dynamics. This gives an exact target-law construction, with practical approximation confined to three explicit and systematically reducible sources: finite terminal noising, Monte Carlo guidance, and time discretisation. This is a fundamental shift in capability. We expect the largest gains for rare, nonlinear, or global conditioning events under non-Gaussian and simulator-defined process priors, including extreme events, pathwise conditions, and constraints on SDE and SPDE driven models. Several limitations remain. LatentFlow requires a tractable latent innovation representation of the discretised prior, and its performance can depend on the dimension, smoothness, and geometry of this representation. It also requires pointwise evaluation of the pulledback likelihood, and ideally that the likelihood is differentiable through the generator. Discontinuous simulators, hard constraints, rare events, sharp likelihoods, multimodal conditionals, or poorly scaled latent parameterisations may therefore require relaxations, surrogates, derivative-free guidance, or larger Monte Carlo budgets. 10

Preprint. Work in progress.

References Idan Achituve, Hai Victor Habi, Amir Rosenfeld, Arnon Netzer, Idit Diamant, and Ethan Fetaya. Inverse problem sampling in latent space using sequential Monte Carlo. In Proceedings of the 42nd International Conference on Machine Learning, pp. 420–443, 2025. Christian Agrell. Gaussian processes with linear operator inequality constraints. Journal of Machine Learning Research, 20(135):1–36, 2019. Ahmad Ajalloeian and Sebastian U Stich. On the convergence of sgd with biased gradients. arXiv preprint arXiv:2008.00051, 2020. Michael S. Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. In Proceedings of the 11th International Conference on Learning Representations, 2023. Michael S. Albergo, Nicholas M. Boffi, and Eric Vanden-Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. Journal of Machine Learning Research, 26 (209):1–80, 2025. Samuel M. Allen and John W. Cahn. A microscopic theory for antiphase boundary motion and its application to antiphase domain coarsening. Acta Metallurgica, 27(6):1085–1095, 1979. Mauricio Álvarez, David Luengo, and Neil D. Lawrence. Latent force models. In Proceedings of The 12th International Conference on Artificial Intelligence and Statistics, pp. 9–16, 2009. Mauricio A. Álvarez, David Luengo, and Neil D. Lawrence. Linear latent force models using Gaussian processes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35 (11):2693–2705, 2013. Brian D. O. Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982. Christophe Andrieu, Nando de Freitas, Arnaud Doucet, and Michael I. Jordan. An introduction to MCMC for machine learning. Machine Learning, 50(1–2):5–43, 2003. Michael Arbel, Alexander G. D. G. Matthews, and Arnaud Doucet. Annealed flow transport Monte Carlo. In Proceedings of the 38th International Conference on Machine Learning, pp. 318–330, 2021. Alexis Arnaudon, Frank van der Meulen, Moritz Schauer, and Stefan Sommer. Diffusion bridges for stochastic hamiltonian systems and shape evolutions. SIAM Journal on Imaging Sciences, 15(1):293–323, 2022. Hossein Askari, Yadan Luo, Hongfu Sun, and Fred Roosta. Latent refinement via flow matching for training-free linear inverse problem solving. In Advances in Neural Information Processing Systems, volume 38, 2025. Lachlan Astfalck, Deborshee Sen, Sayan Patra, Edward Cripps, and David Dunson. Posterior projection for inference in constrained spaces. arXiv preprint arXiv:1812.05741, 2018. Lachlan Astfalck, Cassandra Bird, and Daniel Williamson. Generalised Bayes linear inference. arXiv preprint arXiv:2405.14145, 2024. Elizabeth L. Baker, Alexander Denker, and Jes Frellsen. Supervised guidance training for infinite-dimensional diffusion models. In Proceedings of the 43rd International Conference on Machine Learning, 2026. Elizabeth Louise Baker, Gefan Yang, Michael Lind Severinsen, Christy Anna Hipsley, and Stefan Sommer. Conditioning non-linear and infinite-dimensional diffusion processes. In Advances in Neural Information Processing Systems, volume 37, pp. 10801–10826, 2024. 11

Preprint. Work in progress.

Elizabeth Louise Baker, Moritz Schauer, and Stefan Sommer. Score matching for bridges without learning time-reversals. In Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, pp. 775–783, 2025. Lorenzo Baldassari, Ali Siahkoohi, Josselin Garnier, Knut Sølna, and Maarten V. de Hoop. Conditional score-based diffusion models for Bayesian inference in infinite dimensions. In Advances in Neural Information Processing Systems, volume 36, pp. 24262–24290, 2023. Lorenzo Baldassari, Ali Siahkoohi, Josselin Garnier, Knut Sølna, and Maarten V. de Hoop. Taming score-based diffusion priors for infinite-dimensional nonlinear inverse problems. arXiv preprint arXiv:2405.15676, 2024. Lorenzo Baldassari, Josselin Garnier, Knut Sølna, and Maarten V. de Hoop. Preconditioned Langevin dynamics with score-based generative models for infinite-dimensional linear Bayesian inverse problems. In Advances in Neural Information Processing Systems, volume 38, 2025. Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Roni Sengupta, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Universal guidance for diffusion models. In Proceedings of the 12th International Conference on Learning Representations, 2024. Espen Bernton, Jeremy Heng, Arnaud Doucet, and Pierre E. Jacob. Schrödinger bridge samplers. arXiv preprint arXiv:1912.13170, 2019. Julian Besag. Spatial interaction and the statistical analysis of lattice systems. Journal of the Royal Statistical Society: Series B (Methodological), 36(2):192–225, 1974. Alexandros Beskos, Frank J. Pinski, Jesus Maria Sanz-Serna, and Andrew M. Stuart. Hybrid Monte Carlo on hilbert spaces. Stochastic Processes and their Applications, 121(10): 2201–2230, 2011. Joris Bierkens, Frank van der Meulen, and Moritz Schauer. Simulation of elliptic and hypo-elliptic conditional diffusions. Advances in Applied Probability, 52(1):173–212, 2020. Mogens Bladt and Michael Sørensen. Simple simulation of diffusion bridges with application to likelihood inference for diffusions. Bernoulli, 20(2):645–675, 2014. Mogens Bladt, Samuel Thomas John Finch, and Michael Sørensen. Simulation of multivariate diffusion bridges. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 78(2):343–369, 2016. David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe. Variational inference: A review for statisticians. Journal of the American Statistical Association, 112(518):859–877, 2017. Vanessa Böhm, François Lanusse, and Uroš Seljak. Uncertainty quantification with generative models. arXiv preprint arXiv:1910.10046, 2019. Peter G. Bolhuis and David W. H. Swenson. Transition path sampling as Markov chain Monte Carlo of trajectories: Recent algorithms, software, applications, and future outlook. Advanced Theory and Simulations, 4(4):2000237, 2021. Peter G. Bolhuis, David Chandler, Christoph Dellago, and Phillip L. Geissler. Transition path sampling: Throwing ropes over rough mountain passes, in the dark. Annual Review of Physical Chemistry, 53:291–318, 2002. Zdravko I. Botev and Pierre L’Ecuyer. Sampling conditionally on a rare event via generalized splitting. INFORMS Journal on Computing, 32(4):986–995, 2020. Alberto Cabezas and Christopher Nemeth. Transport elliptical slice sampling. In Proceedings of The 26th International Conference on Artificial Intelligence and Statistics, pp. 3664–3676, 2023. Alberto Cabezas, Louis Sharrock, and Christopher Nemeth. Markovian flow matching: Accelerating MCMC with continuous normalizing flows. In Advances in Neural Information Processing Systems, volume 37, pp. 104383–104411, 2024. 12

Preprint. Work in progress.

Gabriel Cardoso, Yazid Janati El Idrissi, Sylvain Le Corff, and Eric Moulines. Monte Carlo guided denoising diffusion models for Bayesian linear inverse problems. In Proceedings of the 12th International Conference on Learning Representations, 2024. Frédéric Cérou and Arnaud Guyader. Adaptive multilevel splitting for rare event analysis. Stochastic Analysis and Applications, 25(2):417–443, 2007. Frédéric Cérou, Pierre Del Moral, Teddy Furon, and Arnaud Guyader. Sequential Monte Carlo for rare event estimation. Statistics and Computing, 22(3):795–808, 2012. Yifan Chen, Bamdad Hosseini, Houman Owhadi, and Andrew M. Stuart. Solving and learning nonlinear PDEs with Gaussian processes. Journal of Computational Physics, 447: 110668, 2021. Yongxin Chen, Tryphon T. Georgiou, and Michele Pavon. On the relation between optimal transport and Schrödinger bridges: A stochastic control viewpoint. Journal of Optimization Theory and Applications, 169(2):671–691, 2016. Hyungjin Chung, Jeongsol Kim, Michael T. McCann, Marc L. Klasky, and Jong Chul Ye. Diffusion posterior sampling for general noisy inverse problems. In Proceedings of the 11th International Conference on Learning Representations, 2023. Adrien Corenflos, Zheng Zhao, Thomas B. Schön, Simo Särkkä, and Jens Sjölund. Conditioning diffusion models by explicit forward-backward bridging. In Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, pp. 3709–3717, 2025. Simon L. Cotter, Gareth O. Roberts, Andrew M. Stuart, and David White. MCMC methods for functions: Modifying old algorithms to make them faster. Statistical Science, 28(3): 424–446, 2013. Noel A. C. Cressie. Statistics for Spatial Data. John Wiley & Sons, New York, revised edition, 1993. ISBN 9780471002550. Sébastien Da Veiga and Amandine Marrel. Gaussian process modeling with inequality constraints. Annales de la Faculté des Sciences de Toulouse: Mathématiques, 21(3): 529–555, 2012. Andreas Damianou and Neil D. Lawrence. Deep Gaussian processes. In Proceedings of The 16th International Conference on Artificial Intelligence and Statistics, pp. 207–215, 2013. Masoumeh Dashti and Andrew M. Stuart. The Bayesian approach to inverse problems. In Roger Ghanem, David Higdon, and Houman Owhadi (eds.), Handbook of Uncertainty Quantification, pp. 311–428. Springer, Cham, 2017. Valentin De Bortoli, James Thornton, Jeremy Heng, and Arnaud Doucet. Diffusion Schrödinger bridge with applications to score-based generative modeling. In Advances in Neural Information Processing Systems, volume 34, pp. 17695–17709, 2021. Pierre Del Moral and Josselin Garnier. Genealogical particle analysis of rare events. The Annals of Applied Probability, 15(4):2496–2534, 2005. Pierre Del Moral, Arnaud Doucet, and Ajay Jasra. Sequential Monte Carlo samplers. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 68(3):411–436, 2006. Christoph Dellago, Peter G. Bolhuis, Felix S. Csajka, and David Chandler. Transition path sampling and the calculation of rate constants. The Journal of Chemical Physics, 108(5): 1964–1977, 1998. Christoph Dellago, Peter G. Bolhuis, and Phillip L. Geissler. Transition path sampling. In Ilya Prigogine and Stuart A. Rice (eds.), Advances in Chemical Physics, volume 123, pp. 1–78. John Wiley & Sons, 2002. Bernard Delyon and Ying Hu. Simulation of conditioned diffusion and application to parameter estimation. Stochastic Processes and their Applications, 116(11):1660–1675, 2006. 13

Preprint. Work in progress.

Alexander Denker, Francisco Vargas, Shreyas Padhy, Kieran Didi, Simon Mathis, Vincent Dutordoir, Riccardo Barbano, Emile Mathieu, Urszula Julia Komorowska, and Pietro Liò. DEFT: Efficient fine-tuning of diffusion models by learning the generalised h-transform. In Advances in Neural Information Processing Systems, volume 37, pp. 19636–19682, 2024. Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, volume 34, pp. 8780– 8794, 2021. Kieran Didi, Francisco Vargas, Simon V. Mathis, Vincent Dutordoir, Emile Mathieu, Urszula J. Komorowska, and Pietro Liò. A framework for conditional diffusion modelling with applications in motif scaffolding for protein design. In NeurIPS 2023 Workshop AI4D3, 2023. Monroe D. Donsker and Sathamangalam R. S. Varadhan. Asymptotic evaluation of certain Markov process expectations for large time, I. Communications on Pure and Applied Mathematics, 28(1):1–47, 1975. Joseph L. Doob. Conditional Brownian motion and the boundary limits of harmonic functions. Bulletin de la Société Mathématique de France, 85:431–458, 1957. Arnaud Doucet, Nando de Freitas, and Neil Gordon (eds.). Sequential Monte Carlo Methods in Practice. Information Science and Statistics. Springer, New York, NY, 2001. ISBN 978-0-387-95146-1. Vincent Dutordoir, Alan Saul, Zoubin Ghahramani, and Fergus Simpson. Neural diffusion processes. In Proceedings of the 40th International Conference on Machine Learning, pp. 8990–9012, 2023. Tarek A. El Moselhy and Youssef M. Marzouk. Bayesian inference with optimal maps. Journal of Computational Physics, 231(23):7815–7850, 2012. Ola Elerian, Siddhartha Chib, and Neil Shephard. Likelihood inference for discretely observed nonlinear diffusions. Econometrica, 69(4):959–993, 2001. Ruiqi Feng, Chenglei Yu, Wenhao Deng, Peiyan Hu, and Tailin Wu. On the guidance of flow matching. In Proceedings of the 42nd International Conference on Machine Learning, pp. 16993–17029, 2025. Marc Anton Finzi, Anudhyan Boral, Andrew Gordon Wilson, Fei Sha, and Leonardo ZepedaNúñez. User-defined event sampling and uncertainty quantification in diffusion models for physical dynamical systems. In Proceedings of the 40th International Conference on Machine Learning, pp. 10136–10152, 2023. Richard FitzHugh. Impulses and physiological states in theoretical models of nerve membrane. Biophysical Journal, 1(6):445–466, 1961. Hans Föllmer. An entropy approach to the time reversal of diffusion processes. In Michel Métivier and Étienne Pardoux (eds.), Stochastic Differential Systems: Filtering and Control, volume 69 of Lecture Notes in Control and Information Sciences, pp. 156–163. Springer, Berlin, Heidelberg, 1985. Giulio Franzese, Giulio Corallo, Simone Rossi, Markus Heinonen, Maurizio Filippone, and Pietro Michiardi. Continuous-time functional diffusion processes. In Advances in Neural Information Processing Systems, volume 36, pp. 37370–37400, 2023. Roger Frigola, Fredrik Lindsten, Thomas B. Schön, and Carl Edward Rasmussen. Bayesian inference and learning in Gaussian process state-space models with particle MCMC. In Advances in Neural Information Processing Systems, volume 26, pp. 3156–3164, 2013. Roger Frigola, Yutian Chen, and Carl Edward Rasmussen. Variational Gaussian process state-space models. In Advances in Neural Information Processing Systems, volume 27, pp. 3680–3688, 2014. 14

Preprint. Work in progress.

Marta Garnelo, Jonathan Schwarz, Dan Rosenbaum, Fabio Viola, Danilo J. Rezende, S. M. Ali Eslami, and Yee Whye Teh. Neural processes. arXiv preprint arXiv:1807.01622, 2018. Igor V. Girsanov. On transforming a certain class of stochastic processes by absolutely continuous substitution of measures. Theory of Probability & Its Applications, 5(3):285–301, 1960. Andrew Golightly and Darren J. Wilkinson. Bayesian inference for nonlinear multivariate diffusion models observed with error. Computational Statistics & Data Analysis, 52(3): 1674–1693, 2008. Jonathan Gordon, Wessel P. Bruinsma, Andrew Y. K. Foong, James Requeima, Yann Dubois, and Richard E. Turner. Convolutional conditional neural processes. In Proceedings of the 8th International Conference on Learning Representations, 2020. Tobias Grafke. Sampling conditioned diffusions via pathspace projected monte carlo. arXiv preprint arXiv:2506.15743, 2025. Matthew Graham and Amos Storkey. Asymptotically exact inference in differentiable generative models. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, pp. 499–508, 2017. Matthew M. Graham, Alexandre H. Thiery, and Alexandros Beskos. Manifold Markov chain Monte Carlo methods for Bayesian inference in diffusion models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 84(4):1229–1256, 2022. Mamikon Gulian, Ari L. Frankel, and Laura P. Swiler. Gaussian process regression constrained by boundary value problems. Computer Methods in Applied Mechanics and Engineering, 388:114117, 2022. Zhengyi Guo, Wenpin Tang, and Renyuan Xu. Conditional diffusion guidance under hard constraint: A stochastic analysis approach. arXiv preprint arXiv:2602.05533, 2026. Martin Hairer, Andrew M. Stuart, and Sebastian J. Vollmer. Spectral gaps for a Metropolis– Hastings algorithm in infinite dimensions. The Annals of Applied Probability, 24(6): 2455–2490, 2014. Oliver Hamelijnck, Arno Solin, and Theodoros Damoulas. Physics-informed variational state-space Gaussian processes. In Advances in Neural Information Processing Systems, volume 37, pp. 98505–98536, 2024. Jeremy Heng, Valentin De Bortoli, Arnaud Doucet, and James Thornton. Simulating diffusion bridges with score matching. Biometrika, 112(4):asaf048, 2025. James Hensman, Alexander G. de G. Matthews, and Zoubin Ghahramani. Scalable variational Gaussian process classification. In Proceedings of The 18th International Conference on Artificial Intelligence and Statistics, pp. 351–360, 2015. Steven L. Heston. A closed-form solution for options with stochastic volatility with applications to bond and currency options. The Review of Financial Studies, 6(2):327–343, 1993. Jonathan Ho and Tim Salimans. arXiv:2207.12598, 2022.

Classifier-free diffusion guidance.

arXiv preprint

Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33, pp. 6840–6851, 2020. Matthew D. Hoffman and Andrew Gelman. The No-U-Turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo. Journal of Machine Learning Research, 15(47): 1593–1623, 2014. Matthew D. Hoffman, Pavel Sountsov, Joshua V. Dillon, Ian Langmore, Dustin Tran, and Srinivas Vasudevan. NeuTra-lizing bad geometry in hamiltonian Monte Carlo using neural transport. arXiv preprint arXiv:1903.03704, 2019. 15

Preprint. Work in progress.

Matthew Holden, Marcelo Pereyra, and Konstantinos C. Zygalakis. Bayesian imaging with data-driven priors encoded by neural networks: Theory, methods, and algorithms. SIAM Journal on Imaging Sciences, 15(2):892–924, 2022. Bamdad Hosseini, Alexander W. Hsu, and Amirhossein Taghvaei. Conditional optimal transport on function spaces. SIAM/ASA Journal on Uncertainty Quantification, 13(1): 304–338, 2025. Jian Huang, Yuling Jiao, Lican Kang, Xu Liao, Jin Liu, and Yanyan Liu. Schrödinger–Föllmer sampler: Sampling without ergodicity. arXiv preprint arXiv:2106.10880, 2021. Carl Jidling, Niklas Wahlström, Adrian Wills, and Thomas B. Schön. Linearly constrained Gaussian processes. In Advances in Neural Information Processing Systems, volume 30, pp. 1215–1224, 2017. William O. Kermack and Anderson G. McKendrick. A contribution to the mathematical theory of epidemics. Proceedings of the Royal Society of London. Series A, Containing Papers of a Mathematical and Physical Character, 115(772):700–721, 1927. Gavin Kerrigan, Giosue Migliorini, and Padhraic Smyth. Functional flow matching. In Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, pp. 3934–3942, 2024. Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, S. M. Ali Eslami, Dan Rosenbaum, Oriol Vinyals, and Yee Whye Teh. Attentive neural processes. In Proceedings of the 7th International Conference on Learning Representations, 2019. Malte Kuss and Carl Edward Rasmussen. Assessing approximate inference for binary Gaussian process classification. Journal of Machine Learning Research, 6:1679–1704, 2005. Markus Lange-Hegermann. Algorithmic linearly constrained Gaussian processes. In Advances in Neural Information Processing Systems, volume 31, pp. 2137–2148, 2018. Kody J. H. Law, Andrew M. Stuart, and Konstantinos C. Zygalakis. Data Assimilation: A Mathematical Introduction, volume 62 of Texts in Applied Mathematics. Springer, 2015. Christian Léonard. A survey of the Schrödinger problem and some of its connections with optimal transport. Discrete & Continuous Dynamical Systems - A, 34(4):1533–1574, 2014. Jae Hyun Lim, Nikola B. Kovachki, Ricardo Baptista, Christopher Beckham, Kamyar Azizzadenesheli, Jean Kossaifi, Vikram Voleti, Jiaming Song, Karsten Kreis, Jan Kautz, Christopher Pal, Arash Vahdat, and Anima Anandkumar. Score-based diffusion models in function space. Journal of Machine Learning Research, 26(158):1–62, 2025. Lizhen Lin and David B. Dunson. Bayesian monotone regression using Gaussian process projection. Biometrika, 101(2):303–317, 2014. Finn Lindgren, Håvard Rue, and Johan Lindström. An explicit link between Gaussian fields and Gaussian Markov random fields: The stochastic partial differential equation approach. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73(4):423–498, 2011. Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. In Proceedings of the 11th International Conference on Learning Representations, 2023. Guan-Horng Liu, Arash Vahdat, De-An Huang, Evangelos A. Theodorou, Weili Nie, and Anima Anandkumar. I2 SB: Image-to-image Schrödinger bridge. In Proceedings of the 40th International Conference on Machine Learning, pp. 22042–22062, 2023a. Guan-Horng Liu, Yaron Lipman, Maximilian Nickel, Brian Karrer, Evangelos A. Theodorou, and Ricky T. Q. Chen. Generalized Schrödinger bridge matching. In Proceedings of the 12th International Conference on Learning Representations, 2024. 16

Preprint. Work in progress.

Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In Proceedings of the 11th International Conference on Learning Representations, 2023b. Andrés F. López-Lopera, François Bachoc, Nicolas Durrande, and Olivier Roustant. Finitedimensional Gaussian approximation with linear inequality constraints. SIAM/ASA Journal on Uncertainty Quantification, 6(3):1224–1255, 2018. Alfred J. Lotka. Elements of Physical Biology. Williams & Wilkins, Baltimore, 1925. Chao Ma, Yingzhen Li, and Jose Miguel Hernandez-Lobato. Variational implicit processes. In Proceedings of the 36th International Conference on Machine Learning, pp. 4222–4233, 2019. Hassan Maatouk and Xavier Bay. Gaussian process emulators for computer experiments with inequality constraints. Mathematical Geosciences, 49(5):557–582, 2017. Hassan Maatouk, Didier Rullière, and Xavier Bay. Bayesian analysis of constrained Gaussian processes. Bayesian Analysis, 20(3):973–1002, 2025. Andrew J. Majda and Peter R. Kramer. Simplified models for turbulent diffusion: Theory, numerical modelling, and physical phenomena. Physics Reports, 314(4–5):237–574, 1999. Konstantin Mark, Leonard Galustian, Maximilian P.-P. Kovar, and Esther Heid. Feynman– kac-flow: Inference steering of conditional flow matching to an energy-tilted posterior. arXiv preprint arXiv:2509.01543, 2025. Henry B. Moss, Lachlan Astfalck, Thomas Cowperthwaite, Colin Doumont, Sam Willis, Philipp Hennig, Christopher Nemeth, and Andrew Zammit-Mangion. Conditioning Gaussian processes on almost anything. arXiv preprint arXiv:2605.21041, 2026. Jin-Ichi Nagumo, Suguru Arimoto, and Shuji Yoshizawa. An active pulse transmission line simulating nerve axon. Proceedings of the IRE, 50(10):2061–2070, 1962. Radford M. Neal. Annealed importance sampling. Statistics and Computing, 11(2):125–139, 2001. Hannes Nickisch and Carl Edward Rasmussen. Approximations for binary Gaussian process classification. Journal of Machine Learning Research, 9:2035–2078, 2008. Hannes Nickisch, Arno Solin, and Alexander Grigorevskiy. State space Gaussian processes with non-Gaussian likelihood. In Proceedings of the 35th International Conference on Machine Learning, pp. 3789–3798, 2018. Erik Nijkamp, Ruiqi Gao, Pavel Sountsov, Srinivas Vasudevan, Bo Pang, Song-Chun Zhu, and Ying Nian Wu. MCMC should mix: Learning energy-based model with neural transport latent space MCMC. In Proceedings of the 10th International Conference on Learning Representations, 2022. Bernt Øksendal. Stochastic Differential Equations: An Introduction with Applications. Universitext. Springer, Berlin, 6 edition, 2003. George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. Normalizing flows for probabilistic modeling and inference. Journal of Machine Learning Research, 22(57):1–64, 2021. Omiros Papaspiliopoulos, Gareth O. Roberts, and Osnat Stramer. Data augmentation for diffusions. Journal of Computational and Graphical Statistics, 22(3):665–688, 2013. Meet Hemant Parikh, Yaqin Chen, and Jian-Xun Wang. D-Flow SGLD: Source-space posterior sampling for scientific inverse problems with flow matching. arXiv preprint arXiv:2602.21469, 2026. Byoungwoo Park, Jungwon Choi, Sungbin Lim, and Juho Lee. Stochastic optimal control for diffusion bridges in function spaces. In Advances in Neural Information Processing Systems, volume 37, pp. 28745–28771, 2024. 17

Preprint. Work in progress.

Matthew D. Parno and Youssef M. Marzouk. Transport map accelerated Markov chain Monte Carlo. SIAM/ASA Journal on Uncertainty Quantification, 6(2):645–682, 2018. Dhruv V. Patel, Deep Ray, and Assad A. Oberai. Solution of physics-based Bayesian inverse problems with deep generative priors. Computer Methods in Applied Mechanics and Engineering, 400:115428, 2022. Stefano Peluchetti. Diffusion bridge mixture transports, Schrödinger bridge problems and generative modeling. Journal of Machine Learning Research, 24(374):1–51, 2023. Jakiw Pidstrigach, Elizabeth Louise Baker, Carles Domingo-Enrich, George Deligiannidis, and Nikolas Nüsken. Conditioning diffusions using Malliavin calculus. In Proceedings of the 42nd International Conference on Machine Learning, pp. 49292–49315, 2025. Thorben Pieper-Sethmacher, Frank van der Meulen, and Aad van der Vaart. Simulation of infinite-dimensional diffusion bridges. arXiv preprint arXiv:2503.13177, 2025. Renfrey B. Potts. Some generalized order-disorder transformations. Mathematical Proceedings of the Cambridge Philosophical Society, 48(1):106–109, 1952. Vishal Purohit, Matthew Repasky, Jianfeng Lu, Qiang Qiu, Yao Xie, and Xiuyuan Cheng. Consistency posterior sampling for diverse image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 28327– 28336, 2025. Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. In Advances in Neural Information Processing Systems, volume 20, pp. 1177–1184, 2007. Balram S. Rajput and Jan Rosiński. Spectral representations of infinitely divisible processes. Probability Theory and Related Fields, 82(3):451–487, 1989. Carl Edward Rasmussen and Christopher K.I. Williams. Gaussian Processes for Machine Learning. Adaptive Computation and Machine Learning. MIT Press, Cambridge, MA, 2006. Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning, pp. 1530–1538, 2015. Jaakko Riihimäki and Aki Vehtari. Gaussian processes with monotonicity information. In Proceedings of The 13th International Conference on Artificial Intelligence and Statistics, pp. 645–652, 2010. Gareth O. Roberts and Osnat Stramer. On inference for partially observed nonlinear diffusion models using the Metropolis–Hastings algorithm. Biometrika, 88(3):603–621, 2001. Murray Rosenblatt. Remarks on a multivariate transformation. The Annals of Mathematical Statistics, 23(3):470–472, 1952. Håvard Rue and Leonhard Held. Gaussian Markov Random Fields: Theory and Applications. Monographs on Statistics and Applied Probability. Chapman & Hall/CRC, Boca Raton, FL, 2005. Håvard Rue, Sara Martino, and Nicolas Chopin. Approximate Bayesian inference for latent Gaussian models by using integrated nested Laplace approximations. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 71(2):319–392, 2009. Simo Särkkä. Bayesian Filtering and Smoothing. Cambridge University Press, Cambridge, 2013. Moritz Schauer, Frank van der Meulen, and Harry van Zanten. Guided proposals for simulating multi-dimensional diffusion bridges. Bernoulli, 23(4A):2917–2950, 2017. Kiyoung Seong, Seonghyun Park, Seonghwan Kim, Woo Youn Kim, and Sungsoo Ahn. Transition path sampling with improved off-policy training of diffusion path samplers. In Proceedings of the 13th International Conference on Learning Representations, 2025. 18

Preprint. Work in progress.

Amar Shah, Andrew Gordon Wilson, and Zoubin Ghahramani. Student-t processes as alternatives to Gaussian processes. In Proceedings of The 17th International Conference on Artificial Intelligence and Statistics, pp. 877–885, 2014. Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, and Rajesh Ranganath. A general framework for inference-time scaling and steering of diffusion models. In Proceedings of the 42nd International Conference on Machine Learning, pp. 55810–55827, 2025. Marta Skreta, Tara Akhound-Sadegh, Viktor Ohanesian, Roberto Bondesan, Alan AspuruGuzik, Arnaud Doucet, Rob Brekelmans, Alexander Tong, and Kirill Neklyudov. Feynman– Kac correctors in diffusion: Annealing, guidance, and product of experts. In Proceedings of the 42nd International Conference on Machine Learning, pp. 55906–55949, 2025. Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, pp. 2256–2265, 2015. Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In Proceedings of the 9th International Conference on Learning Representations, 2021. Andrew M. Stuart. Inverse problems: A Bayesian perspective. Acta Numerica, 19:451–559, 2010. Laura Swiler, Mamikon Gulian, Ari Frankel, Cosmin Safta, and John Jakeman. A survey of constrained Gaussian process Regression: Approaches and implementation challenges. Journal of Machine Learning for Modeling and Computing, 1(2):119–156, 2020. Kirill Tamogashev and Nikolay Malkin. Data-to-energy stochastic dynamics. In Proceedings of the 14th International Conference on Learning Representations, 2026. Francisco Vargas, Will Sussman Grathwohl, and Arnaud Doucet. Denoising diffusion samplers. In Proceedings of the 11th International Conference on Learning Representations, 2023a. Francisco Vargas, Andrius Ovsianas, David Fernandes, Mark Girolami, Neil D. Lawrence, and Nikolas Nüsken. Bayesian learning via neural Schrödinger–Föllmer flows. Statistics and Computing, 33(3), 2023b. Francisco Vargas, Shreyas Padhy, Denis Blessing, and Nikolas Nüsken. Transport meets variational inference: Controlled Monte Carlo diffusions. In Proceedings of the 12th International Conference on Learning Representations, 2024. Siddarth Venkatraman, Mohsin Hasan, Minsu Kim, Luca Scimeca, Marcin Sendera, Yoshua Bengio, Glen Berseth, and Nikolay Malkin. Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models. In Proceedings of the 42nd International Conference on Machine Learning, pp. 61212–61239, 2025. Vito Volterra. Fluctuations in the abundance of a species considered mathematically. Nature, 118:558–560, 1926. Julia Walchessen, Andrew Zammit-Mangion, Raphaël Huser, and Mikael Kuusela. Neural conditional simulation for complex spatial processes. arXiv preprint arXiv:2508.20067, 2025. Jin Wang, Kun Zhang, Li Xu, and Erkang Wang. Quantifying the waddington landscape and biological paths for development and differentiation. Proceedings of the National Academy of Sciences, 108(20):8257–8262, 2011. Yuanzhe Wang and Alexandre M. Tartakovsky. Latent diffusion posterior sampling with surrogate likelihood guidance for PDE inverse problems. arXiv preprint arXiv:2606.26592, 2026. Zifan Wang, Alice Harting, Matthieu Barreau, Michael M. Zavlanos, and Karl H. Johansson. Source-guided flow matching. In Proceedings of the 14th International Conference on Learning Representations, 2026. 19

Preprint. Work in progress.

Jay Whang, Erik M. Lindgren, and Alexandros G. Dimakis. Composing normalizing flows for inverse problems. In Proceedings of the 38th International Conference on Machine Learning, pp. 11158–11169, 2021. Gavin A. Whitaker, Andrew Golightly, Richard J. Boys, and Christopher G. Sherlock. Bayesian inference for diffusion-driven mixed-effects models. Bayesian Analysis, 12(2): 435–463, 2017. James T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky, and Marc Peter Deisenroth. Efficiently sampling functions from Gaussian process posteriors. In Proceedings of the 37th International Conference on Machine Learning, pp. 10292–10302, 2020. James T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky, and Marc Peter Deisenroth. Pathwise conditioning of Gaussian processes. Journal of Machine Learning Research, 22(105):1–47, 2021. Robert L. Wolpert, Merlise A. Clyde, and Chong Tu. Stochastic expansions using continuous dictionaries: Lévy adaptive regression kernels. The Annals of Statistics, 39(4):1916–1962, 2011. Chen Henry Wu, Saman Motamed, Shaunak Srivastava, and Fernando D. De la Torre. Generative visual prompt: Unifying distributional control of pre-trained generative models. In Advances in Neural Information Processing Systems, volume 35, pp. 22422–22437, 2022. Luhuan Wu, Brian L. Trippe, Christian A. Naesseth, David M. Blei, and John P. Cunningham. Practical and asymptotically exact conditional sampling in diffusion models. In Advances in Neural Information Processing Systems, volume 36, pp. 31372–31403, 2023. Luhuan Wu, Yi Han, Christian Andersson Naesseth, and John P. Cunningham. Reverse diffusion sequential Monte Carlo samplers. In Advances in Neural Information Processing Systems, volume 38, 2025. Yingzhi Xia, Setthakorn Tanomkiattikun, Liangli Zhen, and Zaiwang Gu. Noise-adaptive diffusion sampling for inverse problems without task-specific tuning. In Proceedings of the 14th International Conference on Learning Representations, 2026. Gefan Yang, Elizabeth Louise Baker, Michael Lind Severinsen, Christy Anna Hipsley, and Stefan Sommer. Infinite-dimensional diffusion bridge simulation via operator learning. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics, volume 258, pp. 3556–3564, 2025a. Gefan Yang, Frank van der Meulen, and Stefan Sommer. Neural guided diffusion bridges. In Proceedings of the 42nd International Conference on Machine Learning, pp. 71210–71230, 2025b. Jiachen Yao, Abbas Mammadov, Julius Berner, Gavin Kerrigan, Jong Chul Ye, Kamyar Azizzadenesheli, and Anima Anandkumar. Guided diffusion sampling on function spaces with applications to PDEs. In Advances in Neural Information Processing Systems, volume 38, 2025. Andrew Zammit-Mangion, Matthew Sainsbury-Dale, and Raphaël Huser. Neural methods for amortized inference. Annual Review of Statistics and Its Application, 12(1):311–335, 2025. Qinsheng Zhang and Yongxin Chen. Path integral sampler: a stochastic control approach for sampling. In Proceedings of the 10th International Conference on Learning Representations, 2022. Zheng Zhao, Ziwei Luo, Jens Sjölund, and Thomas B. Schön. Conditional sampling within generative diffusion models. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 383(2299):20240329, 2025.

20

Preprint. Work in progress.

A

Examples of latent innovation representations

This appendix contains examples of stochastic-process priors that admit the latent innovation representation in Assumption 2.1, or its measure-theoretic extension; see Remark B.1 for details. Throughout, s1:m denotes the discretisation inputs, f0 ∈ Em denotes the finitedimensional process realisation, and Tϑ denotes a generator satisfying f0 = Tϑ (ξ0 ). In several examples, the model is most naturally written in terms of a native innovation variable ζ0 , such as a uniform simulator seed, a Gamma scale variable, or a collection of nonGaussian increments. To ensure all examples are consistent with the common latent form used by LatentFlow, we express this native innovation as a deterministic function of a standard Gaussian reference: ζ0 = Hϑ (ξ0 ), with ξ0 ∼ N (0, I). The generator in Assumption 2.1 is then the composite map from the Gaussian innovation to the process realisation. When the native innovation is already Gaussian, Hϑ is simply the identity map. Table 1: Examples of stochastic-process priors and their latent innovation representations. Model class

Innovation representation

Gaussian, Student-t, and basis-expansion priors Gaussian process

A finite-dimensional realisation f0 = (f (s1 ), . . . , f (sm )) of a Gaussian process can be written as f0 = µϑ + Lϑ ξ0 , Lϑ L⊤ ϑ = Kϑ , where (Kϑ )ij = kϑ (si , sj ) is the kernel matrix, and ξ0 ∼ N (0, Im ). The innovation is given by ξ0 ∈ Rm , ξ0 ∼ N (0, Im ). The generator is given by Tϑ (ξ0 ) = µϑ + Lϑ ξ0 . This is the affine whitening map used in finite-dimensional Gaussian process inference.

Student-t process

A finite-dimensional realisation f0 = (f (s1 ), . . . , f (sm )) of a Student-t process can be written as f0 = µϑ + ω −1/2 Lϑ v, Lϑ L⊤ ϑ = Kϑ , where (Kϑ )ij = kϑ (si , sj ) is the scale matrix induced by the Student-t process kernel kϑ , and v and ω are independent, with v ∼ N (0, Im ) and ω ∼ Gamma(ν/2, ν/2) in the shape–rate parameterisation. The natural innovation is given by ζ0 = (v, ω) ∈ Rm+1 . Let w = Φ−1 N (FΓ,ν (ω)), where FΓ,ν is the distribution function of Gamma(ν/2, ν/2), and ΦN is the standard normal distribution function. The Gaussian innovation is then given by ξ0 = (v, w) ∈ Rm+1 , ξ0 ∼ N (0, Im+1 ). The corresponding generator is given by −1 Tϑ (ξ0 ) = µϑ + FΓ,ν (ΦN (w))−1/2 Lϑ v. This gives a heavy-tailed analogue of the Gaussian process while retaining an explicit scale-mixture innovation map.

continued on next page

21

Preprint. Work in progress.

Model class

Innovation representation

Finite-basis, random-feature, or pathwise GP approximation

A truncated Karhunen–Loève expansion, finite basis expansion, random-feature approximation, or pathwise Gaussian process approximation can be written as f0 = µ ϑ + Φ ϑ ξ 0 , where Φϑ ∈ Rm×q contains the basis or feature evaluations at s1:m , and ξ0 ∼ N (0, Iq ). For example,

p

(Φϑ )ik = λϑ,k ϕϑ,k (si ) for a truncated Karhunen–Loève expansion. The innovation is given by ξ0 ∈ Rq , ξ0 ∼ N (0, Iq ). The generator is given by Tϑ (ξ0 ) = µϑ + Φϑ ξ0 . The innovation dimension q need not equal the discretisation dimension m. When q < m, the induced prior on Em is typically supported on a low-dimensional subspace unless an additional residual or nugget term is included. More elaborate pathwise Gaussian process constructions fit the same form after augmenting ξ0 with auxiliary Gaussian variables. Spectral random field

A finite spectral random field can be written as f (s) = µϑ (s) +

q X

Aϑ,k ak cos(ωk⊤ s) + bk sin(ωk⊤ s) ,



k=1 i.i.d.

where Aϑ,k are spectral amplitudes, ωk are frequencies, and ak , bk ∼ N (0, 1). Here the frequencies ωk are treated as fixed, or as part of ϑ. If the frequencies are themselves random, they can be appended to the innovation variable. The innovation is given by ξ0 = (a1 , b1 , . . . , aq , bq ) ∈ R2q , ξ0 ∼ N (0, I2q ). The generator is given by (Tϑ (ξ0 ))i = µϑ (si ) +

q X

Aϑ,k ak cos(ωk⊤ si ) + bk sin(ωk⊤ si ) ,



k=1

for components i = 1, . . . , m. Gaussian Markov random field, or SPDE-defined Matérn field

A Gaussian Markov random field, including an SPDE-defined Matérn field after finite-element discretisation, can be written as Rϑ⊤ Rϑ = Qϑ ,

f0 = µϑ + Aϑ Rϑ−1 ξ0 ,

where xϑ = Rϑ−1 ξ0 ∈ Rq is a vector of GMRF or finite-element coefficients, Qϑ ≻ 0 is its sparse precision matrix, Aϑ : Rq → Em maps the coefficients to the requested process values, and ξ0 ∼ N (0, Iq ). For an SPDE-defined Matérn field, Qϑ is the sparse precision matrix of the finite-element coefficient vector, and Aϑ is the corresponding finite-element evaluation or projection matrix. The innovation is given by ξ0 ∈ Rq , ξ0 ∼ N (0, Iq ). The generator is given by Tϑ (ξ0 ) = µϑ + Aϑ Rϑ−1 ξ0 . In practice, one solves a sparse triangular system rather than forming Rϑ−1 . This includes proper areal Gaussian Markov random fields such as proper CAR and Leroux-type models, as well as SPDE-defined Matérn fields. Intrinsic CAR or intrinsic GMRF precision matrices are singular and instead require an identifying constraint, a reducedcoordinate factorisation on the constrained subspace, or an explicitly specified generalised-inverse construction. continued on next page

22

Preprint. Work in progress.

Model class

Innovation representation

Intensity and convolution-field priors Log-Gaussian Cox intensity

Lévy-driven or convolution random field

A discretised log-Gaussian Cox process intensity can be written as f0 = exp(µϑ + Lϑ ξ0 ), Lϑ L⊤ ϑ = Kϑ , where f0 = (λ(s1 ), . . . , λ(sm )), g0 = µϑ + Lϑ ξ0 is the discretised latent Gaussian process, λ(si ) = exp(g0 (si )), and ξ0 ∼ N (0, Im ). The innovation is given by ξ0 ∈ Rm , ξ0 ∼ N (0, Im ). The generator is given by Tϑ (ξ0 ) = exp(µϑ + Lϑ ξ0 ), where the exponential is applied componentwise. A finite Lévy-driven or stable-convolution random field can be written as f (si ) =

q X

Gϑ (si − ck )Jk ,

i = 1, . . . , m,

k=1

where Gϑ is a convolution kernel, ck are support or quadrature locations, and Jk ∼ νϑ,k are increments of an independently scattered random measure. The natural innovation is given by ζ0 = (J1 , . . . , Jq ). Let Hϑ,k be a map such that Jk = Hϑ,k (wk ), wk ∼ N (0, I). The Gaussian innovation is then given by ξ0 = (w1 , . . . , wq ). The corresponding generator is given by (Tϑ (ξ0 ))i =

q X

Gϑ (si − ck )Hϑ,k (wk ),

i = 1, . . . , m.

k=1

This gives a non-Gaussian random-field prior, including Cauchy, normal-inverse Gaussian (NIG), stable, or other heavy-tailed convolution fields, through independently scattered random-measure increments. continued on next page

23

Preprint. Work in progress.

Model class

Innovation representation

Discrete and extreme-value priors Potts or discrete Markov random field

A Potts or discrete Markov random field on sites s1:m can be written as Pϑ (f0 = x) ∝ exp

m X

hϑ,i (xi ) +

i=1



X

βϑ,ij 1(xi = xj ) ,

(i,j)∈E

where x = (x1 , . . . , xm ) ∈ {1, . . . , K}m , xi = f0 (si ), E is the neighbourhood graph on s1:m , hϑ,i are site potentials, and βϑ,ij are interaction parameters. The natural innovation can be represented by a simulator seed ζ0 = u ∈ (0, 1)q , u ∼ Unif((0, 1)q ), together with an exact sampler Sϑ such that f0 = Sϑ (u) ∼ Pϑ . If Sϑ is an approximate MCMC or sequential sampler, then the corresponding Tϑ represents the algorithmic approximation rather than the exact Potts law. Let ξ0 = Φ−1 N (u), where ΦN is the standard normal distribution function, applied componentwise. The Gaussian innovation is then given by ξ0 ∈ Rq , ξ0 ∼ N (0, Iq ). The corresponding generator is given by Tϑ (ξ0 ) = Sϑ (ΦN (ξ0 )). Exact Potts fields are discrete; gradient-based guidance therefore requires a relaxation, surrogate, or derivative-free guidance estimator. Max-stable process

A finite spectral truncation of a max-stable process can be written as f (si ) = max Pk Wϑ,k (si ), 1≤k≤K

Pk = Γ−1 k ,

Γk =

k X

Eℓ ,

ℓ=1 i.i.d.

where Ek ∼ Exp(1), the spectral functions Wϑ,k are independent and identically distributed, and the sequence (Wϑ,k )k≥1 is independent of (Ek )k≥1 . The spectral functions satisfy the usual normalisation E[Wϑ,k (s)] = 1 for each s. For finite K, this is a spectral truncation of the corresponding infinite max-stable representation. The natural innovation is given by ζ0 = (E1:K , υ1:K ), where υk generates the spectral function Wϑ,k . Let wkE = Φ−1 N (exp(−Ek )), so that, equivalently, Ek = − log ΦN (wkE ), wkE ∼ N (0, 1). If the spectral innovation can be Gaussianised as υk = HϑW (wkW ), wkW ∼ N (0, I), then the Gaussian innovation is given by E W ξ0 = (w1E , . . . , wK , w1W , . . . , wK ).

All variables wkE and wkW are taken to be mutually independent across k and across the two innovation blocks. The corresponding generator is given by W W (Tϑ (ξ0 ))i = max Γ−1 k Wϑ (si ; Hϑ (wk )),

i = 1, . . . , m,

1≤k≤K

Pk

where Γk = − log ΦN (wℓE ). The maximum is non-smooth, so ℓ=1 smooth relaxations may be required for gradient-based guidance. continued on next page

24

Preprint. Work in progress.

Model class

Innovation representation

Dynamical and spatiotemporal priors State-space model or stochastic recurrence

A state-space model or stochastic recurrence can be written as x0 = Ψϑ,0 (η0 ), xn = Ψϑ,n (xn−1 , ηn ), n = 1, . . . , N, ind

where ηn ∼ rn are model innovations and Ψϑ,n is the one-step simulator. The natural innovation is given by ζ0 = (η0 , . . . , ηN ). Let Hϑ,n be a map such that ηn = Hϑ,n (wn ), wn ∼ N (0, I). The Gaussian innovation is then given by ξ0 = (w0 , . . . , wN ). The corresponding generator is given by x0 = Ψϑ,0 (Hϑ,0 (w0 )), xn = Ψϑ,n (xn−1 , Hϑ,n (wn )), n = 1, . . . , N, with output Tϑ (ξ0 ) = (x0 , . . . , xN ). When the original innovations are already Gaussian, Hϑ,n is the identity map. For discrete or mixed innovations, the maps Hϑ,n may be nonsmooth, in which case gradient-based guidance again requires a relaxation, surrogate, or derivative-free estimator. Linear Gaussian and nonlinear Gaussian state-space models correspond to particular choices of Ψϑ,n . Diffusion process

A diffusion process satisfying dxτ = bϑ (τ, xτ ) dτ + Σϑ (τ, xτ ) dwτ can be discretised by Euler–Maruyama as

p

xn+1 = xn + bϑ (τn , xn )∆n + Σϑ (τn , xn )

∆n η n ,

ind

where ∆n = τn+1 − τn , ηn ∼ N (0, I), and Σϑ (τn , xn ) maps Gaussian innovations into diffusion increments. The innovation is given by ξ0 = (η−1 , η0 , . . . , ηN −1 ), where η−1 ∼ N (0, I) generates the initial state if x0 is random, through an initial-state map x0 = χϑ,0 (η−1 ), and is omitted otherwise. The generator is given by applying the Euler–Maruyama recursion above and returning Tϑ (ξ0 ) = (x0 , . . . , xN ). More generally, higher-order numerical solvers correspond to different choices of Tϑ . continued on next page

25

Preprint. Work in progress.

Model class

Innovation representation

Jump diffusion, Lévy-driven SDE, or Lévy process

A jump diffusion or Lévy-driven SDE satisfying dxτ = bϑ (τ, xτ ) dτ + Σϑ (τ, xτ ) dWτ + djτ can be discretised as xn+1 = xn + bϑ (τn , xn )∆n + Σϑ (τn , xn )

p

∆n η n + ℓ n ,

ind

where ∆n = τn+1 − τn , ηn ∼ N (0, I), and ℓn ∼ νϑ,∆n is the jump or Lévy increment over [τn , τn+1 ]. The natural innovation is given by ζ0 = (η−1 , η0 , . . . , ηN −1 , ℓ0 , . . . , ℓN −1 ), where η−1 ∼ N (0, I) generates the initial state if x0 is random, through an initial-state map x0 = χϑ,0 (η−1 ), and is omitted otherwise. Let Hϑ,∆n be a map such that ℓn = Hϑ,∆n (wn ), wn ∼ N (0, I). The Gaussian innovation is then given by ξ0 = (η−1 , η0 , . . . , ηN −1 , w0 , . . . , wN −1 ). The corresponding generator is given by applying the jump–Euler recursion with ℓn = Hϑ,∆n (wn ) and returning Tϑ (ξ0 ) = (x0 , . . . , xN ). A Lévy process or subordinator is obtained as the special case with no state-dependent drift or diffusion term and additive Lévy increments. Time-discretised SPDE

A spatially discretised SPDE of the form duτ = Aϑ (uτ ) dτ + Bϑ (uτ ) dWτ can be time-discretised as un+1 = Ψϑ,n (un , ηn ), n = 0, . . . , N − 1, ind

where ηn ∼ N (0, I) are finite-dimensional Gaussian noise innovations and Ψϑ,n is the numerical SPDE time-stepper. The innovation is given by ξ0 = (η−1 , η0 , . . . , ηN −1 ), where η−1 ∼ N (0, I) generates the initial state if u0 is random, through an initial-state map u0 = χϑ,0 (η−1 ), and is omitted otherwise. The generator is given by applying the numerical SPDE solver above and returning the discretised space-time field. continued on next page

26

Preprint. Work in progress.

Model class

Innovation representation

Implicit and simulator-defined priors Neural or deep implicit process

A neural or deep implicit process can be written as f (si ) = Gϑ (si , h, ϵi ), i = 1, . . . , m, where Gϑ is a neural decoder, h is a global latent variable, and ϵi are optional local noise variables. The innovation is given by ξ0 = (h, ϵ1 , . . . , ϵm ), with a simple reference law, typically standard Gaussian. The generator is given by (Tϑ (ξ0 ))i = Gϑ (si , h, ϵi ), i = 1, . . . , m. If the decoder is differentiable, Dξ Tϑ is available by automatic differentiation.

Black-box simulator

A black-box simulator can be written as f0 = Sϑ (s1:m ; u), where Sϑ is the simulator and u ∼ Unif((0, 1)q ) is the simulator random seed. The natural innovation is given by ζ0 = u ∈ (0, 1)q . Let ξ0 = Φ−1 N (u), where ΦN is the standard normal distribution function, applied componentwise. Then ξ0 ∼ N (0, Iq ). The corresponding generator is given by Tϑ (ξ0 ) = Sϑ (s1:m ; ΦN (ξ0 )). If Sϑ is not differentiable, guidance requires a smooth surrogate, relaxation, or derivative-free estimator.

B

Proofs for the latent conditioning principle

Proof of Proposition 2.2. Let ξ0 ∼ ρϑ (· | C). Then, for any measurable A ⊆ Em , we have that Z  1 P Tϑ (ξ0 ) ∈ A = 1{Tϑ (ξ0 )∈A} LC (Tϑ (ξ0 )) r(ξ0 ) dξ0 Zϑ Rmξ Z Z 1 LC (f0 ) pϑ (f0 ) df0 = πϑ (f0 | C) df0 . = Zϑ A A The third equality follows from the pushforward relation Tϑ# (r(ξ0 ) dξ0 ) = pϑ (f0 ) df0 , applied to the nonnegative measurable function g(f0 ) = 1A (f0 )LC (f0 ). Hence Tϑ (ξ0 ) has density πϑ (· | C), as required. Remark B.1. We use density notation in the main text to simplify the exposition. Our approach, however, admits a much more general measure-theoretic formulation. Let Pϑ denote the prior law of f0 . Assume that LC : Em → [0, ∞) is measurable. The conditional law on Em is then Z LC (f0 )Pϑ (df0 ) Πϑ (df0 | C) = , Zϑ := LC (f0 )Pϑ (df0 ). (8) Zϑ Our assumption now reads as follows. There exist a latent dimension mξ , a reference probability measure R on Rmξ , and a measurable map Tϑ : Rmξ → Em such that Pϑ = (Tϑ )# R. Equivalently, if ξ0 ∼ R, then f0 = Tϑ (ξ0 ) ∼ Pϑ . The corresponding latent target is given by Z LC (Tϑ (ξ))R(dξ) ρϑ (dξ | C) = , Zϑ = LC (Tϑ (ξ))R(dξ). (9) Zϑ We emphasise that the values of Zϑ as defined in (8) and (9) coincide, since Pϑ = Tϑ# R R R implies LC (f0 )Pϑ (df0 ) = LC (Tϑ (ξ))R(dξ) by definition of the pushforward. Now, for any 27

Preprint. Work in progress.

measurable set A ⊆ Em , we have that Z LC (Tϑ (ξ))R(dξ) Tϑ# ρϑ (· | C)(A) = 1{Tϑ (ξ)∈A} Zϑ Z 1 = LC (f0 )Pϑ (df0 ) = Πϑ (A | C). Zϑ A It follows that Tϑ# ρϑ (· | C) = Πϑ (· | C). This is the measure-theoretic form of Proposition 2.2. The density-based statement in the main text is recovered when Pϑ admits a density pϑ . This more general argument also covers singular and discrete laws. Our next result establishes that the latent representation also preserves the relative-entropy cost of conditioning. Proposition B.2. Suppose that Assumption 2.1 holds and that 0 < Zϑ < ∞. Then, in the extended sense,   KL ρϑ (· | C) r = KL πϑ (· | C) pϑ . Proof. By construction, we have Λϑ,C (ξ0 ) LC (Tϑ (ξ0 )) ρϑ (ξ0 | C) = = r(ξ0 ) Zϑ Zϑ

r-a.e.

(10)

Similarly, due to (1), it holds that πϑ (f0 | C) LC (f0 ) = pϑ (f0 ) Zϑ Let ℓ(f0 ) = log



LC (f0 ) Zϑ



pϑ -a.e.

(11)

, with the usual extended-value convention. Using (10), we obtain 

KL ρϑ (· | C) r =

Z Rmξ

ℓ(Tϑ (ξ0 ))ρϑ (ξ0 | C) dξ0 .

By Proposition 2.2, pushing ρϑ (· | C) forward through Tϑ gives πϑ (· | C). We therefore have, Z Z ℓ(Tϑ (ξ0 ))ρϑ (ξ0 | C) dξ0 = ℓ(f0 )πϑ (f0 | C) df0 . (12) Rmξ

Em

 Finally, using (11), the right-hand side of (12) is exactly KL πϑ (· | C) pϑ . This proves the identity. This result suggests that the latent construction is not just a convenient sampling device. Since the additional likelihood depends on the latent variable only through the generated process value Tϑ (ξ0 ), the relative-entropy deformation from the prior pϑ to the conditional πϑ (· | C) is exactly matched by the deformation from the reference r to the latent conditional ρϑ (· | C). In contrast, for a generic latent density η on Rmξ , if pη denotes the density of Tϑ (ξ0 ) when ξ0 ∼ η, then the data-processing inequality only gives KL(pη ∥ pϑ ) ≤ KL(η ∥ r). Thus the latent target introduces no additional information beyond the output-level conditioning likelihood.

C

Exactness of the latent dynamics

We now fix ϑ and C, and write T = Tϑ , Λ = Λϑ,C , Z = Zϑ , and ρ0 (ξ0 ) = ρϑ (ξ0 | C). We also write r(ξ) = N (ξ; 0, Imξ ). Recall that the VP noising schedule is specified by a locally integrable function β : [0, 1) → (0, ∞). We write Z t  1 σ(t)2 = 1 − α(t)2 . B(t) = β(s) ds, α(t) = exp − B(t) , 2 0 28

Preprint. Work in progress.

We now allow either a finite-noising horizon, where B(1) := limt↑1 B(t) < ∞, or an ideal infinite-noising horizon, where B(1) = ∞ and hence α(t) ↓ 0 as t ↑ 1. For t ∈ [0, 1), the VP bridge is ξt = α(t)ξ0 + σ(t)ε, ε ∼ N (0, Imξ ), (13) where ε is independent of ξ0 . We write ρt for the density of the noised latent variable obtained by drawing ξ0 ∼ ρ0 and applying (13). Equivalently, ψ(t, u)r(u) , ψ(t, u) = Eξ0 ∼r [Λ(ξ0 ) | ξt = u] . Z where the conditional expectation is taken under the reference bridge with ξ0 ∼ r. Finally, assuming ψ(t, ·) is differentiable, we write g(t, u) = ∇u log ψ(t, u) for the latent guidance field. ρt (u) =

In Section 3.2, we formally identified the SDE associated with the guided latent marginals (ρt )0≤t<1 . We now recall these dynamics using notation that is convenient for the exactness statements below. The forward noising SDE is given by p 1 ξ0 ∼ ρ0 , 0 ≤ t < 1, dξt = − β(t)ξt dt + β(t) dwt , 2 where w = (wt )0≤t<1 is a standard Rmξ -valued Brownian motion. Fix 0 ≤ t0 < t1 < 1, and set τs = t1 − s, for 0 ≤ s ≤ t1 − t0 . The reverse-time SDE is then given by   p 1 ← ← ← dξs = β(τs )g(τs , ξs ) − β(τs )ξs ds + β(τs ) dws , ξ0← ∼ ρt1 , (14) 2 where w = (ws )0≤s≤t1 −t0 is a standard Brownian motion in reverse time. The same family of marginals can also be generated by a deterministic process known as the probability-flow ODE (e.g., Song et al., 2021). In forward noising time, this ODE is given by dut 1 = − β(t)g(t, ut ), u0 ∼ ρ0 , 0 ≤ t < 1. (15) dt 2 For sampling, we again use the reverse-time parametrisation u← s = uτs , with τs as defined above. In particular, the reverse-time probability-flow ODE is du← 1 s = β(τs )g(τs , u← u← (16) s ), 0 ∼ ρt1 . ds 2 Assumption C.1. For every compact interval [t0 , t1 ] ⊂ (0, 1), ρt is a strictly positive Lebesgue density for all t ∈ [t0 , t1 ]. The map (t, u) 7→ ∇u log ρt (u) is sufficiently regular that the reverse-time SDE (14) and the probability-flow ODE (16) are well posed on the corresponding time intervals, and their associated Fokker–Planck and continuity equations have unique density-valued solutions. Theorem C.2. Suppose that Assumption 2.1 holds, that r = N (0, Imξ ), that 0 < Z < ∞, and that Assumption C.1 holds. Fix 0 < t0 < t1 < 1, and define τs = t1 −s for 0 ≤ s ≤ t1 −t0 . Let ξ0← ∼ ρt1 , and let (ξs← )0≤s≤t1 −t0 solve the reverse-time SDE (14). Then ξs← ∼ ρτs for all 0 ≤ s ≤ t1 − t0 . In particular, ξt← ∼ ρt0 . If the reverse SDE is well defined down to 1 −t0 t0 = 0, then ξt← ∼ ρ0 = ρϑ (· | C), Tϑ (ξt← ) ∼ πϑ (· | C). 1 1 Proof. The proof is standard. Let bt (u) = − 21 β(t)u denote the drift of the forward VP SDE. By the standard time-reversal formula for diffusions (e.g., Anderson, 1982), the reverse-time diffusion with clock τs = t1 − s has drift given by ebs (u) = −bτ (u) + β(τs )∇u log ρτ (u). s s −1

(17)

Moreover, since ρt (u) = Z ψ(t, u)r(u) and r(u) = N (u; 0, Imξ ), it follows straightforwardly that ∇u log ρt (u) = ∇u log ψ(t, u) − u = g(t, u) − u. (18) We thus have, substituting (18) into (17), that ebs (u) = 1 β(τs )u + β(τs ) (g(τs , u) − u) = β(τs )g(τs , u) − 1 β(τs )u. 2 2 29

Preprint. Work in progress.

This is exactly the drift in (14). Equivalently, if qs (u) = ρτs (u), then qs solves the Fokker– Planck equation associated with (14). Indeed, 1 1 ∂s qs = −∂t ρt t=τs = ∇ · (bτs ρτs ) − β(τs )∆ρτs = −∇ · (ebs qs ) + β(τs )∆qs . 2 2 By uniqueness in Assumption C.1, the law of ξs← has density qs = ρτs . The claims at t0 = 0 follow because ρ0 = ρϑ (· | C). Finally, the process-space statement follows from Proposition 2.2. Theorem C.3. Suppose that Assumption 2.1 holds, that r = N (0, Imξ ), that 0 < Z < ∞, and that Assumption C.1 holds. Fix 0 < t0 < t1 < 1, and define τs = t1 − s for 0 ≤ s ≤ t1 − t0 . ← If u← 0 ∼ ρt1 and (us )0≤s≤t1 −t0 solves the reverse-time probability-flow ODE (16), then ← us ∼ ρτs for all 0 ≤ s ≤ t1 − t0 . In particular, u← t1 −t0 ∼ ρt0 . If the flow is well defined down to t0 = 0, then u← Tϑ (u← t1 ∼ ρ0 = ρϑ (· | C), t1 ) ∼ πϑ (· | C). Proof. The marginals (ρt )0≤t<1 are generated by the forward VP-SDE when the initial law is ρ0 = ρϑ (· | C). Indeed, for any bounded measurable test function φ, 1 Eξ ∼r,ε∼r [Λ(ξ0 )φ (α(t)ξ0 + σ(t)ε)] Z 0 Z 1 φ(u)ρt (u) du. = Eu∼r [φ(u)E [Λ(ξ0 ) | ξt = u]] = Z Rmξ

Eξ0 ∼ρ0 [φ(ξt )] =

Thus ρt solves the Fokker–Planck equation 1 ∂t ρt = −∇ · (bt ρt ) + β(t)∆ρt , 2 Writing st (u) = ∇u log ρt (u), we can rewrite (19) as

1 bt (u) = − β(t)u. 2

(19)

   1 1 ∂t ρt = −∇ · (bt ρt ) + β(t)∇ · (ρt st ) = −∇ · bt − β(t)st ρt . 2 2 It follows that ρt also solves the continuity equation ∂t ρt + ∇ · (vt ρt ) = 0,

1 vt (u) = bt (u) − β(t)st (u). 2

Meanwhile, substituting (18), we have that 1 1 1 vt (u) = − β(t)u − β(t) (g(t, u) − u) = − β(t)g(t, u). 2 2 2 This is exactly the vector field in the forward probability-flow ODE (15). By Assumption C.1, the ODE flow is well posed and gives the unique solution of the continuity equation. Therefore the backward flow (16) transports ρt1 to ρt0 . Taking t0 = 0 gives u← t1 ∼ ρ0 = ρϑ (· | C), and the process-space statement follows from Proposition 2.2. Remark C.4. If B(1) = ∞, then α(t) ↓ 0 and the noised target marginals satisfy ρt → r in total variation as t ↑ 1, provided Λ ∈ L1 (r). In this sense, starting from u← 0 ∼ r can be viewed as the limiting infinite-noising version of Theorem C.3. For a finite noising horizon, exact sampling from a terminal time t1 requires initialising from the terminal marginal ρt1 , not from r. If one instead starts from r, the initialisation discrepancy is ∥ρt1 − r∥TV , and the discrepancy after exact reverse probability flow is at most this quantity, with equality when the deterministic flow is invertible.

D

Approximation and stability guarantees

The practical sampler replaces the exact guidance field g(t, u) by an approximate field gbS (t, u), and then solves the resulting ODE or SDE numerically. This introduces two sources of error: the guidance approximation error and the numerical discretisation error. Both errors enter the probability-flow ODE and the reverse-time SDE through the same perturbation mechanism. We thus state the main stability and numerical-error bounds in a unified form. 30

Preprint. Work in progress.

D.1

Common perturbation framework

We work on a truncated forward-time interval [τ, t1 ], where 0 < τ < t1 < 1, and use reverse-time notation throughout this subsection. Set Rτ,t1 = t1 − τ,

κs = t1 − s,

0 ≤ s ≤ Rτ,t1 .

Thus s denotes reverse time and κs denotes the corresponding forward noising time. We index the two reverse dynamics by χ ∈ {ode, sde}, and define χ ode sde

λχ 1 2

ρχ 0

1

Σχ (s) p 0 β(κs )Imξ .

1 2

For an exact guidance field g and an approximate guidance field gb, define the exact and approximate reverse-time drifts Fχ (s, u) = λχ β(κs )g(κs , u) − ρχ β(κs )u,

Fbχ (s, u) = λχ β(κs )b g (κs , u) − ρχ β(κs )u. (20)

In particular, for the probability-flow ODE we have 1 β(κs )g(κs , u), 2 1 b a(s, u) = Fbode (s, u) = β(κs )b g (κs , u). 2 a(s, u) = Fode (s, u) =

Meanwhile, for the reverse-time SDE we have 1 b(s, u) = Fsde (s, u) = β(κs )g(κs , u) − β(κs )u, 2 1 bb(s, u) = Fbsde (s, u) = β(κs )b g (κs , u) − β(κs )u. 2 The exact and approximate dynamics may then be written as dXsχ = Fχ (s, Xsχ ) ds + Σχ (s) dWs ,

b χ = Fbχ (s, X b χ ) ds + Σχ (s) dWs . dX s s

(21)

b ode = u When χ = ode, we write X ode = u← , X b← . When χ = sde, we write X sde = ← b sde ← ξ ,X = ξb . Let Rτ,t1 ⊂ [τ, t1 ] × Rmξ be a region on which the exact and approximate guidance fields are compared, and which contains the exact and approximate trajectories under consideration. Define the realised pointwise guidance error by n o ε(s) = sup ∥g(κs , u) − gb(κs , u)∥ : (κs , u) ∈ Rτ,t1 , δχ (s) = λχ β(κs )ε(s). (22) For the ODE and the SDE, we thus have that δode (s) =

1 β(κs )ε(s), 2

δsde (s) = β(κs )ε(s).

In the common case where ετ,t1 := sup(t,u)∈Rτ,t1 ∥g(t, u) − gb(t, u)∥ < ∞, it follows from (22) that δχ (s) ≤ λχ β(κs )ετ,t1 . Given a time-dependent one-sided Lipschitz rate η ∈ L1 ([0, Rτ,t1 ]), we define the integral Z s  Θη (s, r) = exp η(q) dq , 0 ≤ r ≤ s ≤ Rτ,t1 . r

Finally, throughout the statements below, we treat gb as a realised approximate guidance field. If gb is random, our results are to be read conditionally on this realisation. Probabilistic rates follow by combining these conditional stability bounds with separate probabilistic control of the realised guidance error. 31

Preprint. Work in progress.

D.2

Unified stability of approximate guidance

The following result is a variable-rate version of the deterministic perturbation bound obtained in (Moss et al., 2026, Proposition E.5), restated in our present notation and extended to cover the reverse-time SDE under a synchronous coupling. b χ solve (21) on [0, Rτ,t ]. For Proposition D.1. Fix χ ∈ {ode, sde}, and let X χ and X 1 the SDE, assume that the two equations are driven by the same Brownian motion. Assume bsχ ) ∈ Rτ,t for all s ∈ [0, Rτ,t ]. Suppose that the exact reverse-time drift that (κs , Xsχ ), (κs , X 1 1 satisfies the one-sided Lipschitz condition ⟨x − y, Fχ (s, x) − Fχ (s, y)⟩ ≤ ηχ (s)∥x − y∥2

(23)

whenever (κs , x) and (κs , y) belong to Rτ,t1 . Then, for every 0 ≤ s ≤ Rτ,t1 , the following bound holds: Z s χ χ χ χ b b ∥Xs − Xs ∥ ≤ Θηχ (s, 0)∥X0 − X0 ∥ + Θηχ (s, r)δχ (r) dr. (24) 0

If the exact and approximate solves are coupled by the same reverse-time initial point, equivalently by the same terminal point at forward time t1 , the first term is absent. Moreover, for p ≥ 1, provided the initial laws have finite p-th moments,     Z s bsχ ) ≤ Θη (s, 0)Wp L(X χ ), L(X b χ) + Wp L(Xsχ ), L(X Θηχ (s, r)δχ (r) dr. χ 0 0 0

b χ . When χ = ode, this is an ordinary deterministic difference Proof. Let Ds = Xsχ − X s equation. When χ = sde, the Brownian terms cancel under the synchronous coupling. In either case, we have Z s brχ )} dr. Ds = D0 + {Fχ (r, Xrχ ) − Fbχ (r, X 0

brχ ) = Fχ (r, Xrχ ) − Fχ (r, X brχ ) + Fχ (r, X brχ ) − We can write the integrand as Fχ (r, Xrχ ) − Fbχ (r, X χ b b Fχ (r, Xr ). The first term is controlled by (23). The second term satisfies ∥Fχ (r, u) − Fbχ (r, u)∥ = λχ β(κr )∥g(κr , u) − gb(κr , u)∥ ≤ δχ (r) on the region under consideration. Hence the upper right-hand derivative satisfies d+ ∥Ds ∥ ≤ ηχ (s)∥Ds ∥ + δχ (s). ds The variable-coefficient Gronwall inequality gives (24). The Wasserstein estimate follows by coupling the initial variables optimally and, in the SDE case, using the same Brownian motion for the two SDEs. Corollary D.2. Suppose the conditions of Proposition D.1 hold. Suppose in addition that ηχ (s) ≤ η̄χ,τ,t1 on [0, Rτ,t1 ] and β̄τ,t1 = ess supt∈[τ,t1 ] β(t) < ∞. Then, under common terminal coupling, it holds for all s ∈ [0, Rτ,t1 ] that ( ηr e −1 χ χ η , η ̸= 0, bs ∥ ≤ λχ β̄τ,t ετ,t Ψη̄ ∥Xs − X (s), Ψη (r) = 1 1 χ,τ,t1 r, η = 0. Proof. Under common terminal coupling, the first term in (24) vanishes. Moreover, since ηχ (s) ≤ η̄χ,τ,t1 and δχ (s) ≤ λχ β̄τ,t1 ετ,t1 , we have Θηχ (s, r) ≤ exp{η̄χ,τ,t1 (s − r)}. Therefore Z s χ χ b ∥Xs − Xs ∥ ≤ λχ β̄τ,t1 ετ,t1 exp{η̄χ,τ,t1 (s − r)} dr = λχ β̄τ,t1 ετ,t1 Ψη̄χ,τ,t1 (s). 0

32

Preprint. Work in progress.

D.3

Curvature regimes

The sign of the one-sided Lipschitz rate ηχ determines the nature of the bound in Proposition D.1 and Corollary D.2. This, in turn, is determined by the curvature of the smoothed latent log-likelihood. This is the subject of the following proposition. Proposition D.3. Suppose that, uniformly for t ∈ [τ, t1 ], g is one-sided Lipschitz with constant ℓτ,t1 on the truncated region: ⟨x − y, g(t, x) − g(t, y)⟩ ≤ ℓτ,t1 ∥x − y∥2

(25)

Then Proposition D.1 applies with ηχ (s) = β(κs ) (λχ ℓτ,t1 − ρχ ) .

(26)

Suppose, more strongly, that uniformly for t ∈ [τ, t1 ], g is strongly dissipative with constant µτ,t1 ≥ 0, on the truncated region: ⟨x − y, g(t, x) − g(t, y)⟩ ≤ −µτ,t1 ∥x − y∥2 .

(27)

Then Proposition D.1 applies with ηχ (s) = − (ρχ + λχ µτ,t1 ) β(κs ).

(28)

Proof. Recalling the definition of the reverse-time drift Fχ from (20), it is straightforward to compute ⟨x − y, Fχ (s, x) − Fχ (s, y)⟩ = λχ β(κs )⟨x − y, g(κs , x) − g(κs , y)⟩ − ρχ β(κs )∥x − y∥2 . The rates given in (26) and (28) now follow immediately after substituting (25) and (27), respectively. D.4

Guidance and discretisation error

The previous stability result controls the error between the exact guided dynamics and the exact approximately guided dynamics. It is also necessary to control the discretisation error introduced by a numerical solver. Corollary D.4. Assume that the hypotheses of Proposition D.1 hold for the exact and approximately guided ODE and SDE dynamics. Let 0 = s0 < · · · < sN = Rτ,t1 be a reverse-time grid with maximal step size h. Define Z s Aηχ = sup Θηχ (s, 0), Hηχ ,δχ = sup Θηχ (s, r)δχ (r) dr. 0≤s≤Rτ,t1

0≤s≤Rτ,t1

0

Suppose that the numerical approximations u b←,h and ξb←,h to the approximately guided ODE and SDE satisfy ode q max ∥b u← b←,h sn − u sn ∥ ≤ Cnum,τ,t1 h ,

0≤n≤N

sde max ∥ξbs←n − ξbn←,h ∥ Lp ≤ Cnum,p hϱ .

0≤n≤N

(29)

Then it holds that ← q ode max ∥u← b←,h b← sn − u sn ∥ ≤ Aηode ∥u0 − u 0 ∥ + Hηode ,δode + Cnum,τ,t1 h ,

0≤n≤N

(30)

sde max ∥ξs←n − ξbn←,h ∥ Lp ≤ Aηsde ∥ξ0← − ξb0← ∥Lp + Hηsde ,δsde + Cnum,p hϱ .

0≤n≤N

In both displays, the initial-discrepancy term vanishes under common terminal coupling. Proof. Apply Proposition D.1 at the grid points and take the maximum over n. For the ODE, add the deterministic numerical error by the triangle inequality. For the SDE, first use the pathwise stability bound, then take the Lp -norm and apply Minkowski’s inequality together with (29). 33

Preprint. Work in progress.

Corollary D.5. Assume that the hypotheses of Corollary D.4 hold under a common terminal coupling, and consider a joint limit in which S → ∞, h → 0, and aS → 0. Suppose the realised guidance error satisfies ετ,t1 = OP (aS ) on Rτ,t1 . For χ ∈ {ode, sde}, define Z s Bχ,τ,t1 = sup Θηχ (s, r)λχ β(κr ) dr. 0≤s≤Rτ,t1

0

Assume that the stability and numerical-error constants are uniformly controlled along this ode sde limit, in the sense that Bχ,τ,t1 = OP (1), Cnum,τ,t = OP (1), Cnum,p = OP (1). This is automatic 1 when these quantities are deterministic and bounded independently of S and h. Then q max ∥u← b←,h sn − u sn ∥ = OP (aS + h ),

max ∥ξs←n − ξbn←,h ∥ = OP (aS + hϱ ).

0≤n≤N

0≤n≤N

In particular, if a uniform Monte Carlo guidance bound gives aS = S −1/2 , then the corresponding rates are OP (S −1/2 + hq ) for the ODE and OP (S −1/2 + hϱ ) for the SDE. Proof. Under a common terminal coupling, the initial-discrepancy terms in Corollary D.4 vanish. Moreover, Hηχ ,δχ ≤ Bχ,τ,t1 ετ,t1 = OP (aS ). ode For the ODE, the numerical term satisfies Cnum,τ,t hq = OP (hq ), so the first conclusion follows 1 p sde from (30). For the SDE, the conditional L estimate in (29), together with Cnum,p = OP (1), implies that max ∥ξbs←n − ξbn←,h ∥ = OP (hϱ ). 0≤n≤N

due to Markov’s inequality. The second conclusion follows from the triangle inequality. D.5

Process-space propagation

The latent bounds pass directly to process space whenever the generator is Lipschitz on the region visited by the coupled latent trajectories. Corollary D.6. Assume the hypotheses of Corollary D.4. Suppose that Em is equipped with a norm and that T : Rmξ → Em is LT -Lipschitz on the latent region visited by the exact, approximately guided, and numerical trajectories. Then the ODE and SDE bounds in Corollary D.4 remain valid after applying T , with the right-hand sides multiplied by LT . In particular,  ← ode q max ∥T (u← u←,h b← , sn ) − T (b sn )∥ ≤ LT Aηode ∥u0 − u 0 ∥ + Hηode ,δode + Cnum,τ,t1 h 0≤n≤N

  sde ≤ LT Aηsde ∥ξ0← − ξb0← ∥Lp + Hηsde ,δsde + Cnum,p hϱ .

max ∥T (ξs←n ) − T (ξbn←,h )∥

0≤n≤N

Lp

Moreover, any Wasserstein bound above is inherited under pushforward, since Wp (T# µ, T# ν) ≤ LT Wp (µ, ν) whenever the relevant latent coupling is supported on the Lipschitz region. Proof. The pathwise and strong Lp statements follow from ∥T (x) − T (y)∥ ≤ LT ∥x − y∥. The Wasserstein statement follows by pushing forward the same latent coupling through T . D.6

Additional path-law control for the reverse SDE

The preceding bounds control coupled paths. For the reverse-time SDE, one also obtains a path-law bound by Girsanov’s theorem. b denote Proposition D.7. Assume that the usual Girsanov conditions hold, and let P and P the path laws of the exact and approximately guided reverse SDEs on C([0, Rτ,t1 ]; Rmξ ). Assume also that L(ξ0← ) ≪ L(ξb0← ). Then   1 Z Rτ,t1 ← ← b b KL(P∥P) = KL L(ξ0 ) L(ξ0 ) + EP β(κs )∥g(κs , ξs← ) − gb(κs , ξs← )∥2 ds. (31) 2 0 34

Preprint. Work in progress.

In particular, if the two SDEs start from the same initial law, then the first term vanishes. If the realised guidance error is uniformly bounded by ετ,t1 along the exact paths, then, for 0 ≤ s ≤ Rτ,t1 , Z s   b [0,s] ) ≤ KL L(ξ ← ) L(ξb← ) + 1 ε2 KL(P[0,s] ∥P β(κr ) dr. 0 0 2 τ,t1 0 Consequently, L(ξs← ) − L(ξbs← )

2

≤ TV

Z s   1 1 KL L(ξ0← ) L(ξb0← ) + ε2τ,t1 β(κr ) dr. 2 4 0

Proof. The two SDEs have a common diffusion coefficient. Moreover, we have from the definitions that ∥Σsde (s)−1 (b(s, u) − bb(s, u))∥2 = β(κs )∥g(κs , u) − gb(κs , u)∥2 . The identity in (31) follows immediately from Girsanov’s theorem. The KL bound follows from the uniform guidance-error assumption. The total-variation bound follows by data processing and Pinsker’s inequality. D.7

Derivative-free guidance

When ∇ξ log Λ(ξ) is unavailable, one can estimate the guidance field using a form of Fisher’s identity. Under the tilted bridge conditional density, we have  p(ξ0 | ξt = u, C) ∝ Λ(ξ0 )N ξ0 ; α(t)u, σ(t)2 Imξ . Using Fisher’s identity, it follows that ∇u log ρt (u) = −

1 (u − α(t)E[ξ0 | ξt = u, C]) . σ(t)2

Meanwhile, since ρt (u) ∝ ψ(t, u)r(u), we have ∇u log ρt (u) = g(t, u) − u. Combining this with the previous expression, we arrive at g(t, u) =

α(t) (E[ξ0 | ξt = u, C] − α(t)u) . σ(t)2

This immediately suggests a gradient-free estimator. For a given pair (t, u), draw bridge samples ξ (s) = α(t)u + σ(t)ϵ(s) , ϵ(s) ∼ N (0, Imξ ), s = 1, . . . , S. We can then define the self-normalised approximation ! S α(t) X df (s) gbS (t, u) = w̄s ξ − α(t)u , σ(t)2 s=1

Λ(ξ (s) ) w̄s = PS . (ℓ) ) ℓ=1 Λ(ξ

This estimator only requires likelihood evaluations, but can be less sample-efficient than the gradient-based estimator because it estimates a small conditional-mean difference. Moreover, the prefactor α(t)/σ(t)2 can amplify Monte Carlo noise at small noising times.

E

Additional details for guidance-aware parameter estimation

In this appendix, we fix the discretisation level m and suppress this from the notation. We thus write Zϑ for Zϑ,m , and J (ϑ) = log Zϑ . E.1

A static interpretation via the Gibbs–Donsker–Varadhan formula

The evidence for the condition has a useful interpretation via the Gibbs–Donsker–Varadhan formula (Donsker & Varadhan, 1975). 35

Preprint. Work in progress.

Proposition E.1. Assume that Λϑ,C ≥ 0 and that 0 < Zϑ < ∞. Then the log evidence admits the representation log Zϑ = sup {Eξ0 ∼q [log Λϑ,C (ξ0 )] − KL(q∥r)} . q≪r

The supremum is attained at ρϑ (ξ0 | C), whenever the displayed quantities are finite. Moreover, whenever q ≪ ρϑ (· | C),  Eξ0 ∼q [log Λϑ,C (ξ0 )] − KL(q∥r) = log Zϑ − KL q∥ρϑ (· | C) . (32) (ξ0 |C) 0 Proof. The proof is standard. By definition, ρϑr(ξ = ϑ,C . It follows that, whenever Zϑ 0) q ≪ ρϑ (· | C),   Z  q(ξ0 ) KL q∥ρϑ (· | C) = log q(ξ0 ) dξ0 ρϑ (ξ0 | C) = KL(q∥r) − Eξ0 ∼q [log Λϑ,C (ξ0 )] + log Zϑ . Λ

(ξ )

Rearranging this identity gives (32). The variational formula follows from nonnegativity of KL divergence, with equality at q = ρϑ (· | C). By evaluating the identity in Proposition E.1 at q = ρϑ (· | C), we obtain  log Zϑ = Eξ0 ∼ρϑ (·|C) [log Λϑ,C (ξ0 )] − KL ρϑ (· | C)∥r Combining this with the identity in Proposition B.2, we also have that  log Zϑ = Ef0 ∼πϑ (·|C) [log LC (f0 )] − KL πϑ (· | C)∥pϑ .

(33)

This shows that the evidence represents a balance between satisfaction of C and the relativeentropy cost of deforming pϑ into πϑ (· | C). Corollary E.2. Suppose that the additional condition C is the event that the discretised process lies in a feasible set A ⊆ Em , so that LC (f0 ) = 1A (f0 ). If pϑ (A) > 0, then Zϑ = pϑ (A) and πϑ (· | C) = pϑ (· | A). Thus, in particular,  − log Zϑ = KL pϑ (· | A)∥pϑ . Thus, maximising the evidence is equivalent to minimising the KL deformation from the prior process law to the constrained process law. Proof. Since LC = 1A , Zϑ = Moreover,

R A

pϑ (f0 ) df0 = pϑ (A), and the conditioned law is pϑ (· | A). dpϑ (· | A) 1A (f0 ) . (f0 ) = dpϑ pϑ (A)

It follows that KL pϑ (· | A)∥pϑ



 = Epϑ (·|A) log

 1 = − log pϑ (A) = − log Zϑ . pϑ (A)

Corollary E.3. Suppose that the additional condition is encoded by a likelihood of the form LC (f0 ) = a exp{−ΦC (f0 )}, where ΦC (f0 ) ≥ 0 and a > 0 is independent of ϑ. Then, up to additive constants independent of ϑ,  − log Zϑ ≡ KL πϑ (· | C)∥pϑ + Ef0 ∼πϑ (·|C) [ΦC (f0 )] . (34) Thus evidence maximisation balances a small deformation from the ordinary process law against good satisfaction of the soft condition. Proof. By (33), we have that  log Zϑ = Ef0 ∼πϑ (·|C) [log LC (f0 )] − KL πϑ (· | C)∥pϑ . Since log LC (f0 ) = log a − ΦC (f0 ), we obtain  − log Zϑ = − log a + Ef0 ∼πϑ (·|C) [ΦC (f0 )] + KL πϑ (· | C)∥pϑ . The term − log a is independent of ϑ, giving (34). 36

Preprint. Work in progress.

E.2

A dynamic interpretation via Doob transform and Girsanov’s theorem

The static free-energy identity shows that maximising Zϑ balances condition satisfaction against the KL cost of deforming the ordinary process law. We can also obtain a dynamic interpretation of this deformation cost, using Doob transforms and Girsanov’s theorem (Doob, 1957; Girsanov, 1960; Øksendal, 2003). The result below shows that, under a suitable finite-horizon reference diffusion, the same deformation cost is represented by the energy of the guidance drift. For each ϑ, recall that the latent likelihood is Λϑ,C , the normalising constant is Zϑ , the reference density is r(ξ) = N (ξ; 0, Imξ ), and the latent target is ρϑ (ξ | C) =

Λϑ,C (ξ)r(ξ) . Zϑ

We now give a path-space version of the same construction. Instead of tilting the static latent law r by Λϑ,C (ξ), we consider a stationary reference diffusion in latent space and tilt its path law by the likelihood of its terminal endpoint. This lets us interpret the KL deformation in terms of the energy of the Doob-transform guidance drift. We work on a finite generation horizon, rescaled to [0, 1]. Let β̄ ∈ L1 ([0, 1]) be nonnegative, and let Pϑ be the path law of the stationary reference latent Ornstein–Uhlenbeck process q 1 dUs = − β̄(s)Us ds + β̄(s) dBs , U0 ∼ r, 0 ≤ s ≤ 1. (35) 2 Under Pϑ , the one-time marginal of Us is r for every s ∈ [0, 1]. Thus Pϑ is an unconditioned stationary reference dynamics in latent space. Define the tilted path law Qϑ by dQϑ Λϑ,C (U1 ) = . dPϑ Zϑ Under Qϑ , paths are reweighted according to how well their terminal state satisfies the condition. The terminal marginal under Pϑ is r, while the terminal marginal under Qϑ is precisely ρϑ (· | C). The corresponding Doob potential is hϑ (s, u) = EPϑ [Λϑ,C (U1 ) | Us = u] .

(36)

This is the expected future condition likelihood from the current generation-time state u. Equivalently, hϑ is the same smoothed likelihood object as the noising-time potential ψϑ , written in generation time rather than noising time. Assumption E.4. For the fixed parameter value ϑ,  assume that 0 < Zϑ < ∞, that Λϑ,C > 0 r-almost everywhere, and that KL ρϑ (· | C)∥r < ∞. Assume that hϑ in (36) has a positive, sufficiently regular version for which ∇u log hϑ is well defined, Itô’s formula applies to hϑ (s, Us ), and the backward Kolmogorov drift term vanishes. Assume that the density process hϑ (s, Us ) , 0 ≤ s ≤ 1, Mϑ,s := Zϑ is a true Pϑ -martingale. Further, suppose that this process admits the stochastic exponential 0) representation as Mϑ,s = Mϑ,0 Eϑ,s , where Mϑ,0 = hϑ (0,U and Zϑ Z s q  Z s 1 ⊤ 2 Eϑ,s = exp β̄(r)∇u log hϑ (r, Ur ) dBr − β̄(r)∥∇u log hϑ (r, Ur )∥ dr . 2 0 0 Theorem E.5. Under Assumption E.4, the tilted law Qϑ is the Doob transform of Pϑ associated with hϑ . In particular, under Qϑ , the coordinate process Us solves   q 1 dUs = − β̄(s)Us + β̄(s)∇u log hϑ (s, Us ) ds + β̄(s) dBsQϑ , (37) 2 where B Qϑ is Brownian motion under Qϑ , with initial law U0 ∼ (Qϑ )0 . Moreover, Z 1   1 2 EQϑ β̄(s)∥∇u log hϑ (s, Us )∥ ds = KL ρϑ (· | C)∥r − KL((Qϑ )0 ∥(Pϑ )0 ). 2 0 37

(38)

Preprint. Work in progress.

Proof. By the Markov property and Assumption E.4, hϑ solves the backward Kolmogorov equation associated with (35), with terminal condition hϑ (1, u) = Λϑ,C (u). Using also Itô’s formula, we have that q dhϑ (s, Us ) = β̄(s)∇hϑ (s, Us )⊤ dBs . The density process of Qϑ with respect to Pϑ is   Λϑ,C (U1 ) hϑ (s, Us ) Mϑ,s = EPϑ | Fs = . Zϑ Zϑ It follows that q dMϑ,s = β̄(s)∇u log hϑ (s, Us )⊤ dBs . Mϑ,s By Girsanov’s theorem, Z sq BsQϑ = Bs − β̄(r)∇u log hϑ (r, Ur ) dr 0

is Brownian motion under Qϑ . Substituting this into (35) gives (37). For the energy identity, write Z 1q log Mϑ,1 = log Mϑ,0 + β̄(s)∇u log hϑ (s, Us )⊤ dBsQϑ 0

Z 1 1 β̄(s)∥∇u log hϑ (s, Us )∥2 ds. + 2 0 Taking expectation under Qϑ , the stochastic integral has mean zero, and hence KL(Qϑ ∥Pϑ ) = EQϑ [log Mϑ,1 ] Z 1 1 = KL((Qϑ )0 ∥(Pϑ )0 ) + EQϑ β̄(s)∥∇u log hϑ (s, Us )∥2 ds. (39) 2 0 On the other hand, the likelihood ratio dQϑ / dPϑ depends only on U1 . The left-hand side therefore simplifies to KL(Qϑ ∥Pϑ ) = KL((Qϑ )1 ∥(Pϑ )1 ). Under Pϑ , U1 ∼ r, and under Qϑ , U1 ∼ ρϑ (· | C). Substituting this into the previous display gives  KL(Qϑ ∥Pϑ ) = KL ρϑ (· | C)∥r . Combining this with (39) proves (38). Remark E.6. For a finite stationary OU horizon with finite integrated noise, U0 and U1 are generally correlated. Terminal R 1 reweighting can thus change the law of U0 , resulting in the boundary term in (38). If 0 β̄(s) ds → ∞, then the correlation between U0 and U1 tends to zero. Equivalently, the initial generation-time state becomes independent of the endpoint. In that limit, terminal reweighting does not alter the initial law, and (Qϑ )0 = (Pϑ )0 . We can now provide an interpretation of the evidence in terms of the guidance. Define the guidance energy by Z 1  1 2 β̄(s)∥∇u log hϑ (s, Us )∥ ds . Eguide (ϑ) = EQϑ 2 0 Suppose that we are in the limiting regime where (Qϑ )0 = (Pϑ )0 . Then, by Theorem E.5, the guidance energy coincides with the relative-entropy cost  Eguide (ϑ) = KL ρϑ (· | C)∥r . It follows, combining this with the result in Proposition E.1, that the log evidence can be written as log Zϑ = Eξ0 ∼ρϑ (·|C) [log Λϑ,C (ξ0 )] − Eguide (ϑ). Thus evidence maximisation can be interpreted dynamically as balancing condition satisfaction against the guidance energy required to deform the reference latent process. For hard constraints, interpreted as limits of strictly positive soft constraints, the log-likelihood contribution vanishes on the constrained support, so evidence maximisation reduces, in the boundary-free regime, to minimising guidance energy. For soft constraints, the objective also includes the expected residual loss in Corollary E.3. 38

Preprint. Work in progress.

E.3

Basic theoretical guarantees for parameter estimation

We now record some basic theoretical results relating to our parameter estimation scheme. We begin by establishing that the exact fixed-parameter time slices give unbiased evidence gradients. Proposition E.7. Assume that 0 < Zϑ < ∞, that the reference density r and the noising schedule do not depend on ϑ, and that differentiation may be interchanged with integration. Then, for every t ∈ [0, 1) such that ψϑ (t, ·) is positive and differentiable in ϑ, ∇ϑ J (ϑ) = ∇ϑ log Zϑ = Eu∼ρϑ,t (·|C) [∇ϑ log ψϑ (t, u)] . Thus, if a fixed-parameter LatentFlow solve at ϑ produces particles with selected time-slice (i) marginals uj ∼ ρϑ,tj (· | C), with 0 ≤ tj < 1, then an unbiased estimator of ∇ϑ J (ϑ) is M

b G(ϑ) =

J

1 XX (i) ∇ϑ log ψϑ (tj , uj ). M J i=1 j=1

R Proof. Recall that Zϑ = Rmξ ψϑ (t, u)r(u) du. Taking derivatives, and simplifying, we have Z 1 ∇ϑ log Zϑ = ∇ϑ ψϑ (t, u)r(u) du Zϑ Rmξ Z ψϑ (t, u)r(u) = ∇ϑ log ψϑ (t, u) du = Eu∼ρϑ,t (·|C) [∇ϑ log ψϑ (t, u)] . mξ Zϑ R

We now record an elementary convergence result for our update scheme. In practice, we cannot compute hϑ (t, u) := ∇ϑ log ψϑ (t, u) directly, and it must be replaced by a self-normalised ratio estimator b hϑ,S (t, u) := ∇ϑ log ψbϑ,S (t, u); see Section 3.3 for a related discussion. This estimator is generally biased for finite S, since it is a ratio of Monte Carlo averages. It is, however, consistent as S → ∞, under the usual integrability assumptions and provided that the limiting denominator is positive. In any case, the following result only requires a conditional bias bound. Theorem E.8. Let Θ ⊆ Rq be convex, and suppose that J : Θ → R is continuously differentiable, bounded above by J ⋆ , and LJ -smooth along the iterates. Consider the updates 1 bk , ϑk+1 = ϑk + ηk G 0 < ηk < , 2LJ and assume that the iterates remain in Θ. Let Fk be the filtration generated by the algorithm b k . Suppose, in addition, that G b k admits the decomposition before drawing G b k = ∇ϑ J (ϑk ) + bk + ξk , G where bk is Fk -measurable, and satisfies ∥bk ∥ ≤ βk for some Fk -measurable βk ≥ 0, and where ξk satisfies E[ξk | Fk ] = 0, E[∥ξk ∥2 | Fk ] ≤ σ 2 . Then, for every K ≥ 1, P LJ σ 2 PK−1 2 2   J ⋆ − J (ϑ0 ) + K−1 k=0 ηk k=0 ηk E[βk ] + 2 2 min E ∥∇ϑ J (ϑk )∥ ≤ . (40) PK−1 ηk 0≤k<K k=0 2 (1 − 2LJ ηk ) In particular, if ηk ≡ η < 1/(2LJ ) and βk ≤ β for all k, the limiting stationarity floor is of order  O β 2 + LJ σ 2 η . If, in addition, E[βk2 ] → 0, then the bias contribution to this floor vanishes; the remaining stochastic contribution can be made small by decreasing η, and disappears when σ = 0. Proof. This follows from the standard smooth nonconvex analysis of stochastic gradient methods with biased gradients; see, for instance, Ajalloeian & Stich (2020). Applying the same descent-lemma argument to f = −J , and using Young’s inequality to absorb the bias term bk , gives (40). 39

Preprint. Work in progress.

F

Additional Related Work

Conditioning stochastic-process priors. Perhaps the most well-studied case of stochasticprocess conditioning is inference for Gaussian processes (GPs). Under finitely many linearGaussian observations, the posterior over any finite collection of function values remains Gaussian and can be computed analytically (Rasmussen & Williams, 2006). This conjugacy is generally lost under non-Gaussian likelihoods, motivating a large literature on approximate inference (Kuss & Rasmussen, 2005; Nickisch & Rasmussen, 2008; Hensman et al., 2015; Nickisch et al., 2018). Related work encodes known linear operator or differential-equation constraints directly in the GP prior or regression construction (Jidling et al., 2017; LangeHegermann, 2018; Gulian et al., 2022). By contrast, shape information such as monotonicity, boundedness, positivity, and convexity is commonly incorporated through virtual inequality observations, finite-dimensional coefficient inequalities, truncation, or posterior projection, generally resulting in non-Gaussian or projected posterior laws (Riihimäki & Vehtari, 2010; Da Veiga & Marrel, 2012; Lin & Dunson, 2014; Maatouk & Bay, 2017; López-Lopera et al., 2018; Agrell, 2019; Swiler et al., 2020; Maatouk et al., 2025; Astfalck et al., 2018). An alternative generalised Bayes linear approach enforces domain restrictions through the choice of inferential solution space (Astfalck et al., 2024). The closest GP-specific antecedent to the current work is FlowGP (Moss et al., 2026), which conditions a finite-dimensional whitened GP representation on nonlinear or non-Gaussian information by solving a guided probability-flow ODE. Our framework contains FlowGP as a special case but applies to a substantially broader class of stochastic-process priors; see Appendix A. For SDE priors, LatentFlow is closely connected to diffusion-bridge simulation and conditioned diffusion processes. Classical bridge methods condition a diffusion to hit specified endpoints or endpoint observations, often by adding an approximate guiding term to its drift. Delyon & Hu (2006) construct readily simulated endpoint-guided diffusions and derive an explicit likelihood-ratio correction. Schauer et al. (2017) derive guiding terms from auxiliary tractable diffusions with known transition densities, while Bierkens et al. (2020) extend this guided-proposal framework to partial endpoint observations and elliptic or hypoelliptic settings. Related bridge samplers use coupling or other proposal mechanisms (Bladt & Sørensen, 2014; Bladt et al., 2016). More recent methods learn time reversals using score matching (Heng et al., 2025), estimate bridge scores without first learning the time reversal (Baker et al., 2025), develop infinite-dimensional conditioning constructions (Baker et al., 2024), introduce neural guided bridges (Yang et al., 2025b), or exploit Malliavin-calculus identities (Pidstrigach et al., 2025). Rare-event simulation and transition-path sampling address closely related path ensembles. Transition-path sampling uses path-space Monte Carlo to study rare reactive trajectories (Dellago et al., 1998; 2002), with recent diffusion-based methods also targeting transition-path ensembles (Seong et al., 2025), whereas genealogical particle methods and adaptive multilevel splitting estimate rare-event probabilities by repeatedly selecting and resampling trajectories according to their progress toward a rare set (Del Moral & Garnier, 2005; Cérou & Guyader, 2007; Cérou et al., 2012). LatentFlow targets likelihood-tilted path laws by tilting the law of the underlying innovations and transporting it toward regions of high condition likelihood. It therefore provides a likelihoodguided complement to output-space bridge proposals and splitting methods when suitable pulled-back gradients or guidance estimators are available. Latent and reference-space formulations. Many stochastic-process laws admit representations as pushforwards of a tractable reference distribution rather than through an explicit density on the resulting process space. This perspective is implicit in simulator and filtering representations of state-space models (Doucet et al., 2001; Särkkä, 2013); in GMRF and SPDE constructions of spatial fields (Rue & Held, 2005; Lindgren et al., 2011); and in process priors and finite-dimensional representations built by composing or transforming latent random variables or functions, including random-feature kernel representations, latent-force models, GP state-space models, and deep GPs (Rahimi & Recht, 2007; Álvarez et al., 2009; 2013; Frigola et al., 2013; Damianou & Lawrence, 2013; Frigola et al., 2014). Pathwise GP constructions based on Matheron’s rule similarly represent posterior sample paths as transformations of prior samples and provide scalable samples for downstream Monte Carlo computation (Wilson et al., 2020; 2021). Innovation and non-centred parameterisations make 40

Preprint. Work in progress.

this explicit for diffusion processes by parameterising latent paths through transformed path variables or driving-noise variables together with model parameters (Roberts & Stramer, 2001; Golightly & Wilkinson, 2008; Papaspiliopoulos et al., 2013; Whitaker et al., 2017; Graham et al., 2022). Neural-process and implicit-process families likewise define learned distributions over functions through neural maps (Garnelo et al., 2018; Kim et al., 2019; Ma et al., 2019; Gordon et al., 2020). Similarly, normalising flows push forward simple latent laws, although they usually learn invertible transformations so that output densities remain available by change of variables (Rezende & Mohamed, 2015; Papamakarios et al., 2021). The same reference-law viewpoint is central to Bayesian inverse problems. A posterior on a possibly infinite-dimensional parameter or function space is commonly defined by tilting a tractable reference measure by a likelihood (Stuart, 2010; Dashti & Stuart, 2017). Function-space MCMC methods, including pCN and Hilbert-space HMC, are designed directly for such targets and can remain well behaved under mesh refinement (Beskos et al., 2011; Cotter et al., 2013; Hairer et al., 2014). SMC, annealed importance sampling, and flow-augmented transport samplers construct sequences of intermediate distributions between reference and posterior laws (Neal, 2001; Del Moral et al., 2006; Arbel et al., 2021), while transport-map methods seek explicit reference-to-posterior maps, including in function-space inverse problems (El Moselhy & Marzouk, 2012; Hosseini et al., 2025). Recent score-based methods have likewise emphasised defining and approximating posterior sampling dynamics directly on function spaces (Baldassari et al., 2023; 2024; 2025). Bayesian inference in generator noise space. A related literature performs posterior inference directly in the input-noise coordinates of a fixed generative model. For example, Holden et al. (2022) use parallel-tempered pCN in the latent space of a VAE, while Patel et al. (2022) use HMC together with a GAN. Another early contribution is Böhm et al. (2019). Normalising-flow priors have also been combined with variationally trained conditional flows for amortised inverse-problem inference (Whang et al., 2021). Related early work performs constrained-HMC inference in the input-noise space of a differentiable generator (Graham & Storkey, 2017), while distributional-control methods learn a transformation of the generator’s reference law toward specified output-space constraints (Wu et al., 2022). There is also a substantial literature on combining transport maps with Monte Carlo methods. Transport-map methods construct invertible transformations between tractable reference laws and target laws, either as approximate posterior representations or as preconditioners and proposals for Monte Carlo sampling (El Moselhy & Marzouk, 2012; Parno & Marzouk, 2018; Hoffman et al., 2019; Cabezas & Nemeth, 2023; Cabezas et al., 2024). When an approximate map is embedded within a valid MCMC kernel with the appropriate correction, map approximation error affects sampling efficiency rather than the invariant target (Parno & Marzouk, 2018; Hoffman et al., 2019). NeuTra HMC uses a variationally trained normalising flow followed by HMC in the transported coordinates (Hoffman et al., 2019), whereas transport elliptical slice sampling uses a related nonlinear preconditioning strategy for elliptical slice sampling (Cabezas & Nemeth, 2023). A closely related latent construction is the neural-transport MCMC method of Nijkamp et al. (2022), in which an energy correction exponentially tilts the output law of a learned flow and sampling is performed in its latent coordinates. That work is directed toward learning and sampling energy-based generative models, rather than conditioning a stochastic-process simulator. More recent methods apply the same principle to diffusion and flow-based generators. Outsourced diffusion sampling considers a deterministic representation x = fϑ (z), with z drawn from a Gaussian reference, and trains an amortised diffusion sampler for the latent posterior induced by an output-space likelihood or reward (Venkatraman et al., 2025). Noise-space HMC instead applies HMC directly to the initial noise of a deterministic reversediffusion map (Xia et al., 2026). Similar constructions have appeared for flow-matching models: Source-Guided Flow Matching modifies the source distribution while retaining the learned vector field, whereas D-Flow SGLD samples the source posterior induced by a new measurement operator (Wang et al., 2026; Parikh et al., 2026). Related inverse-problem methods perform SMC in the latent coordinates of diffusions, guide sampling trajectories under pre-trained latent flow priors, or combine latent diffusion posterior sampling with surrogate-likelihood guidance (Achituve et al., 2025; Askari et al., 2025; Wang & Tartakovsky, 2026). 41

Preprint. Work in progress.

These approaches share the central principle of conditioning a generator by changing the distribution of its input noise. LatentFlow differs primarily in the scope of the generator and in how the latent posterior is sampled. In our case, Tϑ may describe an explicit stochastic-process prior, a numerical simulator, or a pre-trained neural generator, and need not be invertible or density-evaluable. Moreover, LatentFlow does not train an auxiliary conditional generator or amortised latent sampler. Instead, its time-dependent guidance is determined directly by the pulled-back likelihood and a known Gaussian noising bridge. Diffusion guidance and finite-time controlled transport. The sampling construction underlying LatentFlow is closely related to diffusion and flow-based generative modelling. Diffusion probabilistic models define a forward noising process from a data distribution to a simple reference law together with a learned reverse process (Sohl-Dickstein et al., 2015; Ho et al., 2020); continuous-time score-based models generate samples by reversing a noising SDE or solving the associated probability-flow ODE (Song et al., 2021). Flow matching, rectified flows, and stochastic interpolants provide related continuous-time transport formulations (Lipman et al., 2023; Liu et al., 2023b; Albergo & Vanden-Eijnden, 2023; Albergo et al., 2025). Recent guided flow-matching methods extend these constructions to conditional generation and energy- or reward-guided sampling (Feng et al., 2025; Mark et al., 2025). A large literature conditions pre-trained diffusion models on labels, measurements, inverse problems, or reward functions by modifying their reverse dynamics through classifier, classifierfree, measurement, reward, or off-the-shelf guidance terms (Dhariwal & Nichol, 2021; Ho & Salimans, 2022; Chung et al., 2023; Bansal et al., 2024). To control bias introduced by approximate guidance, other conditional diffusion samplers use SMC, particle MCMC, Feynman–Kac representations, twisting, or related correction mechanisms (Wu et al., 2023; Cardoso et al., 2024; Corenflos et al., 2025; Singhal et al., 2025; Skreta et al., 2025; Wu et al., 2025). Under their stated assumptions, some of these methods are asymptotically exact as the particle number increases, while others preserve a model-defined conditional target through a valid Monte Carlo correction; see Zhao et al. (2025) for a recent review. Diffusion and flow models have also been developed directly for function-valued data and function-space posterior sampling (Dutordoir et al., 2023; Franzese et al., 2023; Kerrigan et al., 2024; Lim et al., 2025; Yao et al., 2025). These constructions admit a broader interpretation through finite-time stochastic control. Doob h-transforms provide a classical mechanism for conditioning Markov processes; for diffusions, the transformed dynamics acquire a drift correction involving the log-gradient of an appropriate space–time h-function (Doob, 1957; Didi et al., 2023). Related entropyminimisation and Schrödinger-bridge formulations characterise controlled path measures as minimum-relative-entropy perturbations of a reference process subject to marginal or endpoint constraints (Föllmer, 1985; Léonard, 2014; Chen et al., 2016). Modern Schrödinger-bridge, Schrödinger–Föllmer, and controlled-diffusion samplers exploit this finite-time transport perspective for sampling unnormalised targets (Bernton et al., 2019; Huang et al., 2021; Zhang & Chen, 2022; Vargas et al., 2023b;a; 2024; Tamogashev & Malkin, 2026), while diffusion Schrödinger-bridge methods construct transports between prescribed endpoint laws (De Bortoli et al., 2021; Peluchetti, 2023; Liu et al., 2023a; 2024). Several recent conditional-diffusion methods make the connection to Doob transforms explicit (Didi et al., 2023; Denker et al., 2024; Guo et al., 2026). In infinite dimensions, Park et al. (2024) develop a function-space stochastic-control formulation of diffusion bridges; Yang et al. (2025a) learn discretisation-equivariant bridge scores using operator learning; and Pieper-Sethmacher et al. (2025) construct guided bridge measures for semilinear SPDEs conditioned on linear terminal observations. Baker et al. (2026) develop a complementary construction for conditioning trained function-space diffusion priors by supervised guidance. LatentFlow instantiates the finite-time guidance principle in latent innovation space. Rather than learning a noising or transport process from samples, it uses a known referencespace bridge and the stochastic-process law supplied by Tϑ . Its guiding field is the gradient of log ht , where ht is the conditional expectation of the pulled-back likelihood under the latent noising bridge, equivalently the log-gradient of the corresponding smoothed pulled-back likelihood. Thus, LatentFlow combines the pullback formulation used in generator-noise inference with a diffusion-style finite-time transport, yielding a reference-space sampler for a broad class of stochastic-process priors. 42

Preprint. Work in progress.

G

Additional Numerical Experiments

G.1

Sampling Benchmarks

We now evaluate LatentFlow on three canonical benchmark distributions that are challenging for standard inference methods. Each target takes the form p∗ (x) ∝ pprior (x) · p(y | x), with prior pprior (x) = N (x; 0, I) and likelihoods p(y | x) described below. For LatentFlow, we use an Euler–Maruyama discretisation with 100 time steps, and 10 Monte Carlo samples for the guidance approximation. We benchmark against the No-UTurn Sampler (NUTS; Hoffman & Gelman, 2014), mean-field VI (Blei et al., 2017), and a Laplace approximation. Sample quality is measured using the Maximum Mean Discrepancy (MMD2 ) with a multi-scale RBF kernel, with uncertainty reported as standard errors over 5 independent runs, alongside the effective sample size per second (ESS/s). We generate N = 1,000 samples using each method. NUTS thus requires 1, 000 steps to produce its 1, 000 samples, as well as an additional 500 warmup steps to ensure the chains are well mixed. Across all benchmarks, LatentFlow achieves MMD2 scores competitive with or better than NUTS, while producing independent samples. LatentFlow also succeeds on problems where standard MCMC collapses to a single mode. Banana Distribution. The banana (twisted Gaussian) distribution has a curved, parabolic valley that resists random-walk exploration:  x1 ∼ N (0, σ12 ), x2 | x1 ∼ N b(x21 − σ12 ), σ22 , with b = 0.7, σ1 = 1.0, and σ2 = 0.4. Figure 7 shows samples from each method overlaid on the true density. LatentFlow closely follows the curved valley, while VI and Laplace fail to capture the non-linear geometry. NUTS tracks the true distribution but produces far fewer independent samples, reflected in the low ESS of 67; see Table 2. LatentFlow achieves the best MMD2 , and is substantially faster than NUTS, achieving over 100 times greater ESS/s. Direct

8

LatentFlow

NUTS

VI

Laplace

6

x2

4 2 0

−2

−4 −2

0

x1

2

4

−4 −2

0

2

x1

4

−4 −2

0

x1

2

4

−4 −2

0

x1

2

4

−4 −2

0

2

4

x1

Figure 7: Banana distribution. Background: true density. Scatter: 1,000 samples from each method. Direct samples are exact draws. LatentFlow matches the true density well. VI and Laplace fail to capture the curved geometry. Table 2: Banana distribution. MMD2 ± stderr (lower is better), wall-clock time, and ESS/s, all over 5 independent runs (3 d.p.). Method

MMD2

Time (s)

ESS

ESS/s

LatentFlow NUTS VI Laplace

−0.001 ± 0.000 0.004 ± 0.003 0.059 0.044

0.060 ± 0.000 0.470 ± 0.160 0.310 —

1000 67 1000 —

18010.0 141.7 3270.7 —

Ring Distribution. The ring (shell) distribution is concentrated on a thin annulus of radius R, with density given by   (∥x∥2 − R2 )2 ∗ p (x) ∝ exp − , 2σ 2 with R = 2.0 and σ = 0.3. Since gradients point predominantly radially, MCMC chains circulate slowly around the ring and mix poorly. As shown in Figure 8, LatentFlow traces the full annulus accurately and achieves over 20× the ESS/s of NUTS. 43

Preprint. Work in progress.

Direct

LatentFlow

NUTS

VI

Laplace

x2

2 0

−2 −2

0

−2

2

x1

0

−2

2

x1

0

−2

2

x1

0

−2

2

x1

0

2

x1

Figure 8: Ring distribution. Background: true density. Scatter: 1,000 samples from each method. LatentFlow traces the annulus accurately. NUTS also covers the ring but requires more warm-up steps. VI collapses to a central blob. Laplace diverges catastrophically.

Table 3: Ring distribution. MMD2 ± stderr (lower is better), wall-clock time, ESS/s, and radial statistics, all over 5 independent runs (3 d.p.). Ground truth: E[r] = 2, Std[r] ≈ 0.076. Method

MMD2

Time (s)

ESS/s

E[r]

Std[r]

LatentFlow NUTS VI Laplace

0.000 ± 0.000 0.001 ± 0.001 0.453 0.458

0.050 ± 0.000 0.340 ± 0.060 0.290 —

21263.4 897.1 3412.0 —

2.004 1.999 1.960 274.105

0.086 0.075 0.087 209.312

Mixture of Gaussians. We test multi-modal coverage using a mixture of K = 4 equalweight Gaussians centred at (±2, ±2) with σ = 0.5. LatentFlow places samples in all four quadrants, while NUTS concentrates in quadrant 4; see Figure 9. While NUTS achieves a raw ESS of 943 and an ESS/s of 2,191.4 (Table 4), its samples originate from a single mode. Its apparent efficiency is thus misleading: it fails to cover multiple modes, reflected in its MMD2 of 0.401 ± 0.009. LatentFlow achieves the best sample quality with MMD2 = 0.00079 ± 0.00027, recovering the correct global mean of approximately (0, 0) and standard deviation near 2.04 in both dimensions, compared to NUTS’s collapsed mean of (2.016, 1.984) with standard deviation ≈ 0.490. Direct

LatentFlow

NUTS

VI

Laplace

4

x2

2 0

−2 −4

−4 −2

0

x1

2

4

−4 −2

0

2

4

x1

−4 −2

0

2

4

x1

−4 −2

0

x1

2

4

−4 −2

0

2

4

x1

Figure 9: Mixture of Gaussians. Background: true density (four modes). Scatter: 1,000 samples from each method. LatentFlow covers all four modes. NUTS collapses entirely to quadrant 4 (100% of the chain). VI collapses to a single Gaussian centred near (2, 2).

Table 4: Mixture of Gaussians. MMD2 ± stderr (lower is better), wall-clock time, and ESS/s, all over 5 independent runs (3 d.p.). Despite its high ESS/s, NUTS fails to achieve multi-modal coverage (all samples are from one mode). Method

MMD2

Time (s)

ESS

ESS/s

LatentFlow NUTS VI

0.001 ± 0.000 0.401 ± 0.009 0.413

0.050 ± 0.000 0.430 ± 0.130 0.140

1000 943 1000

19865.9 2191.4 7119.4

44

Preprint. Work in progress.

G.2

Spatial Processes

Gaussian process

Samples

Smooth, light tails

3

Prior

0

3 2

Sparse observations

0 2 2

Areal max

0

3 2

Areal mean

0 2 5

Censored observations

1 3

Figure 10: Conditional sampling from a Gaussian process prior. The base process is a smooth, light-tailed Gaussian field (Matérn-3/2 kernel, lengthscale 0.15). The posteriors interpolate smoothly between observations and honour each areal summary while preserving the short-lengthscale texture of the prior.

Student-t process

Samples

Heavy tails, spatially coupled

3

Prior

0 3 2

Sparse observations

0 1 2

Areal max

1

1 1

Areal mean

0 1 4

Censored observations

1

3

Figure 11: Conditional sampling from a Student-t process prior. The base process is a heavy-tailed, spatially coupled Student-t field (νt = 4, Matérn-3/2 kernel, lengthscale 0.5). 45

Preprint. Work in progress.

Cauchy convolution

Samples

Heavy tails

3

Prior

0

3 2

Sparse observations

0 2 2

Areal max

0 2 2

Areal mean

0 2 3

Censored observations

0 2

Figure 12: Conditional sampling from a Cauchy convolution prior. The base process is a very heavy-tailed field formed by convolving a Cauchy noise source with a short-lengthscale kernel (lengthscale 0.05). The infinite-variance source produces sharp, highly localised features; the posteriors reconcile these with the imposed point, areal, and censoring constraints.

Potts model

Samples

Discrete labels, ordered domains

Prior

0

Sparse observations

0

Areal max

0

Areal mean

0

Censored observations

0

Figure 13: Conditional sampling from a Potts model prior. The base process is a discrete-label field with ordered domains (Potts model, q = 3 states mapped to {−1, 0, 1}, inverse temperature β = 1.2). Conditioning reshapes the domain structure to match the observations while retaining the piecewise-constant, label-coherent character of the prior.

46

Preprint. Work in progress.

Temporal Processes

0.0

0.2

0.4

0.6

time

0.8

1.0

1.2 1.0 0.8 0.6 0.4 0.2 0.0 −0.2

γ = 1.7, μ = 1, σ = 1

γ = 0.6, μ = 1, σ = 0.3

γ = 1.7, μ = 0, σ = 0.3

2

0.2

1

0.0

state

prior posterior analytic

state

state

γ = 1.7, μ = 1, σ = 0.3

1.2 1.0 0.8 0.6 0.4 0.2 0.0

state

G.3

0

−0.2

−1 0.0

0.2

0.4

0.6

time

0.8

1.0

−0.4 0.0

0.2

0.4

0.6

0.8

time

1.0

0.0

0.2

0.4

0.6

time

0.8

1.0

Figure 14: Conditional sampling from an OU process prior. The prior is an OU diffusion process with mean-reversion rate γ, long-run mean µ, and diffusion coefficient σ. Each panel shows prior trajectories, posterior bridge samples, and the analytic bridge mean for a different choice of (γ, µ, σ). The red marker denotes the terminal constraint, blue curves denote LatentFlow posterior samples, grey curves denote prior samples, and the green dashed curve denotes the analytic bridge mean. Prior

state

Normal event

x1 x2

2.0

Rare event

1.5

1.5

1.0

1.0

0.5

0.5 0.0

0.0 0

1

2

3

time

4

Multimodal event

1.2 1.0 0.8 0.6 0.4 0.2 0.0

2.0

0

1

2

3

time

4

1.5 1.0 0.5 0.0

−0.5 0

1

2

time

3

4

−1.0

0

1

2

time

3

4

5

Figure 15: Conditional sampling from a cell-differentiation diffusion process prior. The panels show samples from the prior and from three endpoint-conditioned laws: a typical terminal event, a rare terminal event, and a multimodal terminal event. Blue and orange curves denote the two state variables x1 and x2 , black markers denote the initial state, and red markers denote the terminal constraint. The rare and multimodal examples illustrate that the same latent bridge construction can condition nonlinear diffusion priors on very low-probability events. In particular, for the rare event and the multimodal event no valid samples were found after simulating the unconditioned process 100,000 times. prior

reference

posterior

reference

0.5

state

low endpoint

posterior

x1 x2

1.0

0.0

−0.5 −1.0

prior

0.5

state

high endpoint

1.0 0.0

−0.5 −1.0 0.0

0.5

1.0

time

1.5

2.0

0.0

0.5

1.0

time

1.5

2.0

0.0

0.5

1.0

time

1.5

2.0

Figure 16: Conditional sampling from a FitzHugh–Nagumo diffusion process prior. Rows: two endpoint constraints on the first state variable: a low endpoint and a high endpoint. Columns: prior trajectories, LatentFlow posterior bridge samples, and reference bridge trajectories. Blue and orange curves denote the two state variables x1 and x2 , black markers denote the initial state, and red markers and dashed lines denote the terminal constraint. The high-endpoint case illustrates conditioning on a rare excursion of the excitable system. 47

Preprint. Work in progress.

G.4

Spatio-temporal Processes

t = 0.3

t = 0.7

t = 1.1

time

t = 1.6

Prior

t = 2.0

t = 2.2

0.6

no conditioning Sparse sensors

0.4

concentration

8 point obs. × 4 times Obstacle avoidance

plume bypasses

0.2

central zone Detector hit

alarm triggered downstream

0.0

Figure 17: Posterior samples from the advection–diffusion plume model under three conditioning regimes, each showing four independent draws at six time snapshots. Prior: unconditioned samples exhibiting diverse plume trajectories and intensities. Sparse sensors: conditioning on exact concentration values at 8 spatial locations observed at 4 times (32 observations total) recovers the ground-truth plume trajectory. Obstacle avoidance: a binary condition requiring the plume to stay below concentration 0.06 in the central rectangular region (dashed box, ×) at all times steers samples around the obstacle. Detector hit: a binary condition requiring peak concentration to exceed 0.40 in the downstream detector region (dashed box, ✓) concentrates mass on trajectories that reach the top-right corner.

48

Preprint. Work in progress.

t = 0.5

t = 1.2

t = 1.9

time

t = 2.7

Prior

t = 3.4

t = 4.0

1.50

two pacemakers, variable activity

0.75

Source 1 only

0.00

Source 2 only

activator u

left fires early, upper-right silent

right fires late, lower-left silent

−0.75

Both fire

sequential activation: left then right

−1.50

Figure 18: Posterior samples from the FitzHugh–Nagumo excitable medium under three conditioning regimes, each showing four independent draws at six time snapshots. Prior: unconditioned samples show variable activation of two spatially separated pacemakers (green boxes: active; red boxes: silent). Source 1 only: conditioning on the lower-left pacemaker firing early (t ≲ 1.3) while the upper-right remains silent produces a single travelling wavefront originating from the lower-left corner. Source 2 only: the complementary condition suppresses the lower-left source and requires the upper-right to fire late (t ≳ 1.8), yielding a wavefront from the upper-right. Both fire: conditioning on sequential activation of both sources produces two distinct wavefronts propagating across the medium in succession.

49

Preprint. Work in progress.

time

Prior

t = 0.00

t = 0.21

t = 0.41

t = 0.62

t = 0.83

t = 1.00

1.0

no conditioning

Sparse sensors

0.5

36 point obs.

0.0

prescribed mean per quadrant

field value

× 6 times

Quadrant means Annular ring

pos. annulus, neg. core & exterior

0.5

Phase front

neg. left, zero centre, pos. right

1.0

Figure 19: Posterior samples from the Allen–Cahn phase-field model under four conditioning regimes, each showing four independent draws at six time snapshots. Except for the sparse-sensor experiment, all conditions constrain statistics of the final field u(·, t29 ). Prior: unconditioned samples exhibit stochastic domain coarsening toward the ±1 phases. Sparse sensors: conditioning on 36 point observations at each of 6 snapshot times (216 observations total) recovers the ground-truth field trajectory. Quadrant means: prescribing the spatial mean in each quadrant to a checkerboard pattern (+0.10, −0.10, −0.10, +0.10) steers the final phase arrangement into the corresponding diagonal structure. Annular ring: prescribing a positive mean in an annular band (0.30 ≤ r ≤ 0.65) and negative means in the core and exterior produces a ring-shaped positive domain surrounded by negative phase. Phase front: prescribing a left-to-right mean gradient (−0.20, 0.0, +0.20) across three vertical strips (|x| ≤ 0.30 boundaries) induces a diffuse phase interface running vertically through the domain.

50

Record · ID 366245 · SHA-256 1427a4dd052bbf92
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.