Provable diffusion-based posterior sampling for linear inverse problems via DDIM Yuchen Jiao∗†
Na Li∗†
Changxiao Cai‡
Yuxin Chen§
Gen Li†
July 22, 2026
arXiv:2607.19333v1 [cs.LG] 21 Jul 2026
Abstract Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers either lack rigorous theoretical guarantees or incur substantial computational overhead. We propose a simple and efficient algorithm, called Posterior-DDIM, for solving linear inverse problems with diffusion priors via a DDIM-type sampler. Our method requires only lightweight, coordinate-wise modifications to the standard DDIM update, while explicitly incorporating the measurement model. The key idea is to perform posterior sampling separately along each singular direction of the measurement operator: for each direction, the sampler follows the learned diffusion prior when the observation signal-to-noise ratio (SNR) is below the corresponding diffusion SNR, and switches to a calibrated measurement-based predictor otherwise. We prove that the proposed sampler converges to the Bayesian posterior conditioned on the measurements. Empirical results show that the proposed sampler performs favorably against existing diffusion-based posterior samplers across a range of image restoration tasks, achieving the best performance on the majority of evaluation metrics considered. Overall, our results convert posterior sampling for noisy linear inverse problems to simple coordinate-wise DDIM updates, yielding an efficient, easy-to-implement algorithm with provable posterior consistency.
Keywords: diffusion models; linear inverse problems; DDIM; posterior sampling
Contents 1 Introduction 1.1 Motivation: inverse problems with diffusion priors . . . . . . . . . . . . . . . . . . . . . . . . 1.2 An overview of our contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.3 Related work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
2 2 3 3
2 Preliminaries
4
3 Main results 3.1 Motivation: SNR-guided partition of singular directions . . . . . . . . . . . . . . . . . . . . . 3.2 Algorithm: Posterior-DDIM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3 Theoretical guarantees: asymptotic posterior correctness . . . . . . . . . . . . . . . . . . . . .
6 6 7 9
4 Experiments 10 4.1 Choices of η0 and η1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 4.2 Empirical comparisons with reference algorithms . . . . . . . . . . . . . . . . . . . . . . . . . 10 5 Discussion
12
∗ The authors contributed equally. Corresponding author: Gen Li. † Department of Statistics and Data Science, Chinese University of Hong Kong, Hong Kong ‡ Department of Industrial and Operations Engineering, University of Michigan, Ann Arbor, USA § Department of Statistics and Data Science, the Wharton School, University of Pennsylvania, PA, USA
1
A Proof of Theorem 1
13
B Proof of auxiliary lemmas and facts 18 B.1 Proof of (23) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 B.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 B.3 Proof of Lemma 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 B.4 Proof of Lemma 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 B.5 Proof of Lemma 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 B.6 Proof of Lemma 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 B.7 Proof of Lemma 6 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 C Extension to general A 34 C.1 Proof of Theorem 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
1
Introduction
Diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020; Song et al., 2020; Lai et al., 2025) have emerged as one of the most successful paradigms for modern generative modeling, achieving remarkable performance across a wide range of domains, from image synthesis, video generation, to scientific data generation. Beyond sample generation, their ability to capture complex, high-dimensional data distributions makes them particularly well-suited as flexible, expressive, and fully data-driven priors over the signals of interest, offering a versatile framework for Bayesian inference in inverse problems (Song et al., 2022; Chung et al., 2022, 2023a; Kawar et al., 2022; Song et al., 2023a; Rout et al., 2023; Daras et al., 2024).
1.1
Motivation: inverse problems with diffusion priors
In an inverse problem, one is asked to infer an unknown random signal X0 ∈ Rd from noisy observations Y = A(X0 ) + ε,
(1)
where A : Rd → Rn is a known measurement operator, and ε ∈ Rn represents the measurement noise. Such problems are often ill-posed: even in the absence of noise, the observations y alone may be insufficient to uniquely determine the underlying signal X0 , particularly when n < d. A principled approach to resolving this ambiguity and improving statistical performance is provided by Bayesian inference, which combines the measurement model with a prior distribution over the unknown signal X0 . Specifically, given a prior pX0 , the central objective of Bayesian inference is to characterize, or sample from, the posterior distribution pX0 |y . The effectiveness of Bayesian inference therefore depends largely on the choice of prior. Classical approaches typically impose hand-crafted priors, such as sparsity or low-rank structure (Daubechies et al., 2004; Fan and Li, 2001; Candès et al., 2006; Donoho, 2006; Park and Casella, 2008; Ji et al., 2008; Babacan et al., 2009; Donoho et al., 2009; Candes and Recht, 2012; Candès, 2014; Rockova and George, 2018; Chi et al., 2019). While these priors have led to many important advances, they often oversimplify the structure of natural signals. Diffusion models offer a fundamentally different alternative: rather than manually specifying a prior or regularizer, they use a pre-trained generative model as an implicit prior over the signal distribution. This perspective has inspired a rapidly growing line of work on posterior sampling for inverse problems with diffusion priors (Song et al., 2022; Chung et al., 2023a; Kawar et al., 2022; Song et al., 2023a; Xu and Chi, 2024; Wang et al., 2024; Zhang et al., 2025). Moreover, because pre-trained diffusion models can be used as “plug-and-play” priors (Venkatakrishnan et al., 2013; Sreehari et al., 2016; Chan et al., 2017; Xu and Chi, 2024), they often require no task-specific retraining and readily transfer across sensing modalities and application domains. Using diffusion models as priors, however, presents a new suite of algorithmic challenges. Rather than assuming access to an explicit, closed-form expression for the prior distribution, the diffusion-prior framework models the prior implicitly through a pre-trained diffusion model that provides score estimates (i.e., estimates of the Stein score functions) for progressively diffused versions of the prior. A central algorithmic challenge is therefore to combine the learned diffusion prior with the measurement model (2) in order to generate samples 2
from the posterior distribution. Existing diffusion-based methods typically address this through heuristic correction steps, projection operators, likelihood-gradient guidance, approximate posterior updates, and so on (Chung et al., 2022; Chung and Ye, 2022; Choi et al., 2021; Kawar et al., 2022). Although these methods have demonstrated strong empirical performance, many lack rigorous guarantees that the generated samples converge to the true Bayesian posterior. Furthermore, some require computationally intensive modifications to the underlying sampling procedure. For example, Chung et al. (2023a); Song et al. (2023a) required computing the gradient of the score functions, while the sequential Monte Carlo approach of Cardoso et al. (2023) may require a prohibitively large number of particles. These considerations can limit their practical applicability in settings where both computational efficiency and statistical reliability are important. These issues highlight the need for diffusion-based inverse problem solvers that achieve rigorous theoretical guarantees and computational efficiency at once. Such properties are particularly important in high-stakes applications such as medical imaging and scientific discovery. At the same time, practical deployment often favors algorithms that preserve the simplicity and efficiency of standard diffusion samplers.
1.2
An overview of our contributions
In this paper, we make progress towards developing a theoretically principled, and computationally efficient framework for posterior sampling in linear inverse problems with diffusion priors. Specifically, we consider the following noisy linear measurement model: Y = AX0 + ε,
(2)
where X0 ∈ Rd is the unknown signal (generated randomly from a data distribution p0 ), A : Rd → Rn represents a linear measurement operator known a priori, and ε ∼ N (0, σ 2 In ) stands for Gaussian measurement noise with σ > 0 the noise standard deviation. Throughout, we assume that the prior distribution of X0 can be represented implicitly by a pre-trained diffusion model. The objective is to generate samples from the posterior distribution of X0 conditioned on the observation y. Our starting point is the observation that, under a linear measurement model, different singular directions of the measurement operator carry different amounts of statistical information. Rather than treating all directions uniformly, we design a DDIM-type posterior sampler, called Posterior-DDIM, that adaptively determines, for each singular direction, whether the update should be based primarily on the learned diffusion prior or the observed data. This results in a simple, lightweight, coordinate-wise modification of the standard DDIM sampler that explicitly exploits the structure of the measurement operator, while preserving the efficiency of DDIM-type sampling. Along directions where the measurements have sufficient signal-to-noise ratio (SNR), the algorithm injects a carefully calibrated noisy version of the observations so that the resulting coordinates have the correct forward-diffusion marginal distribution. In contrast, directions that are either observed with low SNR or entirely unobserved continue to evolve according to the standard DDIM dynamics induced by the prior. Consequently, posterior sampling reduces to a sequence of simple, direction-wise DDIM updates that are both computationally efficient and easy to implement. Theoretically, our main contribution is to establish posterior consistency of the proposed Posterior-DDIM sampler. Under an exact diffusion prior and a sufficiently fine time discretization, we prove that the sampler converges to the posterior distribution of the unknown signal conditioned on the measurements. This result provides rigorous theoretical support for DDIM-type samplers in linear inverse problems, and suggests a principled approach to integrating measurement information with diffusion priors in noisy, potentially illconditioned or rank-deficient settings. Empirically, we evaluate the proposed sampler on a range of image restoration tasks, including inpainting, super-resolution, Gaussian deblurring, and compressed sensing, using PSNR, SSIM, LPIPS, and FID as evaluation metrics. Across these tasks, our Posterior-DDIM algorithm performs favorably against existing diffusion-based posterior samplers: it achieves the best performance on at least three of the four metrics in most tasks, and on at least two metrics in every task considered.
1.3
Related work
Despite their significant empirical success, existing diffusion-based inverse solvers still leave room for improvement. One prominent class of methods (Song et al., 2020; Chung et al., 2022; Chung and Ye, 2022; 3
Choi et al., 2021; Kawar et al., 2022; Chung et al., 2023b) enforces measurement consistency by incorporating projection steps into unconditional diffusion sampling procedures. While these approaches have demonstrated strong empirical performance, they do not provide theoretical guarantees on the convergence or posterior correctness of the resulting sampling procedures. A second line of work (Chung et al., 2023a; Song et al., 2023a; Kawar et al., 2021; Cardoso et al., 2023; Song et al., 2023b) attempted to approximate the (intractable) posterior score and subsequently apply standard diffusion samplers for posterior sampling. Although empirically effective, it remains unclear whether the resulting sampler faithfully yields the true posterior distribution. In addition, Song et al. (2022) proposed ReSample, which promotes measurement consistency by repeatedly optimizing the clean image estimate at intermediate time steps and then re-noising it to continue the reserve diffusion process. Similarly, Graikos et al. (2022) proposed to iteratively optimize the clean image via gradient descent in the reverse diffusion process. A different perspective was introduced by Mardani et al. (2023), who formulated inverse problems with diffusion priors as a variational inference problem via direct KL minimization rather than reverse diffusion sampling. Like the aforementioned approaches, these methods currently lack rigorous theoretical guarantees. Another strand of work Trippe et al. (2022); Wu et al. (2023); Dou and Song (2024); Cardoso et al. (2023) combined Sequential Monte Carlo (MC) methods (Doucet et al., 2001) with (unconditional) scores to sample from the correct posterior; in particular, Cardoso et al. (2023); Dou and Song (2024) established convergence guarantees as the number of particles tends to infinity. However, these methods incur a prohibitive computational cost. In particular, (Gupta et al., 2024) showed that polynomially many particles are in general insufficient for the above Sequential MC approaches to converge. More recently, Xu and Chi (2024) proposed an alternating framework that accommodates even nonlinear inverse problems by interleaving two samplers: a proximal consistency sampler based on the likelihood function of the forward model and a denoising diffusion sampler driven by the learned score function of the prior. They also established rigorous theoretical guarantees for the resulting algorithm. In contrast, our proposed method is built almost entirely upon DDIM-type sampling. Finally, we note that the convergence theory of DDPM-type samplers has been studied extensively in recent years (see, e.g., Chen et al. (2023); Li et al. (2024); Huang et al. (2024); Liang et al. (2025); Tang and Yan (2026); Cai et al. (2026); Jiao et al. (2025); Li and Jiao (2024); Jiao and Li (2024)), validating the efficacy of DDIM-type sampling and paving the way for more refined theoretical analyses in future work.
2
Preliminaries
In this section, we briefly review the background on diffusion models and their connections to inverse problems, which will be useful throughout the paper. Diffusion models. We begin by reviewing the basics of diffusion models. Let X0 ∼ p0 denote a random sample in Rd drawn from an underlying data distribution p0 . A diffusion model defines a forward stochastic process (Xt )t∈[0,T ] that progressively corrupts X0 by injecting Gaussian noise according to: d
Xt = αt X0 + σt Zt ,
Zt ∼ N (0, Id )
(3)
for all 0 ≤ t ≤ T , where the coefficients αt ≥ 0 and σt ≥ 0 specify the noise schedule. Throughout this paper, we assume that the SNR αt /σt decreasing monotonically and continuously as t grows, and satisfies α0 /σ0 = ∞ and α∞ /σ∞ = 0. As t increases, the signal component gradually diminishes while the noise level increases, so that, under an appropriate noise schedule, the distribution of Xt transitions smoothly from the data distribution toward an approximately Gaussian distribution. A central goal of a diffusion model is to learn the corresponding reverse process, which starts from a Gaussian sample and progressively removes the injected noise to yield a sample from the data distribution of interest. Score functions.
A fundamental object governing the reverse process is the (Stein) score function st (x) := ∇ log pXt (x),
0 ≤ t ≤ T,
(4a)
where pXt denotes the marginal distribution of Xt . The score function specifies the direction in which the log-density of Xt increases most rapidly and plays a central role in reversing the forward process (Anderson, 4
1982; Haussmann and Pardoux, 1986). In practice, the score functions are often parameterized through through equivalent prediction targets, such as the noise predictor ϵt (x) := E[Zt | Xt = x]
(4b)
µt (x) := E[X0 | Xt = x].
(4c)
and the data predictor
These objects are linked through Tweedie’s formula (Efron, 2011), which yields the identities ϵt (x) = −σt st (x)
and
µt (x) =
x − σt ϵt (x) . αt
(5)
Reparameterzed reverse process and DDIM-type sampler. The forward diffusion process (3) admits a reverse-time stochastic differential equation (SDE), originally characterized by Anderson (1982), which forms the basis for diffusion-based generative sampling. It is sometimes convenient to reparameterize time t using the following parameter concerning the log SNR: λt := log
αt . σt
(6)
Under this parameterization, the reverse SDE admits the following simple form using the λ-variable (i.e., the log SNR variable) as the “time variable” in place of t: √ X d (7) = −2e−λ ϵλ (X) dλ + 2e−λ dBλ , α where Bλ denotes a standard Brownian motion indexed by λ, and we abuse the notation by letting ϵλ be a shorthand for ϵt with λ = λt . Practical sampling procedures like DDPM-type samplers (Ho et al., 2020) can be obtained by approximating this reverse SDE (7) through suitable time discretization. More precisely, for a sequence of decreasing time points t0 > t1 > · · · > tM with corresponding log SNR values λt0 , · · · , λtM (cf. (6)), a first-order discretization yields the DDPM-type update q Xti+1 Xti −λti+1 −λti (8) = −2 e −e ϵti (Xti ) + e−2λti − e−2λti+1 Zti , αti+1 αti i.i.d.
where zti ∼ N (0, I). In light of the relation (5) between the noise predictor and the data predictor, the same update can be equivalently expressed in terms of µt . Letting δi+1 := λti − λti+1
(9)
with λt defined in (7), one obtains the first-order SDE-DPMSolver-1 update (Lu et al., 2022) p Xti+1 Xt = e−δi+1 i + eλti+1 1 − e−2δi+1 µti (Xti ) + 1 − e−2δi+1 Zti . σti+1 σti
(10)
More generally, DDIM-type samplers (Song et al., 2021) interpolate between deterministic and stochastic reverse updates. For any given parameter 0 ≤ η ≤ 1, the (general) DDIM-type update can be written as q q Xti+1 Xti (DDIM) : = 1 − η 1 − e−2δi+1 + eλti+1 1 − e−δi+1 1 − η 1 − e−2δi+1 µti (Xti ) σti+1 σ ti q + η 1 − e−2δi+1 Zti . (11) Notably, η controls the stochasticity of the update: the choice η = 0 yields the deterministic DDIM-type update, while η = 1 recovers the stochastic SDE-DPMSolver-1 update in (10). 5
Diffusion priors for inverse problems. Next, we briefly describe how diffusion models are commonly used as priors for linear inverse problems. Given a pre-trained diffusion model representing the prior distribution of X0 , the aim is to generate samples from the posterior distribution pX0 |Y =y . In analogy with the unconditional case, consider the conditional data predictor µt,y (x) := E[X0 | Xt = x, AX0 + ε = y].
(12)
If this conditional predictor were available, then one could simply replace the unconditional predictor µt by µt,y in the DDIM-type sampler (11), thereby directly obtaining samples from the posterior distribution pX0 |y . Consequently, a central challenge in diffusion-based inverse problems is to approximate the conditional predictor µt,y using only the pre-trained unconditional predictor µt together with the measurement y. Many existing methods are built upon this idea. Starting from the unconditional prediction µt (x), they seek to compute a measurement-compatible data predictor µ by balancing two competing objectives: remaining faithful to to the diffusion prior that yields the unconditional data predictor µt (·), while simultaneously enforcing consistency with the observed measurements. To approximate the conditional data predictor, existing approaches often rely on heuristic approximations. A common strategy is to combine the unconditional predictor µt (·) and the measurement Y in a simple — often linear — manner (Wang et al., 2022b; Kawar et al., 2022; Song et al., 2022), while another is to use µt (·) directly as a proxy for the unknown signal X0 (Chung et al., 2023a).
3
Main results
In this section, we develop a posterior-consistent DDIM-type sampler for noisy linear inverse problems, and establish its theoretical guarantees. Before proceeding, we find it convenient to introduce notation associated with the singular value decomposition (SVD) of the measurement matrix A ∈ Rn×d . Suppose that A has rank r, and let (13)
A = U ΣV ⊤ = US ΣS VS⊤
represent the full and compact singular value decompositions (SVDs) of A, where U ∈ Rn×n and V ∈ Rd×d are orthonormal matrices, US ∈ Rn×r and VS ∈ Rd×r consist of the left and right singular vectors associated with the nonzero singular values, and Σ ∈ Rn×d and ΣS ∈ Rr×r are the corresponding rectangular diagonal and diagonal matrices of singular values, respectively. We denote by S := {s ∈ [min{n, d}] : Σss > 0}
(14)
the indices of the nonzero singular values, and hence the notation US , VS and ΣS .We also write US := [us ]s∈S ∈ Rn×r ,
3.1
VS := [vs ]s∈S ∈ Rd×r ,
and
ΣS := diag([Σss ]s∈S ) ∈ Rr×r .
(15)
Motivation: SNR-guided partition of singular directions
The SVD (13) of the measurement matrix A naturally decouples the inverse problem into a sequence of scalar inference problems along its singular directions. Indeed, expressing the measurement model Y = AX0 + ε (i.e., (2)) in the SVD basis yields ⊤ ⊤ ′ ⊤ 2 −2 Σ−1 with ε′ := Σ−1 (16) S US Y = VS X0 + ε S US ε ∼ N 0, σ ΣS ; or equivalently, for every s ∈ S, σ2 with ε′s ∼ N 0, 2 . Σss
⊤ ⊤ ′ Σ−1 ss us Y = vs X0 + εs
(17)
Thus, after transforming into the singular-vector basis, the effective observation SNR along the s-th singular direction is naturally measured by Σss /σ. Meanwhile, projecting the forward diffusion process (3) along the same singular direction gives vs⊤ Xt = αt vs⊤ X0 + σt Zs 6
with Zs ∼ N (0, 1),
(18)
whose SNR — which we often refer to as diffusion SNR — is given by αt /σt . Consequently, the relative magnitudes of these two SNRs dictate whether the posterior update should rely primarily on the measurements or on the diffusion prior. This comparison naturally leads to a partition of the singular directions. For each diffusion time t, we partition the singular directions according to Stmeas := {s ∈ S : αt σ < Σss σt }
Stprior := S \ Stmeas .
and
(19)
For each s ∈ S, we define the threshold time τs ∈ R≥0 ∪ {∞} such that Σss ατs = , σ τs σ namely, the diffusion time at which the diffusion SNR matches the observation SNR along the s-th singular direction. When αt /σt is continuous and strictly decreasing in t, with α0 /σ0 = ∞ and α∞ /σ∞ = 0, the threshold τs is well defined. It can be easily verified that s ∈ Stmeas
⇐⇒
t ≥ τs .
(20)
Therefore, at each diffusion time t, it is natural to partition the singular directions into three groups: • Measurement-dominated directions: s ∈ Stmeas , or equivalently, directions obeying t ≥ τs ; • Prior-dominated observed directions: s ∈ Stprior , or equivalently, directions in S obeying t < τs ; • Unobserved directions: s ∈ S c := [d] \ S. Our sampler treats these three groups differently. As we shall see, each case admits a simple DDIM-type update, resulting in an efficient posterior sampler that adapts to the SNR along each singular direction.
3.2
Algorithm: Posterior-DDIM
We now describe the proposed sampler, based on the partition of singular directions developed in Section 3.1. The central idea is to exploit the measurements whenever the corresponding observation SNR exceeds the diffusion SNR, and otherwise rely on the learned diffusion prior. This idea gives rise to distinct update rules for the three groups of singular directions, which we describe below. The complete algorithm of the proposed DDIM-based posterior sampler, hereafter referred to as Posterior-DDIM, is summarized in Algorithm 1. Measurement-dominated directions. For directions s ∈ Stmeas , the measurements are sufficiently informative to be used directly. To this end, consider the auxiliary forward process q ξt,s := αt ξ0,s + σt2 − αt2 σ 2 Σ−2 Zs ∼ N (0, 1), (21) ss Zs , where ξ0,s is taken to be ⊤ ξ0,s := Σ−1 ss us Y.
(22)
By construction, this auxiliary forward process ξt,s has the same marginal distribution as the forward diffusion process Xt (cf. (3)) projected onto the direction the s-th singular direction vs , namely, d
ξt,s = vs⊤ Xt = vs⊤ (αt X0 + σt Z). bt at the time point ti can be accomplished by generating Consequently, sampling vs⊤ X i q ⊤ bt = αt Σ−1 vs⊤ X σt2i − αt2i σ 2 Σ−2 with Zti ∼ N (0, 1), ss Zti ss us Y + i i provided that s ∈ Stmeas . A formal proof is given in Appendix B.1. i 7
(23)
Algorithm 1: Posterior-DDIM sampler for the linear inverse problem (2) M input: time points t0 > t1 > · · · > tM , learned data predictor {µti (·)}M i=1 , schedule {(αti , σti )}i=1 , stochasticity parameters η0 , η1 . bt ∼ N (0, Id ), Z = [Zs ]1≤s≤d ∼ N (0, Id ). 2 initialize X 0 3 compute the SVD of A (cf. (13)); set S ← {i : Σii > 0}. /* initialize DDPM updates in all directions. */ 4 for s = 1, · · · , d do 5 if s ≤ min{n, d} and αt0 σ < Σssq σt0 then ⊤ b −1 ⊤ 6 set vs Xt ← αt Σss us Y + σt2 − αt2 σ 2 Σ−2 ss Zs . 1
0
7 8
0
0
0
else bt ← Zs . set vs⊤ X 0
for i = 0, . . . , M − 1 do 10 set Stmeas ← {s : αti+1 σ < Σss σti+1 } and Stprior ← S \ Stmeas . i+1 i+1 i+1 11 draw Zti+1 ∼ N (0, I). 12 for s ∈ Stmeas do i+1 /* updates for measurement-dominated directions. q bt ← αt Σ−1 u⊤ Y + σt2 − αt2 σ 2 Σ−2 13 compute v ⊤ X ss Zs . 9
s
14
15 16 17 18
i+1
i+1
ss
s
i+1
*/
i+1
for s ∈ Stprior do i+1 /* DDIM for prior-dominated directions. if all nonzero singular values of A are equal then bt compute vs⊤ X via (24) with η = η0 . i+1
*/
else bt compute vs⊤ X i+1 via (25) with η = η0 . // add correction for general case.
for s ∈ S c do bt via (24) with η = η1 . // DDIM for unobserved directions. 20 compute vs⊤ X i+1 bt = P vs vs⊤ X bt . 21 output: the generated sample X M M s 19
Prior-dominated observed directions. For directions s ∈ Stprior , the measurements are less informative than the diffusion prior. Consequently, rather than directly incorporating the observations, we update these directions using the standard DDIM sampler driven primarily by the learned data predictor. We begin with the special case in which all nonzero singular values of A are identical. In this setting, for bt according to the standard DDIM-type update rule (cf. (24)) as follows s ∈ Stprior , we update vs⊤ X i+1 q bt bt q vs⊤ X vs⊤ X i+1 λti+1 i −δi+1 −2δ −2δ bt ) i+1 i+1 = 1 − η(1 − e )+e 1−e 1 − η(1 − e ) vs⊤ µti (X i σti+1 σ ti q + η(1 − e−2δi+1 ) vs⊤ Zti+1 , Zti+1 ∼ N (0, Id ),
(24)
where the stochasticity parameter is set to η = η0 . The above update extends naturally to the general setting where A has arbitrary nonzero singular values. The only modification is the addition of a correction term. Specifically, for s ∈ Stprior , we replace (24) with i q bt bt q vs⊤ X v⊤ X i+1 = s i 1 − η(1 − e−2δi+1 ) + η(1 − e−2δi+1 ) vs⊤ Zti+1 σti+1 σ ti q λti+1 −δi+1 −2δ bt ) + i+1 +e 1−e 1 − η(1 − e ) vs⊤ µti (X i
! b τ − ατ X bt αti X s s i , ατs αti (e2(λti −λτs ) − 1) | {z } correction term
8
(25)
where Zti+1 ∼ N (0, Id ). The derivation of this modified update rule is deferred to Appendix C. Unobserved directions. For directions s ∈ S c , the measurements contain no information about vs⊤ X0 . Accordingly, these coordinates should be sampled entirely from the learned diffusion prior, given the coordinates that are observed. As in the prior-dominated observed case, we employ the DDIM update (24). However, we choose a larger stochasticity parameter η = η1 to encourage sufficient exploration of the unobserved subspace. This is important because the posterior uncertainty along S c is determined entirely by the learned diffusion prior, rather than by the measurement model. Choices of the stochasticity parameters η0 and η1 . We take a moment to discuss the roles of the two DDIM stochasticity parameters η0 and η1 . Although both parameters control the amount of randomness injected into the DDIM update, they serve fundamentally different purposes because the target conditional distributions differ across the two subspaces. For directions s ∈ Stprior , the corresponding coordinates are observed, although the measurements are not yet sufficiently informative to be imposed directly. The update should therefore approximate the conditional reverse transition given the current diffusion state and the measurements. This motivates using a relatively small value of η0 , so as to keep the update closer to the deterministic DDIM sampler. In contrast, for directions s ∈ S c , no measurements are available. These coordinates must be sampled entirely from the learned conditional prior given the observed coordinates. A larger value of η1 makes the DDIM update more Langevin-like, thereby promoting exploration of the unobserved subspace
3.3
Theoretical guarantees: asymptotic posterior correctness
In this subsection, we establish the posterior correctness of the proposed Posterior-DDIM sampler. We begin with the special case in which all nonzero singular values of the measurement matrix A are identical, as this setting captures the main ideas while allowing for cleaner algorithm design analysis. Here and throughout, TV(p, q) denotes the total-variation (TV) distance between two probability measures p and q. Theorem 1. Consider a sequence of time points t0 > · · · > tM . Recall the definition of δi in (9), and set n o δ := max max δi , αt0 , |1 − σt0 | . (26) 1≤i≤M
Assume that the support of X0 is bounded, and that all nonzero singular values of A are equal. Suppose further that η0 = 1. When η1 = δ −p with 1/2 < p < 2/3, the Posterior-DDIM sampler (see Algorithm 1) achieves (27) lim lim TV pXbt |Y , pX0 |Y = 0. tM →0 δ→0
M
In words, Theorem 1 establishes that, under the assumption that all nonzero singular values of A are identical, the distribution of the output generated by Algorithm 1 converges to the posterior distribution of X0 conditioned on the measurements. This guarantee is asymptotic and holds in the regime where (i) the terminal time tM tends to 0, approaching the initial time of the forward diffusion process; (ii) the initial sampling time t0 is chosen such that αt0 → 0 and σt0 → 1, ensuring that the SNR at the initialization stage vanishes; and (iii) the maximum gap between the log-SNRs of consecutive time points vanishes, which in particular requires the number of sampling steps M to tend to infinity. The proof is deferred to Appendix A. The restriction that all nonzero singular values of A are identical is not essential. The proposed approach extends naturally to the general setting with arbitrary nonzero singular values, requiring only a slight modification of the sampling procedure (as already described in (25)). The corresponding theoretical guarantee is stated below. Compared to Theorem 1, the only additional change is a different choice of η0 . The proof of this theorem is postponed to Appendix C. Theorem 2. Under the assumptions of Theorem 1, except that the nonzero singular values of A are allowed to be arbitrary and η0 = η1 = δ −p for some 1/2 < p < 2/3, the conclusion of Theorem 1, namely (27), continues to hold. 9
Table 1: Empirical results for inpainting on CelebA and ImageNet under different η. CelebA ImageNet η0 η1 PSNR↑ SSIM↑ LPIPS↓ FID↓ PSNR↑ SSIM↑ LPIPS↓ FID↓ 0 27.22 0.79 27.11 61.46 21.24 0.54 46.72 89.56 1 31.32 0.88 16.97 30.24 26.13 0.78 26.17 37.76 2 32.55 0.90 15.44 26.62 28.76 0.86 16.74 18.60 0 4 33.16 0.91 15.07 25.42 30.72 0.89 14.45 15.74 8 33.23 0.91 15.06 25.33 31.06 0.89 14.31 15.86 16 33.23 0.91 15.06 25.33 31.06 0.89 14.31 15.85 32 33.23 0.91 15.06 25.33 31.06 0.89 14.31 15.85 0 33.23 0.91 15.06 25.33 31.06 0.89 14.31 15.85 0.5 33.03 0.90 15.66 26.60 30.86 0.88 15.33 17.89 16 32.84 0.90 16.36 28.42 30.67 0.88 16.61 20.85 1 2 32.51 0.89 17.89 33.21 30.26 0.86 19.52 28.71
4
Experiments
In this section, we conduct extensive empirical evaluations of the proposed Posterior-DDIM algorithm on a diverse range of image restoration tasks, including inpainting, super-resolution, Gaussian deblurring, and compressed sensing. We begin by describing the common experimental setup. Unless otherwise stated, the experiments are run with the NFE (number of function evaluations) set to 100, except for DPS, for which we follow Zhang et al. (2024) and set the NFE to be 1000. To justify our default choice, we also investigate the effect of NFE and find that the reconstruction quality largely saturates once the NFE exceeds 80 (see Figure 2). All experiments are conducted on a server equipped with NVIDIA A100-PCIE-40GB GPUs and Intel Xeon (Cascade Lake) CPUs.
4.1
Choices of η0 and η1
We evaluate the influence of the hyperparameters η0 and η1 in Posterior-DDIM, which govern the DDIM updates on S c and Stprior , respectively. Experiments are conducted on two image restoration tasks: 50% random inpainting and 4× super-resolution with average pooling, using the CelebA (Liu et al., 2015) and ImageNet (Deng et al., 2009) datasets. We set the noise standard deviation to σ = 0.05 (doubled when pixels are rescaled to [−1, 1]). For ImageNet, we use the pretrained 256 × 256 model without classifier guidance from Dhariwal and Nichol (2021). For CelebA, we use the pretrained model released with SDEdit (Meng et al., 2021). Following Wang et al. (2022a); Zhang et al. (2024), we evaluate on the 1K ImageNet sub-test set and the CelebA test set provided by Wang et al. (2022a). Performance is measured using PSNR (peak SNR), SSIM (Wang et al., 2004), LPIPS (Zhang et al., 2018), and FID (Heusel et al., 2017). For readability, all reported LPIPS values are multiplied by 100. We vary η1 from 0 to 32 while fixing η0 = 0, and vary η0 from 0 to 2 while fixing η1 = 16. The results for inpainting and super-resolution are reported in Tables 1 and 2, respectively. For fixed η0 , increasing η1 generally improves performance across most evaluation metrics, although the improvement is not strictly monotone and small fluctuations are observed. This overall trend is consistent with our theoretical analysis that requires η1 to be sufficiently large. Turning to the second set of experiments, where η1 is fixed and η0 is varied, we find that fine-tuning η0 can yield additional, albeit modest, performance gains. Empirically, the best empirical performance is attained at η0 = 0, which corresponds to the deterministic DDIM update.
4.2
Empirical comparisons with reference algorithms
Next, we compare Posterior-DDIM against several representative diffusion-based image restoration algorithms, including DDRM (Kawar et al., 2022), DDNM+ (Wang et al., 2022a), DPS (Chung et al., 2023a), RED-diff
10
Table 2: Empirical results for super-resolution on CelebA and ImageNet under different η. CelebA ImageNet η0 η1 PSNR↑ SSIM↑ LPIPS↓ FID↓ PSNR↑ SSIM↑ LPIPS↓ FID↓ 0 27.77 0.78 22.49 32.97 22.48 0.56 38.50 53.08 1 28.39 0.80 21.09 31.24 24.81 0.70 31.22 39.94 2 28.92 0.82 21.31 32.56 25.67 0.72 32.03 45.74 0 4 29.59 0.83 21.85 33.49 25.93 0.73 32.73 49.06 8 29.83 0.84 22.02 34.13 25.90 0.73 32.84 49.99 16 29.84 0.84 22.02 34.16 25.90 0.73 32.84 49.99 32 29.84 0.84 22.02 34.16 25.90 0.73 32.84 49.99 0 29.84 0.84 22.02 34.16 25.90 0.73 32.84 49.99 0.5 29.48 0.83 22.86 35.35 25.64 0.72 34.11 52.27 16 29.21 0.82 23.57 37.15 25.44 0.70 35.23 55.03 1 2 28.82 0.81 24.89 41.79 25.13 0.69 37.24 61.86 (Mardani et al., 2023), and ProjDiff (Zhang et al., 2024).1 As a classical baseline, we also include the b ls = A† Y . In addition to the inpainting and super-resolution tasks considered in least-squares estimator, X Section 4.1, we further evaluate all methods on Gaussian deblurring and compressed sensing. For Gaussian deblurring, we use a one-dimensional Gaussian blur kernel of size 5 with standard deviation 10. For compressed sensing, we employ a Walsh–Hadamard sampling matrix with a sampling ratio of 0.5. Throughout all experiments, the observation noise standard deviation is fixed at σ = 0.05. In addition, note that the posterior mean can be approximately computed via Monte Carlo sampling, i.e, N
E[X0 | Y ] ≈
1 X b (i) X , n i=1
i.i.d.
b (i) ∼ pX | Y . As is well-known, the posterior mean achieves the optimal reconstruction accuracy where X 0 under the mean squared estimation error metric. Consequently, reconstruction fidelity can be improved by averaging samples generated from independent noise realizations. We observe, however, that this approach may sometimes degrade quality as measured by FID. To illustrate this effect, we report both single-sample results and posterior-mean results, where the latter are obtained by averaging 4 independent trials. The results are reported in Tables 3–6. Boldface indicates the best overall result, while underlining indicates the best sampling-based result. Across all tasks and datasets, the proposed Posterior-DDIM method consistently delivers strong performance. For super-resolution, it achieves the best performance on three and two of the four evaluation metrics on CelebA and ImageNet datasets respectively. For inpainting, it outperforms the reference algorithms on all four metrics for CelebA and on two metrics for ImageNet. For Gaussian deblurring, Posterior-DDIM achieves the best performance on three metrics for CelebA and two metrics for ImageNet. For compressed sensing, Posterior-DDIM outperforms DDNM+ across all of the four evaluation metrics. Overall, whereas existing methods typically achieve the best result on only one or two metrics in most settings, our method attains the best performance on at least two evaluation metrics across all tasks and datasets, demonstrating its efficacy and robustness across a diverse range of inverse problems. Figure 1 shows the performance of Posterior-DDIM, DDNM+, and ProjDiff under different numbers of Monte Carlo samples used to estimate the posterior mean. As discussed previously, increasing the number of samples improves reconstruction fidelity, as reflected by PSNR, SSIM, and LPIPS, but comes at the expense of reduced sample diversity, leading to worse FID scores. Furthermore, among all reference algorithms considered, our method achieves the best trade-off between fidelity and diversity on this task. Figure 2 illustrates the effect of the NFE on the super-resolution task using the ImageNet dataset. As the NFE increases, the performance improves, and becomes stable once NFE exceeds 80, suggesting that our default choice of NFE=100 provides a favorable balance between computational cost and reconstruction quality. 1 In addition to our algorithm, we also implement ProjDiff and DDNM+. The remaining results reported in this section are taken from Zhang et al. (2024). 2 When reproducing DDNM+ on the Gaussian deblurring task, we observe that the obtained results are substantially inferior
11
Table 3: Empirical results for super-resolution on CelebA and ImageNet. Method
Singlesample
Posterior mean
A† y DPS DDRM RED-diff DDNM+ ProjDiff Posterior-DDIM (Ours) DDNM+ ProjDiff Posterior-DDIM (Ours)
PSNR↑ 23.64 27.98 29.20 24.98 29.20 29.49 29.84 30.22 30.31 30.35
CelebA SSIM↑ LPIPS↓ 0.49 68.72 0.78 23.10 0.82 21.92 0.55 50.59 0.82 21.91 0.83 20.89 0.84 22.02 0.85 21.97 0.85 22.55 0.85 22.55
FID↓ 147.89 39.91 40.14 73.89 39.96 36.61 34.16 43.54 38.34 40.62
PSNR↑ 21.85 24.44 25.66 22.74 25.62 25.73 25.90 26.11 26.10 26.13
ImageNet SSIM↑ LPIPS↓ 0.45 65.34 0.67 31.81 0.72 34.88 0.49 53.24 0.72 34.39 0.72 33.03 0.73 32.84 0.73 34.38 0.74 33.13 0.74 33.35
FID↓ 183.32 36.17 55.71 96.26 53.78 49.70 49.99 54.27 50.27 52.91
Table 4: Empirical results for inpainting on CelebA and ImageNet. Method
Singlesample
Posterior mean
5
A† y DPS DDRM RED-diff DDNM+ ProjDiff Posterior-DDIM (Ours) DDNM+ ProjDiff Posterior-DDIM (Ours)
PSNR↑ 13.70 32.80 32.81 7.96 32.81 33.44 33.23 33.98 34.22 34.24
CelebA SSIM↑ LPIPS↓ 0.19 76.07 0.90 16.32 0.90 16.78 0.18 78.27 0.90 16.78 0.91 15.32 0.91 15.06 0.92 16.06 0.92 14.41 0.92 13.81
FID↓ 226.28 32.80 35.28 192.32 35.33 31.13 25.33 37.64 32.40 27.15
PSNR↑ 14.21 30.15 29.99 9.85 30.00 31.32 31.06 31.22 32.07 32.04
ImageNet SSIM↑ LPIPS↓ 0.24 67.42 0.86 17.76 0.87 17.11 0.17 83.92 0.87 17.07 0.89 13.87 0.89 14.31 0.89 16.44 0.90 13.78 0.90 13.67
FID↓ 176.52 22.03 19.88 281.65 19.92 14.97 15.85 22.07 16.57 16.75
Discussion
In this work, we have proposed a novel DDIM-type posterior sampler for solving linear inverse problems with diffusion priors. In contrast to many existing methods that rely on heuristic projection steps or approximate posterior scores, our approach follows a different route by treating different singular directions of the measurement operator differently according to the relative strengths of the observation noise and the diffusion noise. This leads to a simple, easy-to-implement, DDIM-type sampling procedure that is both practically effective and theoretically principled. In particular, we have established that, under mild conditions, the proposed Posterior-DDIM sampler converges to the true posterior distribution of the signal given the measurement. Empirically, our experiments on a range of real-world image restoration tasks have shown that the proposed method consistently performs favorably against existing diffusion-based posterior samplers, achieving the best performance on the majority of evaluation metrics across all tasks. Several directions merit further investigation. For instance, the present work has focused on linear inverse problems, and a fundamentally important next step is to extend the proposed framework to broader nonlinear inverse problems, which naturally arise in wide-ranging real-world applications. In addition, while our theoretical analysis has established asymptotic convergence to the true posterior distribution, it falls short of characterizing the rate of convergence, which would be of interest to quantify. It would also be valuable to understand how the convergence behavior and sampling quality depend on additional factors such as the noise level, the properties of the measurement operator, and the choice of time discretization of the diffusion process. to those reported in Zhang et al. (2024). Consequently, we adopt the results from Zhang et al. (2024); since only the published single-sample results are available, posterior-mean results for DDNM+. 3 Results for the compressed sensing task are not implemented in Zhang et al. (2024); therefore, we compare only with DDNM+. We note that Wang et al. (2022a) reported that DDNM+ outperforms DDRM with PSNR, SSIM, and FID metrics, and our algorithm outperforms DDNM+.
12