ConceptioArchivearXiv CS
arXiv CSopen access

Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models Bartlomiej Sobieski1,2,3,4∗, Matthew Tivnan3,4 , Dawid Płudowski1 , Michał Jan Włodarczyk1 , Pengfei Jin3,4 , Przemyslaw Biecek1,2 , Quanzheng Li3,4

arXiv:2605.05026v1 [cs.CV] 6 May 2026

1

Centre for Credible AI, Warsaw University of Technology, 2 University of Warsaw, 3 Massachusetts General Hospital, 4 Harvard Medical School

Abstract Diffusion models are prone to generating structural hallucinations - samples that match the statistical properties of the training data yet defy underlying structural rules, resulting in anomalies like hands with more than five fingers. Recent research studied this failure mode from several viewpoints, offering partial explanations to their occurrence, such as mode interpolation. In this work, we propose a complementary perspective that treats hallucinations as instabilities on the model-induced manifold. We begin by showing that a hallucination filter based on such instabilities matches or exceeds the performance of the recently proposed temporal one. By tracing the source of these instabilities, we identify local intrinsic dimension (LID) as their primary driver and propose Intrinsic Quenching (IQ), a direct corrective mechanism that deflates it to alleviate hallucinations. IQ consistently outperforms standard hallucination reduction baselines across a wide array of benchmarks and offers a highly promising solution for enforcing anatomical consistency in downstream medical imaging tasks.

1

Introduction

Diffusion models [DMs; Sohl-Dickstein et al., 2015, Song et al., 2021, Lipman et al., 2023, Liu et al., 2023b, Albergo and Vanden-Eijnden, 2023, Song et al., 2023, Geng et al., 2025] drive modern generative modeling across computer vision and other modalities [Dhariwal and Nichol, 2021, Lou et al., 2024, Kong et al., 2021, Kotelnikov et al., 2023, Hoogeboom et al., 2022]. Recent research highlights their critical failure mode termed hallucinations: generated samples that successfully capture the statistical properties of the training data, but fundamentally violate its underlying physical, logical, or morphological patterns [Ramesh et al., 2021, Rombach et al., 2022].

Figure 1: A diffusion model-induced manifold, approximating the data manifold, invents spurious dimensions in unstable regions, leading to structural hallucinations.

In particular, structural hallucinations [Kim et al., 2024] fail to recover correct object forms, generating anomalies like six-fingered hands or misaligned eyes. Lacking a formal definition, identifying them relies strictly on human perception, restricting research to empirical and correlational studies [Oorloff et al., 2025, Triaridis et al., 2025, Tian et al., 2025]. Despite initial explanatory hypotheses [Aithal et al., 2024], understanding their root causes and prevention remains an open question. ∗ Corresponding author at [email protected]

Preprint.

In this work, we propose shifting the analysis of structural hallucinations from temporal generative behavior [Aithal et al., 2024, Tian et al., 2025] to the local instability of the model-induced manifold. Specifically: C1. We demonstrate that a spatial hallucination filter based on local geometric instability matches or exceeds the recently proposed temporal filter [Aithal et al., 2024], identifying the inflation of local intrinsic dimension (LID) as the primary indicator. C2. We theoretically formulate Intrinsic Quenching (IQ), a thermodynamically-inspired corrective mechanism that dynamically deflates LID of the generated sample on the model-induced manifold throughout the generative process, reducing the hallucinatory behavior. C3. Through large-scale, human-annotated evaluations, we show that IQ outperforms all baselines in reducing structural hallucinations and offers a promising solution for enforcing anatomical consistency in downstream medical imaging tasks.

2

Background

DMs can be formulated through the framework of stochastic differential equations (SDEs): dxt = Ft xt dt + Gt dwt ,

(1)

dxt = [Ft xt − Gt G⊤ t ∇xt log p(xt )]dt + Gt dwt , n

(2) n×n

n×n

where Ft xt is a linear drift term with xt ∈ R and time-dependent matrix Ft ∈ R , Gt ∈ R is a matrix-valued diffusion coefficient, wt ∈ Rn and wt ∈ Rn are the Wiener processes running forward and reverse in time respectively, ∇xt log p(xt ) is the score function of the time-parameterized distribution p(xt ), t ∈ [0, 1], where p(x0 ) denotes the data and p(x1 ) a terminal distribution, typically chosen to be Gaussian. A DM, parameterized by θ, is trained to approximate the score function. Equation (1) defines the forward process, which gradually corrupts data samples into Gaussian noise, whereas eq. (2) defines the reverse process, which generates the data through iterative denoising. The Probability Flow ODE (PF-ODE) is the deterministic counterpart of eq. (2) given by FODE (xt , t) ≜ dxt 1 ⊤ dt = [Ft xt − 2 Gt Gt ∇xt log p(xt )]. Substituting the θ-parameterized model’s approximated score, we define the generator as Z 1 G θ (x1 ) = x1 − FθODE (xt , t)dt, (3) 0

denoting the model-based mapping of a sample x1 ∼ p(x1 ) to the approximate data distribution pθ (x0 ) given by the DM. At the center of our reasoning, we rely on the following assumption. Assumption 1. We follow the stratified manifold hypothesis [Goresky and MacPherson, 1988], which posits that high-dimensional data resides on a disjoint collection of low-dimensional submanifolds, and assume that a DM is able to learn a statistical approximation of these strata. Based on assumption 1, we define Mθ = {x0 | ∃x1 ∈Rn x0 = G θ (x1 )}, the set of all possible outputs of the generator, and refer to it as the model-induced manifold for simplicity. This construction is justified by foundational works and recent research on the geometry of DMs [Fefferman et al., 2016, Pidstrigach, 2022, Stanczuk et al., 2024]. Training DMs relies on a tractable formula R for theforward kernel p(xt | x0 ) = N (Ht x0 , Σt ), t where Ht = Φ(t, 0) with Φ(t, s) = exp s Fu du , assuming Ft commutes for all t, and Σt = Rt ⊤ Φ(t, τ )Gτ G⊤ τ Φ(t, τ ) dτ . To train a DM, one can either use ∇xt log p(xt | x0 ) as target or 0 utilize equivalent objectives such as h i 2 LDSM (x0 , t, θ) ≜ Eϵ∼N (0,I) ∥ϵ − ϵθ (xt )∥2 , (4)   1 2 LISM (x0 , t, θ) ≜ Eϵ∼N (0,I) Tr (Σt ∇xt sθ (xt )) + ∥sθ (xt )∥Σt (5) 2 1

where xt = Ht x0 + Σt2 ϵ, in each case marginalizing over x0 and t to learn the underlying score function [Hyvärinen, 2005, Vincent, 2011, Yeats et al., 2025]. We refer to ϵθ as the noise prediction and sθ as the score prediction. Thanks to the Tweedie’s formula [Efron, 2011, Alain and Bengio, 2014], one can directly relate the score function to the mean of the posterior p(x0 | xt ) through x̂0 (xt ) ≜ E[x0 | xt ] = Φ(t, 0)−1 (xt + Σt ∇xt log p(xt )) . 2

(6)

We denote by x̂θ0 (xt ) the result of replacing the true score function in eq. (6) with sθ . Arising from the manifold hypothesis, LID [Levina and Bickel, 2004, Bengio et al., 2013] encodes the local tangent space dimensionality, representing the valid degrees of freedom within a data stratum [Pope et al., 2021]. Historically a diagnostic tool in early deep learning research [Ma et al., 2018, Ansuini et al., 2019], LID can now be natively estimated by DMs. This capability, which justifies assumption 1, relies on x̂θ0 (xt ) acting as an approximate orthogonal projector for small t. While early DM-based estimators were computationally heavy [Tempczyk et al., 2022, Stanczuk et al., 2024, Kamkari et al., 2024], Yeats et al. [2025] demonstrated that standard training losses (eqs. (4) and (5)) provide efficient upper bounds of LID.

3

Empirical Investigation

The initial hypothesis given by Aithal et al. [2024] suggested that DMs’ hallucinations stem from interpolating between the modes of the data distribution and thus amplifying the regions of low probability. To filter out hallucinations post-generation, the authors proposed the Trajectory Variance 2 Rt Filter (TVF) defined as TVF(x0 ) = t12 x̂θ0 (xt ) − x̂θ0,t1 :t2 dt, where x̂θ0,t1 :t2 is the average 2

predicted posterior mean over the [t1 , t2 ] interval. This represents the temporal view, which is used to analyze hallucinations using the intermediate properties of the reverse process. We propose to shift that perspective and instead focus on the spatial view by analyzing the local geometry of the DM on the induced manifold. We define the Local Manifold Instability (LMI) of G θ for an initial point x1 as LMI(x1 ) ≜ Tr(Covε (G θ (x1 + ε))) = Varε (G θ (x1 + ε)),

(7)

where ε ∼ N (0, β 2 I) for some small β > 0, often also referred to as local sensitivity [Cacuci et al., 2005]. Intuitively, LMI measures the total spatial spread of a small spherical region of noise after it is transported to the data space, providing a proxy for exploring the local geometry of a given DM. Geometric perspective. We begin with a core question: If structural hallucinations can be partially captured by temporal variance during the reverse process, do these failure modes also manifest as localized spatial instabilities on the model-induced manifold? To assess it, we compare the performance of TVF and LMI as hallucination filters on a benchmark of hand images from the 11kHands dataset [Afifi, 2019], where they are easily recognizable as hands with deformed, missing or additional fingers. We use the capable EDM model [Karras et al., 2022] trained by Tian et al. [2025] with a deterministic 40-step Euler solver to generate 128 samples. To decide whether a given sample is correct or represents a structural hallucination, we conduct a user study with independent human annotators to label the samples in a binary manner. For TVF, we collect x̂θ0 (xt ) for each t throughout generation and then optimize t1 and t2 for best performance. For LMI, we estimate it with 32 perturbations per sample for a chosen β. For more details, see section B.2. Figure 2 depicts densities for both filters divided between correct and hallucinated samples, as well as classification and separability performance metrics. Interestingly, both filters are highly effective in separating correct and hallucinated samples, with a minor advantage of LMI despite choosing the optimal time interval for TVF. While the practical applicability of LMI is limited due to its computational costs based on running the reverse process multiple times, these results provide a novel identification of hallucinations as unstable states on the model-induced manifold. We extend this experiment to other datasets in section B.3, showing that while the performance of these two filters varies, the relationship between them is preserved. To better understand the predictive ability of LMI, we provide a proposition that connects it with the LID of the generated sample on the model-induced manifold, which we denote as LIDθ . Proposition 1. Let x0 = G θ (x1 ) and assume a sufficiently small perturbation scale β > 0 such that a first-order linear approximation holds. Let σ1 ≥ σ2 ≥ · · · ≥ σn ≥ 0 denote the singular values of the generator’s Jacobian ∇x1 G θ (x1 ), where n is the ambient dimension. Then it holds that LIDθ (x0 )

LMI(x1 ) ≈ β 2

X

σi2 .

i=1

We provide the proofs and full formal statements of all theoretical results in section A. 3

(8)

Figure 2: Left. Estimated densities for the values of each filter across correct and hallucinated samples. Middle. Separability and classification performance metrics for each filter. Right. Evolution of LID filter performance across the time interval. Vertical line indicates the optimal timestep. All results are reported on 11kHands dataset with EDM. Assumption 1 implies that LIDθ (x0 ) measures the effective dimensionality (significantly non-zero singular values) of x0 at Mθ when considering a theoretically full-rank G θ . Proposition 1 hence instantly reveals the reasons for the inflated LMI: it is caused by R1. over-estimating the true singular values (of the true stratified manifold), thus resulting in more rapid but plausible modifications under perturbation, or R2. the invention of spurious, off-manifold directions that influence the generated sample in an implausible or artifactual manner. With proposition 2, we extend the result of Yeats et al. [2025] to show that the losses in eqs. (4) and (5), shown to upper-bound the true LID, also form unbiased estimators of LIDθ . Proposition 2. Let x0 = G θ (x1 ). Under mild regularity conditions, it is true that LDSM (x0 , t, θ) = LIDθ (x0 ), LISM (x0 , t, θ) = − 21 (n − LIDθ (x0 )), where n is the dimension of the ambient space and t is sufficiently small. Hence, proposition 2 provides a tool to investigate the individual influences of R1. and R2. on the inflation of LMI by estimating LIDθ for both correct and hallucinated samples. Crucially, it also results in a continuous relaxation of the problem of estimating LIDθ , a theoretically discrete quantity. Under the same experimental setup, we compare the performance of LIDθ as a hallucination filter with TVF and LMI using LDSM -based estimator (eq. (4)) with 32 noise samples for approximating the expectation. Figure 2 (right) first shows how this performance depends on t, revealing large variability and the importance of choosing a small enough timestep, following the assumption of proposition 2. Crucially, fig. 2 (middle) shows that LIDθ significantly outperforms both LMI and TVF using the optimal t, providing the best separability and classification performance with higher values indicating hallucinations. From a geometric perspective, this result reveals the primary origin of the local instabilities; hallucinations occur when the DM overestimates the complexity of the generated samples by inflating their LIDθ . This result suggests that hallucinations appear when the model-induced manifold possesses unnecessary directions of expansion, which represent illogical degrees of freedom, such as a varying number of fingers or the possibility of more than one thumb (R2.). This also explains the relatively worse performance of LMI, which is simultaneously inflated when the variability along entirely valid directions increases (R1.).

4

Theoretical Foundation

While the empirical evidence justifies a hallucination filter based on LIDθ , filters themselves remain only a partial solution to the hallucination problem, as they inevitably require generating multiple invalid samples. In what follows, we propose a corrective mechanism for the sampling process inspired by the evidence that one should deflate LIDθ for particular samples to avoid hallucinations. >τ Theorem 1. Let x0 = G θ (x1 ). Assume that G θ (x1 ) is decomposed as G θ = G ≤τ θ ◦ G θ for ≤τ sufficiently small time τ > 0, where G >τ θ : x1 7→ xτ and G θ : xτ 7→ x0 . Let the singular values >τ of the macroscopic flow Jacobian ∇x1 G θ (x1 ) be monotonically ordered as σ1>τ ≥ σ2>τ ≥ · · · ≥ σn>τ ≥ 0. For sufficiently small β > 0, it is true that LIDθ (x̂θ 0 (xτ ))

X

LMI(x1 ) ⪅ β 2

i=1

4

(σi>τ )2 .

(9)

Figure 3: Hallucination depicts the initial unconditionally generated sample by EDM on 11kHands. For a set of manually selected timesteps, LIDθ gradients are visualized to highlight their coarse-tofine transition and perceptual connection to hallucinated content. Correction displays the sample generated with IQ, the proposed corrective mechanism applied in the [0.02, 0.05] interval. The contribution of theorem 1 is two-fold. First, it preserves the connection between local instability and the sample’s LIDθ , while only requiring the posterior mean x̂θ0 (xτ ) for some small τ > 0 instead of the final x0 . Second, it replaces the global singular values from proposition 1 with a spectral bottleneck: a summation where the macroscopic variance (t > τ ) is strictly truncated by the intrinsic dimensionality of the terminal projection (t ≤ τ ). Therefore, actively steering the trajectory to minimize LIDθ (x̂θ0 (xτ )) dynamically drops variance terms from this summation, provably decreasing the upper bound of LMI for the resulting sample. We formalize the implied sampling process below. Theorem 2. Assume t ≤ τ , where τ > 0 is sufficiently small. Replacing the standard trained score function sθ (xt ) in the reverse process (eq. (2)) with s̃θ (xt ) = sθ (xt ) − λt ∇xt E(xt ),

(10)

where the energy E(xt ) = LIDθ (x̂θ0 (xt )), is equivalent to sampling from the ideal Boltzmann t distribution over the true terminal states, pθ,λ (xt ) ∝ pθt (xt )Epθ (x0 |xt ) [exp(−λt LIDθ (x0 ))]. t Theorem 2 guarantees that, for small enough t, correcting the reverse process with the gradient of LIDθ for the posterior mean x̂θ0 (xt ) moves the resulting sample along Mθ towards terminal states of lower LIDθ , i.e., a lower-dimensional stratum. We obtain the gradient through ∇xt LIDθ (x̂θ0 (xt )) ≈ ∇xt LDSM (x̂θ0 (xt ), t, θ), where the approximation comes from estimating the DSM expectation with k i.i.d noise samples. Probabilistic perspective. The resulting correction represents a projected gradient descent scheme on Mθ . However, utilizing the LDSM estimator does not immediately reveal the probabilistic effect of that correction on the generated sample x0 . To address that, we provide corollary 1 showing that one may interchangeably use either the LDSM - or LISM -based estimator. Corollary 1. The gradients ∇xt LDSM (x̂θ0 (xt ), t, θ) and ∇xt LISM (x̂θ0 (xt ), t, θ) are exactly collinear and ∇xt LDSM (x̂θ0 (xt ), t, θ) = 2∇xt LISM (x̂θ0 (xt ), t, θ). This collinearity allows us to reason about the probabilistic effect of the stable and computationally cheap LDSM gradient through the theoretical lens of LISM . Corollary 2. Assuming sθ h(xt ) = ∇xt log pθt (xt ), eq. (5) can be iequivalently reformulated as 2

LISM (x0 , t, θ) = Eϵ∼N (0,I) Tr (Σt ∇2xt log pθt (xt )) + 12 ∥sθ (xt )∥Σt . Thus, as t → 0, minimizing

the energy from eq. (10) steers towards stationary points (∇xt log pθt (xt ) → 0) with maximal negative curvature (Tr (Σt ∇2xt log pθt (xt )) ≪ 0) of the model-induced probability distribution pθt (xt ), inducing a mode-seeking behavior. By mitigating mode interpolations, corollary 2 directly connects our optimization scheme with the work of Aithal et al. [2024]. As the proposed correction essentially aims at eliminating unstable, low-probability states by ‘cooling’ them and shifting towards stable local maxima, we term our approach as Intrinsic Quenching (IQ). From theory to practice. Both the theoretical assumptions and the separability performance from fig. 2 indicate that IQ should be only used for small times t < τ . We thus apply it solely within a narrow interval [t1 , t2 ] of the reverse process, where t2 is small. This focused intervention is also empirically justified when inspecting the energy gradients ∇xt E(xt ) (fig. 3). By overlaying them over the final generated image, we discover a smooth transition from global, structural attributes 5

through localized, highly-specific features correlated with the hallucinatory content, up to dispersed but cohesive states. This observation naturally aligns with multiple prior studies that decompose the diffusion process into semantically interpretable phases [Park et al., 2023, Sclocchi et al., 2025, Wang and Pehlevan, 2025]. In this case, however, LIDθ provides an inherent attribution map [Lundberg and Lee, 2017, Sundararajan et al., 2017] that allows quantifying such effects, as it indicates pixels with the highest contribution to the model’s error at time t. To ensure that the modified updates behave stably, we compute the λt scale dynamically for each t by ensuring that the magnitude of the energy term projected into data space ( λt Φ(t, 0)−1 Σt ∇xt E(xt ) 2 ) always equates to a fixed ratio λ of the natural update. For example, in the EDM framework, where the network is trained directly to predict x̂θ0 (xt ) (and Φ(t, 0) = I), ∥x̂θ (xt )−xt ∥ we use λt = λ Σ ∇0 E(x ) 2+ϵ , where ϵ ensures numerical stability. ∥ t xt t ∥2 The DM’s score function naturally interacts with the energy gradient from IQ; the former pushes the sample towards Mθ with decreasing t, while the latter shifts the sample along the manifold to regions of lower LIDθ , suppressing the most volatile directions. To ensure that this intervention selectively targets hallucinations rather than uniformly restricting all samples, we employ LIDθ as a dynamic, mid-generation filter. Specifically, we disable the correction (λt = 0) for stable samples satisfying E(xt ) < qt , where qt is a time-dependent threshold. To calibrate qt , we unconditionally generate a reference set of samples from the DM, precompute E(xt ) for each xt across the active interval [t1 , t2 ], and set qt to the q-th percentile of these empirical energy values at each timestep. We provide the pseudocode of IQ in algorithm 1.

5

Related Works

Hallucination reduction. Observed since early natural image DMs [Ramesh et al., 2021, Rombach et al., 2022], hallucinations were initially hypothesized by Aithal et al. [2024] as effects of mode interpolations, exacerbating errors during iterative DM training. Subsequent mitigations [Fu et al., 2025, Cho et al., 2025, Bhosale et al., 2026] include Adaptive Attention Modulation [AAM; Oorloff et al., 2025], which optimizes attention temperatures [Vaswani et al., 2017] via an anomaly detector [Roth et al., 2022], requiring the original training data. Dynamic Guidance [DG; Triaridis et al., 2025] applies Classifier Guidance [CG; Dhariwal and Nichol, 2021] toward the most probable class at each step, necessitating an independently trained noisy classifier. More recently, Tian et al. [2025] (RODS) formulate DMs as a continuation method [Allgower and Georg, 2012], mitigating hallucinations by intervening intrinsically upon detecting vector field instabilities. Task-specific methods. Because deformed hands are a prominent failure mode, several methods aim to correct or filter them specifically [Narasimhaswamy et al., 2024, Lu et al., 2024, Shi et al., 2026]. Similarly, Lu et al. [2025] identify local generation bias causing visual text failures. Hallucinations pose critical risks in image translation and reconstruction by affecting real-world decisions. Addressing this, Kim et al. [2024] propose local diffusion for conditional translation, partitioning inand out-of-distribution regions. In reconstruction, several works introduce quantitative hallucination scores [Tivnan et al., 2024, Ren et al., 2025, Ku et al., 2025]. Distinctively, Cao et al. [2025] employ temporal score analysis, akin to Aithal et al. [2024] and Tian et al. [2025], to eliminate related out-of-distribution artifacts.

6

Experiments

Datasets. We perform a large-scale quantitative comparison of IQ with recent methods for hallucination reduction by unifying several existing benchmarks. As a tractable toy problem, we use the GaussianGrid dataset proposed by Aithal et al. [2024], comprising a 2D mixture of uniformly spaced Gaussians in a form of rotated grid with 25 modes. For semantic structural hallucinations, we evaluate on SimpleShapes and MNIST for synthetic images following Aithal et al. [2024], Triaridis et al. [2025], Oorloff et al. [2025], and animal faces [AFHQV2, 64 × 64 resolution; Choi et al., 2020], human faces [FFHQ, 64 × 64 resolution Karras et al., 2019] and hand images [11kHands, 128 × 128 resolution; Afifi, 2019] entirely following Tian et al. [2025] for natural images. Models. For GaussianGrid, we train a 3-layer MLP with hidden dimension and time embedding of size 256 based on DDPM [Ho et al., 2020] with a linear schedule and 1000 steps. For SimpleShapes, 6

Table 1: Quantitative comparison of all sampling methods across six datasets. We highlight the primary performance indicators, based on human evaluation, regarding perceived image quality ( UP ) and hallucination ratio ( HR ), reported as averages across annotators. For 95% CIs, see table 5. 11kHands

FFHQ

AFHQV2

Method

FID ↓

IV ↑

DSV ↑ UP ↑ HR ↓ FID ↓

IV ↑

DSV ↑ UP ↑ HR ↓ FID ↓

IV ↑

DSV ↑ UP ↑ HR ↓

Baseline DG AAM RODSCAS RODSSAS IQ

16.3 16.2 16.4 15.8 16.2 16.6

0.032 0.033 0.033 0.033 0.033 0.030

0.15 0.15 0.15 0.15 0.15 0.13

0.067 0.070 0.068 0.066 0.066 0.065

0.40 0.40 0.40 0.40 0.40 0.39

0.054 0.054 0.055 0.054 0.054 0.054

0.48 0.48 0.49 0.48 0.48 0.48

Method

FID ↓

IV ↑

DSV ↑

HR ↓

FID ↓

IV ↑

DSV ↑

HR ↓

MMD ↓

HR ↓

Baseline DG AAM RODSCAS RODSSAS IQ

32.3 32.1 32.3 35.4 34.8 31.8

0.050 0.050 0.049 0.049 0.047 0.051

0.13 0.14 0.13 0.13 0.13 0.14

37.3 23.6 34.8 41.8 55.3 10.2

27.5 27.8 27.2 28.1 27.8 23.3

0.031 0.031 0.025 0.031 0.031 0.033

0.11 0.11 0.06 0.10 0.11 0.13

25.8 27.3 55.5 30.5 28.9 9.4

0.010 0.080 0.016 0.016 0.016

20.2 11.1 10.4 19.4 8.9

39.8 39.5 40.6 40.2 41.0 68.0

29.3 29.7 29.3 25.8 29.7 9.0

13.5 13.8 13.5 13.6 13.6 13.9

MNIST

45.3 44.1 45.7 45.3 45.3 46.1

8.2 9.8 10.2 7.7 8.2 4.2

16.7 16.7 17.6 16.9 16.8 17.0

SimpleShapes

41.8 40.2 42.2 42.2 42.2 42.6

6.9 8.6 14.3 6.9 6.6 5.9

GaussianGrid

we use a UNet [Ronneberger et al., 2015] with 3 convolutional blocks and an attention layer for both encoder and decoder, train it using 1000-step DDPM with squared cosine scheduler [Nichol and Dhariwal, 2021] and use 250-step sampling for inference. For MNIST, we use a pretrained UNet2 from HuggingFace [von Platen et al., 2022] also based on 1000-step DDPM and 250-step inference. For AFHQV2 and FFHQ, we use unconditional VE EDM checkpoints from the original paper [Karras et al., 2022] and sample with a deterministic 40-step Euler solver, following Tian et al. [2025]. For 11kHands, we reuse the VE EDM network pretrained by Tian et al. [2025], also with 40-step sampling for inference. Baselines. We compare IQ against recent methods for hallucination reduction: DG [Triaridis et al., 2025], AAM [Oorloff et al., 2025] and RODS [Tian et al., 2025]. Importantly, DG requires labeled training data of the DM of interest and assumes access to a pretrained noisy classifier for CG. Similarly, AAM also assumes access to training data to obtain the anomaly detection model and is limited to architectures with attention layers. In this context, RODS and IQ remain the only methods with no additional assumptions about the DM of interest. Our comparisons also include the standard unmodified sampling as a default baseline. For details regarding the unification of methods and adaptation to our codebase, see section B.1. Metrics. On GaussianGrid, we measure hallucination ratio (HR) as the percentage of samples exceeding 3 standard deviations from the closest mode [Aithal et al., 2024], alongside Maximum Mean Discrepancy (MMD) for distributional similarity. For images, we evaluate 2048 samples per method using proxy metrics: Fréchet Inception Distance [FID; Heusel et al., 2017] for quality, Inception Variance [IV; Miao et al., 2024] and DreamSim Variance [DSV; Domingo-Enrich et al., 2025] for diversity [Tian et al., 2025]. Because automated evaluations struggle with semantic perception, we quantify HR on image datasets via a blinded user study (section B.2). Annotators first complete a calibration phase viewing 128 true images to grasp dataset variability. In the subsequent labeling phase, they identify structural hallucinations in a blinded manner outputs generated by competing methods from identical latent points x1 . Furthermore, feature-based metrics notoriously misalign with human perception, often favoring texture over structure [Geirhos et al., 2018], a flaw amplified in DMs where human preference remains the ultimate ground truth [Stein et al., 2023, Jayasumana et al., 2024]. Addressing this, we expand our study to evaluate User Preference (UP) on natural images. For each latent point, annotators select the highest-quality generations, with UP reported as the overall selection frequency per method. Results. Quantitative and qualitative results are summarized in table 1 and fig. 4. On GaussianGrid, IQ confirms its theoretical validity: it minimally deviates from the baseline while maximally reducing 2 https://huggingface.co/1aurent/ddpm-mnist

7

Figure 4: Qualitative comparison of all sampling methods across six datasets. AAM is omitted on GaussianGrid due to incompatible network architecture. IQ∗ denotes a variant of IQ without filtering (q = 0). For high-resolution results on GaussianGrid, see fig. 16. hallucinations by eliminating mode interpolations. Figure 4 also highlights the role of dynamic filtering; without it, IQ over-squeezes samples into modes, whereas filtering halts correction appropriately. Crucially, IQ succeeds even when true data LID equals the ambient dimension, confirming our probabilistic interpretation. Conversely, DG achieves comparable HR but samples unevenly due to classifier bias, while RODSSAS fails to meaningfully reduce hallucinations despite matching RODSCAS in MMD. On SimpleShapes, IQ reduces over 60% of hallucinations. We hypothesize the concurrent improvement in feature-based metrics stems from the specific nature of these failures: generating extra shapes creates natural outliers that inflate FID. Since these anomalies dominate baseline generation, their removal naturally increases valid diversity. This is corroborated on MNIST, where FID and diversity strongly correlate with HR. Despite extensive tuning, baselines struggle on these simple sets. This validates our LID formulation: an optimal DM must bound its factors of variation, and while IQ directly suppresses unnecessary degrees of freedom, other methods cannot effectively enforce this constraint. On natural images, where human-centric UP and HR are the primary performance indicators, differences in feature-based metrics shrink. Crucially, on 11kHands, IQ reduces hallucinations by almost 70% and achieves almost 70% UP, vastly outperforming baselines. The accompanying minor regressions in FID and diversity corroborate findings [Stein et al., 2023, Jayasumana et al., 2024] that such metrics fail to accurately capture human preference. On FFHQ and AFHQV2, performance gaps narrow. While users preferred IQ slightly more, overlapping confidence intervals (CIs) indicate no statistically significant UP difference, which annotators also attributed to the datasets’ lower resolutions. Nonetheless, IQ successfully cuts hallucinations by around 50% on FFHQ and 15% on the highly challenging AFHQV2, where IQ’s HR CI overlaps with RODS and baseline sampling despite a better average. Hallucinations in medical diagnosis. Generative hallucinations pose severe risks in real-world decision-making, particularly in medical imaging where models restore high-frequency details to enable low-dose radiation [Wang et al., 2025]. To demonstrate IQ’s practical capabilities, we evaluate hallucination reduction in low-dose computed tomography (LDCT) reconstruction. We formulate the inverse problem y = Ax + σϵ, where x represents ground truth RSNA brain CT scans [Anouk Stein et al., 2019], A is the forward projection matrix (X-ray line integrals), ϵ ∼ N (0, I), and y is the measurement sinogram. By zeroing the singular values of A below a threshold ξ and adding noise, we simulate low-dose sparse-view CT. We employ System-Embedded Diffusion Bridges [SDB; Sobieski et al., 2026], which solve linear Gaussian problems by embedding measurement parameters into the diffusion coefficients: Ht x = A+ Ax + αt (I − A+ A)x and ⊤ Σt = γt A+ ΣA+ + βt (I − A+ A), utilizing the Moore-Penrose pseudoinverse A+ . Because SDB maps measurements y to reconstructions x via a matrix-valued diffusion process, it allows for direct integration of the evaluated hallucination reduction methods. Throughout this experiment, we strictly adhere to the original SDB experimental setup on RSNA [Sobieski et al., 2026]. Reducing misdiagnosis. In this context, structural hallucinations manifest as mischaracterizations of the patient condition, e.g., fabricated synthetic pathology. To quantify diagnostic safety, we trained a ResNet50 observer to detect five hemorrhage types using 17,948 ground-truth RSNA slices. On a 8

Table 2: Quantitative comparison of all sampling methods applied to SDB on the RSNA dataset. We include standard reconstruction metrics regarding perception and distortion. We highlight mAP and mROC , performance measures of the ResNet50 observer, as indicators of reconstruction errors and hallucinations, where higher values suggest improved anatomical consistency. Method

FID ↓

IV ↑

DSV ↑

PSNR ↑

SSIM ↑

LPIPS ↓

mAP ↑

mROC ↑

Baseline DG AAM RODSCAS RODSSAS IQ

33.3 35.6 35.6 33.3 33.3 33.3

0.066 0.066 0.066 0.067 0.066 0.066

0.264 0.264 0.264 0.264 0.263 0.264

33.54 33.55 33.54 33.54 33.55 33.54

0.898 0.899 0.898 0.898 0.898 0.901

0.0387 0.0388 0.0388 0.0388 0.0387 0.0388

0.27 0.31 0.29 0.29 0.29 0.31

0.85 0.86 0.86 0.86 0.86 0.86

512-slice validation split, the observer achieves 0.38 macro-averaged multi-label Average Precision (mAP) and 0.91 macro-averaged multi-label Area Under the ROC (mROC). When evaluated on baseline SDB reconstructions [Sobieski et al., 2026], performance drops to 0.27 mAP and 0.85 mROC, reflecting quality degradation and generative hallucinations. We adapt all baseline mitigations to the SDB framework, tracking image quality and diversity (FID. IV, DSV), distortion (PSNR, SSIM, LPIPS) [Zhang et al., 2018, Blau and Michaeli, 2018], and diagnostic safety (mAP, mROC) (table 2). Because SDB strictly preserves the measurement range space, image and reconstruction quality metric variations are heavily constrained to the null space and remain mostly flat, although DG and AAM noticeably degrade FID. Diagnostic performance, however, reveals a sharp distinction. While all methods marginally raise mROC to 0.86, IQ and DG increase mAP by roughly 15% (reaching 0.31). Crucially, DG requires an independently trained classifier with direct access to ground-truth hemorrhage labels, granting it an inherent advantage. Despite operating entirely zero-shot, IQ matches DG’s diagnostic gains while strictly preserving reconstruction quality. This underscores IQ’s universal applicability and viability for broader inverse problem frameworks [Chung et al., 2023, Liu et al., 2023a, Luo et al., 2023, Garber and Tirer, 2024, Yue et al., 2024, Zhou et al., 2024]. Ablation studies. Due to limited space, we continue the experimental evaluation in section B, providing ablation studies for each component of IQ (section B.6), runtime comparison (section B.7), qualitative examples (section B.11) and more.

7

Discussion and Limitations

Our work serves as a direct counterpart to Ross et al. [2025], who demonstrate how the collapse of LID values provides a direct indicator of memorized samples. These complementary findings offer empirical evidence for an intuitive understanding of LID: it serves as a proxy for the creativity of a DM, where hallucinations (the model being overly creative) and memorization (a lack of creativity) span the full generative spectrum. Evaluation persists as the primary limitation of IQ and other hallucination reduction methods. While human annotators generally agree when identifying structural failures, disagreements remain common, highlighting the subjective nature of the task. Constructing fully objective, large-scale benchmarks with automated evaluation would streamline the development of more effective methods. Currently, deciding whether the observed drops in DSV and IV, proxy metrics for diversity, in datasets like 11kHands and FFHQ are strictly caused by the removal of hallucinated samples remains highly probable but difficult to definitively isolate. Furthermore, methods intrinsic to the model, such as RODS and IQ, still induce a moderate computational overhead. Discovering cheaper estimators for LID would lead to immediate practical improvements for IQ. Finally, while unconditional generation is the cornerstone of generative modeling, hallucinations pose the highest risk in conditional settings involving real-world decision-making and diagnosis [Antun et al., 2020], highlighting a critical direction for future investigation. 9

Acknowledgments and Disclosure of Funding Work on this project is financially supported by the Polish National Science Centre PRELUDIUM BIS grant No. 2023/50/O/ST6/00301, the Foundation for Polish Science (FNP) grant ‘Centre for Credible AI’ No. FENG.02.01-IP.05-0058/24. The computational resources for this work were provided by the Laboratory of Bioinformatics and Computational Genomics and the High Performance Computing Center of the Faculty of Mathematics and Information Science, Warsaw University of Technology. We also gratefully acknowledge Poland’s High-performance Infrastructure PLGrid ACC Cyfronet AGH for providing computer facilities and support within computational grant no. PLG/2025/018330.

References Mahmoud Afifi. 11k hands: gender recognition and biometric identification using a large dataset of hand images. In Multimedia Tools and Applications, 2019. Sumukh K Aithal, Pratyush Maini, Zachary Lipton, and J Zico Kolter. Understanding hallucinations in diffusion models through mode interpolation. In Advances in neural information processing systems, 2024. Guillaume Alain and Yoshua Bengio. What regularized auto-encoders learn from the data-generating distribution. In The Journal of Machine Learning Research. JMLR, 2014. Michael Samuel Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. In The Eleventh International Conference on Learning Representations, 2023. Eugene L Allgower and Kurt Georg. Numerical continuation methods: an introduction. Springer Science & Business Media, 2012. MD Anouk Stein, Carol Wu, Chris Carr, George Shih, Jayashree Kalpathy-Cramer, Julia Elliott, kalpathy, Luciano Prevedello, MD Marc Kohli, Matt Lungren, Phil Culliton, Robyn Ball, and Safwan Halabi MD. Rsna intracranial hemorrhage detection. https://kaggle.com/competitions/ rsna-intracranial-hemorrhage-detection, 2019. Kaggle. Alessio Ansuini, Alessandro Laio, Jakob H Macke, and Davide Zoccolan. Intrinsic dimension of data representations in deep neural networks. In Advances in Neural Information Processing Systems, 2019. Vegard Antun, Francesco Renna, Clarice Poon, Ben Adcock, and Anders C Hansen. On instabilities of deep learning in image reconstruction and the potential costs of ai. In Proceedings of the National Academy of Sciences. National Academy of Sciences, 2020. Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. In IEEE transactions on pattern analysis and machine intelligence. IEEE, 2013. Mahesh Bhosale, Naresh Kumar Devulapally, Abdul Wasi, Chau Pham, Vishnu Suresh Lokhande, and David Doermann. Variance-guided score regularization for hallucination mitigation in diffusion models, 2026. URL https://openreview.net/forum?id=nY4nULFzDP. Yochai Blau and Tomer Michaeli. The perception-distortion tradeoff. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6228–6237, 2018. Dan G Cacuci, Mihaela Ionescu-Bujor, and Ionel Michael Navon. Sensitivity and uncertainty analysis, volume II: applications to large-scale systems. CRC press, 2005. Yu Cao, Zengqun Zhao, Ioannis Patras, and Shaogang Gong. Temporal score analysis for understanding and correcting diffusion artifacts. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 7707–7716, 2025. Hyunmin Cho, Donghoon Ahn, Susung Hong, Jee Eun Kim, Seungryong Kim, and Kyong Hwan Jin. Tag: Tangential amplifying guidance for hallucination-resistant diffusion sampling. In arXiv, 2025. 10

Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha. Stargan v2: Diverse image synthesis for multiple domains. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020. Hyungjin Chung, Jeongsol Kim, Michael Thompson Mccann, Marc Louis Klasky, and Jong Chul Ye. Diffusion posterior sampling for general noisy inverse problems. In The Eleventh International Conference on Learning Representations, 2023. Prafulla Dhariwal and Alex Nichol. Diffusion models beat gans on image synthesis. In Advances in neural information processing systems, 2021. Carles Domingo-Enrich, Michal Drozdzal, Brian Karrer, and Ricky T. Q. Chen. Adjoint matching: Fine-tuning flow and diffusion generative models with memoryless stochastic optimal control. In The Thirteenth International Conference on Learning Representations, 2025. URL https: //openreview.net/forum?id=xQBRrtQM8u. Bradley Efron. Tweedie’s formula and selection bias. In Journal of the American Statistical Association. Taylor & Francis, 2011. Charles Fefferman, Sanjoy Mitter, and Hariharan Narayanan. Testing the manifold hypothesis. In Journal of the American Mathematical Society, 2016. Shuai Fu, Jian Zhou, Qi Chen, Huang Jing, Huy Anh Nguyen, Xiaohan Liu, Zhixiong Zeng, Lin Ma, Quanshi Zhang, and Qi Wu. Counting hallucinations in diffusion models. In arXiv, 2025. Tomer Garber and Tom Tirer. Image restoration by denoising diffusion models with iteratively preconditioned guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness. In International conference on learning representations, 2018. Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, and Kaiming He. Mean flows for one-step generative modeling. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. Mark Goresky and Robert MacPherson. Stratified morse theory. In Stratified Morse Theory. Springer, 1988. Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in neural information processing systems, 2017. Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in neural information processing systems, 2020. Emiel Hoogeboom, Vıctor Garcia Satorras, Clément Vignac, and Max Welling. Equivariant diffusion for molecule generation in 3d. In International conference on machine learning, pages 8867–8887. PMLR, 2022. Aapo Hyvärinen. Estimation of non-normalized statistical models by score matching. In Journal of Machine Learning Research, 2005. Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar. Rethinking fid: Towards a better evaluation metric for image generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9307–9315, 2024. Hamidreza Kamkari, Brendan L Ross, Rasa Hosseinzadeh, Jesse C Cresswell, and Gabriel LoaizaGanem. A geometric view of data complexity: Efficient local intrinsic dimension estimation with diffusion models. In Advances in Neural Information Processing Systems, 2024. 11

Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4401–4410, 2019. Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusionbased generative models. In Advances in neural information processing systems, 2022. Seunghoi Kim, Chen Jin, Tom Diethe, Matteo Figini, Henry FJ Tregidgo, Asher Mullokandov, Philip Teare, and Daniel C Alexander. Tackling structural hallucination in image translation with local diffusion. In European Conference on Computer Vision. Springer, 2024. Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=a-xFK8Ymz5J. Akim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, and Artem Babenko. Tabddpm: Modelling tabular data with diffusion models. In International conference on machine learning. PMLR, 2023. Alice Ku, Matthew Tivnan, and Dufan Wu. Hallucination analysis of score-based diffusion models for ct denoising through spatial frequency decomposition. In 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI), pages 1–5, 2025. doi: 10.1109/ISBI60581.2025.10981211. Elizaveta Levina and Peter Bickel. Maximum likelihood estimation of intrinsic dimension. In Advances in neural information processing systems, 2004. Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In International Conference on Learning Representations, 2023. Guan-Horng Liu, Arash Vahdat, De-An Huang, Evangelos Theodorou, Weili Nie, and Anima Anandkumar. I2SB: Image-to-Image Schrödinger Bridge. In International Conference on Machine Learning, 2023a. Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Representations, 2023b. Aaron Lou, Chenlin Meng, and Stefano Ermon. Discrete diffusion modeling by estimating the ratios of the data distribution. In Forty-first International Conference on Machine Learning, 2024. Rui Lu, Runzhe Wang, Kaifeng Lyu, Xitai Jiang, Gao Huang, and Mengdi Wang. Towards understanding text hallucination of diffusion models via local generation bias. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/ forum?id=SKW10XJlAI. Wenquan Lu, Yufei Xu, Jing Zhang, Chaoyue Wang, and Dacheng Tao. Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 7085–7093, 2024. Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017. Ziwei Luo, Fredrik K Gustafsson, Zheng Zhao, Jens Sjölund, and Thomas B Schön. Image restoration with mean-reverting stochastic differential equations. In International Conference on Machine Learning, 2023. Xingjun Ma, Bo Li, Yisen Wang, Sarah M Erfani, Sudanthi Wijewickrema, Grant Schoenebeck, Dawn Song, Michael E Houle, and James Bailey. Characterizing adversarial subspaces using local intrinsic dimensionality. In International Conference on Learning Representations, 2018. Zichen Miao, Jiang Wang, Ze Wang, Zhengyuan Yang, Lijuan Wang, Qiang Qiu, and Zicheng Liu. Training diffusion models towards diverse image generation with reinforcement learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10844–10853, 2024. 12

Supreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta, Saayan Mitra, and Minh Hoai. Handiffuser: Text-to-image generation with realistic hand appearances. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2468–2479, 2024. Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International conference on machine learning, pages 8162–8171. PMLR, 2021. Trevine Oorloff, Yaser Yacoob, and Abhinav Shrivastava. Mitigating hallucinations in diffusion models through adaptive attention modulation. In arXiv, 2025. Yong-Hyun Park, Mingi Kwon, Jaewoong Choi, Junghyo Jo, and Youngjung Uh. Understanding the latent space of diffusion models through the lens of riemannian geometry. In Advances in Neural Information Processing Systems, volume 36, pages 24129–24142, 2023. Jakiw Pidstrigach. Score-based generative models detect manifolds. In Advances in Neural Information Processing Systems, 2022. Phil Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein. The intrinsic dimension of images and its impact on learning. In International Conference on Learning Representations, 2021. Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International conference on machine learning. PMLR, 2021. Weiming Ren, Raghav Goyal, Zhiming Hu, Tristan Ty Aumentado-Armstrong, Iqbal Mohomed, and Alex Levinshtein. Hallucination score: Towards mitigating hallucinations in generative image super-resolution. In arXiv, 2025. Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. Highresolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, pages 234–241. Springer, 2015. Brendan Leigh Ross, Hamidreza Kamkari, Tongzi Wu, Rasa Hosseinzadeh, Zhaoyan Liu, George Stein, Jesse C. Cresswell, and Gabriel Loaiza-Ganem. A geometric framework for understanding memorization in generative models. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=aZ1gNJu8wO. Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf, Thomas Brox, and Peter Gehler. Towards total recall in industrial anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14318–14328, 2022. Antonio Sclocchi, Alessandro Favero, and Matthieu Wyart. A phase transition in diffusion models reveals the hierarchical nature of data. In Proceedings of the National Academy of Sciences. National Academy of Sciences, 2025. Chuancheng Shi, Shiming Guo, Ke Shui, Yixiang Chen, and Fei Shen. Sgmhand: Structure-guided modulation for structure-aware hand inpainting. In Proceedings of the AAAI Conference on Artificial Intelligence, 2026. Bartlomiej Sobieski, Matthew Tivnan, Yuang Wang, Siyeop yoon, Pengfei Jin, Dufan Wu, Quanzheng Li, and Przemyslaw Biecek. System-embedded diffusion bridge models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. URL https://openreview.net/ forum?id=cipx3rwfWp. Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, 2015. 13

Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021. Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In International Conference on Machine Learning, 2023. Jan Pawel Stanczuk, Georgios Batzolis, Teo Deveney, and Carola-Bibiane Schönlieb. Diffusion models encode the intrinsic dimension of data manifolds. In Proceedings of the 41st International Conference on Machine Learning, 2024. George Stein, Jesse Cresswell, Rasa Hosseinzadeh, Yi Sui, Brendan Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L Caterini, Eric Taylor, and Gabriel Loaiza-Ganem. Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models. In Advances in Neural Information Processing Systems, 2023. Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In International Conference on Machine Learning, pages 3319–3328. PMLR, 2017. Piotr Tempczyk, Rafał Michaluk, Lukasz Garncarek, Przemysław Spurek, Jacek Tabor, and Adam Golinski. Lidl: Local intrinsic dimension estimation using approximate likelihood. In International Conference on Machine Learning, pages 21205–21231. PMLR, 2022. Yiqi Tian, Pengfei Jin, Mingze Yuan, Na Li, Bo Zeng, and Quanzheng Li. RODS: Robust optimization inspired diffusion sampling for detecting and reducing hallucination in generative models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=fhuqIxoPcr. Matthew Tivnan, Siyeop Yoon, Zhennong Chen, Xiang Li, Dufan Wu, and Quanzheng Li. Hallucination index: An image quality metric for generative reconstruction models. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 449–458. Springer, 2024. Kostas Triaridis, Alexandros Graikos, Aggelina Chatziagapi, Grigorios G Chrysos, and Dimitris Samaras. Mitigating diffusion model hallucinations with dynamic guidance. In arXiv, 2025. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, 2017. Pascal Vincent. A connection between score matching and denoising autoencoders. In Neural computation. MIT Press, 2011. Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, Dhruv Nair, Sayak Paul, William Berman, Yiyi Xu, Steven Liu, and Thomas Wolf. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022. Binxu Wang and Cengiz Pehlevan. An analytical theory of spectral bias in the learning dynamics of diffusion models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. Shanshan Wang, Xuanru Zhou, Cheng Li, Shuqiang Wang, Ye Li, Tao Tan, and Hairong Zheng. Generative artificial intelligence in medical imaging: Foundations, progress, and clinical translation. In Research, volume 8, page 1029. AAAS, 2025. Eric Yeats, Aaron Jacobson, Darryl Hannan, Yiran Jia, Timothy Doster, Henry Kvinge, and Scott Mahan. A connection between score matching and local intrinsic dimension. In NeurIPS 2025 Workshop on Structured Probabilistic Inference & Generative Modeling, 2025. Conghan Yue, Zhengwei Peng, Junlong Ma, Shiyan Du, Pengxu Wei, and Dongyu Zhang. Image restoration through generalized ornstein-uhlenbeck bridge. In International Conference on Machine Learning, 2024. 14

Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018. Linqi Zhou, Aaron Lou, Samar Khanna, and Stefano Ermon. Denoising diffusion bridge models. In International Conference on Learning Representations, 2024.

15

Appendix A Theoretical results

16

A.1 Proposition 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

16

A.2 Proposition 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

17

A.3 Theorem 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

18

A.4 Theorem 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

20

B Extended experiments

21

B.1 Experimental setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

21

B.2 User study . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

21

B.3 Filter performance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

21

B.4 LID performance curves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

22

B.5 Extended results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

23

B.6 Ablation studies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

23

B.7 Runtime comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

25

B.8 Hyperparameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

26

B.9 Pseudocode . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

26

B.10 Resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

27

B.11 Qualitative examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

27

A

Theoretical results

A.1

Proposition 1

Proposition 1. Let G θ be the deterministic generative mapping of a diffusion model from the latent noise space to the induced data manifold Mθ , evaluated at an initial state x1 . Assume the perturbation scale β > 0 is sufficiently small such that a first-order approximation holds. Let σ1 ≥ σ2 ≥ · · · ≥ σn ≥ 0 denote the singular values of the generator’s Jacobian ∇x1 G θ (x1 ). Let LIDθ (x0 ) denote the local intrinsic dimensionality of the generated data point x0 = G θ (x1 ) with respect to the induced manifold Mθ . The Local Manifold Instability (LMI) is approximately equal to the scaled sum of its top ⌊LIDθ (x0 )⌋ squared singular values: ⌊LIDθ (x0 )⌋

LMI(x1 ) ≈ β

2

X

σi2 .

(11)

i=1

Proof. By definition, the Local Manifold Instability (LMI) at an initial state x1 is given by LMI(x1 ) ≜ Tr(Covϵ (G θ (x1 + ϵ))), where ϵ ∼ N (0, β 2 I) for a perturbation scale β > 0. Applying a first-order Taylor expansion to the generative mapping G θ around x1 yields: 2

G θ (x1 + ϵ) = G θ (x1 ) + ∇x1 G θ (x1 )ϵ + O(∥ϵ∥2 ).

(12)

Substituting this expansion into the covariance expression and recognizing that G θ (x1 ) is a deterministic constant with respect to ϵ, it vanishes from the computation. Letting J = ∇x1 G θ (x1 ), we obtain: Covϵ (G θ (x1 + ϵ)) = J Covϵ (ϵ)J⊤ + O(β 3 ) = β 2 JJ⊤ + O(β 3 ). (13) Applying the trace operator to calculate the total variance (LMI): 2

LMI(x1 ) = Tr(β 2 JJ⊤ ) + O(β 3 ) = β 2 ∥J∥F + O(β 3 ). 16

(14)

For a sufficiently small noise scale β, the linear regime dominates and the higher-order remainder O(β 3 ) can be safely neglected, yielding the first-order approximation: 2

LMI(x1 ) ≈ β 2 ∥J∥F .

(15)

Our next objective is to relate this squared Frobenius norm to the effective spectral properties of the mapping. The squared Frobenius norm of any matrix is algebraically equivalent to the sum of all its squared singular values: n X 2 ∥J∥F = σi2 . (16) i=1

We address the geometric dimensionality of the generation process. The generator G θ continuously maps the n-dimensional ambient noise space into a lower-dimensional induced data manifold Mθ . The local intrinsic dimensionality at the specific generated data point x0 , denoted as LIDθ (x0 ), represents the effective number of independent, structurally significant directions in which the model’s output varies. While deep neural networks practically possess full-rank Jacobians (meaning all σi > 0), assumption 1 implies that the mapping highly contracts the space along off-manifold directions. Consequently, the singular values corresponding to the remaining n − ⌊LIDθ (x0 )⌋ directions are infinitesimally small, representing negligible off-manifold numerical noise rather than structurally meaningful degrees of freedom. By treating these tail singular values as negligible (σi ≈ 0 for i > ⌊LIDθ (x0 )⌋), we can approximate the full spectral sum by truncating it at the effective dimensionality: n X

⌊LIDθ (x0 )⌋

X

σi2 ≈

i=1

σi2 .

(17)

i=1

Finally, substituting this spectral truncation back into eq. (15) yields the completed theoretical approximation: ⌊LIDθ (x0 )⌋ X 2 LMI(x1 ) ≈ β σi2 . (18) i=1

This demonstrates that, under a localized linear regime, the spatial instability of the generator is closely approximated by the variance of the injected noise scaled by the structurally significant spectral mass of the local manifold. A.2

Proposition 2

Proposition 2. Let Mθ be the LIDθ (x0 )-dimensional manifold induced by a fully trained diffusion model parameterized by θ, embedded in an n-dimensional ambient space, with probability density pθ (x0 ). Assume pθ (x0 ) is locally uniform on Mθ with negligible curvature at the limit t → 0. Let the forward process be defined by an arbitrary drift matrix Ht and a strictly positive-definite covariance matrix Σt . As t → 0, the expected DSM loss evaluates exactly to the local intrinsic dimension LIDθ (x0 ), and the expected ISM loss evaluates exactly to − 21 (n − LIDθ (x0 )). 1

Proof. Let the forward diffusion state be xt = Ht x0 + Σt2 ϵ, where ϵ ∼ N (0, I) and x0 ∈ Mθ . By Tweedie’s formula applied to the model’s induced distribution, the predicted score function sθ (xt ) defines the model’s posterior mean estimate x̂θ0 (xt ) via: θ sθ (xt ) = −Σ−1 t (xt − Ht x̂0 (xt )).

(19)

1 2

The predicted noise is defined as ϵθ (xt ) ≜ −Σt sθ (xt ). Substituting the predicted score yields: −1

ϵθ (xt ) = Σt 2 (xt − Ht x̂θ0 (xt )). (20) h i 2 To evaluate the DSM expectation Eϵ ∥ϵ − ϵθ (xt )∥2 , we express the target forward noise as ϵ = −1

Σt 2 (xt − Ht x0 ). The residual error inside the norm becomes: −1

ϵ − ϵθ (xt ) = Σt 2 Ht (x̂θ0 (xt ) − x0 ). 17

(21)

As t → 0, the model’s posterior mean x̂θ0 (xt ) acts as the optimal projection of the noisy state back onto the locally flat model-induced manifold Mθ . Because the forward process introduces spatial 2 drift Ht and covariance Σt , the projected state minimizes the Mahalanobis distance ∥xt − Ht x∥Σ−1 t for x ∈ Mθ . Let Tx0 Mθ be the tangent space of the model’s manifold at x0 . The projection onto this mapped tangent space Ht Tx0 Mθ under the Mahalanobis inner product ⟨u, v⟩Σ−1 = u⊤ Σ−1 t v is governed t

t ,Ht by an oblique projection matrix PΣ . Therefore, the deviation of the model’s posterior mean T mapped through the drift is exactly the oblique projection of the forward perturbation: 1

t ,Ht t ,Ht Ht (x̂θ0 (xt ) − x0 ) ≈ PΣ (xt − Ht x0 ) = PΣ Σt2 ϵ. T T

(22)

Substituting this geometric relation back into eq. (21), the DSM expectation becomes:   2 1 −1 t ,Ht 2 Eϵ [LDSM (x0 , t, θ)] = Eϵ Σt 2 PΣ Σ ϵ . t T 2

(23)

1

−1

t ,Ht Σt2 . This operator represents the tangent space Define the transformed operator MT ≜ Σt 2 PΣ T projection mapped into the whitened isotropic noise space. We prove MT is a standard Euclidean t ,Ht orthogonal projector by checking symmetry. By definition, PΣ is self-adjoint with respect to the T Σt ,Ht ⊤ −1 −1 Σt ,Ht Mahalanobis inner product, meaning (PT ) Σt = Σt PT . Thus: 1

−1

1

−1

−1

1

Σt ,Ht ⊤ Σt ,Ht t ,Ht 2 M⊤ ) Σt 2 = Σt2 (Σ−1 Σt )Σt 2 = Σt 2 PΣ Σt2 = MT . T = Σt (PT t PT T

(24)

2 Because MT is symmetric (M⊤ T = MT ) and idempotent (MT = MT ), it is a valid Euclidean orthogonal projector. Its rank is strictly preserved under similarity transformations and the nont ,Ht singular linear mapping Ht , meaning rank(MT ) = rank(PΣ ) = LIDθ (x0 ). T

Applying the standard identity for the expected squared norm of an orthogonally projected standard Gaussian, we evaluate the trace:   Eϵ [LDSM (x0 , t, θ)] = Eϵ ϵ⊤ M⊤ (25) T MT ϵ = Tr(MT ) = LIDθ (x0 ). Finally, by Stein’s Lemma, the generalized score matching objectives satisfy the exact algebraic identity Eϵ [LDSM (x0 , t, θ)] = n+2Eϵ [LISM (x0 , t, θ)]. Rearranging this relationship and substituting the evaluated DSM expectation directly yields: 1 Eϵ [LISM (x0 , t, θ)] = − (n − LIDθ (x0 )), 2

(26)

proving geometric and algebraic consistency for arbitrary drift and covariance schedules exclusively within the model-induced geometry. A.3

Theorem 1

We include a more general formulation of theorem 1 that is not limited to isotropic noise as t → 0. This allows us to provide theoretical guarantees also for frameworks like SDB. >τ Theorem 1. Let x0 = G θ (x1 ). Assume that G θ (x1 ) is decomposed as G θ = G ≤τ θ ◦ G θ for ≤τ >τ sufficiently small time τ > 0, where G θ : x1 7→ xτ and G θ : xτ 7→ x0 . Let the singular values >τ >τ of the macroscopic flow Jacobian ∇x1 G >τ θ (x1 ) be monotonically ordered as σ1 ≥ σ2 ≥ · · · ≥ 2

σn>τ ≥ 0. Let K = ∇xτ G ≤τ ≥ 1 represent the squared spectral norm of the terminal θ (xτ ) 2 oblique projection. For sufficiently small β > 0, it is true that LMI(x1 ) ⪅ Kβ 2

(x̂θ ⌊LIDθX 0 (xτ ))⌋ i=1

18

(σi>τ )2 .

(27)

>τ Proof. Applying the multivariate chain rule to the decomposed generative flow G θ = G ≤τ θ ◦ Gθ , the total generator Jacobian J = ∇x1 G θ (x1 ) is the matrix product of the stage-wise Jacobians: >τ J = ∇xτ G ≤τ θ (xτ )∇x1 G θ (x1 ) = J≤τ J>τ .

(28)

Recall from proposition 1 the first-order approximation for LMI for sufficiently small β, which evalu2 ates to the scaled squared Frobenius norm of the total Jacobian: LMI(x1 ) ≈ β 2 ∥J∥F . Substituting the chain rule expansion into this norm yields: 2

2

∥J∥F = ∥J≤τ J>τ ∥F .

(29)

For sufficiently small τ , the terminal integration step G ≤τ θ behaves as a projection onto the induced manifold Mθ . In the general case of anisotropic terminal noise (such as the embedded covariance structure in SDB), this operation acts as an oblique projector rather than a standard Euclidean orthogonal projector. Consequently, its Jacobian J≤τ possesses exactly ⌊LIDθ (x0 )⌋ non-zero singular values. We denote these terminal singular values as σi≤τ and bound them by the squared spectral norm of the oblique 2 projector: (σi≤τ )2 ≤ ∥J≤τ ∥2 ≜ K. The remaining n − ⌊LIDθ (x0 )⌋ singular values are strictly equal to 0 (annihilating variance along normal directions). To decouple the two generative regimes, we apply the singular value product inequality (von Neumann trace inequality) to bound the Frobenius norm of the product. The sum of the squared singular values of the product is bounded by the sum of the products of their individual squared singular values: 2

∥J≤τ J>τ ∥F ≤

n X (σi≤τ )2 (σi>τ )2 .

(30)

i=1

√ Substituting the bounded spectrum of the terminal oblique projector (σi≤τ ≤ K for i ≤ ⌊LIDθ (x0 )⌋ and 0 otherwise) acts as a strict algebraic cutoff for the summation, scaled by the projection constant K: ⌊LIDθ (x0 )⌋ n X X (σi≤τ )2 (σi>τ )2 ≤ K (σi>τ )2 . (31) i=1

i=1

Substituting this strict upper bound back into the LMI approximation provides an intermediate theoretical guarantee: ⌊LIDθ (x0 )⌋ X LMI(x1 ) ⪅ Kβ 2 (σi>τ )2 . (32) i=1

Next, we address the topological equivalence of evaluating the intrinsic dimensionality at the unobserved terminal state x0 versus the predicted posterior mean x̂θ0 (xτ ). At any intermediate time τ > 0, the state xτ is corrupted by forward process variance. Consequently, xτ has full support in the n-dimensional ambient space, meaning xτ ∈ / Mθ . To evaluate the topological complexity of the target manifold from xτ , we compute the Tweedie estimate derived from the predicted score function, x̂θ0 (xτ ). By the fundamental theorem of estimation, this conditional expectation acts as the Minimum Mean Square Error (MMSE) estimator. Because Mθ generally possesses curvature, the expectation of points on the manifold strictly resides within the convex hull of Mθ , meaning x̂θ0 (xτ ) ∈ / Mθ in the general case. However, we establish asymptotic equivalence between this estimation and the terminal generative mapping. The terminal state is obtained by integrating the Probability Flow ODE from τ to 0: x0 = G ≤τ θ (xτ ). As the integration interval τ → 0, the distribution concentrates locally, the exact ODE integration converges to the analytic score-based projection, and the Euclidean distance to the manifold vanishes: lim dist(x̂θ0 (xτ ), Mθ ) = 0. (33) τ →0

This implies that for a sufficiently small τ , the Tweedie estimate x̂θ0 (xτ ) and the true terminal state x0 reside within the exact same infinitesimal local neighborhood. 19

Under the assumption that Mθ is a smooth manifold, the local intrinsic dimensionality LIDθ (x) is a locally constant topological invariant. Because x̂θ0 (xτ ) approaches x0 sufficiently close, their geometric dimensionalities strictly align: LIDθ (x̂θ0 (xτ )) = LIDθ (x0 ).

(34)

Substituting this topological equivalence into the intermediate bound (eq. (32)) yields the completed operational approximation: LMI(x1 ) ⪅ Kβ 2

(x̂θ ⌊LIDθX 0 (xτ ))⌋

(σi>τ )2 .

(35)

i=1

A.4

Theorem 2

Theorem 2. Let pθt (xt ) be the model-induced marginal density at time t ≤ τ . Assume the network parameterizes the exact score of its induced distribution, such that sθ (xt ) = ∇xt log pθt (xt ). We define the ideal geometrically-guided distribution by re-weighting the marginal over the true terminal t states: pθ,λ (xt ) ∝ pθt (xt )Epθ (x0 |xt ) [exp(−λt LIDθ (x0 ))]. Assume τ is sufficiently small such that t the posterior variance σt2 → 0. Under these conditions, guiding the reverse diffusion with the modified score s̃θ (xt ) = sθ (xt ) − λt ∇xt LIDθ (x̂θ0 (xt )) is mathematically equivalent to sampling t from the ideal distribution pθ,λ (xt ), shifting probability mass toward lower-dimensional target t strata. Proof. We define the ideal guided target density by re-weighting the model-induced marginal pθt (xt ) with the expected terminal energy: 1 t pθ,λ (xt ) = θ pθt (xt )Epθ (x0 |xt ) [exp(−λt LIDθ (x0 ))] , (36) t Zt where Ztθ is the partition function. Taking the logarithm and differentiating with respect to xt yields the ideal modified score: t (xt ) = ∇xt log pθt (xt ) + ∇xt log Epθ (x0 |xt ) [exp(−λt LIDθ (x0 ))] . ∇xt log pθ,λ t

(37)

To render this tractable, we analyze the domain t ≤ τ . As the noise scale diminishes (σt2 → 0), the model’s posterior distribution pθ (x0 |xt ) concentrates into a Dirac delta distribution centered at the Tweedie estimate x̂θ0 (xt ). Consequently, the log-expectation of the energy converges exactly to the energy of the expectation: lim log Epθ (x0 |xt ) [exp(−λt LIDθ (x0 ))] = −λt LIDθ (x̂θ0 (xt )).

σt2 →0

(38)

While the theoretical intrinsic dimension of a smooth manifold is a locally constant integer (which would yield a trivial zero gradient), our operational definition of LIDθ relies on the generalized score-matching estimator LDSM (x, t, θ). Because this estimator evaluates the model’s spatial error, it is a continuous, non-linear, and differentiable function of its inputs, parameterized by the neural network θ. Therefore, applying the spatial gradient operator yields a well-defined, non-zero vector field: ∇xt E(xt ) = ∇xt LIDθ (x̂θ0 (xt )) ≈ ∇xt LDSM (x̂θ0 (xt ), t, θ). (39) Substituting this limit and the assumed model score function (sθ (xt ) = ∇xt log pθt (xt )) into the target gradient yields the operational guidance rule: s̃θ (xt ) = sθ (xt ) − λt ∇xt LIDθ (x̂θ0 (xt )).

(40)

By substituting s̃θ (xt ) into the reverse probability flow ODE, the generated trajectory simulates t sampling from the ideal terminal-weighted distribution pθ,λ (xt ). Because this guided distribution is t θ multiplicatively bounded by the original prior pt (xt ), the integration strictly preserves the topological validity dictated by the model-induced distribution while explicitly minimizing the local intrinsic dimension of the expected final sample via the continuous gradient signal. 20

B

Extended experiments

B.1

Experimental setup

As the domain of hallucination reduction currently lacks a single unified evaluation scheme, we reimplement all of the methods within our codebase. For DG, we train 6 noisy classifiers from scratch, one for each dataset, using ResNets of different sizes adapted to the complexity of the data for image datasets and a 3-layer MLP for GaussianGrid. For AAM, we adapt the temperature gradient descent scheme for all DMs and train a separate PatchCore anomaly detection model [Roth et al., 2022] for each dataset. To find optimal hyperparameters, we perform a large grid search for AAM, DG and RODS on all datasets, except 11kHands, FFHQ and AFHQV2 for RODS, where we use the originally proposed ones [Tian et al., 2025]. We report the hyperparameters of all methods in section B.8.

B.2

User study

For each of the five image datasets, the user study consisted of two subsequent phases: the calibration phase and the labeling phase. The former involved presenting 128 images from the original dataset in batches, aiming to calibrate the user’s perception of structures and correlations that appear in the true data distribution. Although this phase might initially seem redundant, fig. 5 presents several examples of hand images from the 11kHands dataset that contain atypical hand and finger placements, highlighting the diversity and complex relationships that should not be labeled as hallucinations. Hence, this phase ensures that annotators are aware of what a structural hallucination is in each considered case. The labeling phase involved presenting the annotator with a grid of images, each generated with a different sampling method from the same starting point x1 of the diffusion process. The placement of all methods was fixed across grids, and the study was blinded; i.e., the annotators were not given the names of the methods, which were replaced by Image n captions, with n being the method index (see fig. 5). The user study relied on two independent annotators, both of whom were machine learning researchers. The task was first briefly described to them to ensure that the term structural hallucinations was understood correctly. At each iteration of the labeling phase, the annotators first had to indicate the indices of images that contained a structural hallucination (which could also be none) and then (for natural image datasets) the indices of images that were of the highest quality according to their own perception. Because the labeling process was highly time-consuming, the study utilized a different number of images for each dataset, chosen a priori based on our visual assessment and the subjective advantage of IQ. This approach ensured that statistically significant results would be obtained (with a high probability) using a minimal number of images to reduce the burden on the annotators, which could have otherwise negatively impacted the study. The numbers of starting points (i.e., the number of grids presented to each annotator) were as follows: 128 for 11kHands, 1024 for AFHQV2 and FFHQ each, 512 for MNIST and SimpleShapes each. Each annotator spent approximately 8 hours in total (across two sessions) labeling all of the provided samples. At last, we note that the initial empirical investigation (section 3) uses the labels provided by one of the annotators on the 11kHands dataset.

B.3

Filter performance

We compare three filters (TVF, LMI, and LID) in terms of separability and classification performance of correct and hallucinated samples in table 3. For hyperparameters of each filter, see table 4. LID consistently outperforms both TVF and LMI in almost all cases. LMI outperforms TVF in all cases, except for PR AUC and Cohen’s d on SimpleShapes, and outperforms LID in some cases on AFHQV2 and MNIST. 21

Figure 5: Examples of the user study interface from the 11kHands dataset, depicting calibration and labeling phases.

Table 3: Extended comparison of three filters (TVF, LMI, and LID) in terms of separability and classification performance on natural and synthetic datasets.

AFHQV2

FFHQ

Method PR AUC ↑ ROC AUC ↑ Cohen’s d ↑ PR AUC ↑ ROC AUC ↑ Cohen’s d ↑ TVF LMI LID

0.12 0.16 0.22

0.55 0.66 0.57

0.18 0.31 0.41

0.16 0.17 0.20

MNIST

0.53 0.60 0.60

0.17 0.17 0.43

SimpleShapes

Method PR AUC ↑ ROC AUC ↑ Cohen’s d ↑ PR AUC ↑ ROC AUC ↑ Cohen’s d ↑ TVF LMI LID

B.4

0.58 0.62 0.63

0.60 0.67 0.66

0.40 0.57 0.42

0.67 0.60 0.67

0.58 0.59 0.61

0.34 0.23 0.34

LID performance curves

We visualize how the classification and separability performance of the LID filter changes across time on all image datasets in fig. 6. Crucially, the optimal timestep is very small in all cases, highlighting the importance of the theoretical assumptions. 22

Table 4: Summary of hyperparameters for the three evaluated filters (LID, LMI, and TVF) across natural and synthetic datasets. Time variables (t, t1 , t2 ) are normalized to the [0, 1] interval as decimal fractions. LID LMI TVF Dataset t β t1 t2

AUC Score (PR & ROC)

0.80 0.64 0.48 0.32 0.16 0.00

0.0

0.2

0.4

0.6 0.8 1.0 t Separability across time (MNIST) ROC AUC

0.0

0.2

PR AUC

0.4

t

0.6

Cohen's d

0.8

1.0

0.70 0.56 0.42 0.28 0.14 0.00

0.60 0.48 0.36 0.24 0.12 0.00

0.80 0.64 0.48 0.32 0.16 0.00

0.05 0.05 0.9 0.155 0.063

0.125 0.1 0.95 0.158 0.222

Separability across time (FFHQ) ROC AUC PR AUC Cohen's d

0.0

0.2

0.6 0.8 1.0 t Separability across time (SimpleShapes)

0.4

ROC AUC PR AUC Cohen's d

0.0

0.2

0.4

t

0.6

0.70 0.56 0.42 0.28 0.14 0.00

Cohen's d

ROC AUC PR AUC Cohen's d

0.65 0.52 0.39 0.26 0.13 0.00

Cohen's d AUC Score (PR & ROC)

Separability across time (AFHQV2)

Cohen's d AUC Score (PR & ROC)

AUC Score (PR & ROC)

0.65 0.52 0.39 0.26 0.13 0.00

2.0 2.0 2.0 0.1 5.0

0.60 0.48 0.36 0.24 0.12 0.00

Cohen's d

0.05 0.00625 0.0125 0.001 0.37

11kHands FFHQ AFHQV2 MNIST SimpleShapes

0.8

1.0

Figure 6: Performance curves of the LID filter across diffusion timesteps. The optimal timestep for classification and separability metrics is indicated by a vertical line in each case. B.5

Extended results

We provide the average HR and UP in all considered cases, together with 95% CIs, in table 5. These results are based on the annotations of two independent experts (as mentioned in section B.2) using varying sample sizes for each dataset (for reasons previously described): 128 for 11kHands, 1024 for AFHQV2 and FFHQ each, and 512 for MNIST and SimpleShapes each. The results for GaussianGrid are based on 16384 samples for each method. We also note that we generate the SimpleShapes dataset procedurally by using the implementation of Aithal et al. [2024], which divides a black 64×64 plane into three vertical bands of equal width and then independently inserts a triangle (leftmost band), a square (middle band) and a rhomb (rightmost band), each with 50% probability, into a random position within its band. Due to this procedural approach, we define hallucinations as improper placement of an object (e.g., a triangle in the middle band) or the generation of more than 3 objects. B.6

Ablation studies

In the following subsections, we provide an ablation study of each hyperparameter of IQ on the 11kHands dataset, where hallucinations are easiest to judge subjectively. For each hyperparameter combination, we manually labeled 128 samples, resulting in a total of 4096 labeled cases, and reported feature-based metrics using 2048 samples. Because the original annotators did not take part in these studies, the HR differs from the main results in table 1. In each study, we freeze all of the 23

Table 5: Detailed 95% confidence intervals for User Preference (UP) and Hallucination Ratio (HR) evaluated on high-resolution and synthetic datasets. Values are formatted as Mean [Lower Bound, Upper Bound]. 11kHands FFHQ AFHQV2 UP (%) ↑

HR (%) ↓

39.8 [34.8, 44.9] 39.5 [34.6, 44.3] 40.6 [35.7, 45.5] 40.2 [35.3, 45.1] 41.0 [35.8, 46.2] 68.0 [63.1, 72.8]

29.3 [24.1, 35.1] 29.7 [24.4, 35.6] 29.3 [24.1, 35.1] 25.8 [20.8, 31.5] 29.7 [24.4, 35.6] 9.0 [6.1, 13.1]

Method

HR (%) ↓

UP (%) ↑

HR (%) ↓

45.3 [42.8, 47.9] 8.2 [7.0, 9.4] 44.1 [41.3, 47.0] 9.8 [8.6, 11.1] 45.7 [43.2, 48.2] 10.2 [9.0, 11.6] 45.3 [42.8, 47.9] 7.7 [6.6, 8.9] 45.3 [42.8, 47.9] 8.2 [7.0, 9.4] 46.1 [43.7, 48.4] 4.2 [3.4, 5.1]

41.8 [38.5, 45.0] 6.9 [5.9, 8.1] 40.2 [36.4, 44.0] 8.6 [7.5, 9.9] 42.2 [39.0, 45.4] 14.3 [12.9, 15.9] 42.2 [39.0, 45.4] 6.9 [5.9, 8.1] 42.2 [39.0, 45.4] 6.6 [5.6, 7.7] 42.6 [39.5, 45.7] 5.9 [4.9, 7.0]

MNIST

SimpleShapes

GaussianGrid

Method

HR (%) ↓

HR (%) ↓

HR (%) ↓

Baseline DG AAM RODS-CAS RODS-SAS IQ

37.3 [33.2, 41.6] 23.6 [20.2, 27.5] 34.8 [30.8, 39.0] 41.8 [37.6, 46.1] 55.3 [50.9, 59.5] 10.2 [7.8, 13.1]

25.8 [19.0, 34.0] 27.3 [20.4, 35.6] 55.5 [46.8, 63.8] 30.5 [23.2, 38.9] 28.9 [21.8, 37.3] 9.4 [5.4, 15.7]

20.2 [19.5, 20.8] 11.1 [10.6, 11.6] 10.4 [9.9, 10.9] 19.4 [18.8, 20.0] 8.9 [8.4, 9.3]

1.0 0.8

IV vs HR over λ 0.055 0.050

0.8

0.2

0.4

λ

0.6

0.8

DSV HR

0.22

1.0 0.8

0.20

0.6

0.045

0.6

0.4

0.040

0.4

0.16

0.4

0.2

0.14

0.2

1.00.0

0.12

0.2 0.0

DSV vs HR over λ

1.0

IV HR

HR IV

FID

22 21 20 19 18 17 16

FID HR

1.00.0

0.6

0.035 0.030 0.0

0.2

0.4

λ

0.6

0.8

0.18

HR

FID vs HR over λ

HR DSV

Baseline DG AAM RODS-CAS RODS-SAS IQ

UP (%) ↑

0.0

0.2

0.4

λ

0.6

0.8

1.00.0

Figure 7: Ablation of the adaptive step size magnitude λ on the 11kHands dataset.

hyperparameters and vary only the one of interest. We use the settings from table 11 as the optimal configuration. We note that while feature-based metrics are not the optimal choice for evaluating image quality when hallucinations are present, they serve as a useful proxy; large collapses or expansions in these metrics indicate an inevitable change in quality. Crucially, in the following experiments, we also observe that diversity metrics (IV and DSV) consistently inflate (suggesting improvements) even when FID grows exponentially and HR achieves almost 100%. This observation further highlights the misalignment between feature-based metrics (for diversity in this case) and human perception. B.6.1

Ablating λ

Figure 7 provides the results of varying the adaptive step size magnitude λ across the [0, 1] range. Across all metrics, we observe that [0.05, 0.2] is its optimal range, where HR plateaus around 0.3, while other metrics start to grow exponentially. B.6.2

Ablating time interval

To ablate the choice of the time interval for IQ, we pick t1 from an approximately log-uniform range of the EDM σ scale and choose the corresponding t2 so that guidance is applied at exactly 4 consecutive timesteps. Figure 8 reveals a sharp transition from high HR values into a clear, optimal point, after which the hallucinations once again increase, confirming the theoretical results. Moreover, a clear trade-off between feature-based metrics and HR is observed, once again showing that FID, as 24

17 16 15

0.45 0.40 0.35 0.30 0.25 0.20 0.15

IV HR

0.032 0.030

HR IV

FID

18

IV vs HR over t1

0

10

20

t1

30

40

50

0.028 0.026

DSV vs HR over t1

0.45 0.40 0.35 0.30 0.25 0.20 0.15

DSV HR

0.14 0.13

0

10

20

t1

30

40

50

HR

0.45 0.40 0.35 0.30 0.25 0.20 0.15

FID HR

HR DSV

FID vs HR over t1

19

0.12 0

10

20

t1

30

40

50

Figure 8: Ablation of the time interval bounds for IQ on the 11kHands dataset. For each t1 , its corresponding t2 was chosen so that IQ is applied at exactly 4 consecutive timesteps.

Figure 9: Ablation of the filtering parameter q on the 11kHands dataset. well as other metrics, improve when hallucinations occur frequently, and start to worsen once they decrease. B.6.3

Ablating filtering

To ablate the filtering parameter q, we sweep it across the [0, 1] range, where q = 0 implies that IQ is applied unconditionally across all samples, while q = 1 means that it is never applied. Figure 9 visualizes the resulting dependency between q, feature-based metrics, and HR, leading to two conclusions. First, there is an evident correlation between q and the metrics (negative for FID, positive for IV and DSV). Second, there appears to be an optimal choice of q at around 0.4, where the HR achieves satisfactory values while not significantly altering feature-based metrics. This once again provides evidence that IQ might be trading the inevitable cost in feature-based metrics to decrease hallucinations. The effect of filtering is also clear when analyzing the GaussianGrid results (fig. 16). With q = 0 (depicted as IQ∗ ), all of the generated samples squeeze at the Gaussian modes, almost entirely zeroing out the variance of each Gaussian component. By applying filtering (IQ), this effect is eliminated and the original variance is almost matched. B.6.4

Ablating constant scaling

To assess the validity of adaptive step size λt , we replace it with constant scaling and sweep its values between 0.0 and 10.0 (based on the range of values we observed from the adaptive approach). Figure 10 visualizes the results with an additional dimension for the presence of artifacts in the generated samples, i.e., oversaturated colors and severe structural distortions in more extreme cases. The results confirm the instability of constant scaling, showing that within the optimal range of the trade-off between HR and other metrics, artifacts start to appear. This behavior is not observed when using the adaptive size (fig. 7), confirming its validity. B.7

Runtime comparison

We compare the average sampling runtime of IQ with default unconditional sampling and both RODS variants as the only unconstrained methods for hallucination reduction. The comparison is performed on the most practical and high-resolution RSNA dataset, with results presented in table 6. While all methods add a significant overhead to baseline sampling, IQ remains around 20% faster than 25

0.27 0.26

IV HR Artifact Present

0.03050 0.03025

16.7

0.25 0.24

16.6 2

4

6

8

Guidance Scale

0.26

0.03000

0.25

0.02975 0.02950

0.24

0.02925

10

DSV vs HR over Guidance Scale

0.27

HR IV

FID

16.8

IV vs HR over Guidance Scale

DSV HR Artifact Present

0.132

0.27 0.26

0.130

HR DSV

FID HR Artifact Present

2

4

6

8

Guidance Scale

HR

FID vs HR over Guidance Scale

16.9

0.25

0.128

0.24

0.126

10

2

4

6

8

Guidance Scale

10

Figure 10: Ablation of constant scaling on the 11kHands dataset. both RODSCAS and RODSSAS . These results confirm the superiority of IQ but also point toward an important research direction of reducing the computational cost while preserving the method’s performance. Table 6: Runtime comparison of unconstrained methods on the RSNA dataset. Values are reported as the mean sampling loop time (in seconds) over 16 batches of size 8, alongside the 95% confidence interval. Method Runtime (s) Baseline IQ RODSCAS RODSSAS

B.8

11.49 ± 0.63 49.83 ± 0.40 63.86 ± 2.27 60.33 ± 2.37

Hyperparameters

We include the hyperparameters used for DG in table 7, AAM in table 8, RODSCAS in table 9, RODSSAS in table 10, and IQ in table 11. These were obtained by running a grid search over approximately 500 configurations across all datasets for each baseline method. For IQ, we emphasize that k, the number of i.i.d noise samples used to estimate LIDθ , is kept constant at k = 32 across all datasets. We observed that this value universally ensures the stability of the estimator, but did not experiment further with lowering it. To compute qt , we utilize 2048 reference samples for each dataset. For RSNA, IQ is applied between t = 0.7 and t = 0.75, which corresponds to low noise levels within the SDB framework that utilizes an inverted U-shape noise schedule. For SimpleShapes, we followed Aithal et al. [2024], which noted that the typical behavior observed in DMs shifts to larger timesteps and hence apply IQ at relatively higher times. However, for timesteps smaller than the indicated interval, the image is already fully formed and hence guidance has no practical effect. Table 7: Hyperparameters for Dynamic Guidance (DG). The names of hyperparameters follow their respective names in our implementation. For EDM datasets, dg_start and dg_end are normalized to the [0, 1] interval by dividing the σ boundaries by σmax = 80. For DDPM datasets, they are normalized by dividing the timesteps by Tmax = 1000. For SDB, timesteps natively reside in the [0, 1] interval.

B.9

Parameter

11kHands

FFHQ

AFHQV2

MNIST

SimpleShapes

GaussianGrid

RSNA

dg_scale dg_start dg_end

15.0 0.125 0.00125

10.0 1.0 0.0125

1.0 0.125 0.0125

10.0 0.5 0.25

10.0 0.25 0.0

10.0 0.25 0.0

105 0.7 0.4

Pseudocode

We provide detailed pseudocode of IQ in algorithm 1. 26

Table 8: Hyperparameters for Adaptive Attention Modulation (AAM). The names of hyperparameters follow their respective names in our implementation. Parameter 11kHands FFHQ AFHQV2 MNIST SimpleShapes RSNA γ β N η

2.5 1.25 15 0.1

1.5 1.3 15 0.05

2.5 1.3 15 0.05

2.5 0.1 15 0.1

2.5 0.3 5 0.05

2.0 0.01 15 0.01

Table 9: Hyperparameters for RODSCAS . The names of hyperparameters follow their respective names in our implementation. Parameter perturb_scale cor_threshold num_adv_steps

B.10

11kHands

FFHQ

AFHQV2

MNIST

SimpleShapes

GaussianGrid

RSNA

30.0 0.014 1

8.0 0.09 1

1.0 0.1 1

5.0 0.005 1

1.0 0.005 1

0.1 0.0001 1

0.1 0.002 1

Resources

All of the experiments were performed on NVIDIA A100 40GB and H100 80GB GPUs, with runtime (table 6) reported on the former. The entire project required around 1000 GPU hours in total (this includes any preliminary or grid search experiment). The GPUs were part of NVIDIA DGX nodes. B.11

Qualitative examples

We include more qualitative comparisons on 11kHands (fig. 11), AFHQV2 (fig. 13), FFHQ (fig. 12), MNIST (fig. 14), SimpleShapes (fig. 15), high-resolution version of GaussianGrid results (fig. 16), ground truth scans and their reconstructions from RSNA (fig. 17) and the corresponding qualitative comparison (fig. 18).

27

Figure 11: Qualitative comparison of all sampling methods on 11kHands.

Figure 12: Qualitative comparison of all sampling methods on FFHQ.

28

Figure 13: Qualitative comparison of all sampling methods on AFHQV2.

Figure 14: Qualitative comparison of all sampling methods on MNIST.

29

Table 10: Hyperparameters for RODSSAS . The names of hyperparameters follow their respective names in our implementation. Parameter perturb_scale cor_threshold num_adv_steps

11kHands

FFHQ

AFHQV2

MNIST

SimpleShapes

GaussianGrid

RSNA

30.0 0.014 1

8.0 0.09 1

1.0 0.1 1

5.0 0.01 1

20.0 0.05 1

0.1 0.0001 1

0.1 0.002 1

Table 11: Hyperparameters for IQ. The names of hyperparameters follow their respective names in our implementation. For EDM datasets, t1 and t2 are normalized to the [0, 1] interval by dividing the guidance σ boundaries by σmax = 80. For DDPM datasets, they are normalized by dividing the timesteps by Tmax = 1000. For SDB, timesteps natively reside in the [0, 1] interval. Parameter t1 t2 λ q k

11kHands

FFHQ

AFHQV2

MNIST

SimpleShapes

GaussianGrid

RSNA

0.025 0.0625 0.08 0.4 32

0.025 0.0375 0.035 0.05 32

0.0088 0.025 0.022 0.2 32

0.24 0.40 0.5 0.0 32

0.48 0.96 0.08 0.05 32

0.0 0.72 0.9 0.65 32

0.70 0.75 0.8 0.0 32

Algorithm 1 Intrinsic Quenching (IQ): sampling step. Note: autograd returns the gradient of the energy function w.r.t the input. # # # #

net : EDM denoiser predicting clean data x_0 x : noisy input batch at time t lambda : target magnitude ratio for guidance q : target quantile for filtration ( e . g . , 0.4)

# Apply Intrinsic Quenching only within the window [ t_1 , t_2 ] if t < t_1 or t > t_2 : return stopgrad ( net (x , t )) x . requires_grad_ ( True ) # Predict posterior mean x_0_hat and evaluate DSM energy ( LID ) x_0_hat = net (x , t ) LID = dsm_loss ( x_0_hat , x , t ) # Compute energy gradient w . r . t the input state x grad = autograd ( LID . sum () , x ) # Project guidance into data space ( Sigma_t = t ^2 * I for EDM ) raw_guidance = ( t * * 2) * grad # Calculate adaptive scale to maintain constant magnitude ratio nat_update = x_0_hat - x scale = ( lambda * norm ( nat_update )) / ( norm ( raw_guidance ) + 1 e - 8) # Filtration strategy # q_t is the q - th quantile of the baseline LID distribution at time t q_t = quantile ( baseline_LID ( t ) , q ) mask = where ( LID > = q_t , 1.0 , 0.0) # Apply masked adaptive EDM update guided_x_0_hat = x_0_hat - ( mask * scale * raw_guidance ) return stopgrad ( guided_x_0_hat )

30

Figure 15: Qualitative comparison of all sampling methods on SimpleShapes.

31

Figure 16: High-resolution visualization of samples generated by each method on the GaussianGrid dataset. Each plot depicts 16384 samples.

32

Figure 17: Ground truth brain CT scans from the RSNA dataset and their corresponding reconstructions obtained with SDB using 100-step sampling.

Figure 18: Qualitative comparison of all sampling methods on RSNA.

33

Record · ID 158560 · SHA-256 ecef34c61629eec6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.