CORRECTIVE FORCING: UNIFIED POST-TRAINING FOR DIFFUSIONS AND FLOWS IN GENERATIVE SPEECH ENHANCEMENT Qing Yao, Lijian Gao, and Qirong Mao†
arXiv:2609.24651v1 [cs.LG] 21 Sep 2026
School of Computer Science and Communication Engineering, Jiangsu University Jiangsu Engineering Research Center of Big Data Ubiquitous Perception and Intelligent Agriculture Applications Provincial Key Laboratory of Computational Intelligence and New Technologies in Low-Altitude Digital Agriculture Zhenjiang, China
ABSTRACT Diffusion and flow models, as promising generative paradigms for speech enhancement, face a training–inference mismatch: training uses analytical path states, whereas inference recursively evaluates models on self-generated rollout states along discretized sampling trajectories. This mismatch causes prediction and discretization errors to accumulate. To address it, we introduce Corrective Forcing (CoF), a post-training paradigm that forces diffusion and flow models to learn from self-generated rollouts and correct their predictions. CoF corrects clean-speech predictions on rollout states toward the ground truth under dynamic sampling schedules, exposing the model to varying inference conditions. It further regularizes local evolution using locally corrected counterfactual transitions as references for factual transitions. By expressing model outputs through a shared clean-speech prediction parameterization, CoF applies the same post-training objective across diffusion and flow formulations. Experiments with SB-VE and OT-CFM demonstrate improvements in perceptual quality and reconstruction fidelity, together with robust performance across different numbers of sampling steps. Index Terms— Generative speech enhancement, Schrödinger bridge, flow matching, post-training, exposure bias 1. INTRODUCTION Speech enhancement (SE) aims to recover clean speech from degraded recordings, and generative SE approaches model the underlying conditional distribution of clean speech [1], achieving high perceptual naturalness in reconstructed speech. Recent studies have adapted diffusion and flow models to learn conditional stochastic or deterministic dynamics that transport between degraded and clean speech distributions. Specifically, diffusion-based SE methods [1, 2] define probability paths through stochastic differential equations (SDEs) and train a generative model to learn score fields for constructing reverse-time stochastic dynamics. Schrödinger bridge (SB) [3], formulated as entropy-regularized optimal transport, further provides a principled stochastic formulation bridging degraded and clean speech distributions, achieving high-quality enhancement with efficient sampling. In parallel, flow matching (FM) methods optimize a generative model to learn time-dependent vector fields, enabling deterministic dynamics via ordinary differential equations (ODEs) and facilitating efficient few-step enhancement [4, 5]. † represents the corresponding author. This work was supported in part by the National Natural Science Foundation of China (Grant Nos. 62576155, 62506144, and 62176106), the Natural Science Foundation of Jiangsu Province (Grant No. SH2025108), and the Postgraduate Research & Practice Innovation Program of Jiangsu Province (Grant No. KYCX25 4234). Code and demo are available at https://yorch233.github.io/CoF.
Despite these different formulations, both families face a fundamental training–inference mismatch: training optimizes models on the analytical states sampled from prescribed probability paths, whereas inference recursively evaluates them on self-generated rollout states produced from previous predictions along a discretized sampling trajectory [6]. This mismatch, referred to as exposure bias [7, 8, 9], arises as prediction errors move rollout states away from analytical training states, inducing a progressive distribution shift and error accumulation [10, 11]. Thus, while standard training can learn powerful generative dynamics, the model is never exposed to its own inference behavior, limiting generalization across rollout distributions induced by different sampling schedules. To address this, Regularized Schrödinger Bridge (RSB) perturbs training states and targets to simulate inference deviations [12], but does not expose the model to its actual rollout distributions. In contrast, Correcting the Reverse Process (CRP) uses empirical rollouts [13], but optimizes only the final model call under a fixed number of function evaluations (NFE), thereby restricting both the optimization scope and the rollout conditions encountered during training. In effect, these SE attempts mainly improve predictions over simulated or partial rollout states, while errors can still propagate through local transitions, i.e., numerical solver updates. Robust inference therefore requires not only accurate predictions but also transition behavior that limits error accumulation across sampling steps [14]. Accordingly, these limitations motivate us to (1) force the models to learn from their own rollout distributions under dynamic sampling schedules [11, 15], thereby generalizing across different inference conditions; and (2) regularize their local evolution under transitioninduced deviations before errors accumulate along the trajectory. In this paper, we introduce Corrective Forcing (CoF), a posttraining paradigm for generative SE that forces pretrained models to learn from their own rollout distributions across sampling schedules and correct predictions on them, with corrective supervision at both rollout-state and local-transition levels. First, dynamic rollout correction (DRC) corrects clean-speech predictions on rollout states under dynamic sampling schedules, thereby exposing the model to varying inference conditions. Furthermore, counterfactual transition consistency (CTC) regularizes local evolution under state deviations by aligning predictions at states reached through factual and locally corrected counterfactual transitions, with an exponential moving average (EMA) teacher providing the corrective reference. We validate CoF on Schrödinger bridge with variance-exploding diffusion (SB-VE) and optimal transport conditional flow matching (OTCFM), where a shared clean-speech prediction parameterization enables a unified post-training objective across formulations. Experiments demonstrate improvements in perceptual quality and reconstruction fidelity, together with generalization across different NFEs.
2. BACKGROUND 2.1. Schrödinger Bridge for Speech Enhancement Among diffusion-based approaches to SE, SB [3] models probability transport between paired degraded speech y and clean speech x as an entropy-regularized optimal transport problem [16], equivalently governed by a pair of forward–backward SDEs dXt = [ft + gt2 ∇ log Ψt (Xt )] dt + gt dWt , dXt = [ft − gt2 ∇ log Ψ̂t (Xt )] dt + gt dW̄t ,
(1)
where Xt ∈ CF ×L represents a complex short-time Fourier transform (STFT) spectrogram at time t with F frequency bins and L time frames, where X0 = x and X1 = y. ft and gt are the drift and diffusion coefficients. Ψt and Ψ̂t denote potential functions. Wt and W̄t are Brownian motions. For SB-VE considered in this work, we use the Gaussian marginal pt (Xt | X0 , X1 ) defined in [3]. Training. We sample t ∼ U(tmin , 1) and Xt from the corresponding Gaussian marginal, where tmin denotes the minimum training time. A generative model Xθ is trained to predict X0 by minimizing LSB (θ) = Et, Xt ∥Xθ (Xt , t, y) − X0 ∥22 . (2) 2.2. Flow Matching for Speech Enhancement FM learns a time-dependent vector field that parameterizes deterministic dynamics between two endpoints through an ODE [4, 5]: dXt = ut (Xt | x, y) dt,
(3)
where ut denotes the vector field. In this work, we adopt the OTCFM formulation used in FlowSE [5] and construct the Gaussian path Xt = (1 − t)X0 + tX1 , where X0 = x, X1 = y + σmax z, and z ∼ NC (0, I). σmax specifies the Gaussian perturbation scale at the degraded endpoint. The corresponding vector field is ut (Xt | x, y) = X1 − X0 .
(4)
Training. We sample t ∼ U(tmin , 1) and Xt from the Gaussian path. A generative model vθ is trained to predict ut by minimizing LFM (θ) = Et, Xt ∥vθ (Xt , t, y) − ut ∥22 . (5) 2.3. A Unified View of Diffusions and Flows Despite their distinct formulations, diffusion and flow models can be described under a common framework of Gaussian probability paths [6]. Specifically, for a formulation G and paired data (x, y), the intermediate state Xt can be sampled from the Gaussian marginal pGt (Xt | x, y) through the reparameterization Xt = αtG x + βtG y + γtG z,
z ∼ NC (0, I),
(6)
where (αtG , βtG , γtG ) are the formulation-specific coefficients. Under this framework, standard training optimizes a generative model Gθ on analytical states sampled from pGt to predict clean speech, scores, or vector fields. During inference, a formulationspecific solver evaluates the model recursively along a discretized schedule TNG = {ti }N i=0 , where 1 = t0 > t1 > · · · > tN −1 > tN = 0.
(7)
Here, N denotes the number of sampling steps, which equals the NFE during inference for both SB-VE and OT-CFM. Both formulations start from the endpoint state at t0 = 1 and recursively move
the state from the degraded endpoint toward the clean endpoint, although SB-VE usually uses stochastic SDE updates defined in [3] whereas OT-CFM uses deterministic ODE updates defined in [5]. This recursive evaluation induces an empirical rollout distribution qtG,N that can deviate from the analytical training distribution pGt . Consequently, this mismatch exposes a limitation of conventional training: the model is optimized on analytical states but deployed on states generated by its own numerical solver. 3. CORRECTIVE FORCING To address the training–inference mismatch, we formalize generative SE as a two-stage paradigm: pre-training establishes foundational generative dynamics on analytical distributions, while CoF adapts the pretrained models for inference on empirical rollout states. 3.1. From Pre-Training to Post-Training Pre-Training on Analytical Distributions. During pre-training, Gθ is optimized on analytical states sampled from pGt to learn the formulation-specific generative dynamics. Specifically, pre-training of SB-VE [3] utilizes the data-prediction loss in Eq. (2), and that of OT-CFM [5] utilizes the vector-field prediction loss in Eq. (5). Post-Training on Rollout Distributions. During post-training, Gθ is adapted directly to its empirical rollout distributions. We first define the one-step transition from time ti to ti+1 as G (Xti ) = ΦG (Xti , Gθ (Xti , ti , y), ti , ti+1 ; y, ξ) , Sθ,t i →ti+1
(8)
where ΦG is the solver update rule of G. Here, ξ ∼ NC (0, I) denotes the Gaussian noise used by an SDE solver and is omitted for an ODE solver. We retain it as an implicit argument of S. During inference, given an inference schedule TNG = {ti }N i=0 , its corresponding rollout states form a sampling trajectory, i.e., {X̃tk }N k=0 . This trajectory is recursively generated by the one-step transitions as G X̃tk+1 = Sθ,t X̃tk , k = 0, . . . , N − 1, (9) k →tk+1 where X̃t0 = X1 and X̃tN is the final estimate of X0 . The rollout procedure induces an empirical state distribution qtG,N at each time t, from which states are drawn during post-training. 3.2. Dynamic Rollout Correction (DRC) Given the rollout distribution defined above, DRC directly corrects the model prediction at each self-generated rollout state toward the ground-truth clean-speech endpoint. Rather than forcing rollout states back onto the prescribed probability path, DRC encourages predictions along the rollout to remain consistent with the cleanspeech target, facilitating reconstruction faithfulness. Rollout Generation on Dynamic Schedules. Different numbers of sampling steps produce different transition intervals and consequently distinct rollout distributions. Therefore, DRC exposes the model to rollout states under varying sampling schedules, enabling it to generalize across different inference conditions. Specifically, at each post-training iteration, we uniformly sample Nd ∼ U{1, . . . , Nmax } and construct the corresponding sampling schedule TNGd following the standard discretization adopted by formulation G [3, 5]. Here, varying Nd across iterations induces a dynamic family of rollout distributions. We then uniformly sample an
index i ∈ {0, . . . , Nd − 1}, corresponding to time ti in TNGd . Starting from X̃t0 = X1 , we apply the one-step transitions in Eq. (9) for G,N i steps to obtain the rollout state X̃ti ∼ qti d . Unified Clean-Speech Prediction for Correction. To apply DRC across formulations, we express their original model outputs through a common clean-speech prediction parameterization. Specifically, the output of Gθ is mapped to a clean-speech estimate X̂θ (X̃t , t, y): • Data prediction (SB-VE): Gθ = Xθ directly predicts X0 , so the clean-speech estimate is X̂θ (X̃t , t, y) = Gθ (X̃t , t, y). • Vector-field prediction (OT-CFM): Gθ = vθ predicts the vector field. Since Xt = X0 + t ut , the estimate is given by X̂θ (X̃t , t, y) = X̃t − t Gθ (X̃t , t, y).
(10)
Beyond these, the clean-speech prediction parameterization can also be extended to score or noise predictions [11]. With this shared parameterization, we introduce a reconstruction loss ℓrec to supervise clean-speech predictions on rollout states, aiming to preserve both spectral accuracy and waveform fidelity. Accordingly, DRC is formulated as h i LDRC = ENd ,ti ,X̃t ℓrec X̂θ (sg[X̃ti ], ti , y), X0 , (11) i
where sg[·] denotes the stop-gradient operator. ℓrec is a composite reconstruction criterion following [6], consisting of magnitudespectrum mean squared error (MSE), complex-spectrum MSE, and negative scale-invariant signal-to-noise ratio computed on reconstructed waveforms, weighted by 0.7, 0.3, and 0.01, respectively. 3.3. Counterfactual Transition Consistency (CTC) Although DRC improves predictions at rollout states, residual errors may still affect subsequent model inputs through numerical transitions. To provide such transition-level supervision, CTC constructs factual and counterfactual transitions from the same rollout state, where the counterfactual branch represents a locally corrected evolution and provides a corrective reference for the factual branch. Factual vs. Counterfactual Transitions. To stabilize the factual transition and provide a reliable counterfactual reference, we maintain an EMA teacher Gθ̄ [17], initialized from the pretrained model as θ̄ ← θ and updated after each post-training step via θ̄ ← µθ̄ + (1 − µ)θ, where µ denotes the EMA decay rate. Starting from X̃ti , we sample a continuous time s ∼ U(tmin , ti ) and construct the factual transition to capture the local evolution during inference. In parallel, the counterfactual transition starts from the same state over the same time interval, but replaces the factual model prediction with the ground-truth counterpart GGGT . This yields a locally corrected evolution toward the clean-speech endpoint. The resulting factual and counterfactual states are given by X̃sF = ΦG X̃ti , Gθ̄ (X̃ti , ti , y), ti , s; y, ξ , (12) X̃sCF = ΦG X̃ti , GGGT , ti , s; y, ξ , with shared ξ. Here, GGGT specifies the corrective intervention on the factual prediction, given by X0 for SB-VE and (X̃ti − X0 )/ti for OT-CFM, which is obtained by rearranging Eq. (10). Transition Consistency as Regularization. Given the paired transition states, CTC uses the teacher prediction at the counterfactual state as a corrective reference for the online prediction at the factual state. Specifically, the online prediction is
Method NFE PESQ↑ ESTOI↑ SI-SDR↑ UTMOS↑ Unprocessed – 1.97 0.79 8.5 2.85 SEMamba [18] 1 3.54 0.89 19.7 3.52 MP-SENet [19] 1 3.61 0.89 19.4 3.52 SGMSE+ [1] 100 2.92 0.86 17.5 3.68 0.87 18.6 3.54 StoRM [2] 100+1 2.93 RSB [12] 50 3.04 0.87 18.8 3.64 CRP [13] 5 3.08 0.88 19.3 – 1 2.96 0.88 19.7 3.43 3.07 0.88 18.9 3.60 4 SB-VE [3] 16 3.08 0.87 18.1 3.63 1 3.17 0.88 19.4 3.58 4 3.21 0.88 19.4 3.63 CoF (SB-VE) 16 3.21 0.88 19.4 3.63 1 2.89 0.87 19.7 3.46 OT-CFM [5] 4 3.08 0.87 19.1 3.61 1 3.07 0.88 19.7 3.61 CoF (OT-CFM) 4 3.18 0.88 19.8 3.63 Table 1. Performance comparison on VoiceBank+DEMAND. s = X̂θ (sg[X̃sF ], s, y), while the counterfactual target is X̂0,F s X̂0,CF = sg[X̂θ̄ (X̃sCF , s, y)]. The consistency objective is then: h i s s LCTC = ENd ,ti ,X̃t ,s,ξ ℓrec X̂0,F , X̂0,CF . (13) i
Here, CTC acts as a regularizer that encourages consistent predictions under factual and locally corrected counterfactual transitions, aiming to reduce sensitivity to transition-induced state deviations. 3.4. Overall Objectives Overall, the post-training objective is defined as LCoF = LDRC + λCTC LCTC .
(14)
4. EXPERIMENTS 4.1. Experimental Setup Datasets. We evaluate denoising performance on the VoiceBank+DEMAND [20] and WSJ0+WHAM datasets, and dereverberation performance on the WSJ0+REVERB dataset. Specifically, WSJ0+WHAM is constructed by mixing clean WSJ0 speech [21] with WHAM noise [22], while WSJ0+REVERB is generated by convolving WSJ0 utterances with simulated room impulse responses. The data splits and construction procedures for these two datasets follow those used in RSB [12]. Additionally, VoiceBank+DEMAND is a publicly available benchmark widely used for speech denoising. For all datasets, audio is sampled at 16 kHz and converted into complex spectrograms using an STFT with a squareroot Hann window of length 510 and a hop size of 128. We then apply square-root magnitude warping to the complex spectrograms. Evaluation Metrics. We assess performance with reference-based and reference-free metrics. Specifically, the reference-based metrics include PESQ [23] for perceptual quality, ESTOI [24] for speech intelligibility, SI-SDR [25] for signal fidelity, and the corpus-level word error rate (WER) of transcriptions generated by Whisper base.en [26]. We further employ the reference-free UTMOSv2 (denoted as UTMOS) from VoiceMOS Challenge 2024 [27] to assess overall speech perceptual quality. Best results are shown in bold.
WSJ0+REVERB
WSJ0+WHAM
PESQ ↑
ESTOI ↑
0.81
2.75
SI-SDR (dB) ↑ 15
0.78
2.50 2.25
unproc. 0.55
SGMSE+ 0.72 2.14
2.7
SGMSE+ 0.71
0.80
2.4 2.1 1
4
8
16
1
4
SGMSE+ 0.71 8
NFE
SB-VE
12 7.5
16
OT-CFM
2.60 SGMSE+ −3.2 2.40 16 24 1
unproc. 3.0 −9.1
24
1
4
8
SGMSE+ (NFE=100)
clean 5.7
9 12
unproc. 13.5
SGMSE+ 11.9
10 8
unproc. 1.61 4
NFE
CoF (OT-CFM)
12
unproc. 2.69
2.80 3.20 3.00 2.80
NFE
CoF (SB-VE)
SGMSE+ 10.3
4.5 unproc. 0.42
0.72 24
unproc. 3.9
SGMSE+ 19.7
15
3.00
13
6.0
0.76 unproc. 1.38
WER (%) ↓ 18
3.20
14
0.75 unproc. 1.32
UTMOS ↑
3.40
8
16
24
6
NFE
StoRM (NFE=100+1)
clean 5.8 1
4
8
16
24
NFE
unproc. (unprocessed)
clean
Fig. 1. Performance comparison with and without CoF at different NFEs on WSJ0+WHAM (top) and WSJ0+REVERB (bottom). Compared Baselines. We consider SB-VE [3] and OT-CFM [5] as the baseline generative models and apply CoF to both through posttraining. For broader comparison, we include representative predictive methods, namely, SEMamba [18] and MP-SENet [19], together with generative SE methods, including SGMSE+ [1], StoRM [2], RSB [12], and CRP [13]. SGMSE+, StoRM, and RSB are evaluated using their recommended samplers with 50 sampling steps, corresponding to 100, 101, and 50 NFEs, respectively. Implementation Details. On VoiceBank+DEMAND, results for the baselines are obtained by inference with their released checkpoints, except for CRP, whose results are taken from [5]. On WSJ0+WHAM and WSJ0+REVERB, we reproduce SGMSE+ and StoRM using their official code for reference. We reimplement SB-VE and OT-CFM using the same NCSN++M architecture [2], following the default implementations and parameters of SB [3] and FlowSE [5], respectively. The best pretrained checkpoint is used for CoF post-training. We set λCTC = 0.1 in Eq. (14) and use an EMA decay rate of µ = 0.999. We optimize LCoF using Adam with a learning rate of 10−4 and a total batch size of 8 on two NVIDIA RTX 5090 GPUs. We set (Nmax , tmin ) to (16, 10−4 ) for SB-VE and (8, 0.03) for OT-CFM. For both SB-VE and OT-CFM, the sampling times in TNG are uniformly spaced over the time interval. CoF uses 4,000 post-training steps and requires on average (Nmax − 1)/4 + 4 model evaluations per step, corresponding to 7.75 for SB-VE and 5.75 for OT-CFM. We evaluate the EMA checkpoint every 200 steps on 50 validation utterances from each dataset and select the EMA checkpoint with the best validation PESQ for final inference. 4.2. Experimental Results Comparison with Baselines. Table 1 compares CoF with the baselines on VoiceBank+DEMAND. Overall, CoF improves PESQ for both SB-VE and OT-CFM and substantially improves multi-step SI-SDR, while generally preserving ESTOI and UTMOS. Specifically, CoF (SB-VE) reaches PESQ 3.21 at NFE = 4, exceeding SB-VE at 3.07, RSB at 3.04, and CRP at 3.08, while requiring fewer NFEs than the latter two. Moreover, CoF stabilizes multi-step reconstruction fidelity. In particular, CoF (SB-VE) maintains SISDR at 19.4 dB across NFEs, and CoF (OT-CFM) raises SI-SDR to 19.8 dB at NFE = 4. The resulting SI-SDR is also comparable to that of the predictive baselines MP-SENet and SEMamba. Fig. 1 shows that CoF consistently improves most metrics across NFEs on the more challenging WSJ0+WHAM and WSJ0+REVERB benchmarks. At NFE = 4, CoF (OT-CFM) raises PESQ to 2.85 on WSJ0+WHAM and 2.89 on WSJ0+REVERB, with SI-SDR gains of 1.0 and 1.1 dB, respectively. CoF (SB-VE) also mitigates fidelity degradation under deeper rollouts. On WSJ0+WHAM, as NFE increases from 1 to 16, its SI-SDR improves from 14.9 to 15.2 dB,
Variants SB-VE pre-training Fine-tuning with ℓrec
PESQ↑ 3.07 3.16
SI-SDR↑ 18.9 18.6
UTMOS↑ 3.60 3.64
DRC 3.19‡ 19.2‡ 3.61 + CTC (CoF), λCTC = 0.1 3.21* 19.4* 3.63* + CTC (CoF), λCTC = 0.25 3.20 19.5 3.63 CoF with consistency target X0 3.20 19.3 3.61 CoF with fixed Nd = 8 3.16 19.3 3.62 3.03 19.6 3.58 CoF with plain MSE (‡) and (*) indicate significant improvements over fine-tuning and DRC, respectively (utterance-level paired t-test, p < 0.05). Table 2. Ablation study on VoiceBank+DEMAND with SB-VE. whereas SB-VE degrades from 14.8 to 13.6 dB. Overall, CoF improves PESQ and ESTOI at matched NFEs and enhances multi-step SI-SDR, while its modest and inconsistent UTMOS gains indicate that improvements in reference-free perceptual quality remain limited. CoF also maintains low WERs across NFEs, largely avoiding the degradation observed in pretrained models under deeper rollouts. Ablation Study. Table 2 examines the contributions of CoF with SB-VE at NFE = 4. Compared with fine-tuning on analytical states using ℓrec for 4,000 steps, DRC significantly improves PESQ from 3.16 to 3.19 and SI-SDR from 18.6 to 19.2 dB, supporting the benefit of corrective supervision on rollout states, although UTMOS decreases from 3.64 to 3.61. Adding CTC with λCTC = 0.1 yields further significant improvements in PESQ, SI-SDR, and UTMOS. Increasing λCTC to 0.25 slightly favors SI-SDR but lowers PESQ. We therefore use λCTC = 0.1 as the default. Replacing the counterfactual teacher prediction with X0 as the consistency target yields weaker overall performance, supporting the use of a transition-aware counterfactual reference rather than an additional reconstruction target. Using a fixed Nd = 8 also degrades performance, supporting dynamic sampling schedules that expose the model to rollouts induced by diverse inference conditions. Finally, replacing the composite ℓrec with plain MSE further raises SI-SDR to 19.6 dB but substantially reduces PESQ and UTMOS, highlighting the tradeoff introduced by the training objective. 5. CONCLUSION In this paper, we propose CoF to mitigate the training–inference mismatch in generative SE by adapting pretrained models to selfgenerated sampling trajectories. DRC corrects predictions on rollout states, while CTC regularizes local transitions through counterfactual references. Experiments show that CoF improves referencebased perceptual quality and reconstruction fidelity for both SB-VE and OT-CFM, particularly under multi-step sampling.
6. COMPLIANCE WITH ETHICAL STANDARDS This study used existing speech datasets and involved no new collection of human-subject data. No ethical approval was required. 7. REFERENCES [1] J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech enhancement and dereverberation with diffusion-based generative models,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 31, pp. 2351–2364, 2023. [2] J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “StoRM: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 31, pp. 2724–2737, 2023. [3] A. Jukić, R. Korostik, J. Balam, and B. Ginsburg, “Schrödinger bridge for generative speech enhancement,” in Proc. Interspeech, 2024, pp. 1175–1179. [4] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in Proc. Int. Conf. Learn. Represent., 2023. [5] S. Lee, S. Cheong, S. Han, and J. W. Shin, “FlowSE: Flow matching-based speech enhancement,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process., 2025, pp. 1–5. [6] D. Wang, J. Gao, T. Lei, Y. Hu, C. Zhu, K. Chen, and J. Lu, “Rethinking flow and diffusion bridge models for speech enhancement,” Proc. AAAI Conf. Artif. Intell., vol. 40, no. 39, pp. 33431–33439, 2026. [7] S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Adv. Neural Inf. Process. Syst., 2015. [8] M. Ning, M. Li, J. Su, A. A. Salah, and I. Onal Ertugrul, “Elucidating the exposure bias in diffusion models,” in Proc. Int. Conf. Learn. Represent., 2024. [9] M. Ning, E. Sangineto, A. Porrello, S. Calderara, and R. Cucchiara, “Input perturbation reduces exposure bias in diffusion models,” in Proc. Int. Conf. Mach. Learn., 2023, vol. 202, pp. 26245–26265. [10] S. Ross, G. J. Gordon, and J. A. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proc. Int. Conf. Artif. Intell. Statist., 2011, vol. 15, pp. 627–635. [11] C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu, “DPMSolver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps,” in Adv. Neural Inf. Process. Syst., 2022, vol. 35, pp. 5775–5787. [12] Q. Yao, L. Gao, Q. Mao, and M. Dong, “Regularized Schrödinger bridge via distortion-perception perturbation for high-fidelity speech enhancement,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 34, pp. 3886–3900, 2026. [13] B. Lay, J.-M. Lemercier, J. Richter, and T. Gerkmann, “Single and few-step diffusion for generative speech enhancement,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process., 2024, pp. 626–630. [14] D. Kim, C.-H. Lai, W.-H. Liao, N. Murata, Y. Takida, T. Uesaka, Y. He, Y. Mitsufuji, and S. Ermon, “Consistency trajectory models: Learning probability flow ODE trajectory of diffusion,” in Proc. Int. Conf. Learn. Represent., 2024.
[15] X. Huang, Z. Li, G. He, M. Zhou, and E. Shechtman, “Self Forcing: Bridging the train-test gap in autoregressive video diffusion,” in Adv. Neural Inf. Process. Syst., 2025, vol. 38, pp. 167283–167308. [16] G.-H. Liu, A. Vahdat, D.-A. Huang, E. A. Theodorou, W. Nie, and A. Anandkumar, “I2SB: Image-to-image Schrödinger bridge,” in Proc. Int. Conf. Mach. Learn., 2023, vol. 202, pp. 22042–22062. [17] A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistency targets improve semisupervised deep learning results,” in Adv. Neural Inf. Process. Syst., 2017, vol. 30. [18] R. Chao, W.-H. Cheng, M. La Quatra, S. M. Siniscalchi, C.H. H. Yang, S.-W. Fu, and Y. Tsao, “An investigation of incorporating Mamba for speech enhancement,” in Proc. IEEE Spoken Lang. Technol. Workshop (SLT), 2024, pp. 302–308. [19] Y.-X. Lu, Y. Ai, and Z.-H. Ling, “Explicit estimation of magnitude and phase spectra in parallel for high-quality speech enhancement,” Neural Netw., vol. 189, pp. 107562, 2025. [20] C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust Text-to-Speech,” in Proc. SSW, 2016, pp. 146– 152. [21] J. S. Garofolo, D. Graff, J. M. Baker, D. Paul, and D. Pallett, “CSR-I (WSJ0) Complete,” Linguistic Data Consortium, 1993, LDC93S6A. [22] G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, “WHAM!: Extending speech separation to noisy environments,” in Proc. Interspeech, 2019, pp. 1368–1372. [23] A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process., 2001, vol. 2, pp. 749–752. [24] J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 24, no. 11, pp. 2009–2022, 2016. [25] J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR– half-baked or well done?,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process., 2019, pp. 626–630. [26] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. Int. Conf. Mach. Learn., 2023, vol. 202, pp. 28492–28518. [27] K. Baba, W. Nakata, Y. Saito, and H. Saruwatari, “The T05 system for the VoiceMOS Challenge 2024: Transfer learning from deep image classifier to naturalness MOS prediction of high-quality synthetic speech,” in Proc. IEEE Spoken Lang. Technol. Workshop (SLT), 2024, pp. 818–824.