ConceptioArchivearXiv CS
arXiv CSopen access

Diffusion Language Models for Speech Recognition

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Diffusion Language Models for Speech Recognition Davyd Naveriani1,∗ , Albert Zeyer ID 1,2,∗ , Ralf Schlüter ID 1,2 , Hermann Ney1,2 1

Machine Learning and Human Language Technology Group, RWTH Aachen University, Germany 2 AppTek, Germany [email protected], {zeyer,schlueter,ney}@cs.rwth-aachen.de

arXiv:2604.14001v1 [cs.CL] 15 Apr 2026

Abstract Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models (MDLM) and uniformstate diffusion models (USDMs) for rescoring ASR hypotheses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise probability distributions derived from CTC with the labelwise probability distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong language knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We publish all our code and recipes. Index Terms: Speech Recognition, Diffusion Language Model

1. Introduction Autoregressive language models (LMs) are commonly used to improve automatic speech recognition (ASR) systems due to their strong linguistic capabilities and the ability to incorporate external textual knowledge [1–3]. However, applying traditional autoregressive LMs in joint decoding inherently limits the speed due to their strictly left-to-right decoding structure. Non-autoregressive language models [4] can operate in a parallel manner, potentially resolving this bottleneck and enabling faster decoding. Recently, discrete diffusion models, such as large language diffusion with masking (LLaDA) [5] and masked diffusion language models (MDLM) [6] have emerged as powerful nonautoregressive alternatives [5–10]. Prior work has explored diffusion-based models in ASR as audio-conditioned decoders [11–13]. Other non-autoregressive speech recognition models have been proposed [14–16]. Here, we focus on keeping a separate language model to better leverage large amounts of text data and to allow more flexible integration of the language model into the ASR system [17, 18]. An investigation of diffusion language models as standalone models for joint ASR decoding has not yet been conducted and their integration into token-level joint decoding remains unexplored. In this work, we systematically investigate diffusion language models for ASR rescoring, comparing masked diffusion language models (MDLM) and uniform-state diffusion models * These authors contributed equally.

(USDM). Furthermore, we develop a novel token-level decoding method that allows for the integration of USDM with a non-autoregressive connectionist temporal classification (CTC) speech recognition model [19]. Because USDM corrupts sequences using uniform transitions without artificial mask tokens, it provides a full vocabulary probability distribution for every token at each denoising step, which enables a direct combination of framewise CTC probabilities and labelwise diffusion distributions during hypothesis construction, as detailed in Section 2 and Section 3. To the best of our knowledge, this is the first work that (i) systematically compares masked and uniform-state diffusion language models for ASR rescoring, and (ii) integrates a uniform-state diffusion language model into CTC-based token-level joint decoding.

2. Diffusion Language Models Masked diffusion language model. MDLM corrupts text by randomly masking tokens and learns to reconstruct the sequence during the reverse generative pass. During the forward process, tokens are independently masked based on a monotonically decreasing noise schedule αt ∈ [0, 1]. Essentially, αt represents the probability of a token retaining its original value, while (1 − αt ) is the probability of it being masked. As the process reaches the final step T , αT approaches 0, meaning the sequence becomes completely masked with probability 1. The marginal distribution of this forward process is defined as: q(zt | w) = Cat(zt ; αt 1w + (1 − αt )1m )

(1)

where Cat is the categorical distribution, w is the original clean token, zt ∈ V is the token at diffusion step t, 1w ∈ {0, 1}|V | denotes the one-hot vector with a 1 at position w ∈ V and indicates a clean token, and m is the index of the [MASK] token. During the reverse process, starting from a fully masked sequence, an MDLM iteratively denoises the text to recover the original tokens. The theoretical objective is to align the reverse transition with the true posterior distribution, which is defined as: q(zs | zt , w)  Cat(z  s ; 1zt ) = (1 − αs )1m + (αs − αt )1w Cat zs ; 1 − αt

zt ̸= m, zt = m. (2)

However, because the target token w is unknown during generation, a parameterized model wθ (zt , t) is trained to directly predict the best estimate of the unmasked tokens from the noisy

state zt . This implies that the training objective is computed exclusively over the masked tokens, taking the form of a crossentropy loss weighted by the noise schedule [5, 6]. The bidirectional context modeling of MDLM makes it particularly wellsuited for ASR hypothesis rescoring [11, 12].

Sample-level mask normalization. We normalize each sample by its own mask count before averaging over K: log PDiffLM (ãS̃ 1) ≈

K X 1 X 1 log Pθ (ãj | zt,k ) (5) K |Mk | j∈M k=1

Uniform-state diffusion model. USDM works similarly to MDLM but uses a different corruption strategy. During the forward process, tokens are replaced with random samples from the vocabulary rather than a mask token. The marginal distribution is defined as q(zt | w) = Cat(zt ; αt 1w + (1 − αt )π), where π = |V1 | 1 is the uniform distribution over the vocabulary V [20]. During the reverse process, these forward dynamics allow for continual token updates. Since corrupted tokens are indistinguishable from clean ones, the model must re-evaluate every position at each denoising step, enabling a self-correcting property where previously predicted tokens can be updated or fixed [7, 21–23]. Consequently, at each denoising step, the model wθ (zt , t) produces a full probability distribution over the entire vocabulary for every token in the sequence, regardless of its current state. For ASR, this dense output is particularly advantageous as it provides a continuous stream of vocabularywide probabilities that can be directly aligned and combined with frame-wise CTC scores during joint decoding.

k

Global mask normalization. We pool all masked predictions across all K samples and divide by the total mask count: PK P k=1 j∈Mk log Pθ (ãj | zt,k ) S̃ log PDiffLM (ã1 ) ≈ (6) PK k=1 |Mk | Sample-level normalization weights each Monte Carlo sample equally, whereas global normalization weights each predicted token equally. As shown in Figure 3, they significantly improve WER over sequence-length normalization. Coupled scoring. Additionally, inspired by the coupledsampling scheme from [26], we construct K pairs of comple(1)

(2)

(1)

mentary masks Mk and Mk = Mk , such that every token is masked in exactly one of the two forward passes per pair: log PDiffLM (ãS̃ 1) K

1 X1 K S̃ k=1

3. Methodology

X

(1)

log Pθ (ãj | zt,k )

(1)

j∈Mk

! +

3.1. Rescoring

X

(2) log Pθ (ãj | zt,k )

(7)

(2)

We rescore n-best CTC hypotheses ãS̃ 1 = (ã1 , . . . , ãS̃ ) by combining the CTC log-probability with a diffusion language model (DiffLM) score and a prior correction term: S̃ T S(ãS̃ 1 ) = λCTC log PCTC (ã1 | x1 )

+ λDiffLM log PDiffLM (ãS̃ 1) − λprior log Pprior (ãS̃ 1)

(3)

where xT1 denotes the sequence of T acoustic feature frames, λCTC , λDiffLM , and λprior are tunable interpolation weights, and the prior term compensates for the implicit language model bias learned by the CTC model [24, 25]. Since log PDiffLM (ãS̃ 1 ) is intractable, we approximate it via K Monte Carlo samples by applying forward noise and computing the reconstruction loglikelihood. MDLM. A naive sequence-length normalization estimate averages over the full sequence length S̃: log PDiffLM (ãS̃ 1) ≈

K 1 X1 X log Pθ (ãj | zt,k ) K S̃ j∈M k=1

(4)

k

where t is a fixed noise step that determines the masking probability, Mk is the set of randomly masked positions in the kth sample, ãj is the j-th token of the hypothesis, and zt,k denotes the noisy sequence obtained by masking positions Mk with masking probability (1 − αt ). This overweights heavily corrupted samples—steps with many masks contribute many log-prob terms, dominating the score regardless of model confidence. We propose three alternative scoring strategies to address this issue.

j∈Mk

This guarantees that every token contributes to the score. USDM. We adopt the ELBO from [7] as the scoring objective. Unlike MDLM, where only masked tokens contribute to the score, USDM corrupts positions with uniform noise and every token participates in the score regardless of the noise level. 3.2. Joint-Decoding USDM exhibits several properties that make it particularly suitable for integration with CTC in a joint decoding framework (see Figure 1): its self-correcting nature allows the model to continuously refine all positions rather than committing to predictions early; and since tokens are corrupted with uniform noise rather than a mask token, the model maintains a welldefined probability distribution over the full vocabulary at every position throughout the entire denoising process. These properties motivate us to extend USDM to a novel joint CTC-USDM decoding framework. We initialize the denoising process from the CTC greedy sequence at noise level tstart [13], which we treat as a tunable hyperparameter. Since CTC operates on frames while USDM operates on tokens, we align the two by extracting the logprobability distribution of the first frame corresponding to each collapsed token and renormalizing it over the non-blank vocabulary. At each denoising step l, and each token position i, USDM produces a token-level distribution Pθ (· | ztl ), which we combine with the CTC distribution: log Pcomb,i (vj ) = λCTC log PCTC,τi (vj | xT1 ) + λDiffLM log Pθ,i (vj | ztl )

(8)

Here, vj ∈ V denotes the j-th token from the vocabulary V . τi denotes the first CTC frame aligned with token position i in

140

𝑦!"#$%

120

log 𝑃./01,3 𝑣4 = 𝜆CTC log 𝑃$%$,5! 𝑣4 𝑥6% + 𝜆73889: log 𝑃;,3 𝑣4 𝑧#" log 𝑃$%$,5$ 𝑣3 𝑥6%

𝒕𝒐𝒌𝒆𝒏𝒎

𝒕𝒐𝒌𝒆𝒏𝒎

𝒃𝒍𝒂𝒏𝒌

100

𝒕𝒐𝒌𝒆𝒏𝒌

𝐷𝑖𝑓𝑓𝑢𝑠𝑖𝑜𝑛 𝑀𝑜𝑑𝑒𝑙 𝑈𝑛𝑖𝑓𝑜𝑟𝑚 − 𝑆𝑡𝑎𝑡𝑒 Audio

𝑙 =𝑙−1 𝑧#"%#,3 ∼ 𝑃./01,3

PPL

log 𝑃$%$,5# 𝑣3 𝑥6%

USDM 5ep MDLM 5ep

𝐶𝑜𝑛𝑓𝑜𝑟𝑚𝑒𝑟 𝐸𝑛𝑐𝑜𝑑𝑒𝑟

80 60

𝒕𝒐𝒌𝒆𝒏𝒎

𝑧# =𝑦$%$&'())*+ 𝑡 = 𝑡,#-(#

40

𝒕𝒐𝒌𝒆𝒏𝒌

𝐼𝑛𝑝𝑢𝑡

0

𝑆𝑡𝑒𝑝 0

Figure 1: Overview of the proposed joint CTC-USDM decoding. At each denoising step, the USDM token-level distribution is combined with the CTC frame-level distribution to sample input to the next denoising step.

the collapsed greedy sequence. The term PCTC,τi (vj | xT1 ) corresponds to the frame-level CTC probability of token vj at frame τi , obtained from the encoder output and renormalized over the non-blank vocabulary. The term Pθ,i (vj | ztl ) denotes the token-level probability predicted by the USDM at position i given the current noisy sequence ztl . The combined score Pcomb,i (vj ) therefore integrates acoustic evidence from the CTC model with contextual information from the diffusion language model during each denoising step. The next input ztl−1 is then sampled from this combined distribution. We use ancestral sampling from [20, 21, 27], which draws the next state directly from the USDM reverse posterior.

4. Experiments 4.1. Experimental Setup We trained MDLM and USDM on a combined corpus of normalized LibriSpeech LM data and train-other transcriptions [28]. For our experiments, we leveraged the training frameworks from [6, 7]. Models were trained for 5, 10 and 25 epochs using AdamW (0.1 weight decay) [29], a piecewise linear LR scheduler, and a 20,000 token batch size. Our primary "medium" architecture for both models is a 24-layer Diffusion Transformer (DiT) [30] with 16 attention heads and a 1024dimensional hidden state. We also evaluated a "small" variant (12 layers, 12 heads, 768-dim, 0.1 dropout [31]). Both configurations used 128-dimensional diffusion time embeddings. Text was tokenized via SentencePiece into 10,240 subwords [32]. 4.2. Results Language model training. Table 1 shows the perplexity upper bounds for USDM and MDLM trained with the same configuration. MDLM achieves lower PPL at 5 and 10 epochs (see Figure 2, e.g. 36.6 vs. 40.2 on dev at 5 epochs), but USDM surpasses it at 25 epochs (34.0 vs. 32.3 on dev). This can be explained by the fact that USDM corrupts tokens with uniform noise rather than explicit mask tokens, making the task inherently harder since the model must evaluate every position. Both models show improvement with longer training. Rescoring. Figure 3 compares MDLM and USDM rescoring strategies across varying numbers of Monte Carlo samples K. As illustrated in Figure 3, MDLM rescoring consistently outperforms the CTC baseline (5.08% WER) and USDM. While stan-

20

40 60 Subepoch

80

100

Figure 2: Dev PPL learning curves for MDLM and USDM trained for 5 full epochs. USDM learning curve is shown after the curriculum learning phase. Table 1: Perplexity (PPL) upper bounds on train, dev, and devtrain splits for MDLM and USDM trained for 5, 10 and 25 epochs on LibriSpeech LM data. Diffusion LM MDLM

USDM

Number of Epochs 5 10 25 5 10 25

train ≤ 36.2 ≤ 32.7 ≤ 30.0 ≤ 40.5 ≤ 37.0 ≤ 33.5

PPL dev ≤ 36.6 ≤ 37.0 ≤ 34.0 ≤ 40.2 ≤ 39.4 ≤ 32.3

devtrain ≤ 47.9 ≤ 44.6 ≤ 42.2 ≤ 59.5 ≤ 48.0 ≤ 44.1

Table 2: MDLM rescoring WER [%] on dev-other across 5, 10 and 25 training epochs, varying the number of Monte Carlo samples (K). Standard deviations computed over 5 random seeds. K 1 2 16 32 64 128 256

MDLM 5 ep 4.97 ± .03 4.94 ± .01 4.79 ± .03 4.72 ± .03 4.65 ± .04 4.60 ± .02 4.59 ± .02

WER [%] MDLM 10 ep 4.94 ± .01 4.94 ± .02 4.78 ± .02 4.67 ± .02 4.60 ± .02 4.58 ± .02 4.56 ± .03

MDLM 25 ep 4.95 ± .01 4.91 ± .02 4.75 ± .03 4.65 ± .03 4.56 ± .02 4.55 ± .02 4.52 ± .01

dard sequence-length normalization achieves a WER of 4.73% (K = 256), our proposed global mask normalization yields a further reduction to 4.59%. Additionally, coupled scoring improves results with fewer samples, reaching 4.71% WER at K = 46. Furthermore, as shown in Table 2, extending MDLM training to 10 and 25 epochs provides additional gains, reaching a rescoring performance of 4.56% and 4.52% WER at K = 256. While USDM rescoring does not match the performance of MDLM, it still yields improvements over the CTC baseline. As shown in Table 3, USDM reduces the WER from 5.08% to 4.82% at K = 256. With an increased number of epochs, the WER decreased to 4.80% for 25 epochs. CTC-USDM Joint Decoding. As shown in Table 4, joint decoding yields better results than USDM rescoring, suggesting that the active participation of USDM in hypothesis construc-

MDLM Seq Level MDLM Global Mask USDM MDLM Coupled MDLM Sample Mask

5.0 WER [%]

4.9

Epochs 5 10 25

4.8 4.7 4.6 0

50

100

K

150

200

250

Figure 3: WER [%] on dev-other comparing MDLM rescoring (sequence-level, global-mask and sample-mask score normalization), MDLM (5 ep) coupled scoring, and USDM (5 ep) rescoring, across different numbers of Monte Carlo samples (K). Stars mark the best WER for each method. Table 3: USDM rescoring WER [%] on dev-other across 5, 10 and 25 training epochs, varying the number of Monte Carlo samples (K). Standard deviations computed over 5 random seeds. K 1 2 16 32 64 128 256

USDM 5 ep 4.99 ± .03 4.96 ± .03 4.94 ± .03 4.91 ± .02 4.87 ± .04 4.85 ± .03 4.82 ± .02

WER [%] USDM 10 ep 4.97 ± .02 4.99 ± .03 4.92 ± .02 4.89 ± .04 4.85 ± .04 4.83 ± .02 4.82 ± .02

USDM 25 ep 4.97 ± .01 4.99 ± .02 4.92 ± .02 4.89 ± .03 4.86 ± .02 4.80 ± .01 4.80 ± .02

Table 4: WER [%] on dev-other for CTC+USDM joint decoding with different initial noise level tstart , and denoising steps K (for all experiments, λDiffLM = 0.3 shows the best performance).

tstart 0.3 0.5 0.8

Table 5: WER [%] on dev-other for CTC+USDM joint decoding comparing models trained on different numbers of epochs with λDiffLM = 0.3 and tstart = 0.3.

1 4.79 4.81 4.82

8 4.78 4.81 4.79

WER [%] K 12 16 32 4.79 4.77 4.78 4.80 4.78 4.77 4.78 4.78 4.77

48 4.77 4.77 4.77

64 4.77 4.77 4.77

tion produces better hypotheses. A lower initial noise level (tstart = 0.3) allows the model to reach optimal WER with fewer denoising steps, while all configurations converge to the same final WER. As shown in Table 5, extending USDM training to 10 and 25 epochs further improves joint decoding, reaching a peak WER of 4.73% and 4.71%, respectively. Comparison with Autoregressive LMs. Table 6 provides an overview of the results across the different modeling approaches. As expected, the autoregressive language models achieve the lowest overall WER, with 4.19% for rescoring and 3.86% for joint decoding (Table 6). However, scaling behavior differs across architectures: for MDLM, increasing from

1 4.79 4.74 4.74

8 4.78 4.73 4.72

WER [%] K 12 16 32 4.79 4.77 4.78 4.74 4.76 4.74 4.74 4.72 4.73

48 4.77 4.74 4.72

64 4.77 4.73 4.71

Table 6: WER [%] on dev-other comparing CTC without LM, with an autoregressive LM, with MDLM (10 ep, global mask norm., K = 256) and USDM (10 ep, K = 256 rescoring and K = 64 joint-decoding). Model Dim –

Num Layers –

768

12

1024

24

MDLM

768 1024

12 24

USDM

1024

24

LM None Autoregressive LM

WER [%] dev-other greedy 5.08 rescoring 4.10 joint-decoding 3.98 rescoring 4.19 joint-decoding 3.86 4.64 rescoring 4.56 rescoring 4.82 joint-decoding 4.73

12 to 24 layers reduces rescoring WER from 4.64% to 4.56%, whereas for the autoregressive model, the same scaling degrades rescoring performance from 4.10% to 4.19%, while still improving joint decoding from 3.98% to 3.86%.

5. Conclusions In this work, we systematically explored the integration of discrete diffusion language models into ASR systems. While traditional autoregressive models are constrained by a strictly sequential, left-to-right decoding structure, diffusion LMs leverage bidirectional context and parallel generation, offering a more flexible and theoretically faster alternative for ASR. We introduced new methods to rescore ASR hypotheses using MDLM, namely Global Mask Normalization and Sample Mask Normalization. By utilizing the mask length for normalization, these methods significantly improved performance in comparison to standard sequence-level normalization. Most notably, we noticed unique properties of Uniform-State Diffusion Models, specifically their lack of artificial mask tokens and their fullvocabulary probability distribution for each position, and developed a CTC-USDM joint decoding framework, which successfully outperformed static rescoring with USDM. Evaluation shows MDLM achieves better rescoring accuracy on limited data than USDM. This is likely because MDLM’s explicit mask tokens provide a clearer reconstruction signal, whereas USDM’s uniform noise forces the model to implicitly distinguish clean from noisy tokens at every position. Consequently, USDM’s more complex objective may require additional data or scaling to reach peak performance. Additionally, we compared our methods with standard autoregressive language models. While the autoregressive baselines achieved better results, increasing the model capacity significantly improved MDLM’s rescoring performance, whereas it degraded the rescoring re-

sults of the autoregressive models. In future work, we plan to evaluate these models on larger datasets and scale the model capacities further to close the performance gap with autoregressive baselines. We also plan to investigate the joint decoding framework in greater detail, including extending it to MDLM.

6. Acknowledgements This work was partially supported by NeuroSys, which as part of the initiative “Clusters4Future” is funded by the Federal Ministry of Education and Research BMBF (funding IDs 03ZU2106DA and 03ZU2106DD), and by the project RESCALE within the program AI Lighthouse Projects for the Environment, Climate, Nature and Resources funded by the Federal Ministry for the Environment, Nature Conservation, Nuclear Safety and Consumer Protection (BMUV), funding ID: 67KI32006A. The authors gratefully acknowledge the computing time provided to them at the NHR Center NHR4CES at RWTH Aachen University (project number p0023565 and p0023999). This is funded by the Federal Ministry of Education and Research, and the state governments participating on the basis of the resolutions of the GWK for national high performance computing at universities (www.nhr-verein. de/unsere-partner).

7. Generative AI Use Disclosure We use LLMs to improve the formulations and grammar of the paper.

8. References [1] F. Jelinek, Statistical Methods for Speech Recognition. press, 1998.

MIT

[2] K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Language Modeling with Deep Transformers,” in Interspeech, Graz, Austria, Sep. 2019, pp. 3905–3909, iSCA Best Student Paper Award. [slides]. [Online]. Available: http://arxiv.org/pdf/1905.04226.pdf [3] R. Prabhavalkar, T. Hori, T. N. Sainath, R. Schlüter, and S. Watanabe, “End-to-End Speech Recognition: A Survey,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 325–351, 2023. [4] Y. Su, D. Cai, Y. Wang, D. Vandyke, S. Baker, P. Li, and N. Collier, “Non-Autoregressive Text Generation with Pre-trained Language Models,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 2021, pp. 234–243.

V. Kuleshov, “Mercury: Ultra-Fast Language Models Based on Diffusion,” arXiv preprint arXiv:2506.17298, 2025. [11] M. Wang, Z. Liu, Z. Jin, G. Sun, C. Zhang, and P. C. Woodland, “Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing,” arXiv preprint arXiv:2509.16622, 2025. [12] T. Kwon, J. Ahn, T. Yun, H. Jwa, Y. Choi, S. Park, N.-J. Kim, J. Kim, H. G. Ryu, and H.-J. Lee, “Whisfusion: Parallel ASR Decoding via a Diffusion Transformer,” arXiv preprint arXiv:2508.07048, 2025. [13] W. Tian, B. Mu, G. Ma, X. Geng, Z. Zhao, and L. Xie, “dLLMASR: A Faster Diffusion LLM-based Framework for Speech Recognition,” arXiv preprint arXiv:2601.17902, 2026. [14] Y. Higuchi, N. Chen, Y. Fujita, H. Inaguma, T. Komatsu, J. Lee, J. Nozaki, T. Wang, and S. Watanabe, “A Comparative Study on Non-autoregressive Modelings for Speech-to-Text Generation,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, pp. 47–54. [15] K. Deng, Z. Yang, S. Watanabe, Y. Higuchi, G. Cheng, and P. Zhang, “Improving Non-Autoregressive End-to-End Speech Recognition with Pre-trained Acoustic and Language Models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 8522–8526. [16] A. Navon, A. Shamsian, N. Glazer, Y. Segal-Feldman, G. Hetz, J. Keshet, and E. Fetaya, “Drax: Speech Recognition with Discrete Flow Matching,” arXiv preprint arXiv:2510.04162, 2025. [17] C. Gulcehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, F. Bougares, H. Schwenk, and Y. Bengio, “On Using Monolingual Corpora in Neural Machine Translation,” Computer Speech & Language, vol. 45, pp. 137–148, 2015. [18] S. Toshniwal, A. Kannan, C.-C. Chiu, Y. Wu, T. N. Sainath, and K. Livescu, “A Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognition,” in IEEE Spoken Language Technology Workshop (SLT), 2018, pp. 369– 375. [19] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,” in Twenty-third International Conference on Machine Learning, 2006, pp. 369– 376. [20] J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. van den Berg, “Structured Denoising Diffusion Models in Discrete StateSpaces,” in The Thirty-fifth Annual Conference on Neural Information Processing Systems, 2021. [21] J. Deschenaux, C. Gulcehre, and S. S. Sahoo, “The Diffusion Duality, Chapter II: Ψ-Samplers and Efficient Curriculum,” in The Fourteenth International Conference on Learning Representations, 2026.

[5] S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J.-R. Wen, and C. Li, “Large Language Diffusion Models,” in The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.

[22] D. von Rütte, A. Orvieto, J. Fluri, O. Pooladzandi, B. Schölkopf, and T. Hofmann, “Scaling Behavior of Discrete Diffusion Language Models,” in The Fourteenth International Conference on Learning Representations, 2026.

[6] S. S. Sahoo, M. Arriola, A. Gokaslan, E. M. Marroquin, A. M. Rush, Y. Schiff, J. T. Chiu, and V. Kuleshov, “Simple and Effective Masked Diffusion Language Models,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

[23] S. S. Sahoo, J.-M. Lemercier, Z. Yang, J. Deschenaux, J. Liu, J. Thickstun, and A. Jukic, “Scaling Beyond Masked Diffusion Language Models,” arXiv preprint arXiv:2602.15014, 2026.

[7] S. S. Sahoo, J. Deschenaux, A. Gokaslan, G. Wang, J. T. Chiu, and V. Kuleshov, “The Diffusion Duality,” in Forty-second International Conference on Machine Learning, 2025. [8] T. Li, M. Chen, B. Guo, and Z. Shen, “A Survey on Diffusion Language Models,” arXiv preprint arXiv:2508.10875, 2025. [9] J. Ni, Q. Liu, L. Dou, C. Du, Z. Wang, H. Yan, T. Pang, and M. Q. Shieh, “Diffusion Language Models are Super Data Learners,” arXiv preprint arXiv:2511.03276, 2025. [10] S. Khanna, S. Kharbanda, S. Li, H. Varma, E. Wang, S. Birnbaum, Z. Luo, Y. Miraoui, A. Palrecha, S. Ermon, A. Grover, and

[24] A. Zeyer, A. Merboldt, W. Michel, R. Schlüter, and H. Ney, “Librispeech Transducer Model with Internal Language Model Prior Correction,” in Interspeech, 2021, pp. 2052–2056. [25] Z. Meng, S. Parthasarathy, E. Sun, Y. Gaur, N. Kanda, L. Lu, X. Chen, R. Zhao, J. Li, and Y. Gong, “Internal Language Model Estimation for Domain-Adaptive End-to-End Speech Recognition,” in IEEE Spoken Language Technology Workshop (SLT), 2021, pp. 243–250. [26] S. Gong, R. Zhang, H. Zheng, J. Gu, N. Jaitly, L. Kong, and Y. Zhang, “DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation,” in The Fourteenth International Conference on Learning Representations, 2026.

[27] A. Campbell, J. Benton, V. D. Bortoli, T. Rainforth, G. Deligiannidis, and A. Doucet, “A Continuous Time Framework for Discrete Denoising Models,” in The Thirty-sixth Annual Conference on Neural Information Processing Systems, 2022. [28] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR Corpus Based on Public Domain Audio Books,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5206–5210. [29] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in The Seventh International Conference on Learning Representations, 2019. [30] W. Peebles and S. Xie, “Scalable Diffusion Models with Transformers,” in IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 4172–4182. [31] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Research, vol. 15, no. 56, pp. 1929–1958, 2014. [32] T. Kudo and J. Richardson, “SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2018, pp. 66–71.

Record · ID 14019 · SHA-256 da3c7eed6aec338b
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.