Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Abstract. Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout—particularly layer dropout—has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, we observe that at the same training FLOPs layer dropout leads to lower loss. For a given number of training steps, LLMs can achieve lower or similar validation loss while saving upto 25% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.5× inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.
Training
Inference
Self-Speculative Inference
with Layer Dropout
with Static Early Exit and Layer Skipping
→ Up to 25% Faster
→ Elastic Depth
→ Upto 1.5x Speedup
Layer Layer Layer Layer
Classifier
Layer Layer
Layer Skipping
Classifier Early Exit
arXiv:2609.05275v1 [cs.AI] 4 Sep 2026
Mostafa Elhoushi† , Alex Pretko‡ , Nolan Dey† , Bin Claire Zhang† , Gavia Gray† , Gurpreet Gosal† , Abdulrahman Mahmoud‡ , Shane Bergsma† , Joel Hestness† † Cerebras Systems, ‡ MBZUAI [email protected], [email protected]
Classifier
Classifier
Classifier
Layer
Layer
Layer
Layer
Layer
Layer
Layer
Layer Layer
Adapter
Layer
Layer
Layer
Layer
Layer
Layer
Embedding
Embedding
Embedding
Embedding
Embedding
Draft
Draft
Verify
Verify
Figure 1: Layer dropout as a unified mechanism for efficient LLM training and inference. (Left) Layer dropout skips layers stochastically during pre-training, leading to faster training, and with our proposed configuration, does so without sacrificing validation loss. (Center) Trained models gain zero-shot “elastic depth,” degrading gracefully under early exit and layer skipping. (Right) This robustness carries over to post-training adapters and self-speculative decoding for lossless inference speedup.
© 2026 Cerebras Systems Inc. All Rights Reserved.
1
Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
1
Introduction
Pretraining large language models (LLMs) demands extraordinary computational resources (Narayanan et al., 2021; Meng et al., 2025), where small improvements in time-to-accuracy can save millions of dollars (Coleman et al., 2019; Shen et al., 2024) and reduce carbon emissions (Acun et al., 2023; Wang et al., 2025). Historically, regularization techniques improved validation accuracy for given training budgets by reducing overfitting and stabilizing optimization (Moradi et al., 2020; Wang & Manning, 2013; Murugan & Durairaj, 2017). Dropout was widely adopted in convolutional networks (Hinton et al., 2012) and early transformers (Vaswani et al., 2017). However, as LLMs scaled to billions of parameters and trillions of tokens, dropout has been largely abandoned (Raschka, 2025). Models trained for a single epoch over massive datasets have little opportunity for classical overfitting, and empirical evidence suggests activation dropout degrades performance under these conditions (Liu et al., 2025). One type of dropout, Layer Dropout, also known as Stochastic Depth (Huang et al., 2016), can provide benefits beyond regularization. Unlike activation dropout or unstructured sparsity, which typically do not translate into wall-clock speedups due to sparse-kernel overheads, skipping entire transformer blocks yields structured sparsity that can reduce active training FLOPs almost linearly with the dropout rate (Zhang & He, 2020; Elkerdawy et al., 2021). It also encourages robustness to reduced-depth execution at inference time, enabling a single pretrained model to dynamically adapt to different latency and compute budgets without retraining. This supports zero-shot depth-wise optimizations such as elastic depth (Fan et al., 2020), early exit (Elhoushi et al., 2024), and intermediate layer skipping (Huang et al., 2016; Cai et al., 2021). Despite layer dropout’s promise, its role in state-of-the-art LLM pretraining has never been established through a comprehensive evaluation at scale. Existing evidence is fragmented across model families, dataset sizes, and implementation conventions. Many reported degradations may reflect suboptimal schedules or hyperparameters, rather than fundamental limitations. This leaves a basic unresolved question: should layer dropout be used in modern large-scale LLM training, and if so, how should it be configured to preserve accuracy while delivering training and deployment benefits? We provide the first unified experimental study of layer dropout in LLMs, systematically varying (i) optimizer hyperparameters, (ii) depth-wise distribution and granularity of layer sparsity, and (iii) temporal dropout schedules, across fixed architecture and data. Across 2400+ training runs spanning 271M to 3.9B parameters and up to 116B tokens, we identify configurations that reliably improve training and inference efficiency. Our contributions are: 1. Improved Compute–Accuracy Trade-offs: Properly configured layer dropout reduces training FLOPs while achieving validation loss competitive with, and in several cases superior to, dense baselines at scale. 2. Joint Optimization Framework: We identify key interactions between dropout configurations, schedules, and optimizer hyperparameters that mitigate degradations observed in prior work. 3. Depth-Elastic Inference: The average training dropout rate predicts zero-shot robustness to early exit and layer skipping without retraining. 4. Scaling Analysis and Best Practices: We analyze performance across model and data scales, recommending a progressively increasing distribution across depth paired with a decreasing schedule across steps, yielding up to 25% training FLOPs savings and up to 1.5× inference speedup.
2
Related Work
Dropout granularity and scope. Dropout encompasses a family of techniques that differ in the granularity at which stochastic sparsity is applied. Prior work distinguishes activation-level dropout, weight-level dropout (e.g., DropConnect (Wan et al., 2013)), and structured dropout that operates on groups of parameters such as channels, layers, or blocks (Salehin & Kang, 2023). In this paper, we focus exclusively on structured, depth-wise dropout—i.e., stochastic removal of entire transformer blocks during training—commonly referred to as layer dropout or stochastic depth (Huang et al., 2016). We do not study neuron-level or weight-level © 2026 Cerebras Systems Inc. All Rights Reserved.
2
Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference dropout, which induce fine-grained sparsity and are known to interact differently with hardware efficiency and optimization dynamics. Dropout in large-scale LLM pretraining. As language models scaled to billions of parameters and trillion-token datasets, explicit regularization techniques—including activation dropout—have largely disappeared from state-of-the-art pretraining recipes. Early decoder-only models such as GPT-3 (Brown et al., 2020) and OPT (Zhang et al., 2022) retained the dropout settings inherited from Vaswani et al. (2017), while later models such as PaLM (Chowdhery et al., 2022) applied dropout only during finetuning, and LLaMA-style models no longer explicitly document its use. Recent empirical studies further suggest that activation dropout can degrade performance in single-epoch, large-data regimes (Liu et al., 2025), reinforcing the prevailing view that dropout is unnecessary or harmful at scale. At the same time, dropout has been shown to remain beneficial in multi-epoch or data-limited settings (Xue et al., 2023), indicating that its utility is highly regime-dependent. Importantly, techniques originally introduced as regularizers may persist in modern LLM training for reasons unrelated to overfitting prevention. Weight decay, for example, has been shown to primarily influence optimization dynamics rather than classical generalization in large-scale pretraining (D’Angelo et al., 2024). This motivates re-examining dropout—particularly structured variants—without assuming that its value must stem from regularization in the traditional sense. Layer dropout across scale. Layer dropout was originally proposed to stabilize optimization in very deep residual networks (Huang et al., 2016) and later become standard in large-scale vision models. However, its optimal strength has been observed to diminish as dataset scale increases: for example, ConvNeXt models trained on ImageNet-22K require substantially lower dropout rates than those trained on ImageNet-1K (Liu et al., 2022). A similar pattern appears in language modeling. Progressive Layer Dropout (Zhang & He, 2020) and LayerDrop (Fan et al., 2020) reported improved convergence and robustness in BERT-era, multi-epoch settings on relatively small corpora. In contrast, more recent work applying layer dropout to decoder-only LLMs trained on large token budgets has reported non-negligible accuracy degradation (Elhoushi et al., 2024), suggesting that naive extensions of earlier recipes may not transfer to modern regimes. To date, the literature lacks a controlled, large-scale evaluation that reconciles these conflicting findings by systematically varying dropout configurations and optimizer settings. Training-aware approaches to depth elasticity. Layer dropout is closely related to a broader class of training-aware methods designed to enable inference-time efficiency. In compression, approaches such as Quantization-Aware Training (QAT) consistently outperform post-training quantization by exposing the model to reduced precision during optimization (Stock et al., 2021). Analogously, depth-aware training aims to make models robust to reduced depth at inference time. Prior work has explored achieving depth elasticity via auxiliary losses, routers, or adapters added during or after pretraining, including early-exit models (Jamialahmadi et al., 2025), routing-based skipping (Jiang et al., 2024; Raposo et al., 2024), and hybrid speculative decoding schemes (Zhang et al., 2024a). Other approaches train elastic architectures explicitly, such as Once-for-All (Cai et al., 2020), MatFormer (Devvrit et al., 2024), and Nemotron-Elastic (Taghibakhshi et al., 2025). While effective, these methods typically introduce architectural changes, additional parameters, or auxiliary objectives. Layer dropout occupies a distinct position within this landscape: it induces robustness to depth-wise inference optimizations directly during pretraining, without modifying the model architecture or introducing additional losses. Prior work demonstrated that this can enable elastic inference at small scales (Fan et al., 2020), but whether similar benefits can be realized at modern LLM scales without sacrificing base-model accuracy has remained unresolved.
3
Methodology
In our experiments, we train decoder-only transformers following the architecture of Celerity models (Bergsma et al., 2025b): ALiBi position embeddings (Press et al., 2022), squared ReLU activations (Zhang et al., 2024b), © 2026 Cerebras Systems Inc. All Rights Reserved.
3
Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference and Llama3 vocabulary (Grattafiori et al., 2024). Specific architectural dimensions for all model sizes are detailed in the Appendix. Our datasets are obtained from a diverse corpus of natural language text and code. To develop best practices and quantify the effects of layer dropout, we first identify optimal hyperparameters for each dropout rate (Sec. 5). We then determine optimal granularity (Sec. 6) and configuration (Sec. 7). Following Hoffmann et al. (2022), these experiments utilize a compute-optimal budget of 20 tokens-perparameter (TPP) at each model size. Subsequently, we evaluate benefits across various depth-wise inference optimizations (Sec. 8), then quantify accuracy as training scales to larger datasets (Sec. 9). We conclude with larger-scale runs with aggressive dropout rate to demonstrate its final performance and inference advantages.
4
Preliminary
General Formulation of Layer Dropout We start by denoting residual layer ℓ ∈ {0, . . . , L − 1}, of an L layer neural network at training step t ∈ {0, . . . , T − 1}, as: Hℓ+1,t = Hℓ,t + f ℓ (Hℓ,t )
(1)
where, in the domain of natural language processing, activation tensor H ∈ RB×S×d , B is batch size, S is sequence length, d is hidden dimension. When layer dropout is applied with rate pl,t , the operation of the layer during training at step t becomes: ℓ,t Hℓ+1,t = Hℓ,t + rtrain Mℓ,t f l (Hℓ,t )
(2)
where mask Mℓ,t ∈ {0, 1}B ∼ Bernoulli(1 − pℓ,t ) is a Bernoulli random vector, and rtrain is a scaling factor applied during training. rtrain is defined differently in different layer dropout literature, and we will discuss our choice later. The bth sequence of H during training is now equal to1 : ℓ,t H [b, :, :], with probability p, Hℓ+1,t [b, :, :] = ℓ,t Hℓ,t [b, :, :] + rtrain f l Hℓ,t [b, :, :] , with probability 1 − p.
(3)
While layer dropout could be implemented during training by executing f (Hℓ,t ) on all sequences b ∈ {0, 1, ..., B − 1} of H, and multiplying its output by Mℓ,t , a more efficient implementation would be to only execute f (Hℓ,t ) on sequences b ∈ { bi | Mℓ,t [bi ] = 1 }. This leads to a saving a portion p of training FLOPs of the layer. ℓ During inference, dropout is typically disabled and a distinct scaling factor, reval , is applied: ℓ Hℓ+1 = Hℓ + reval f ℓ (Hℓ )
(4)
Layer Dropout for a Transformer We denote the operation of layer ℓ ∈ {0, . . . , L − 1} of an L, layer transformer model, at time step t ∈ {0, . . . , T − 1}, during training as: l Zℓ,t = Xℓ,t + fattn (Xℓ,t ) l Xℓ+1,t = Zℓ,t + fffn (Zℓ,t )
(5)
ℓ ℓ where X, Z ∈ RB×S×d , fattn is the attention layer and fffn is the feed-forward network (FFN).2 1 For neuron dropout, i.e., the default variant of dropout introduced by
(Hinton et al., 2012), M ∈ {0, 1}{B×S×d} .
2 This is a simplified form that does not refer to layer normalization or different variants of attention and FFN, but the
subsequent formulation generalizes to different transformer variants that include pre-, post-, layer normalization, different variants or alternatives to attention, FFNs, and mixture of experts, as long as residual connection exists.
© 2026 Cerebras Systems Inc. All Rights Reserved.
4
Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference When layer dropout is applied with rate pℓ,t , the operation at transformer layer, ℓ, step, t, during training becomes: ℓ,t ℓ ℓ,t Zℓ,t = Xℓ,t + rtrain Mℓ,t attn fattn (X ) (6) ℓ,t ℓ ℓ,t Xℓ+1,t = Zℓ,t + rtrain Mℓ,t ffn fffn (Z ) and during inference becomes: ℓ,t ℓ Zℓ,t = Xℓ,t + reval fattn (Xℓ,t ) ℓ,t ℓ Xℓ+1,t = Zℓ,t + reval fffn (Zℓ,t )
5
(7)
Hyperparameters
Background To avoid the “hyperparameter lottery” phenomenon (Dey et al., 2024), and to ensure we compare against a strong baseline, we systematically optimize learning rate, batch size, and weight decay for each dropout rate before evaluating configurations. Prior literature offers varying strategies—from coupling dropout with max-norm regularization (Hinton et al., 2012) to using learning rates 10× larger than baselines (Zhang & He, 2020)—yet systematic consensus remains elusive. To our knowledge, this is the first study to perform joint optimization of these hyperparameters for layer-wise dropout. We tune a small model with dimensions depth Lbase , width dbase on dataset Dbase to determine learning rate ηbase , weight decay λbase , initialization σbase , ans batch size Bbase , then scale via µP (Yang & Hu, 2021), CompleteP (Dey et al., 2025), and Power Lines (Bergsma et al., 2025a). Dropout Scale The scaling parameters rtrain and reval from Equations 2 and 4 require careful consideration. We define layer density ρ = 1 − p. The choice of scaling parameters varies across different research work and frameworks. Moreover, they are occasionally left implicit in published papers, and we often need to inspect their source code to specify which scaling they use. The original dropout paper (Hinton et al., 2012) used rtrain = 1, reval = ρ. Standard libraries like PyTorch and TensorFlow use rtrain = 1/ρ for dropout. For layer dropout, the first stochastic depth paper (Huang et al., 2016) used rtrain = 1, reval = p; DINOv2 (Oquab et al., 2024)3 used rtrain = 1/ρ, reval = 1; fairseq4 (that implemented Fan et al. (2020)) and torchtune5 (that implemented Elhoushi et al. (2024)) set both to 1. We demonstrate that selecting rtrain = 1/ρ is critical for stable hyperparameter transfer. To determine the optimal scale factor rtrain , we follow CompleteP’s Maximal Residual Stream Update Desideratum (Dey et al., 2025), which facilitates hyperparameter transfer across model depths L. Desideratum 1 (Maximal Residual Stream Update). Each residual block’s weights should contribute order 1/L to feature movements, and each non-residual block should contribute constant order. More precisely, for all ℓ ∈ [L − 1], each block’s parameter update θ ℓ 7→ θ ℓ + ∆θ ℓ should contribute the change 1 ℓ+1 2 ∥2 ∈ Θ(1/L). Moreover, for the embedding and unembedding layers we require d1 ∥∆W0 X∥22 ∈ Θ(1) d ∥∆θ ℓ H 1 and d ∥∆WL HL ∥22 ∈ Θ(1). Since layer dropout reduces the effective depth of the network during training, we treat models with different dropout rates ρ as having different effective depths, and apply this desideratum to ensure stable initialization across these effective depths. Our coordinate checks in Fig. 2 empirically evaluate which scaling factor better satisfies stable initialization across dropout rates: rtrain = 1 fails, necessitating per-rate tuning, whereas rtrain = 1/ρ largely satisfies these checks, enabling optimal hyperparameters transfer across many layer dropout rates. Transfer Test Fig. 3 verifies that rtrain = 1/ρ enables hyperparameter transfer: optimal η, λ, and B remain constant across dropout rates. Hence, we adopt Table A.2’s transfer rules with rtrain = 1/ρ for all our ℓ+1 upcoming experiments. We set reval = 1 to ensure that Hℓ+1 eval = E[Htrain ]. 3 https://github.com/facebookresearch/dinov2/blob/main/dinov2/layers/drop_path.py 4 https://github.com/facebookresearch/fairseq/blob/main/fairseq/modules/layer_drop.py 5 https://github.com/meta-pytorch/torchtune/blob/main/torchtune/modules/layer_dropout.py
© 2026 Cerebras Systems Inc. All Rights Reserved.
5
Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference