ConceptioArchivearXiv CS
arXiv CSopen access

Post-Training Pruning for Diffusion Transformers

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Post-Training Pruning for Diffusion Transformers Chengzhi Hu1,2 , Xuewen Liu1,2 , Jing Zhang1,2 , Mengjuan Chen1 , Zhikai Li1,* , Qingyi Gu1,* 1

Institute of Automation, Chinese Academy of Sciences School of Artificial Intelligence, University of Chinese Academy of Sciences {huchengzhi2024, liuxuewen2023, zhangjing2024, chenmengjuan2016, zhikai.li, qingyi.gu}@ia.ac.cn 2

Magnitude

Wanda

DiT-Pruning

Magnitude

Wanda

DiT-Pruning

20% 40% 50%

arXiv:2607.00927v1 [cs.CV] 1 Jul 2026

Pretrained

Prompt: A giraffe looking over the corral fence in his zoo habitat. / A yellow and white train traveling down train tracks. Image quality: Comparison of different pruning methods for DiTs. We evaluate magnitude, Wanda, and ours (DiT-Pruning) on PixArtΣ model. Our method preserves modeling capability for both foreground and background of images, even under high sparsity.

Abstract Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhead and resource consumption. Posttraining pruning offers a promising solution; however, due to DiTs’ unique architectural design and parameter distribution, traditional pruning methods are inapplicable, leading to significant performance degradation. Specifically, prior methods developed for LLMs, which derive metrics through a series of approximations, amplify the relative contribution of weights in the saliency metric. In addition, weights in DiTs exhibit significantly larger magnitudes than those in LLMs. Moreover, existing pruning granularity overlooks variations in model structures. In this paper, we propose DiT-Pruning, which improves pruning performance by introducing customized saliency criteria and pruning granularity. We design a novel metric that balances the contributions of weights and activations from an energy-based perspective, enabling more effective identification of important elements. Furthermore, we observe distinct clustering patterns in the two-dimensional weight space. Accordingly, we adopt a clustering-aware pruning granularity, enabling effective sparse allocation. Extensive evaluations on various DiTs show that our method consistently preserves image quality, especially under high sparsity. For *

Corresponding author: {zhikai.li, qingyi.gu}@ia.ac.cn.

FLUX.1-dev at 512×512 resolution on MJHQ, DiT-Pruning achieves only a 0.001 loss in CLIP score at 50% sparsity, dramatically outperforming recent pruning methods.

1

Introduction

Diffusion Transformers (Peebles and Xie 2023) (DiTs) exhibit remarkable performance in image generation, fully leveraging their advantages in scalability and expressive modeling capabilities. Recent models, such as PixArt (Chen et al. 2023) and FLUX (Labs et al. 2025), have demonstrated the ability to handle complex image distributions and generate high-quality samples. However, their numerous number of parameters and prolonged inference process result in substantial computational overhead and high resource consumption, posing significant challenges for deployment in resource-constrained environments. Neural network pruning (LeCun, Denker, and Solla 1989; Hassibi, Stork, and Wolff 1993; Han et al. 2015; Frantar and Alistarh 2022) reduces computational and storage overhead by eliminating redundant weights. Post-training pruning avoids additional training and typically adopts unstructured approaches treating individual weights as sparse units. For DiTs, the massive parameter scale and complex architecture make training-based methods (Liu et al. 2018; Blalock et al.

(a) DiT transformer block’s weight distribution.

Figure 2: The distribution characteristics of parameter importance in DiTs, which exhibits distinct clustering patterns in the two-dimensional weight space of DiTs.

(b) LLM transformer block’s weight distribution.

Figure 1: Weight distributions of DiT and LLM transformer layers. Weights in DiT are predominantly distributed within the range of 0 to 1 and show generally larger magnitudes, while LLM weights are concentrated in the range 0 to 0.1. 2020) prohibitively costly, whereas post-training pruning is training-free and offers strong scalability and promising potential. Its effectiveness depends on two aspects: the saliency criteria used to measure parameter importance, and the pruning granularity that dictates sparsity allocation across different structural components like channels or layers. The proper design of both factors is intrinsically associated with the architectural design and the parameter distribution. Early CNN pruning (Molchanov et al. 2016) employs Taylor-based saliency metrics. Similar criteria (Michel, Levy, and Neubig 2019) have been adopted to prune attention heads in ViTs. Recently, several methods have been proposed for LLM pruning. For instance, SparseGPT (Frantar and Alistarh 2023) based on Optimal Brain Surgeon(OBS) theory (Hassibi, Stork, and Wolff 1993) prunes weights that minimally impact output perturbations, while Wanda (Sun et al. 2023) introduces second-order approximations for computational efficiency. However, applicable pruning methods for Diffusion models remain largely unexplored. Some efforts have targeted U-Net-based diffusion models, for example, DiffPruning (Fang, Ma, and Wang 2023), which proposes dynamic pruning strategies incorporating timestep information and relies on retraining to restore performance. Nevertheless, due to the unique inference paradigm, architectural design, and parameter distribution of DiTs, existing pruning methods are inapplicable, leading to significant performance degradation, especially at high sparsity rates. To this end, we thoroughly analyze the characteristics of DiTs from the perspectives of saliency criteria and pruning granularity. Regarding saliency criteria, many pruning methods developed for LLMs, represented by Wanda (Sun et al. 2023), construct efficient sensitivity estimations based on second-order approximations, achieving remarkable performance. However, these methods have the following lim-

itations: Theoretically, the approximations adopted amplify weight contributions, disrupting the balance with activations in sensitivity estimation; empirically, we observe that DiT and LLM exhibit distinct statistical properties in the relative magnitude of weights and activations. As shown in Figure 1, weights in DiTs reach significantly larger magnitudes, leading to a mismatch between weight and activation scales. This discrepancy further exacerbates errors introduced by the approximations. Furthermore, as shown in Figure 2, we observe that the metric demonstrates distinct clustering patterns in the two-dimensional weight space, indicating that DiTs are incompatible with uniform pruning granularity. In this work, we propose DiT-Pruning, an efficient posttraining pruning method for DiTs, which improves pruning performance by introducing customized saliency criteria and pruning granularity. Specifically, we apply a squared transformation on weights to mitigate the amplification effect dominated by weight magnitudes in sensitivity estimation, which alleviates the scale disparity, facilitating accurate identification of the important elements. To address the clustering patterns observed in saliency criteria, we exploit the distributional structure to dynamically adjust pruning granularity. Per-layer and per-channel strategies each demonstrate advantages under different clustering patterns, enabling more effective and robust sparsity allocation. Our main contributions are summarized as follows: • We propose DiT-Pruning, a post-training pruning method for DiTs, which effectively improves pruning performance at high sparsity utilizing parameter saliency estimation and pruning granularity. • We modulate the saliency metric from an energy-based modeling perspective, providing a stronger theoretical foundation and better alignment with DiTs’ properties. Moreover, we propose a clustering-aware pruning granularity to facilitate alignment with distribution patterns. • We conduct extensive evaluations on DiTs, demonstrating that our method consistently preserves image quality under high sparsity. On FLUX.1-dev at 512×512 resolution on MJHQ, we achieve only a 0.001 loss in CLIP score at 50% sparsity, outperforming exsisting methods.

2

Related Works

Diffusion Transformers. Diffusion Transformers (Peebles and Xie 2023) (DiTs) have attracted increasing attention and gained great popularity by replacing convolutional U-NeTs (Ronneberger, Fischer, and Brox 2015) with transformer-based architectures (Vaswani et al. 2017) for image generation. DiTs demonstrate state-of-the-art performance and are gradually emerging as a mainstream backbone for generative applications. Models like Pixart (Chen et al. 2023), Flux (Labs et al. 2025) and SANA (Xie et al. 2025) further extend DiTs to text-to-image tasks, highlighting their strong capability and scalability to model high-resolution images under complex semantic conditions. However, these advances come at the cost of rapidly growing parameter sizes and computational complexity, leading to substantial resource consumption. Post-Training Pruning. Post-training pruning is a widely adopted model compression technique (Han, Mao, and Dally 2015; Kwon et al. 2022) that aims to directly operate on pre-trained models without introducing additional training or fine-tuning costs. By identifying and removing redundant or less important weight parameters, post-training pruning produces a compact model representation, thereby reducing model sizes and computational overhead in practical deployment, while maintaining competitive performance. Magnitude pruning (Han et al. 2015; Park et al. 2020) is a classical model compression method and a strong baseline (Blalock et al. 2020) in network pruning. It prunes parameters with small absolute values by ranking weights based on magnitude. However, in complex network architectures, parameter importance is influenced by multiple factors, further undermining the reliability of magnitude-based criteria. Prior works like CNN pruning (Molchanov et al. 2016) combine gradients when designing saliency metrics. Similar criteria (Michel, Levy, and Neubig 2019) have been adopted to prune attention heads in ViTs. Recent advanced methods have been proposed for LLMs, such as SparseGPT (Frantar and Alistarh 2023) and Wanda (Sun et al. 2023) , which focus on efficient unstructured pruning. Furthermore, Llmpruner (Ma, Fang, and Wang 2023), SlimGPT (Ling, Wang, and Liu 2024) and SoBP (Wei et al. 2024) explore structured pruning. In contrast, pruning methods for diffusion models remain underexplored. Diff-Pruning (Fang, Ma, and Wang 2023) target U-Net-based (Ronneberger, Fischer, and Brox 2015) diffusion models, which introduces a timestepaware dynamic pruning but requires retraining. Due to DiTs’ unique architectural design and parameter distribution, these methods are inapplicable, and resulting in significant performance degradation. This highlights the need for efficient post-training pruning methods specifically designed for DiTs.

3

Method

In this section, we thoroughly analyze the characteristics of DiTs from the perspectives of saliency criteria and pruning granularity and propose a novel post-training pruning method which effectively improves pruning performance. The overview of our method is shown in Figure 3.

Analysis of the OBD-based Criteria Classical post-training pruning methods primarily stems from Optimal Brain Damage (OBD) (LeCun, Denker, and Solla 1989) and Optimal Brain Surgeon (OBS) (Hassibi, Stork, and Wolff 1993), which employ second-order sensitivity analysis to estimate the loss variation caused by pruning individual weights, thereby achieving unstructured sparsity. 1 hii wi2 (1) 2 OBD are shown in Eq. (1), where Ii denotes the importance of the ith weight parameter, defined as the approximate increase in the loss function when this weight is set to zero. wi denotes the magnitude of pruned weights, hii denotes the ith diagonal element of the Hessian matrix(H) of the loss function. OBS are shown in Eq. (2), which further extends OBD. The formulations explicitly show that pruning sensitivity is jointly determined by both the weight magnitude and local curvature information, providing a principled second-order criterion for network sparsification. Ii ≈

Ii ≈

1 wi2 2 (H−1 )ii

(2)

Based on these principles, a variety of pruning methods have been proposed. However, most of them are tailored to LLM architectures and become ineffective when applied to DiTs. In the following, we provide a detailed analysis from both theoretical and empirical perspectives to illustrate. Subsequent works, such as SparseGPT (Frantar and Alistarh 2023) and Wanda (Sun et al. 2023), inherit from OBD and OBS framework and extend to the post-training pruning for LLMs by approximating the Hessian with activationbased second-order statistics under strong independence assumptions. This formulation enables the pruning metric to explicitly account for the combined influence of weight terms and activation terms on pruning sensitivity estimation. To achieve computational efficiency, these methods further adopt approximations, which reduce the accuracy of sensitivity estimation and alter the relative contributions of weight and activation terms. # " |W|2  (3) Iij ≈ diag (X⊤ X + λI)−1 ij

Specifically, methods like SparseGPT formulates the pruning sensitivity as shown in Eq. (3) by employing a diagonal Hessian approximation with reconstruction compensation mechanism. In this equation, W ∈ Rdout ×din denotes the weight matrix, X ∈ RN ×din denotes the input activation matrix, where each row corresponds to a token embedding, and X⊤ X serves as an empirical Hessian approximation. The diagonal operator diag(·) retains only the diagonal elements, and the damping term λI is introduced for numerical stability and partially mitigates the bias caused by the diagonal approximation, but it remains imperfect.  diag approx |W|2 2 λ→0 Iij ≈ = (|Wij | · ∥Xj ∥2 ) diag((X⊤ X)−1 ) ij (4)

Relative contribution

The series of approximations amplify the relative contribution of weights in the saliency metric.

activation activation

DiT Block ×N

weight

compact, isotropic groups

approx

weight

Scale

balance

weight

activation

Pointwise Feedforward

apparent imbalance

adaLN

DiTs exhibit markedly different parameter characteristics compared to LLMs. weight magnitude

1e-3

1e-2

DiT peak~1e-1

1e-1

per-layer pruning

activation norm latent space

density

LLM peak~1e-2

Metric values show clustering patterns in the two-dimensional weight space.

1e0

DiTs show larger weight magnitude.

log

𝑥11

𝑥1𝑗

𝑥𝑖1

𝑥𝑖𝑗 DiTs

Scale

𝑥11 … 𝑥1𝑗

elongated, channel-aligned cluster

Multi-Head Self-Attention

𝑥21 … 𝑥2𝑗 ... … …

adaLN

𝑥𝑖1 … 𝑥𝑖𝑗 LLMs

Input Tokens

per-channel pruning

Fewer tokens in DiTs lead to smaller norms.

Figure 3: Overview. (Left) Existing saliency metrics amplify the relative contribution of weights due to a series of approximations. We introduce a squared transformation on the weights (STW) to balance them. (Right) The importance scores exhibit clear clustering patterns in the two-dimensional weight space, enabling our clustering-aware granularity (CAG) for sparse allocation.

Whereas Wanda in Eq. (4) further simplifies the estimation by discarding the reconstruction-aware optimization and taking the limit λ → 0, leading to a diagonally approximated, first-order surrogate. Here, |Wij | denotes the weight connecting the j-th input channel to the i-th output channel, and ||Xj ||2 represents the ℓ2 norm of the j-th input activation across tokens. The diagonal approximation of (X⊤ X)−1 collapses the curvature information into a channel-wise scaling factor, yielding a multiplicative form composed of weight magnitude and activation norm. Consequently, these approximations operate primarily on the activation terms, reducing their contribution to a bounded scaling factor. Meanwhile, the weight term is no longer modulated by the inverse Hessian and instead directly dominates sensitivity estimation, thus amplifying its contribution.

Pruning Metric We revisit the design of importance metrics from an energybased perspective, as illustrated in Figure 4(Right). We start from the observation that parameter pruning can be interpreted as introducing a perturbation to the model, where the resulting loss increase can be formalized via a Taylor expansion. In classical mechanics, elastic potential energy grows quadratically with displacement, where the stored energy is determined by second-order interactions of the system, which motivates modeling the loss increase as a quadratic energy form governed by the Hessian. In this view, weight importance is treated as an energy allocation problem, where each parameter contributes proportionally to its squared magnitude, weighted by local curvature. This reveals that existing importance metrics, which rely on linear estimations of weight perturbations, inherently ignore second-order effects.

Actually, as shown in Figure 4(Left), such linearization fails to capture the intrinsic quadratic energy structure of the loss landscape, leading to systematic underestimation of relative large-magnitude weights. In particular, linear estimations remain locally accurate in the small-weight regime but increasingly deviate as weight magnitude grows. Moreover, this discrepancy becomes more pronounced as weight magnitude increases, reflecting the important role of secondorder effects in large parameter regimes. This issue is particularly severe in DiTs due to their distinct scale characteristics compared to LLMs. As shown in Figure 3(Left), DiT weights exhibit a wider range and overall larger magnitudes, which leads to the breakdown of linear estimations. Besides, X ∈ RN ×din denotes the input activation matrix, N denotes the number of tokens, since DiTs operate in the latent space with significantly fewer tokens (e.g., N = 256 or 512), compared to LLMs with long token sequences (e.g., N = 2048, 4096 or 8192), the resulting ∥Xj ∥2 values are substantially smaller. Consequently, the scale imbalance further amplifies the relative dominance of weight terms, making the approximation imbalance more apparent and structurally biased in DiTs. To retain the computational efficiency of existing approximations, we adopt |Wij | and |Xj |2 as the fundamental weight and activation components. Following classical second-order pruning criteria, we preserve the separable multiplicative structure and introduce exponents (m, n) to parameterize contributions, as shown in Eq. (5). Since the scaling factors a and b do not affect within-layer ranking, they are omitted, resulting in the simplified form in Eq. (6). m

n

Iij = (a · |Wij |) · (b · ∥Xj ∥2 ) m

n

Iij = (|Wij |) · (∥Xj ∥2 )

(5) (6)

Energy inherently follows a quadratic relationship.

Quadratic Linear

Hooke’s Law

perputation loss

1e-2

OBD Theory

1e-4

Energy is stored under perturbation. large weights:Q grows faster L becomes inaccurate small weights:Q≈L

1e-6

1e-8

𝑤11

𝑤1𝑗

0

𝑤𝑖𝑗

𝑤𝑖1

… 1e-3

1e-2

1e-1

magnitude

Quadratic dependence on parameter magnitude.

𝑤𝑖1

𝑤1𝑗 …

Physic System

Parameter Pruning

displacement:Δx

perputation weight change:Δw

Spring stiffness:k

modulator Hessian matrix:H

potential energy:E system energy loss increase:Δl

𝑤𝑖𝑗

restore the energy

Parameter pruning causes loss increase.

Figure 4: (Left) Pruning loss increase grows quadratically with parameter magnitude. (Right) Analogous to elastic potential energy, loss increase can be interpreted as perturbation energy, we square the weights to restore the quadratic energy structure. 2

1

Iij = (|Wij |) · (∥Xj ∥2 ) (7) Motivated by this observation, we instantiate our parametric formulation by applying a squared transformation to the weight term (STW), corresponding to (m = 2) and (n = 1) in Eq. (6), as defined in Eq. (7). This design restores the quadratic energy structure of perturbation-induced loss variations and rebalances the relative contributions of weights and activations, while preserving the original power relationship, yielding a more faithful and reliable sensitivity estimation. Meanwhile, STW compresses the dominant DiT weight scale from 10−1 to 10−2 , bringing it closer to the regime where Wanda was originally developed for LLMs (Appendix A shows detailed statistics).

Pruning Granularity In general, pruning removes parameters with low importance scores within predefined structural groups. As shown f denote the original and pruned weight in Eq. (8), W and W matrices, respectively, while MG ∈ 0, 1|W| is a group-wise binary mask determined by the predefined pruning groups G. These groups define the comparison scope for ranking importance scores, making the choice of pruning granularity critical to the final sparsification performance. f = W ⊙ MG W (8) Existing post-training pruning methods typically adopt a uniform granularity, overlooking variations in weight distributions and model structures. In contrast, we observe that our metric exhibits distinct clustering patterns in the twodimensional weight space of DiTs, as illustrated in Figure 3(Right). Specifically, some weights form elongated, channel-aligned clusters, while others aggregate into compact, isotropic groups. These observations reveal substantial heterogeneity in importance distributions across layers and channels, suggesting that a globally uniform pruning granularity is inherently suboptimal. C

in 1 X Iij I¯i = Cin j=1

(9)

  Cout Var I¯i i=1 Gout =  2  Cout Mean I¯i i=1

(10)

Motivated by the insight, we propose a clustering-aware pruning granularity (CAG) that dynamically adjusts comparison groups according to the underlying importance distribution. Specifically, we use the averaged channel importance I¯i and the normalized heterogeneity metric Gout in Eq. (9) and Eq. (10) to characterize layer-wise importance concentration. Layers with small Gout , such as MLP and self-attention layers, exhibit compact importance patterns and are pruned at the per-layer level. In contrast, layers with large Gout , notably adaLN layers, show strong channel-wise heterogeneity and are pruned at the per-channel level. This strategy enables more effective sparsity allocation and better preserves model performance under high sparsity.

4

Experiments

Experimental Settings Models. We evaluate our method primarily focus on three types of diffusion models: the classical pre-trained DiTXL/2 model, the state-of-the-art PixArt-Σ model (Chen et al. 2023), and the FLUX.1-dev model (Labs et al. 2025). All evaluated diffusion models adopt transformer-based backbones (Vaswani et al. 2017). Datasets. Following previous works (Li et al. 2023; Zhao et al. 2024), we randomly sample prompts from COCO (Lin et al. 2014) for calibration. We evaluate DiT-XL/2 on ImageNet (Russakovsky et al. 2015), generating 10,000 images at 256×256 resolution using the official DDPM sampler, consistent with prior studies (Nichol and Dhariwal 2021; Shang et al. 2023). Robustness is further assessed with 250, 100, and 50 diffusion steps (Wu et al. 2024). For PixArt-Σ and FLUX.1-dev, we evaluate on COCO and MJHQ, generating images at 256×256 and 512×512 resolutions with default samplers. We use the first 1,024 COCO annotation prompts and randomly sample 1,024 prompts from MJHQ-5K, a subset of the MJHQ-30K dataset (Li et al. 2024).

Table 1: Quantitative comparison on DiT-XL/2 at 256×256 resolution on ImageNet under different timesteps and sparsity levels. The best result per metric is highlighted in bold.

Table 2: Image quality and text-image alignment results of PixArt-Σ on COCO across different resolutions and sparsity. Method

Params FID ↓ (M)

IS ↑

CLIP ↑

IR ↑

Dense

612.08

70.31 33.33

0.265

0.807 249.59

0.49

20%

Magnitude 493.42 Wanda 493.64 DiT-Pruning 493.46

69.44 32.92 68.60 30.62 69.01 31.84

0.266 0.264 0.265

0.821 249.16 0.727 248.03 0.793 249.52

0.49 0.50 0.51

0.543 0.631 0.688

40%

Magnitude 374.48 Wanda 374.87 DiT-Pruning 374.55

69.87 31.36 69.32 28.86 67.50 31.06

0.258 0.256 0.260

0.498 239.25 0.244 235.79 0.635 244.21

0.44 0.47 0.51

0.407 0.456 0.484

50%

Magnitude 315.02 Wanda 315.10 DiT-Pruning 315.10

92.70 21.03 89.66 20.47 69.79 29.76

0.231 0.229 0.251

-0.728 240.16 -0.923 234.46 0.183 237.60

0.23 0.27 0.42

0.312 0.381 0.385

Resolution Sparsity

Timesteps Sparsity

Method

Params FID ↓ (M)

IS ↑

sFID ↓ PR ↑ SSIM ↑

675.13

5.33

273.62 17.84

0.84

20%

Magnitude 541.35 Wanda 541.55 DiT-Pruning 541.43

5.17 5.57 5.46

272.27 17.35 276.77 18.38 276.26 17.83

0.84 0.84 0.84

0.788 0.785 0.825

40%

Magnitude 407.58 9.39 163.42 19.81 Wanda 407.94 13.09 152.66 41.74 DiT-Pruning 407.73 6.53 224.27 23.39

0.70 0.68 0.79

0.545 0.537 0.617

50%

Magnitude 340.69 42.06 51.27 34.91 Wanda 340.69 45.97 49.20 97.17 DiT-Pruning 340.69 13.86 134.94 36.18

0.41 0.62 0.66

0.474 0.461 0.534

0%

250

Dense

675.13

5.53

261.87 18.93

0.81

Magnitude 541.35 Wanda 541.55 DiT-Pruning 541.43

5.41 5.72 5.56

260.92 18.15 263.46 19.06 263.78 18.61

0.82 0.82 0.82

0.803 0.796 0.835

40%

Magnitude 407.58 12.16 150.41 22.02 Wanda 407.94 15.34 142.27 42.19 DiT-Pruning 407.73 7.19 213.22 23.58

0.66 0.65 0.77

0.559 0.541 0.627

50%

Magnitude 340.69 50.25 46.79 40.82 Wanda 340.69 50.96 43.54 97.97 DiT-Pruning 340.69 16.28 124.69 36.86

0.36 0.37 0.64

0.485 0.463 0.541

0% 20% 100

Dense

675.13

6.68

236.27 21.62

0.78

20%

Magnitude 541.35 Wanda 541.55 DiT-Pruning 541.43

6.38 6.57 6.41

237.56 20.43 238.33 20.78 242.23 20.68

0.79 0.79 0.79

0.823 0.813 0.853

40%

Magnitude 407.58 17.29 131.89 26.50 Wanda 407.94 19.20 122.57 43.60 DiT-Pruning 407.73 8.78 194.21 24.50

0.62 0.62 0.75

0.584 0.555 0.646

50%

Magnitude 340.69 62.40 37.98 50.50 0.32 Wanda 340.69 57.32 38.84 101.30 0.34 DiT-Pruning 340.69 20.29 111.95 38.95 0.60

0.508 0.473 0.559

0%

50

Dense

Baselines. We compare our method with two representative approaches. Magnitude pruning (Han et al. 2015; Park et al. 2020) is a classical pruning strategy and serves as a strong baseline in our experiments. Wanda (Sun et al. 2023) is a representative second-order pruning method originally designed for LLMs. To construct the calibration data, we collect 128 input samples from the model at randomly sampled diffusion timesteps (Chen et al. 2025). Metrics. To comprehensively assess generated image quality, we employ the following metrics: We choose Fréchet Inception Distance (FID) (Heusel et al. 2017) for fidelity evaluation, spatial FID (sFID) (Nash et al. 2021; Salimans et al. 2016) for spatial consistency, Inception Score (IS) (Barratt and Sharma 2018) for visual quality and diversity, Clipscore (CLIP) (Hessel et al. 2021) for text-image alignment, ImageReward (IR) (Xu et al. 2023) for human preference, Precision (PR) (Kynkäänniemi et al. 2019) for sample realism, SSIM (Wang et al. 2004) for structural coherence. All experiments are conducted under identical experiment settings on NVIDIA RTX A6000 GPUs.

Pruning Performance We present a comprehensive assessment of our method against prevalent baseline methods in various settings.

0%

256×256

sFID ↓ PRE ↑ SSIM ↑

612.08

62.36 36.12

0.256

0.929 274.47

0.59

20%

Magnitude 494.29 Wanda 494.51 DiT-Pruning 494.34

61.25 37.76 61.22 36.92 61.71 35.78

0.257 0.255 0.257

0.968 268.88 0.879 269.71 0.942 272.10

0.59 0.60 0.60

0.484 0.674 0.693

40%

Magnitude 375.35 Wanda 375.75 DiT-Pruning 375.43

61.12 34.87 61.26 33.33 56.74 35.98

0.252 0.247 0.253

0.587 257.66 0.360 256.87 0.803 264.47

0.53 0.57 0.62

0.378 0.475 0.514

50%

Magnitude 315.90 108.86 19.66 Wanda 315.97 93.97 19.81 DiT-Pruning 315.97 61.78 33.27

0.214 0.219 0.244

-1.142 272.80 -1.143 263.23 0.285 254.82

0.18 0.29 0.51

0.261 0.373 0.439

0%

512×512

Dense

DiT-XL/2 on ImageNet. Table 1 reports the performance of pruned DiT-XL/2 model across various pruning settings and timesteps. At 20% sparsity, a mild pruning, all methods yield similar outcomes. As sparsity increase to 40%, our method consistently outperforms the baselines and remains close to the dense model, suggesting the existence of highly effective sparse sub-networks within DiTs. The advantage becomes more pronounced at 50% sparsity, where baseline methods suffer substantial degradation while our method maintains strong generative performance. For example, under 250 sampling steps and 50% sparsity, magnitude pruning and Wanda increase FID by 36.73 and 40.64, respectively, whereas our method incurs only an 8.53 increase, demonstrating the effectiveness under aggressive sparse. PixArt-Σ on COCO. Table 2 presents the results of pruning PixArt-Σ on COCO. Our method consistently achieves competitive or superior performance in both image quality and text-image alignment. At 40% sparsity, it attains the best FID and CLIP scores, indicating strong preservation of visual quality and semantic fidelity. The advantage becomes more pronounced at 50% sparsity, where baseline methods suffer substantial degradation, while our approach maintains stable performance with lower FID and higher CLIP scores. FLUX.1-dev on COCO and MJHQ. Table 3 and 4 report the results on FLUX.1-dev across COCO and MJHQ. Similar to the observations on PixArt-Σ, all methods perform comparably at 20% and 40% sparsity, with only minor differences in image quality and text-image alignment metrics. As sparsity increases to 50%, the performance gap becomes substantially larger, where our method consistently achieves superior results and demonstrates stronger robustness.

Ablation Study To validate the effectiveness of our method, we conduct ablation studies on DiT-XL/2 over ImageNet at 256×256 resolution under different sparsity and sampling steps. We further evaluate FLUX.1-dev under structured sparsity on MJHQ at different resolutions to verify the effectiveness of STW.

Table 3: Evaluation of FLUX.1-dev on COCO and MJHQ at 256×256 resolution under increasing sparsity.

Table 5: Ablation study of each component on DiT-XL/2. Timesteps Sparsity

Method

Params FID ↓ (B)

IS ↑

CLIP ↑

IR ↑

0%

Dense

11.9

54.94 26.11

0.260

0.686 259.18

0.75

20%

Magnitude Wanda DiT-Pruning

9.53 9.54 9.53

57.05 24.32 54.48 25.61 54.54 26.37

0.257 0.259 0.259

0.607 262.01 0.681 259.97 0.689 260.67

0.73 0.74 0.74

0.692 0.786 0.836

40%

Magnitude Wanda DiT-Pruning

7.16 7.17 7.17

64.56 20.00 57.28 23.83 54.83 23.64

0.246 0.257 0.259

0.251 267.40 0.582 256.94 0.647 255.83

0.63 0.69 0.71

0.499 0.584 0.663

50%

Magnitude Wanda DiT-Pruning

5.97 5.98 5.98

74.38 16.30 62.43 19.00 56.78 23.51

0.235 0.256 0.257

-0.384 273.47 0.281 255.13 0.517 256.29

0.49 0.58 0.69

0.422 0.477 0.567

Dataset Sparsity

MJHQ

Dense

11.9

66.01 38.78

0.154

-0.964 295.08

0.73

20%

Magnitude Wanda DiT-Pruning

9.53 9.54 9.53

66.54 37.26 65.62 37.57 65.31 37.69

0.154 0.154 0.154

-1.037 296.29 -0.991 293.87 -0.965 294.94

0.68 0.72 0.72

0.675 0.784 0.828

40%

Magnitude Wanda DiT-Pruning

7.16 7.17 7.17

71.11 30.49 68.27 36.15 64.96 37.71

0.155 0.155 0.156

-1.227 308.19 -1.049 294.71 -0.992 292.77

0.57 0.61 0.69

0.493 0.561 0.645

50%

Magnitude Wanda DiT-Pruning

5.97 5.98 5.98

78.46 24.14 74.85 27.87 67.56 35.91

0.159 0.155 0.161

-1.476 317.77 -1.179 287.41 -1.089 295.79

0.42 0.44 0.60

0.434 0.452 0.537

COCO

100

40%

50%

40%

50%

0% 256×256

Method

IS ↑

CLIP ↑

IR ↑

sFID ↓ PRE ↑ SSIM ↑

0%

Dense

11.9

48.64 29.56

0.272

0.899 266.47

0.78

20%

Magnitude Wanda DiT-Pruning

9.53 9.54 9.53

50.76 26.74 48.82 26.98 48.93 28.75

0.271 0.272 0.273

0.867 270.15 0.921 265.42 0.913 265.72

0.76 0.77 0.78

0.731 0.801 0.847

40%

Magnitude Wanda DiT-Pruning

7.16 7.17 7.17

57.36 22.32 51.98 23.66 49.56 27.07

0.262 0.271 0.272

0.635 269.75 0.895 267.40 0.899 265.03

0.71 0.75 0.77

0.561 0.624 0.692

50%

Magnitude Wanda DiT-Pruning

5.97 5.98 5.98

66.03 19.77 61.44 18.73 52.42 24.10

0.249 0.265 0.271

0.133 269.02 0.684 270.97 0.828 264.81

0.57 0.62 0.76

0.493 0.504 0.603

MJHQ

0%

Dense

11.9

63.41 39.13

0.154

-0.801 298.68

0.75

20%

Magnitude Wanda DiT-Pruning

9.53 9.54 9.54

63.82 36.45 63.58 37.37 63.12 37.73

0.155 0.155 0.155

-0.811 300.94 -0.812 299.05 -0.813 298.59

0.74 0.72 0.73

0.717 0.793 0.839

40%

Magnitude Wanda DiT-Pruning

7.16 7.17 7.17

66.58 29.94 65.73 34.38 64.14 36.53

0.157 0.156 0.158

-0.916 306.87 -0.836 305.19 -0.824 299.66

0.71 0.66 0.71

0.546 0.616 0.684

50%

Magnitude Wanda DiT-Pruning

5.97 5.98 5.98

69.18 26.77 74.19 26.63 66.15 33.29

0.156 0.158 0.158

-1.139 308.31 -0.915 310.66 -0.863 304.46

0.58 0.49 0.64

0.495 0.495 0.588

COCO

Starting from Wanda as the baseline, we incrementally incorporate each component, STW and CAG, and evaluate their individual contributions. As shown in Tables 5 and 6, each component improves performance and their combination achieves the best. Appendix E shows more results. Parameter Importance Evaluated by STW. Under 50% sparsity with 50 denoising steps, STW substantially improves performance over the baseline, reducing FID and sFID by 10.21 and 8.73, respectively, under the same pruning granularity, demonstrating its effectiveness in identifying important parameters. Similar gains are observed under 2:4 and 4:8 structured sparsity on FLUX.1-dev. For instance, under 2:4 sparsity at 256×256 resolution, STW reduces FID from 79.35 to 69.21 while improving CLIP, IR, and SSIM. These results demonstrate that the proposed importance criterion generalizes to both unstructured and structured pruning.

5.53

261.87

18.93

0.81

/

42.19 42.53 23.58

0.65 0.74 0.77

0.541 0.561 0.627

Wanda 50.96 43.54 + STW 43.47 54.34 + STW + CAG 16.28 124.69

97.97 89.22 36.86

0.37 0.41 0.64

0.463 0.477 0.541

236.27

21.62

0.78

/

19.20 122.57 12.86 143.31 8.78 194.21

43.60 43.61 24.50

0.62 0.73 0.75

0.555 0.576 0.646

Wanda 57.32 38.84 101.30 + STW 47.11 47.01 92.57 + STW + CAG 20.29 111.95 38.95

0.34 0.39 0.60

0.473 0.488 0.559

Dense

6.68

Wanda + STW + STW + CAG

Resolution Sparsity Method FID ↓

Dataset Sparsity

sFID ↓ PRE ↑ SSIM ↑

Table 6: Ablation results of structured pruning on FLUX.1dev at 256×256 and 512×512 resolutions on MJHQ.

Table 4: Evaluation of FLUX.1-dev on COCO and MJHQ at 512×512 resolution under increasing sparsity. Params FID ↓ (B)

IS ↑

15.34 142.27 11.79 161.01 7.19 213.22

Wanda + STW + STW + CAG

0%

50

FID ↓

Dense

0%

sFID ↓ PRE ↑ SSIM ↑

0%

Ablation

4:8 2:4 0%

512×512

4:8 2:4

IS ↑

CLIP ↑

IR ↑

SFID ↓ PRE ↑ SSIM ↑

Dense

54.94 26.11

0.260

0.686

259.18

0.75

/

Wanda

71.56 16.38

0.243

-0.206 260.76

0.49

0.441 0.489

+ STW 63.95 20.51

0.246

0.189

257.96

0.61

Wanda

79.35 14.63

0.235

-0.613 263.35

0.42

0.412

+ STW 69.21 19.98

0.245

-0.078 254.42

0.55

0.444

Dense

48.64 29.56

0.272

0.899

266.47

0.78

/

Wanda

61.31 20.09

0.259

0.414

256.52

0.64

0.522 0.548

+ STW 57.58 23.27

0.261

0.575

254.47

0.71

Wanda

63.07 18.36

0.251

0.054

242.66

0.56

0.466

+ STW 59.89 23.32

0.255

0.340

241.22

0.65

0.497

Sparse Allocation by CAG. Furthermore, we compare fixed pruning granularity with CAG under the same importance metric. By dynamically adapting pruning granularity to the underlying importance distribution, CAG consistently improves performance across sparsity levels and denoising steps. For example, as shown in Table 5, at 40% sparsity with 100 denoising steps, FID decreases from 11.79 to 7.19, while at 50% sparsity with 50 denoising steps, it further drops from 47.11 to 20.29. These results suggest that adaptive granularity allocation enables more effective sparse allocation.

5

Conclusion

In this paper, we propose DiT-Pruning, a post-training pruning method for DiTs. Existing pruning methods suffer substantial performance degradation on DiTs due to their unique architecture, parameter distribution, and approximations that amplify weight contributions. To address this issue, we introduce a customized saliency criterion and pruning granularity to improve pruning performance. Specifically, we balance weight and activation contributions by squaring the weights, enabling more accurate identification of important elements. Furthermore, we observe that the proposed metric exhibits distinct clustering patterns in the two-dimensional weight space. Accordingly, we adopt a clustering-aware pruning granularity for effective sparse allocation. Extensive experiments show that our method significantly outperforms existing approaches, especially under high sparsity.

References Barratt, S.; and Sharma, R. 2018. A note on the inception score. arXiv preprint arXiv:1801.01973. Blalock, D.; Gonzalez Ortiz, J. J.; Frankle, J.; and Guttag, J. 2020. What is the state of neural network pruning? Proceedings of machine learning and systems, 2: 129–146. Chen, J.; Yu, J.; Ge, C.; Yao, L.; Xie, E.; Wu, Y.; Wang, Z.; Kwok, J.; Luo, P.; Lu, H.; et al. 2023. Pixart-α: Fast training of diffusion transformer for photorealistic text-toimage synthesis. arXiv preprint arXiv:2310.00426. Chen, L.; Meng, Y.; Tang, C.; Ma, X.; Jiang, J.; Wang, X.; Wang, Z.; and Zhu, W. 2025. Q-dit: Accurate post-training quantization for diffusion transformers. In Proceedings of the Computer Vision and Pattern Recognition Conference, 28306–28315. Fang, G.; Ma, X.; and Wang, X. 2023. Structural pruning for diffusion models. In Advances in Neural Information Processing Systems. Frantar, E.; and Alistarh, D. 2022. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35: 4475–4488. Frantar, E.; and Alistarh, D. 2023. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International conference on machine learning, 10323–10337. PMLR. Han, S.; Mao, H.; and Dally, W. J. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149. Han, S.; Pool, J.; Tran, J.; and Dally, W. 2015. Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28. Hassibi, B.; Stork, D. G.; and Wolff, G. J. 1993. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, 293–299. IEEE. Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021. Clipscore: A reference-free evaluation metric for image captioning. In Proceedings of the 2021 conference on empirical methods in natural language processing, 7514– 7528. Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30. Kwon, W.; Kim, S.; Mahoney, M. W.; Hassoun, J.; Keutzer, K.; and Gholami, A. 2022. A fast post-training pruning framework for transformers. Advances in Neural Information Processing Systems, 35: 24101–24116. Kynkäänniemi, T.; Karras, T.; Laine, S.; Lehtinen, J.; and Aila, T. 2019. Improved precision and recall metric for assessing generative models. Advances in neural information processing systems, 32. Labs, B. F.; Batifol, S.; Blattmann, A.; Boesel, F.; Consul, S.; Diagne, C.; Dockhorn, T.; English, J.; English, Z.; Esser,

P.; et al. 2025. FLUX. 1 Kontext: Flow Matching for InContext Image Generation and Editing in Latent Space. arXiv preprint arXiv:2506.15742. LeCun, Y.; Denker, J.; and Solla, S. 1989. Optimal brain damage. Advances in neural information processing systems, 2. Li, D.; Kamko, A.; Akhgari, E.; Sabet, A.; Xu, L.; and Doshi, S. 2024. Playground v2. 5: Three insights towards enhancing aesthetic quality in text-to-image generation. arXiv preprint arXiv:2402.17245. Li, X.; Liu, Y.; Lian, L.; Yang, H.; Dong, Z.; Kang, D.; Zhang, S.; and Keutzer, K. 2023. Q-diffusion: Quantizing diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 17535–17545. Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, 740–755. Springer. Ling, G.; Wang, Z.; and Liu, Q. 2024. Slimgpt: Layer-wise structured pruning for large language models. Advances in Neural Information Processing Systems, 37: 107112– 107137. Liu, Z.; Sun, M.; Zhou, T.; Huang, G.; and Darrell, T. 2018. Rethinking the value of network pruning. arXiv preprint arXiv:1810.05270. Lu, Y.; and Liu, W. 2023. Dasp: Specific dense matrix multiply-accumulate units accelerated general sparse matrixvector multiplication. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 1–14. Ma, X.; Fang, G.; and Wang, X. 2023. Llm-pruner: On the structural pruning of large language models. Advances in neural information processing systems, 36: 21702–21720. Macko, V.; and Boža, V. 2025. MACKO: Sparse MatrixVector Multiplication for Low Sparsity. arXiv preprint arXiv:2511.13061. Michel, P.; Levy, O.; and Neubig, G. 2019. Are sixteen heads really better than one? Advances in neural information processing systems, 32. Molchanov, P.; Tyree, S.; Karras, T.; Aila, T.; and Kautz, J. 2016. Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440. Nash, C.; Menick, J.; Dieleman, S.; and Battaglia, P. W. 2021. Generating images with sparse representations. arXiv preprint arXiv:2103.03841. Nichol, A. Q.; and Dhariwal, P. 2021. Improved denoising diffusion probabilistic models. In International conference on machine learning, 8162–8171. PMLR. Park, S.; Lee, J.; Mo, S.; and Shin, J. 2020. Lookahead: A far-sighted alternative of magnitude-based pruning. arXiv preprint arXiv:2002.04809. Peebles, W.; and Xie, S. 2023. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, 4195–4205.

Ronneberger, O.; Fischer, P.; and Brox, T. 2015. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, 234–241. Springer. Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3): 211–252. Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; Chen, X.; and Chen, X. 2016. Improved Techniques for Training GANs. In Lee, D.; Sugiyama, M.; Luxburg, U.; Guyon, I.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc. Shang, Y.; Yuan, Z.; Xie, B.; Wu, B.; and Yan, Y. 2023. Posttraining quantization on diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 1972–1981. Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2023. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30. Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4): 600–612. Wei, J.; Lu, Q.; Jiang, N.; Li, S.; Xiang, J.; Chen, J.; and Liu, Y. 2024. Structured optimal brain pruning for large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 13991–14007. Wu, J.; Wang, H.; Shang, Y.; Shah, M.; and Yan, Y. 2024. Ptq4dit: Post-training quantization for diffusion transformers. Advances in neural information processing systems, 37: 62732–62755. Xie, E.; Chen, J.; Chen, J.; Cai, H.; Tang, H.; Lin, Y.; Zhang, Z.; Li, M.; Zhu, L.; Lu, Y.; et al. 2025. SANA: Efficient high-resolution text-to-image synthesis with linear diffusion transformers. In The Thirteenth International Conference on Learning Representations. Xu, J.; Liu, X.; Wu, Y.; Tong, Y.; Li, Q.; Ding, M.; Tang, J.; and Dong, Y. 2023. Imagereward: Learning and evaluating human preferences for text-to-image generation. Advances in Neural Information Processing Systems, 36: 15903–15935. Zhao, T.; Ning, X.; Fang, T.; Liu, E.; Huang, G.; Lin, Z.; Yan, S.; Dai, G.; and Wang, Y. 2024. Mixdq: Memoryefficient few-step text-to-image diffusion models with metricdecoupled mixed precision quantization. In European Conference on Computer Vision, 285–302. Springer.

A

Weights and Activations

Weights. Due to space limitations, we present the weight visualizations of only two representative layers in the main text. Similar distribution patterns are consistently observed across all layers of the DiTs, indicating that the observed phenomenon is a general characteristic of the model rather than an isolated case. For completeness, the weight distributions of all remaining model layers are presented below. Figure 5 shows detailed information, weights are predominantly distributed around the 10−1 scale, which is consistent with the distribution described before.

Activations. Table 7 and Table 8 report the activation statistics and weight-to-activation ratios of representative DiT and LLM layers, respectively. Owing to their much longer token sequences, LLMs exhibit significantly larger activation outliers. Despite this difference, the relative scales of weights and activations remain substantially more balanced in LLMs than in DiTs. STW reduces the dominant DiT weight scale from 10−1 to 10−2 , yielding a quantitative relationship between weights and activations that is closer to the regime where Wanda has been shown to be effective. Table 8: Statistics of activation norms of LLM layers. Layer

(a) DiT transformer block’s weight distribution.

Figure 5: Remaining layers’ weight distribution of DiT transformer block. The weights in DiT are predominantly distributed within the range of 0 to 1, all the layers show the same characteristics. Table 7: Statistics of activation norms of DiT layers. Layer

Metric

block5

block10 block15 block20 block25

|X| mean 1.1012 1.6378 |X| max 22.3234 10.8295 atten.qkv |W |/|X|mean 0.1052 0.0979 |W |2 /|X|mean 0.0122 0.0157

1.7370 9.2846 0.0972 0.0164

1.8147 9.0306 0.0905 0.0149

1.9152 9.4552 0.0832 0.0133

|X| mean 4.9624 8.5670 11.7076 9.9487 13.3197 |X| max 16.7251 16.3386 36.9063 27.2713 36.1129 atten.proj |W |/|X|mean 0.0283 0.0213 0.0159 0.0193 0.0149 |W |2 /|X|mean 0.0039 0.0039 0.0037 0.0037 0.0029

mlp.fc1

mlp.fc2

adaLN

|X| mean 1.5179 1.5725 |X| max 17.1799 16.0704 |W |/|X|mean 0.0936 0.1047 |W |2 /|X|mean 0.0133 0.0172

1.1350 8.7672 0.1453 0.0239

1.1803 1.4208 22.1027 41.0000 0.1373 0.1098 0.0223 0.0171

|X| mean 2.7627 3.4125 2.9495 2.8912 3.0111 |X| max 16.0696 25.4478 22.8027 27.8377 36.2336 |W |/|X|mean 0.0589 0.0532 0.0647 0.0609 0.0646 |W |2 /|X|mean 0.0095 0.0097 0.0127 0.0129 0.0126 |X| mean |X| max |W |/|X|mean |W |2 /|X|mean

0.0084 0.3409 11.45 1.1017

0.0084 0.3409 16.42 2.2638

0.0084 0.3409 16.00 2.1504

0.0084 0.3409 13.76 1.5908

0.0084 0.3409 13.07 1.4274

Metric

block5

block10

block15

block20

block25

atten.q

|X| mean 11.4072 14.1838 15.8599 17.7827 19.9309 |X| max 214.7673 229.1075 210.9702 185.4212 189.7775 |W |/|X|mean 0.0019 0.0015 0.0013 0.0011 0.0009

atten.k

|X| mean 11.4072 14.1838 15.8599 17.7827 19.9309 |X| max 214.7673 229.1075 210.9702 185.4212 189.7775 |W |/|X|mean 0.0019 0.0009 0.0013 0.0011 0.0009

atten.v

|X| mean 11.4072 14.1838 15.8599 17.7827 19.9309 |X| max 214.7673 229.1075 210.9702 185.4212 189.7775 |W |/|X|mean 0.0012 0.0009 0.0013 0.0009 0.0009

atten.o

|X| mean |X| max |W |/|X|mean

1.7841 14.9358 0.0073

5.1910 37.6035 0.0027

6.5438 26.8657 0.0023

6.4441 43.6976 0.0027

6.3077 36.2537 0.0029

mlp.up

|X| mean |X| max |W |/|X|mean

6.9415 95.0688 0.0024

8.6943 83.8602 0.0020

10.9269 53.8245 0.0016

14.2518 43.4382 0.0012

16.3169 58.0632 0.0011

mlp.gate

|X| mean |X| max |W |/|X|mean

6.9415 95.0688 0.0027

8.6943 83.8602 0.0021

10.9269 53.8245 0.0016

14.2518 43.4382 0.0013

16.3169 58.0632 0.0011

mlp.down

|X| mean |X| max |W |/|X|mean

1.7290 16.3258 0.0096

2.9034 18.8423 0.0058

4.5845 31.4849 0.0038

7.4560 65.0971 0.0024

8.5851 55.4523 0.0021

B

Compression and Speedup Effects

Compression and Acceleration. DiT-Pruning introduces no additional parameters and therefore preserves the inherent benefits of sparsification. As shown in Table 9, both parameter count and computational cost decrease consistently with increasing sparsity. TFLOPs is used to measure the acceleration effect, Params is used to measure the storage reduction. Increasing compression sparsity results in significant reduction in inference cost, demonstrating its effectiveness. Table 9: Model compression and inference efficiency of DiTPruning on DiT-XL/2 under different sparsity levels. Model

Method

DiT-XL/2 DiT-Pruning

Sparsity Params↓ TFLOPs↓ TF/50steps↓ Speedup↑ 0% 20% 40% 50% 2:4

675.13M 541.43M 407.73M 340.69M 340.68M

0.114 0.092 0.069 0.058 0.058

5.723 4.589 3.456 2.888 2.888

/ ∼1.2× ∼1.4× ∼1.6× ∼1.6×

Acceleration Analysis. Existing sparse inference frameworks (Lu and Liu 2023; Macko and Boža 2025) have demonstrated substantial acceleration for unstructured sparse computation, while prior work reports approximately (1.6×) runtime speedup for Transformer layers under 2:4 structured

sparsity (Sun et al. 2023). Since DiT-Pruning follows the same unstructured and 2:4 sparsity patterns without introducing additional modules, the achieved reductions in model size and FLOPs can be directly translated into practical efficiency gains under these sparse execution frameworks.

Table 10: Effect of calibration set size under different sparsity. Num. Calib.

128

C

Experiment Settings

Parameter Settings. For DiT-XL/2, we follow the standard ImageNet generation setup with a classifier-free guidance (CFG) scale of 1.5, a fixed random seed of 42, and the MSE VAE for latent coding. Evaluations are conducted under different timesteps. For PixArt-Σ, we use DPM-Solver with 20 denoising steps and a CFG scale of 4.5. For FLUX.1-dev, we adopt 50 inference steps with a guidance scale of 3.5. For fair comparison across resolutions, all reference images are resized to the target resolution when necessary. Evaluations. We evaluate model performance using generation quality, text-image alignment, and structural consistency metrics. Generation quality is assessed by FID, sFID, Precision, and Inception Score (IS), following the official evaluation pipeline of OpenAI’s guided-diffusion repository with the corresponding reference batches. Text-image alignment is measured using CLIP with the official pretrained weights. For FLUX.1-dev, we additionally report Image Reward (IR), following the official evaluation protocol. Structural consistency is evaluated using SSIM. All evaluation settings follow the corresponding official implementations.

D

Calibration Samples

Number of Calibration Samples. To evaluate the sensitivity to calibration set size, we progressively increase the number of calibration samples while keeping all other settings fixed. Experiments are conducted on DiT-XL/2 at 256×256 resolution with 50 denoising steps and 5,000 generated samples. As shown in Table 10, the performance remains stable across a wide range of calibration sizes, demonstrating strong robustness to calibration data selection. Moreover, increasing the calibration set beyond 128 samples yields only marginal improvements at all sparsity levels, indicating that 128 samples are sufficient for reliable pruning. Selection of Calibration Samples. We further investigate the impact of calibration data selection. Specifically, we fix the calibration set size to 128 and sample calibration data from different diffusion timesteps. As shown in Figure 6, pruning performance is highly sensitive to timestep coverage. Calibration data drawn exclusively from early or late stages consistently lead to inferior results, whereas samples covering a broad timestep range achieve substantially better performance. This observation suggests that different timesteps exhibit distinct activation distributions and parameter sensitivity patterns in DiTs. Therefore, effective calibration data should provide diversity not only in sample content but also across diffusion timesteps, ensuring representative activation statistics and more reliable pruning decisions.

E

More Ablation Materials

Due to space limitations, the main text reports ablation results only on DiT-XL/2 at 40% and 50% sparsity and N:M

256

512

1024

Sparsity

FID ↓

sFID ↓

IS ↑

PRE ↑

SSIM ↑

20%

10.32

250.38

37.52

0.79

0.853

40%

12.75

194.92

40.46

0.74

0.649

50%

24.14

111.01

53.50

0.59

0.560

20%

10.31

251.53

37.59

0.78

0.853

40%

12.76

198.04

40.36

0.74

0.649

50%

23.99

112.32

53.86

0.59

0.562

20%

10.47

250.96

37.45

0.77

0.854

40%

12.76

197.51

40.33

0.73

0.651

50%

23.98

111.94

53.82

0.59

0.561

20%

10.45

251.55

37.32

0.78

0.853

40%

12.75

197.33

40.32

0.75

0.649

50%

24.01

112.03

53.01

0.58

0.560

Quality Perception Metrics and Structural Consistency Metrics

Calibration samples are sensitive to early, middle, and late timesteps.

Figure 6: Calibration timestep evaluation.

structured pruning results on FLUX. Table 11 extends the study to PixArt-Σ across different resolutions and sparsity levels. Similar to the observations on DiT-XL/2, STW consistently improves the baseline, and CAG provides additional gains when combined with STW. The resulting indicates that the two components are complementary and generalize well across different DiT-based architectures. Effect of CFG Scale. We evaluate pruning performance under different classifier-free guidance (CFG) scales while keeping the sparsity ratio and denoising steps fixed. Following the official DiT setup, CFG=1.5 is used as the default configuration, while CFG=1 and CFG=2 are adopted to examine weaker and stronger guidance strengths, respectively. As shown in Table 12, varying CFG changes the trade-off between conditional guidance and generation diversity. Nevertheless, the proposed method consistently achieves the best performance across all CFG settings, maintaining its advantage under both weaker and stronger guidance. These results demonstrate the robustness of our method to variations in guidance strength and sampling configuration. Effect of Optional (m, n) Settings. To validate the choice of the proposed squared transformation, we perform an ablation study by varying the exponents (m, n) in Eq. (6), while keeping the CAG strategy fixed. As shown in Table 13, the proposed setting (m = 2, n = 1) consistently achieves the

Pretrained

Table 11: Ablation study of each component on PixArt-Σ. Method

FID ↓

IS ↑

CLIP ↑

IR ↑

Dense

70.31 33.33

0.265

0.807 249.59

0.49

/

40%

Wanda 69.32 28.86 + STW 67.88 30.12 + STW + CAG 67.5 31.06

0.256 0.260 0.260

0.244 235.79 0.593 239.47 0.635 244.21

0.47 0.48 0.51

0.456 0.468 0.484

50%

Wanda 89.66 20.47 + STW 75.7 26.52 + STW + CAG 69.79 29.76

0.229 0.247 0.251

-0.923 234.46 -0.122 231.97 0.183 237.6

0.27 0.37 0.42

0.381 0.382 0.385

62.36 36.12

0.256

0.929 274.47

0.59

/

40%

Wanda 61.26 33.33 + STW 59.63 35.29 + STW + CAG 56.74 35.98

0.247 0.252 0.253

0.360 256.87 0.754 259.69 0.803 264.47

0.57 0.61 0.62

0.475 0.483 0.514

50%

Wanda 93.97 19.81 + STW 73.27 28.47 + STW + CAG 61.78 33.27

0.219 0.242 0.244

-1.143 263.23 -0.154 253.74 0.285 254.82

0.29 0.41 0.51

0.373 0.428 0.439

Resolution Sparsity 0%

256×256

0%

512×512

Dense

sFID ↓ PRE ↑ SSIM ↑

Magnitude

Table 12: Effect of CFG scale on DiT-XL/2 pruning. Timesteps Sparsity CFG Scale

50

50%

Method

FID ↓

IS ↑

sFID ↓ PRE ↑ SSIM ↑

2.0

Magnitude 38.14 72.04 Wanda 33.39 81.99 DiT-Pruning 10.82 199.70

40.08 77.78 34.02

0.44 0.49 0.73

0.516 0.491 0.565

1.5

Magnitude 62.40 37.98 50.50 Wanda 57.32 38.84 101.30 DiT-Pruning 20.29 111.95 38.95

0.32 0.34 0.60

0.508 0.473 0.559

1.0

Magnitude 97.76 Wanda 90.50 DiT-Pruning 46.54

0.22 0.21 0.42

0.220 0.316 0.519

18.62 17.23 43.67

Wanda

67.32 135.69 48.98

DiT-Pruning

best overall performance, yielding the lowest FID and highest IS, sFID, PRE, and SSIM. In contrast, both smaller and larger weight exponents lead to inferior results. These observations suggest that (m = 2, n = 1) provides a more suitable balance between weight and activation contributions, supporting the effectiveness of the proposed STW. Table 13: Effect of (m, n) settings under CAG on DiT-XL/2. Timesteps Sparsity

Settings m = 1, n = 1

50

50%

F

FID ↓

IS ↑

sFID ↓ PRE ↑ SSIM ↑

22.23 100.91 41.71

m = 0.5, n = 1 41.92

53.64

57.71

0.57

0.557

0.42

0.521

m = 3, n = 1

22.76 104.36 42.27

0.57

0.547

m = 2, n = 1

20.29 111.95 38.95

0.60

0.559

Additional Visualization Results

Figure 7, Figure 8-9, and Figure 10-13 provide qualitative comparisons between DiT-Pruning, Wanda, and the dense models under 50% sparsity level. Across all settings, our method generates images that remain visually closer to the dense model, preserving both structural details and semantic consistency. In contrast, Wanda exhibits noticeable degradation as sparsity increases, including distorted structures, blurred content, and weakened semantic alignment. These results demonstrate that DiT-Pruning effectively preserves generation quality under aggressive sparsification and generalizes consistently across different DiT-based architectures.

Figure 7: Random samples generated by the pruned DiTXL/2 model at 50% sparsity and a resolution of 256×256. Our method preserves image fidelity, fine-grained details and structural integrity, while baseline methods exhibit substantial degradation in visual quality.

G

Limitations and Broader Impacts

Our method demonstrates effective pruning of DiTs, achieving favorable generation quality at different sparsity levels (e.g., 20%, 40%, 50%). However, several limitations remain. Under higher sparsity ratios (e.g., 70% and 80%), generated images exhibit noticeable degradation, including artifacts, reduced fidelity, and inaccurate semantics. This indicates that our pruning strategy, while efficient, cannot fully preserve model performance under extreme sparsity and suggests the trade-off between compression ratio and generation quality. Moreover, DiTs appear inherently more sensitive to pruning than LLMs, highlighting the need for strategies that better exploit their denoising dynamics and timestep-dependent characteristics. Finally, our experiments focus on standard image generation benchmarks, and the robustness of our approach in more diverse visual domains remains unexplored.

Pretrained

Magnitude

Wanda

DiT-Pruning

Pretrained

Magnitude

Wanda

DiT-Pruning

Prompt : A man riding down a snow covered slope in the snow.

Prompt : A motorbike sitting in front of a wine display case.

Prompt : A pastry meal sitting on a trey next to a bottle of orange soda.

Prompt : An elephant walks down a dirt road with brush on either side.

Prompt : A little girl holding a blow dryer next to her head.

Prompt : The van is driving down the street in traffic.

Prompt : A group of people flying kites in a blue cloudy sky.

Prompt : A big commercial plane flying high in the sky.

Prompt : THERE IS AN ADULT CAT THAT IS LOOKING AT SOMETHING .

Prompt : A red fire hydrant on the side of a street.

Prompt : a large cow stairs across the snowy fields.

Prompt : A large doughnut sign above a shop for doughnuts.

Prompt : A baseball game being played before a crowd.

Prompt : A wine bottle and glass sit on a table in front of a couch .

Prompt : A bus and car wait at an intersection on a city street.

Figure 8: Samples generated by pruned PixArt model at 50% sparsity and 256x256 resolution on Coco datasets.

Prompt : A wooden clock sitting up against a white wall.

Figure 9: Samples generated by pruned PixArt model at 50% sparsity and 512x512 resolution on Coco datasets.

Pretrained

Magnitude

Wanda

DiT-Pruning

Prompt : A table holds a cheese pizza and condiments on a checkered tablecloth.

Pretrained

Magnitude

Wanda

DiT-Pruning

Prompt : An orange and white cat laying on top of black shoes.

Prompt : A beautiful girl sitting at a table with an orange.

Prompt : a number of people on bikes under a traffic light

Prompt : A photo of a man swinging a tennis racket.

Prompt : A train traveling down train tracks next to a building.

Prompt : A black and white dog is catching a red frisbee.

Prompt : a public transit bus parked with its doors open.

Prompt : THERE IS A WALL WITH FLOWERS ON IT AND A BIRD .

Prompt : A toddler laying in a bed with their head on the pillow.

Prompt : Several motorcycles that are parked on the side of the street.

Prompt : there are many pieces of broccoli and vegetables here.

Prompt : The young girl runs toward the net to meet the tennis ball.

Prompt : A plate of food with shrimp, pasta, and salad.

Prompt : a big body of water with a freeze be next to it.

Prompt : A european city in nice a sunny bright day.

Figure 10: Samples generated by pruned Flux model at 50% sparsity and 256x256 resolution on Coco datasets.

Figure 11: Samples generated by pruned Flux model at 50% sparsity and 512x512 resolution on Coco datasets.

Pretrained

Magnitude

Wanda

DiT-Pruning

Pretrained

Magnitude

Wanda

DiT-Pruning

Prompt : tograph of a 19yearold girl, with long dark blonde hair and striking green eyes.

Prompt : Countryside with farms and houses and abandoned roadstelepoles.

Prompt : a group of diverse students working together.

Prompt : still life photo, 32k uhd, extreme detailed, open refrigerator, fridge, fresh meat.

Prompt : Design a photorealistic village night landscape set in Bangladesh in 1970s.

Prompt : XXL wooden barrel to be converted into a summer bar table.

Prompt : For years, the squirrel had watched the humans playing golf from a safe distance.

Prompt : a german empire millitary parade in streets of a french city.

Prompt : space, stars and litell planet, create 9d styl.

Prompt : fresh popcorn, isolate with 0e4321 color,

Prompt : a girly sticker design that represents being a lucky woman lucky.

Prompt : Portrait of a woman by Leonardo da Vinci, 15th Century,

Prompt : a dog laying on his back smiling surrounded by tennis balls,

Prompt : Take a delightful photograph of a woman working diligently in the garden.

Prompt : hazel eye close up with a reflection of flames of fire in the iris.

Prompt : a macro shot of an eyeball, with the iris in focus, reveals a rainbow of colored fibers.

Figure 12: Samples generated by pruned Flux model at 50% sparsity and 256x256 resolution on MJHQ datasets.

Figure 13: Samples generated by pruned Flux model at 50% sparsity and 512x512 resolution on MJHQ datasets.

Record · ID 329137 · SHA-256 bb22bbb5de2ea9f0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.