IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
1
MMOE: Modernizing Diffusion Transformers with Efficient Expert Design
arXiv:2607.24665v1 [cs.CV] 27 Jul 2026
Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, Fellow, IEEE, Xuelong Li, Fellow, IEEE
Abstract—Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that made LLM scaling practical, so generation quality is seldom balanced against training and deployment cost. This raises a natural question: can the architectural principles behind efficient LLM scaling be adapted to AFMs in a more balanced way? We introduce ModernMOE (MMOE), a modernization of SiT-style diffusion transformers that systematically adapts routed experts, shared and lightweight experts, gate-residual routing, and attention-residual information reuse to AIGC generation. Rather than treating MoE as a single plug-in replacement, MMOE studies how different modern expert components affect convergence, efficiency, and generation quality when composed inside a diffusion transformer. Every experiment in this paper is trained on a single eight-GPU H100 node with batch size 256 for 400k steps, an accessible singlemachine budget. Under matched training and sampling protocols and at this budget, MMOE reaches lower FID at every recorded checkpoint, that is, it converges faster per training step, than dense and intermediate sparse-expert baselines, and among the sparse variants it attains the best quality-cost balance. Routing analysis further shows stable expert specialization across depth, substantial use of lightweight routes, and modest step-to-step routing changes during denoising. These results suggest that AFMs can follow the balanced scaling path of LLMs by importing proven efficiency designs, rather than by simply increasing total parameters and sparsity ratios. Index Terms—AIGC Foundation Models, Diffusion Transformers, Mixture of Experts, Efficient Generative Models, Expert Routing.
I. I NTRODUCTION
T
RANSFORMER scaling has reshaped both language modeling and visual generation, but the two fields have followed different architectural trajectories. Modern large language models increasingly rely on sparse and high-throughput computation: sparsely gated experts decouple total capacity from activated computation [1], large-scale systems such as Yanhao Jia and Jiepeng Wang contributed equally to this work. This work was done during Yanhao’s internship at TeleAI. Corresponding authors: Erik Cambria and Xuelong Li. Yanhao Jia and Erik Cambria are with the College of Computing and Data Science, Nanyang Technological University, Singapore, 639798 (E-mail: [email protected], [email protected]). Jiepeng Wang, Haibin Huang, Chi Zhang and Xuelong Li are with the Institute of Artificial Intelligence of China Telecom (TeleAI), Shanghai, China, 200232 (E-mail: [email protected], [email protected], [email protected], [email protected]).
GShard and Switch Transformers make conditional computation practical at scale [2], [3], and recent LLMs further popularize top-k expert routing as a standard way to enlarge model capacity while controlling per-token cost [4]. More recent MoE designs go beyond simply adding routed FFN experts. For example, MoE++ introduces zero-computation experts such as zero, copy, and constant experts, together with gate-residual routing, to improve both efficiency and effectiveness [5]. These developments suggest that the progress of foundation models is no longer only a matter of increasing parameter count, but also a matter of designing how computation is activated, skipped, reused, and routed. AIGC Foundation Models (AFMs), especially diffusiontransformer backbones, have also benefited from transformer scaling. DiT demonstrates that replacing convolutional denoisers with transformers yields a scalable generative backbone [6], while SiT provides a closely related scalable interpolant-transformer formulation for flow and diffusion generation [7]. However, scaling AFMs remains expensive because generation requires repeated denoising, and each denoising step processes many visual tokens through deep transformer blocks. Recent work has therefore begun to explore MoE for diffusion transformers. DiT-MoE studies shared expert routing and expert-level balancing for sparse diffusion transformers [8]. EC-DIT uses expert-choice routing to allocate computation according to heterogeneous generation complexity [9]. Diff-MoE introduces time-aware and spaceadaptive experts for denoising stages and spatial tokens [10]. Other studies investigate routing competition, explicit routing guidance, and practical recipes for training diffusion MoE models [11]–[13]. Together, these works show that MoE is a promising direction for AFMs. A recurring pattern, however, is to scale mainly by enlarging the total expert pool and the sparsity ratio: DiT-MoE trains a 4.1B-parameter model for up to 7M iterations at batch size 1024 [8], and Race-DiT scales to 2.79B parameters over 1.7M iterations [11]. These efforts inherit the bigger-pool mindset while largely leaving aside the efficiency mechanisms, such as zero-computation experts, gate-residual routing, and explicit control of activated computation, that let LLM MoE trade capacity against pertoken and deployment cost. This leaves an unresolved design question: which of these LLM efficiency designs actually transfer to diffusion-transformer generation, and how should they be combined when training cost and deployment efficiency matter? This paper takes a modernization perspective. ConvNeXt revisited a classical vision backbone and showed that carefully
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
2
Output
Output AttnRes Ops (α) ω Q
MoE
K
α
V
ω
ω
MoE++
α
ω
Attention
Output
Attention
α
ω
MoE
MoE++ FFN1 ···· Shared
Const
Zero
Copy
Attention
ω
Attention Router
MoE++ Structure
α
Block n-1
Input
Embedding
Vanilla MoE Block
α
Block n-2 Embedding
Our MMoE Block
Fig. 1. Architecture of MMOE. Left: a vanilla MoE block that routes each token to a small pool of feed-forward experts. Right: the MMOE block, which adds MoE++ lightweight experts (copy, zero, and constant), gate-residual routing, and attention-residual aggregation over previously completed block states before both the attention and the expert sub-layer.
importing modern design choices can close much of the gap to a newer architectural paradigm [14]. We ask an analogous question for AFMs: can a standard SiT-style diffusion transformer be modernized with recent LLM MoE designs to improve generation quality without increasing the effective computation and training burden? This question is practically important because simply increasing total parameters and sparsity ratios can make AFMs difficult to train, route, communicate, and deploy on resource-constrained devices. Instead of treating MoE as a single replacement for dense FFNs, we study a sequence of recent expert designs, including routed experts, shared experts, lightweight zero-computation experts, gateresidual routing, and attention-residual information reuse. We test these components under a controlled diffusion-transformer setting and integrate the effective parts into a unified architecture. We propose ModernMOE (MMOE), a modernized sparseexpert diffusion transformer for AFMs, illustrated in Figure 1. MMOE keeps the denoising objective, conditioning interface, latent representation, and sampler unchanged, and modifies
only the internal computation of transformer blocks. Starting from a SiT-style backbone, we progressively introduce standard routed MoE, shared-expert computation, MoE++-style lightweight experts, and attention-residual aggregation. This progressive design allows us to separate the effect of each component rather than attributing all gains to a monolithic architecture. The resulting model allocates expensive MLP computation selectively, provides cheap expert routes when a heavy update is unnecessary, and reuses cross-depth information without adding a new generative objective or sampling procedure. Every experiment in this paper is trained on a single eightGPU H100 node with batch size 256 for 400k steps, an accessible single-machine budget. At this budget, MMOE reaches lower FID at every recorded checkpoint, that is, it converges faster per training step, than the dense and intermediate sparseexpert baselines, and among the sparse variants it attains the best quality-cost balance; a direct empirical comparison against prior diffusion-MoE methods, which are trained at far larger batch sizes and iteration counts, is left as ongoing
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
work. Ablation studies further indicate that each modern LLMinspired MoE design contributes to performance, convergence, or speed when adapted to AFMs. Beyond aggregate generation quality, we analyze routing behavior during denoising and observe strong block-wise specialization, depth-dependent lightweight-expert use, and low adjacent-step routing switch rates. These observations support the view that sparse expert modernization can serve not only as a way to increase nominal capacity, but also as a mechanism for stable adaptive computation allocation in generative foundation models. Our contributions are summarized as follows: • We formulate AFM scaling as a modernization problem and systematically examine whether recent LLM MoE designs can improve diffusion-transformer generation without simply increasing total parameters, sparsity ratios, or training overhead. • We propose ModernMOE, a SiT-style sparse-expert architecture that integrates routed experts, shared experts, lightweight zero-computation experts, gate-residual routing, and attention-residual aggregation into a unified diffusion-transformer backbone. • We provide controlled comparisons, ablations, efficiency analysis, and routing analysis, all within a single eightGPU H100 node, showing that the proposed modernization path improves AFM performance and convergence at an accessible budget while revealing how modern expert components affect convergence, speed, and specialization. II. R ELATED W ORK A. Diffusion Models and Flow Matching Diffusion models define generation as the reverse of a gradual noising process and have become a central paradigm for high-fidelity image synthesis [15]. Score-based generative models further formulate this process through stochastic differential equations, connecting denoising, score estimation, stochastic samplers, and probability-flow ODEs under a unified continuous-time view [16]. A systematic study of the design space of these models further disentangles noise scheduling, network preconditioning, and sampling to markedly improve sample quality [17]. Latent diffusion models reduce the cost of this iterative generation process by moving denoising from pixel space to a compressed autoencoder latent space, enabling high-resolution generation with a more tractable token budget [18]. Flow-based formulations provide a closely related perspective. Flow Matching trains continuous normalizing flows by directly regressing vector fields along probability paths, avoiding simulation of a reverse diffusion process during training [19]. Stochastic interpolants further unify flow and diffusion models by describing generative paths that bridge noise and data distributions through continuous-time interpolating processes [20]. Recent transformer-based generative backbones build on these ideas: U-ViT casts all inputs as tokens with a ViT backbone and long skip connections [21], DiT replaces convolutional denoisers with transformer blocks over latent patches [6], and SiT studies scalable interpolant transformers for both flow- and diffusion-based generation [7];
3
complementary training techniques such as aligning intermediate features with pretrained visual representations further ease diffusion-transformer training [22]. These transformer backbones now reach well beyond class-conditional image synthesis: DiT-based systems unify multimodal image generation and understanding [23], extend to controllable and multi-shot video generation [24]–[26], and support neural 3D geometry and appearance reconstruction [27]. The broader push toward general multimodal foundation models also drives progress in cross-modal understanding and generation [28]– [33]. MMOE is orthogonal to the choice between diffusion and flow matching: it keeps the generative objective, conditioning interface, and sampler fixed, and instead modernizes the internal computation of the diffusion-transformer backbone. B. MoEs in LLMs and Diffusion Models Mixture-of-experts models increase total model capacity by activating only a subset of parameters for each token. The sparsely-gated MoE layer established this conditionalcomputation paradigm [1], and large-scale systems such as GShard and Switch Transformers made sparse expert routing practical for foundation-model training [2], [3]. GLaM shows that sparsely activated experts can match dense models at a fraction of the training cost [34], while ST-MoE improves the training stability and transferability of sparse expert models [35]. Recent LLMs further demonstrate that top-k routed FFN experts can provide a favorable capacitythroughput tradeoff [4]. Beyond standard routed FFNs, modern MoE designs increasingly explore more fine-grained efficiency mechanisms: DeepSeekMoE finely segments experts and isolates shared experts to reduce redundant computation [36], expert-choice routing lets experts select tokens to balance load [37], and large-scale systems such as DeepSeek-V3 adopt auxiliary-loss-free load balancing for efficient training [38]. MoE++ introduces zero-computation experts, including zero, copy, and constant experts, so that not every selected route needs to execute a full MLP [5], and mixture-of-depths applies a related activate-or-skip principle across transformer layers [39]. Sparse experts have likewise been scaled to vision transformers [40]. Attention Residuals studies selective reuse of previous layer states, offering another way to improve information flow and memory behavior in deep transformer models [41]. MoE has also started to appear in diffusion-transformer research. Large sparse diffusion transformers have been explored as a scaling path for high-capacity generative models [8], [42]. EC-DIT uses expert-choice routing for adaptive computation in text-to-image diffusion transformers [9]. Diff-MoE introduces time-aware and space-adaptive experts to match denoising stages and spatial token complexity [10], while DiffMoE (distinct from Diff-MoE above) proposes dynamic token selection with a batch-level global token pool and a capacity predictor [43]. Other works investigate routing competition, explicit routing guidance, and practical architecture choices for training diffusion MoE models [11]–[13]. These studies show that sparse experts are promising for AIGC generation, but most of them focus on a particular routing strategy or scaling
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
recipe. MMOE instead asks a complementary question: which modern MoE components developed in LLMs remain effective when transferred to diffusion transformers, and how can they be combined into a practical AFM backbone? C. Scaling Law with AIGC Foundation Models Scaling has been a dominant driver of progress in AIGC Foundation Models. Latent diffusion reduces the spatial cost of image generation [18], and transformer denoisers such as DiT and SiT show that generative quality can improve with larger transformer backbones and longer training [6], [7]. Sparse diffusion transformers push this trend further by increasing total parameter capacity while activating only part of the network for each token [8]. From a scaling-law perspective, this suggests that AFMs should be evaluated not only by nominal parameter count, but also by activated parameters, routing overhead, memory footprint, training throughput, and deployment efficiency. The contrast with LLM scaling is instructive. In LLMs, sparse-expert scaling became practical because capacity growth was paired with mechanisms that hold per-token and deployment cost roughly constant: routing decouples activated computation from total capacity, while zero-computation experts and gate-residual routing let the model skip or reuse computation instead of always executing a full expert. Current AFM scaling instead tends to enlarge the total expert pool and the sparsity ratio while importing few of these efficiency mechanisms. The asymmetry is also one of budget: prior sparse AFMs are trained with billions of parameters for 1.7M to 7M iterations at batch size 1024 [8], [11], whereas the study in this paper fits on a single eight-GPU H100 node at batch size 256 for 400k steps. This matters because sampling repeatedly applies the backbone over many denoising steps, so routing overhead, expert dispatch and communication, load balancing, and deployment execution all recur at every step. Serving foundation models efficiently under edge and resource-constrained conditions is itself an active research direction [44]–[46], which reinforces the need to control activated computation and communication rather than only total capacity. A useful AFM scaling path should therefore improve quality under realistic compute budgets rather than merely enlarge the total expert pool. MMOE follows this principle by modernizing the blocklevel computation: it allocates expensive MLP computation selectively, provides cheaper expert paths when possible, and reuses cross-depth information to improve the performancecost tradeoff. III. M ETHOD MMOE imports four efficiency mechanisms proven in LLM MoE, namely routed experts, shared and lightweight zero-computation experts, gate-residual routing, and attentionresidual reuse, into the SiT backbone, changing only how block-level computation is activated and reused so that capacity can grow without a proportional increase in activated computation. Figure 1 contrasts the resulting MMOE block with a vanilla MoE block. MMOE is built on the SiT backbone and keeps the external SiT interface unchanged: the model
4
receives latent tokens x, timestep t, and class label y, embeds the latent patches with a fixed positional embedding, forms the conditioning vector c = et (t) + ey (y), and returns the unpatchified denoising prediction. The difference from the dense SiT backbone is entirely inside the transformer stack. Each block is split into an attention sub-layer and a MoE++ MLP sub-layer, and both sub-layers can receive an attentionresidual aggregation over previously completed block states. The full forward pass maintains three recurrent states: a list of completed block states, a current partial block state, and a gate residual passed across expert layers. A. MoE++ expert layer The MLP sub-layer in MMOE is implemented by a MoE++ expert layer. Given hidden states H ∈ RB×T ×D , the layer first flattens the batch and token dimensions into N = BT token representations. A router produces expert logits for every token. The default router is a two-layer gate, g(h) = W2 tanh(W1 h), D
(1) 8E×D
where h ∈ R is a token feature, W1 ∈ R and W2 ∈ RE×8E are learned gate weights, D is the model width, and E is the number of experts. If a previous gate residual is available, the router maps it through a learned linear transformation and adds it to the current logits: ℓi = g(hi ) + Ariprev .
(2)
Here A ∈ RE×E maps the previous gate residual riprev ∈ RE into the current logit space, and the resulting logits satisfy ℓi ∈ RE . The updated logits ℓi are returned as the next gate residual. This implementation makes routing decisions depend not only on the current token state, but also on the routing tendency inherited from the previous expert layer. The router selects top-2 experts for each token. Under the default non-Mixtral gating mode, probabilities are computed by softmax over all experts, the top-2 probabilities are selected and renormalized, and the zero expert receives zero gate weight when selected. The MoE output is accumulated by dispatching only the selected tokens to each expert: X MMOEmlp (hi ) = pi,e fe (hi ), (3) e∈T2 (i)
where T2 (i) denotes the selected expert indices and pi,e denotes the post-renormalization routing weight.1 In code, this is implemented with per-expert token indices, expert-specific forward calls, and an index_add accumulation into the output tensor.2 The expert pool contains both heavy and lightweight experts. If the number of experts is E, the implementation constructs E − 4 standard MLP experts, two constant experts, one copy expert, and one zero expert. A heavy expert is a 1 When the zero expert is among the top-2 for a token, its weight is set to 0 before renormalization, so the remaining weight is assigned entirely to the other selected expert and that token is effectively processed by a single expert. 2 In the default routing mode the dispatched token feature is pre-scaled by the top-k factor, so each expert is evaluated at 2hi rather than hi before its output is weighted by pi,e .
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
standard transformer MLP with hidden width determined by the MLP ratio. The copy and zero experts are fcopy (h) = h,
fzero (h) = 0.
(4)
Each constant expert learns a vector aj and predicts an inputdependent mixture between the token feature and this constant vector: fconst,j (h) = πj,0 (h)h + πj,1 (h)aj , (5) πj (h) = softmax(Wj h). Here aj ∈ RD is a learned constant vector and Wj ∈ R2×D produces the two-way mixture weights πj (h) ∈ R2 . Thus, selected routes do not always trigger a full MLP. Some tokens can use cheap identity, zero, or constant transformations, which is the core implementation mechanism for reducing unnecessary heavy expert computation. B. MMOE transformer block Each MMOE block follows the adaLN-Zero structure of SiT but separates attention and MLP computation into two functions. Given a block input z and conditioning vector c, the block first produces six modulation vectors: (∆msa , Smsa , Gmsa , ∆mlp , Smlp , Gmlp ) = MLPadaLN (c).
(6)
The attention sub-layer applies layer normalization, adaLN modulation, self-attention, and gated residual addition: zattn = z + Gmsa ⊙ Attn (7) mod LN(z), ∆msa , Smsa . The block caches (∆mlp , Smlp , Gmlp ) so that the MLP sublayer does not recompute the adaLN projection. The MLP sublayer then applies the MoE++ expert layer to the modulated token features and updates the gate residual: m, rnext = MMOEmlp (8) mod LN(zmlp ), ∆mlp , Smlp , rprev , zout = zmlp + Gmlp ⊙ m.
(9)
Here zmlp is the input to the MLP sub-layer, which may be either the post-attention state or an attention-residual aggregation of previous block states, as described next. C. Attention-residual forward process MMOE maintains a list of completed block states C and a current partial block state u. Before each attention sub-layer and before each MLP sub-layer, the model optionally aggregates information from C and u. If no completed state exists, the current partial state is used directly. Otherwise, the implementation stacks the completed block states s1 , . . . , sM ∈ C and the current partial state u, V = [s1 , . . . , sM , u],
(10)
5
Algorithm 1 MMOE forward process Require: latent input x, timestep t, label y 1: u ← PatchEmbed(x) + PosEmbed 2: c ← et (t) + ey (y) 3: C ← [ ], r ← None, h ← u 4: for block i = 1, . . . , L do 5: zattn ← u if C is empty, else AttnResattn (C, u) i 6: (u, cache) ← AttentionSubLayeri (zattn , c) 7: zmlp ← u if C is empty, else AttnResmlp (C, u) i 8: (u, r) ← MMOESubLayeri ( zmlp , c, r, cache) 9: h←u 10: if i reaches a block-group boundary then 11: append u to C 12: reset u to zeros 13: end if 14: end for 15: return Unpatchify(FinalLayer(h, c))
normalizes them with an RMSNorm, and applies a learned pseudo-query q along the state dimension: αi = softmaxi q ⊤ RMSNorm(Vi ) , X (11) AttnRes(V ) = αi Vi . i
The model uses separate projection and RMSNorm parameters for attention-residual aggregation before attention and before MLP at every layer. This gives the block two opportunities to reuse previous completed states: one before self-attention and another before expert computation. The complete forward process is summarized in Algorithm 1. After every block group, the current partial state is appended to the completed-state list and the partial state is reset to zeros.3 In the implementation, the boundary interval is block_size / 2. This grouped update avoids attending to every intermediate layer state while still allowing later blocks to reuse representations from earlier completed groups. The implementation keeps the latest block output as the final transformer representation, so the reset operation at a block boundary does not erase the output used by the final prediction head. After the final transformer block, MMOE applies the standard SiT final layer and unpatchifies the result. If sigma prediction is enabled, the output channels are split and only the denoising prediction is returned. No additional denoising objective, conditioning branch, or sampler is introduced. IV. E XPERIMENTS A. Experiment Settings We evaluate MMOE on class-conditional ImageNet-256 generation [47]. All models are trained in the latent space 3 Immediately after a boundary the partial state is reset to zeros but is still included as the last entry of V in the next aggregation, so the first sub-layer of each new block group attends over the completed states together with a zero placeholder.
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
Dense SiT Standard MoE SMOE MoE++ AMOE MMOE
9
FID
8 7 6 5
0.79
SiT-XL/2 MMOE-XL/2 0.78
Denoising Loss
10
6
0.77
0.76
4 0.75
200
300
400
Training Steps (K) Fig. 2. FID convergence on ImageNet-256 for XL/2 models along the modernization path from dense SiT to MMOE. MMOE reaches the lowest FID at every checkpoint, indicating faster convergence per training step. Lower is better.
of a Stable-Diffusion-style VAE [18]. Generated images are decoded with the same VAE and evaluated with FID [48] using pytorch-fid; lower FID is better in all tables. Unless otherwise specified, we report EMA checkpoints, ODE sampling, 50 denoising steps, 50k generated samples, and classifier-free guidance [49]. The purpose of this setup is to isolate architectural changes: dataset preprocessing, latent representation, denoising objective, sampler, and evaluation metric are kept fixed across variants. All variants use the same SiT training interface. The VAE downsampling factor maps ImageNet-256 images to 32 × 32 latent grids with four channels, and the ImageNet-512 runs in Table VI map 512 × 512 images to 64 × 64 latents through the same pipeline. Training uses the linear interpolant, velocity prediction, uniform timestep sampling, and class-label dropout with probability 0.1 for classifier-free guidance. Every run in this paper is trained on a single eight-GPU H100 node with batch size 256, which keeps the entire modernization study within an accessible single-machine budget. Unless otherwise noted, models are trained for 400k optimization steps with AdamW, learning rate 1 × 10−4 , betas (0.9, 0.999), no weight decay, fp16 mixed precision, maximum gradient norm 1.0, and EMA decay 0.9999. For MoE variants that expose a loadbalancing auxiliary loss, the loss is added with coefficient 0.01. The evaluation script samples labels uniformly, uses the same VAE decoder for all models, and computes FID against precomputed ImageNet reference statistics. The main modernization comparison in Table I uses 50-step ODE sampling at guidance scale 1.5 with 50k generated samples; the modelscale study in Table III and the guidance study in Table IX use 250-step ODE sampling; and the sample-count and samplingstep sweeps in Table VII and Table VIII vary those two axes as indicated. We compare a controlled sequence of architectures. Each step introduces one design family from modern sparse expert models while keeping the surrounding SiT training recipe
0
0
100
200
300
400
Training Steps (K) Fig. 3. Denoising-loss convergence of dense SiT-XL/2 and MMOE-XL/2 on ImageNet-256 (0–400k steps). The broken y-axis zooms into the converged range: MMOE tracks the dense baseline early in training and then settles at a consistently lower loss, corroborating its faster FID convergence.
unchanged: • Dense SiT: the baseline with dense MLP blocks. • Standard MoE: all MLP blocks are replaced by routed MLP experts, with a top-k router and load-balancing auxiliary loss. • Shared MoE (SMOE): one shared dense expert is always active, and additional routed experts provide tokendependent specialization. • MoE++: routed MLP experts are combined with copy, zero, and constant experts, together with gate-residual routing. • AttnRes-MoE (AMOE): standard routed MLP experts are combined with attention-residual aggregation. • MMOE: MoE++-style lightweight experts are combined with attention-residual aggregation; this is the final modernized AIGC backbone described in Section III. B. Comparison Results Table I and Figure 2 show the main modernization path, in which dense SiT is progressively modernized with routed experts, shared-expert computation, lightweight experts, and attention-residual expert aggregation; the training-time row reports wall-clock time on a single eight-GPU H100 node. Dense SiT provides the reference point. Replacing dense MLPs with standard routed experts improves FID at all recorded checkpoints, indicating that sparse expert capacity is useful even without changing the diffusion objective or sampler. SMOE further stabilizes the expert architecture by preserving an always-active common computation path. MoE++ introduces lightweight copy, zero, and constant experts, which help at the early and middle training stages by allowing the router to select cheaper transformations for some tokens. Attentionresidual aggregation gives the largest single jump among the measured components, suggesting that cross-depth information reuse is particularly beneficial for sparse expert diffusion transformers.
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
7
TABLE I M ODERNIZATION PATH ON I MAGE N ET-256 CLASS - CONDITIONAL GENERATION FOR XL/2 MODELS , FROM DENSE S I T TO MMOE. FID USES EMA CHECKPOINTS , 50- STEP ODE SAMPLING , GUIDANCE SCALE 1.5, AND 50 K SAMPLES ; THE LAST ROW REPORTS WALL - CLOCK TRAINING TIME ON A SINGLE EIGHT-GPU H100 NODE . L OWER FID IS BETTER . Training steps
SiT
Standard MoE
SMOE
MoE++
AMOE
MMOE
200k 300k 400k
9.80 6.50 5.20
9.14 6.00 4.65
9.01 5.91 4.62
8.53 5.63 4.64
7.13 4.74 3.85
6.88 4.54 3.75
Training time
23h
54h
45h
46h
120h
67h
TABLE II MMOE ALONGSIDE DENSE AND SPARSE - EXPERT DIFFUSION TRANSFORMERS ON CLASS - CONDITIONAL I MAGE N ET-256. T HE S I T-XL/2 AND MMOE-XL/2 ROWS ARE OUR OWN SINGLE - NODE RUNS ; THE OTHERS ARE QUOTED FROM THE CITED PAPERS AND ARE NOT DIRECTLY COMPARABLE ( SEE TEXT ). Method
Total params
Activated params
Train iters
FID
SiT-XL/2 (ours) [7] DyDiT-XL/2 ICLR2025 [50] JiT-L/16 CVPR2026 [51] DiT-MoE-XL/2 arXiv2024 [8] DiffMoE-L-E8 arXiv2025 [43] Race-DiT-XL/2 ICML2025 [11] DSMoE-3B-E16 arXiv2025 [13] ProMoE-XL/2 ICLR2026 [12]
676M 678M 953M 4.1B 1.18B 2.79B 2.96B 1.57B
1.5B 458M 710M 965M 675M
400k 7M + Finetune 1M 7M 7M 1.7M 1M 500k
4.91 2.07 2.29 1.72 2.13 2.06 2.39 4.11
MMOE-XL/2 (ours)
1.57B
≤770M
400k
3.60
The final MMOE variant combines lightweight expert routes with attention-residual aggregation. It obtains the best FID throughout the recorded training trajectory while requiring less wall-clock training time than the AMOE variant in our measured setting. This supports the central hypothesis of the paper: the best performance-cost tradeoff does not come from sparse routing alone, but from composing routing with cheap expert paths and representation reuse. Figure 3 corroborates this from the optimization side: MMOE-XL/2 tracks dense SiT-XL/2 early in training and then reaches a consistently lower denoising loss for most of the 0–400k trajectory, so the FID gains are accompanied by faster convergence at the same single-node budget. The comparisons in Table I are internal: they isolate the contribution of each modern expert component while holding the SiT training recipe fixed. To place MMOE in the broader landscape, Table II places it alongside dense baselines and prior diffusion-MoE methods as reported in their papers, all using classifier-free guidance; the SiT-XL/2 and MMOE-XL/2 rows are our own single-node runs, while the others are quoted from the cited sources. These figures are not directly comparable to ours. The baselines are trained far longer and at much larger batch sizes: DiT-MoE and Race-DiT use batch size 1024 for 1.7M to 7M iterations, whereas our two rows use batch size 256 for 400k iterations on a single eight-GPU H100 node, roughly an order of magnitude less training. They also use their own samplers and sample counts, with DiT-MoE quoting FID50k under 250-step DDPM sampling, DiffMoE-L-E8 FID-50k under 250-step flow sampling, and Race-DiT not stating its sampler; our two rows use 250-step ODE sampling at guidance scale 1.5. EC-DIT [9] is omitted because it reports only text-to-image MS-COCO FID rather than class-conditional
TABLE III M ODEL - SIZE COMPARISON OF DENSE S I T AND MMOE AT THE S/2, B/2, AND L/2 SCALES (250- STEP ODE, GUIDANCE SCALE 1.0, 400 K STEPS ). MMOE IMPROVES FID OVER THE CORRESPONDING DENSE S I T MODEL AT EVERY SCALE , SO THE MODERNIZATION IS NOT SPECIFIC TO THE XL SETTING . Scale
SiT
MMOE
S/2 B/2 L/2
60.0 37.2 22.3
51.0 26.9 14.7
ImageNet. Under these as-reported protocols, several methods trained far longer reach lower FID, including DiT-MoE-XL/2 (1.72) and DiffMoE-L-E8 (2.13) after up to 7M iterations, Race-DiT-XL/2 (2.06), DyDiT-XL/2 (2.07), DSMoE-3B-E16 (2.39), and the dense JiT-L/16 (2.29), all far beyond our 400k-iteration budget. At a comparable 1.57B total-parameter budget, however, MMOE (3.60) already improves on ProMoEXL/2 (4.11) despite fewer training iterations. We therefore do not claim state-of-the-art FID; the contribution of MMOE is the controlled modernization study and its quality-cost balance at a much smaller, single-machine training budget. A matchedprotocol comparison that trains these baselines and MMOE to the same budget is left as ongoing work, and the present empirical claims are scoped to the dense and intermediate sparse-expert baselines in Table I. Table III evaluates whether the modernization remains useful below the XL setting. Under the 250-step, CFG-1.0 protocol, MMOE improves over the corresponding dense SiT model at S, B, and L scales. This suggests that the benefit of the proposed sparse expert modernization is not restricted to a single model width or depth.
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
Training steps
200k
300k
400k
400k (250 steps)
FID
11.69
7.74
5.81
5.47
TABLE V S EED - ROBUSTNESS OF THE XL/2 MODERNIZATION PATH : FID MEAN ± STANDARD DEVIATION OVER THREE RANDOM SEEDS (0, 1, 2) AT 400 K STEPS . T HE BETWEEN - VARIANT GAPS ARE MUCH LARGER THAN THE SEED - INDUCED DEVIATION , SO THE SINGLE - RUN RANKINGS ARE STABLE .
FID (mean ± std)
SiT
SMOE
MMOE
5.16±0.04
4.62±0.06
3.72±0.05
Table IV reports the training convergence of a one-billionparameter MoE++ model at the L/2 scale, using eight experts (four standard MLP experts plus two constant, one copy, and one zero expert) with top-2 routing and the default 50step evaluation protocol. The lightweight-expert design keeps reducing FID throughout training at this larger total-parameter budget, reaching 5.81 at 400k steps. For this checkpoint, increasing the number of ODE steps from 50 to 250 further lowers FID to 5.47, in contrast to the MMOE-XL/2 checkpoint in Table VIII, where additional steps did not help. This indicates that sampling-step sensitivity is checkpoint dependent and should be tuned per model rather than assumed constant. We further check two robustness aspects of the XL/2 comparison, both at 400k steps under the 50-step ODE, guidance-scale-1.5, 50k-sample protocol. Table V reports seed robustness over three random seeds (0, 1, 2): dense SiT, SMOE, and MMOE reach FID 5.16 ± 0.04, 4.62 ± 0.06, and 3.72 ± 0.05, so the gaps between variants are far larger than the seed-induced standard deviation. These three-seed means agree, within that variance, with the single-run (seed-0) values reported elsewhere in the paper, such as the MMOE and SiT entries in Table I (3.72 versus 3.75 and 5.16 versus 5.20), so the single-run rankings are stable. Table VI extends the denseversus-MMOE comparison to ImageNet-512: MMOE again improves over dense SiT, reaching 11.72 versus 12.72 at 200k steps and 6.03 versus 6.58 at 400k steps. C. Ablation Study The modernization sequence in Table I also serves as the component ablation. Standard MoE isolates sparse expert TABLE VI I MAGE N ET-512 COMPARISON BETWEEN DENSE S I T AND MMOE, USING THE SAME VAE PIPELINE THAT MAPS 512 × 512 IMAGES TO 64 × 64 LATENTS . MMOE IMPROVES OVER DENSE S I T AT BOTH THE 200 K AND 400 K CHECKPOINTS . Model
FID (200k)
FID (400k)
SiT MMOE
12.72 11.72
6.58 6.03
1400
Activation memory (MB)
TABLE IV T RAINING CONVERGENCE OF A ONE - BILLION - PARAMETER M O E++ MODEL AT THE L/2 SCALE , WITH EIGHT EXPERTS ( FOUR MLP PLUS TWO CONSTANT, ONE COPY, AND ONE ZERO ) AND TOP -2 ROUTING . FID KEEPS DECREASING THROUGHOUT TRAINING , AND FOR THIS CHECKPOINT 250- STEP SAMPLING FURTHER IMPROVES THE 50- STEP RESULT.
1200 1000 800
8
AMOE MMOE
−32% 1203
−20% 988 787
823
600 400 200 0
Forward pass
Backward pass
Fig. 4. Per-block activation memory (MB) for an XL/2 block: attentionresidual MoE (AMOE) versus the MMOE lightweight-expert design at matched 42.5M parameters. Routing some tokens to copy, zero, or constant experts cuts forward memory by about 20% and backward memory by about 32%.
capacity, SMOE tests the benefit of an always-active common path, MoE++ evaluates lightweight copy/zero/constant experts and gate-residual routing, and AMOE isolates attentionresidual aggregation. Each component contributes to the final performance-cost balance, although the progression is not strictly monotonic at every checkpoint: MoE++ helps mainly at the early and middle checkpoints and is essentially tied with SMOE at 400k (4.64 versus 4.62), whereas attention-residual aggregation and lightweight expert routing give the largest and most consistent gains and are especially important for MMOE. The goal of MMOE is not only to reduce FID, but also to rebalance model capacity, activated computation, memory footprint, and training time. The training-time row in Table I shows that architectural choices have very different system costs. Standard MoE increases training time substantially relative to dense SiT because routing and expert dispatch introduce overhead. AMOE achieves strong quality but is the most expensive measured variant. MMOE retains the quality benefit of attention-residual information reuse while reducing the cost by replacing some heavy expert routes with lightweight expert functions. Figure 4 provides a block-level memory comparison between AMOE and the MMOE-style lightweight expert design. Both configurations use an XL/2 block with hidden width 1152 and the same 42.5M parameters, but the lightweight expert design reduces single-block activation memory from 988 MB to 787 MB in the forward pass (about 20%) and from 1203 MB to 823 MB in the backward pass (about 32%). Because a substantial fraction of route slots select the copy, zero, or constant experts (44.6% in early blocks, decreasing to 23.4% in late blocks; Table X), the corresponding routes skip the activation storage and gradient computation of a full MLP. Aggregated over the 28 transformer blocks, this amounts to approximately 5.5 GB of forward and over 10 GB of backward activation memory saved. This supports the design motivation
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
TABLE VII FID SENSITIVITY TO THE NUMBER OF GENERATED EVALUATION SAMPLES FOR MMOE-XL/2 (400 K STEPS , 50- STEP ODE). T HE ESTIMATE DROPS SHARPLY FROM 10 K TO 50 K SAMPLES , SO LOW- SAMPLE FID SHOULD BE READ AS A DIAGNOSTIC RATHER THAN A HEADLINE NUMBER . Samples
10k
20k
30k
40k
50k
FID
6.64
4.85
4.04
3.82
3.75
TABLE VIII S AMPLING - STEP SENSITIVITY FOR MMOE-XL/2 (400 K STEPS , 10 K SAMPLES ). I NCREASING ODE STEPS BEYOND 50 DOES NOT IMPROVE FID FOR THIS CHECKPOINT; THE SDE COLUMN IS A REFERENCE POINT AT MATCHED STEP COUNTS . Inference Steps
ODE FID
SDE FID
50 100 150 200 250
6.64 6.66 6.63 6.62 6.63
6.49 6.33 6.48 6.41 6.58
behind copy, zero, and constant experts: a sparse model should not force every selected route to execute a full MLP when cheaper transformations are sufficient. Beyond per-block memory, we profiled the multi-GPU training of the sparse expert models to locate where wall-clock time is spent. A single profiling run indicates that collective communication dominates: ncclAllReduce accounts for almost all of the profiled communication overhead, and device synchronization (cudaStreamSynchronize) accounts for roughly half of the total profiled time. The current training is therefore communication bound rather than compute bound, so the wall-clock figures in Table I reflect a communicationlimited regime. This also implies that part of the cost gap between dense and sparse variants stems from expert dispatch and gradient synchronization rather than from raw expert computation, and that a more efficient distributed implementation is a concrete avenue for reducing the training cost of sparse AFMs. Table VII studies the effect of the number of generated samples used for FID estimation. The 10k-sample estimate is directionally consistent with the 50k-sample estimate for the checked checkpoint, but the large gap between the 10k estimate (6.64) and the 50k estimate (3.75) shows that lowsample FID should be treated as a diagnostic rather than a headline metric. Therefore, the main comparison in Table I keeps the evaluation protocol fixed across models, and 50ksample FID remains the preferred setting whenever compute allows. Table VIII evaluates the number of inference steps for the MMOE-XL/2 checkpoint. Increasing ODE steps beyond 50 does not improve the result under the current configuration (ODE FID stays within 6.62–6.66), and the SDE sampler reaches only slightly lower FID than the 50-step ODE setting (6.33–6.58 versus 6.64) rather than a large gain. The absence of improvement with more ODE steps is mildly counterintuitive, because additional integration steps should reduce discretization error; this flat trend more likely reflects a confound
9
TABLE IX C LASSIFIER - FREE GUIDANCE SENSITIVITY (250- STEP ODE, 400 K STEPS ). L OWERING THE GUIDANCE SCALE FROM 1.5 TO 1.0 DEGRADES ALL MODELS , BUT MMOE REMAINS THE STRONGEST AT BOTH SETTINGS . CFG scale
SiT
SMOE
MMOE
1.5 1.0
4.91 18.90
4.39 17.00
3.60 13.40
with the guidance scale or the 10k-sample FID estimate used in this sweep than a genuine property of the sampler, and it warrants a controlled re-run at a fixed guidance scale and 50k samples. We therefore keep 50-step ODE sampling as the default protocol for the main comparison for consistency, and treat this table as a sampling-protocol observation for the measured checkpoint rather than a general claim about all samplers. Table IX evaluates classifier-free guidance sensitivity. Reducing the guidance scale from 1.5 to 1.0 degrades all compared models under this 250-step setting, showing that absolute FID remains sensitive to guidance and sampler choices. Nevertheless, MMOE remains the strongest among the compared models at both guidance scales, suggesting that the architectural improvement is not tied to a single guidance setting.
D. Visualization Figure 5 provides a qualitative comparison in which the five variants use the same initial noise and class labels, 250-step ODE sampling, and classifier-free guidance scale 4.0, with samples decoded from VAE latents by the shared decoder; this guidance scale is chosen for visualization and differs from the scales used in the quantitative tables. The qualitative samples are intended to complement FID rather than replace it: they help inspect whether the quantitative gains along the modernization path correspond to visible improvements in image structure and fidelity, while the controlled tables above provide the primary evidence for architectural comparison. V. E XPERT ROUTING A NALYSIS We analyze MMOE routing with forward hooks on the expert routers during denoising. The analysis records selected experts for each step, block, image, token, and top-k route. Table X reports statistics for the converged MMOE-XL/2 checkpoint at 400k training steps, measured over 10k randomly sampled class-conditional denoising trajectories with 50 ODE steps and no classifier-free guidance. The expert pool contains four standard MLP experts and four lightweight experts: two constant experts, one copy expert, and one zero expert. In the table, routing entropy is computed from harddispatch expert loads and normalized by log E with E = 8; the lightweight route fraction counts dispatch slots assigned to the two constant, copy, or zero experts; and the switch rate is the fraction of tokens whose top-2 expert set changes between adjacent denoising steps.
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
10
Fig. 5. Qualitative comparison along the modernization path (top to bottom: Dense SiT, SMOE, MoE++, AMOE, MMOE) on five ImageNet classes. All rows share the same initial noise and class label and use 250-step ODE sampling at guidance scale 4.0, so differences reflect the architecture rather than the sample.
a) Denoising-time routing stability.: The measured stepto-step switch rate is low. The top-k expert set changes for only 1.43%, 1.78%, and 2.71% of tokens in early, middle, and late block groups, respectively. Thus, the learned routing is not highly volatile across adjacent denoising steps. Instead, the model tends to keep stable expert sets along the denoising trajectory, with slightly more temporal adaptation in later blocks. This behavior is consistent with gate-residual routing: routing decisions can remain coherent across depth and denoising time while still allowing local changes when the token representation evolves. b) Aggregate expert utilization.: The load statistics indicate strong block-wise specialization rather than uniform
expert usage. Across block groups, the load Gini coefficient is about 0.71–0.72 and the normalized load entropy is about 0.43–0.45. Under top-2 routing with eight experts, this corresponds to an effective usage of roughly 2.5 experts per block on average. This is not a single-expert collapse, because the dominant expert typically receives at most about half of all dispatches, but it does show that each block learns a small preferred expert subset. The lightweight route fraction also changes with depth: lightweight experts account for 44.6% of route slots in early blocks, 34.8% in middle blocks, and 23.4% in late blocks. This depth trend supports the intended role of lightweight experts: early computation often selects cheap copy, zero, or constant transformations, while later blocks rely
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
11
TABLE X Q UANTITATIVE ROUTING STATISTICS FOR THE CONVERGED MMOE-XL/2 MODEL , MEASURED OVER 10 K DENOISING TRAJECTORIES (50 ODE STEPS , NO GUIDANCE ). T HE HIGH LOAD G INI WITH MODERATE ENTROPY INDICATES BLOCK - WISE EXPERT SPECIALIZATION , AND THE LIGHTWEIGHT ROUTE FRACTION DECREASES WITH DEPTH . Block group
Load Gini
Routing Entropy
Lightweight Route Frac.
Switch Rate
Early (0–8) Middle (9–18) Late (19–27)
0.7206 ± 0.0450 0.7198 ± 0.0302 0.7115 ± 0.0414
0.4324 ± 0.1118 0.4390 ± 0.0789 0.4532 ± 0.0839
0.4457 0.3482 0.2341
0.0143 0.0178 0.0271
more heavily on standard MLP experts for refinement. VI. D ISCUSSION AND L IMITATIONS MMOE improves the quality-efficiency tradeoff of diffusion transformers by combining sparse routing with lightweight expert functions and attention-residual information reuse. The empirical gains are strongest on ImageNet-256, where MMOE consistently outperforms the dense SiT baseline and the intermediate sparse-expert variants we study under a matched evaluation protocol. We re-implement all compared variants ourselves, so these are controlled internal comparisons rather than a benchmark against published systems. The controlled modernization path also clarifies which components matter: sparse expert capacity helps, but the larger gain appears when lightweight routes and cross-depth aggregation are introduced. More broadly, by importing efficiency mechanisms proven in LLM MoE rather than enlarging the total expert pool, MMOE improves generation quality over the dense and intermediate sparse-expert baselines it studies within a single-node eightGPU H100 budget (batch size 256, 400k steps), which suggests that AFMs can follow the balanced scaling path taken by LLMs instead of inflating total parameters and sparsity ratios. The current evidence has several limitations. First, the main experiments are class-conditional ImageNet experiments, so the results do not establish general text-to-image superiority. Second, all reported comparisons are internal: we reimplement the dense and intermediate sparse-expert baselines under a matched recipe, and a head-to-head comparison against published diffusion-MoE methods (Table II) is still in progress. Third, most FID values are single-run estimates; the seed-robustness study in Table V shows that for the XL/2 path the between-variant gaps exceed the seed-induced standard deviation, but per-seed variance for the other scales and studies is not measured, so small gaps elsewhere should still be read with caution. Fourth, the resolution study covers ImageNet256 and, for the dense-versus-MMOE comparison, ImageNet512 (Table VI), but the full modernization path and the modelscale study are established at 256, so conclusions at higher resolutions are limited. Fifth, the routing analysis is measured on the converged XL checkpoint under the 50-step no-guidance analysis protocol; extending it to other model sizes, guidance scales, and training stages would further clarify how stable the observed specialization pattern is. Finally, sparse expert models introduce distributed-training overheads that blocklevel memory profiling and single-run wall-clock time do not fully capture; our profiling indicates that the current multiGPU training is communication bound, so the reported training times reflect a communication-limited regime.
Future work includes extending MMOE to large-scale textto-image training, studying routing under multimodal conditioning, and improving distributed training efficiency for sparse expert diffusion models. VII. C ONCLUSION We presented MMOE, a controlled modernization of SiTstyle diffusion transformers for AIGC Foundation Models. Rather than following current AFM scaling by inflating total parameters and sparsity ratios, MMOE imports efficiency mechanisms proven in LLM MoE, combining routed MLP experts, lightweight copy/zero/constant experts, gateresidual routing, and attention-residual aggregation. On classconditional ImageNet-256 generation, and within a singlenode eight-GPU H100 budget (batch size 256, 400k steps), the modernization path improves FID and convergence over dense SiT and the intermediate MoE variants under matched training and sampling protocols; among the sparse variants MMOE gives the best quality-cost balance, cutting the training time of the most expensive variant (AMOE) from 120h to 67h, though it remains more expensive than the dense baseline. The routing analysis further suggests that MMOE learns stable block-wise expert specialization, uses lightweight routes substantially in early and middle blocks, and changes its top-2 routing set only modestly between adjacent denoising steps. These findings support sparse expert modernization as a practical direction for improving the quality-cost balance of diffusion transformers while preserving the standard training and sampling interface. ACKNOWLEDGMENTS This research is supported by the Institute of Artificial Intelligence of China Telecom (TeleAI). This research is supported by the RIE2025 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) (Award I2301E0026), administered by A*STAR, as well as supported by Alibaba Group and NTU Singapore through Alibaba-NTU Global eSustainability CorpLab (ANGEL). The work is also supported by the Ministry of Education, Singapore under its MOE Academic Research Fund Tier 2 (MOE-T2EP20123-0005).
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
R EFERENCES [1] N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in International Conference on Learning Representations (ICLR), 2017. [2] D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen, “GShard: Scaling giant models with conditional computation and automatic sharding,” in International Conference on Learning Representations (ICLR), 2021. [3] W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research, vol. 23, no. 120, pp. 1–39, 2022. [4] A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. Singh Chaplot, D. de las Casas et al., “Mixtral of experts,” 2024. [5] P. Jin, B. Zhu, L. Yuan, and S. Yan, “MoE++: Accelerating mixtureof-experts methods with zero-computation experts,” in International Conference on Learning Representations (ICLR), 2025. [6] W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 4172–4182. [7] N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie, “SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers,” in European Conference on Computer Vision (ECCV), 2024. [8] Z. Fei, M. Fan, C. Yu, D. Li, and J. Huang, “Scaling diffusion transformers to 16 billion parameters,” 2024. [9] H. Sun, T. Lei, B. Zhang, Y. Li, H. Huang, R. Pang, B. Dai, and N. Du, “EC-DIT: Scaling diffusion transformers with adaptive expertchoice routing,” in International Conference on Learning Representations (ICLR), 2025. [10] K. Cheng, X. He, L. Yu, Z. Tu, M. Zhu, N. Wang, X. Gao, and J. Hu, “Diff-MoE: Diffusion transformer with time-aware and spaceadaptive experts,” in Proceedings of the 42nd International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 267, 2025, pp. 10 010–10 024. [11] Y. Yuan, Z. Wang, Z. Huang, D. Zhu, X. Zhou, J. Yu, and Q. Min, “Expert race: A flexible routing strategy for scaling diffusion transformer with mixture of experts,” 2025. [12] Y. Wei, S. Zhang, H. Yuan, Y. Han, Z. Chen, J. Wang, D. Zou, X. Liu, Y. Zhang, Y. Liu, and H. Shan, “Routing matters in MoE: Scaling diffusion transformers with explicit routing guidance,” 2025. [13] Y. Liu, Y. Yue, J. Zhang, C. Sun, Y. Zhou, W. Zeng, R. Tang, and G. Zhou, “Efficient training of diffusion mixture-of-experts models: A practical recipe,” 2025. [14] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 11 976– 11 986. [15] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 6840–6851. [16] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in International Conference on Learning Representations (ICLR), 2021. [17] T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” in Advances in Neural Information Processing Systems (NeurIPS), 2022. [18] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “Highresolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10 684–10 695. [19] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in International Conference on Learning Representations (ICLR), 2023. [20] M. S. Albergo, N. M. Boffi, and E. Vanden-Eijnden, “Stochastic interpolants: A unifying framework for flows and diffusions,” Journal of Machine Learning Research, vol. 26, no. 209, pp. 1–80, 2025. [21] F. Bao, S. Nie, K. Xue, Y. Cao, C. Li, H. Su, and J. Zhu, “All are worth words: A ViT backbone for diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. [22] S. Yu, S. Kwak, H. Jang, J. Jeong, J. Huang, J. Shin, and S. Xie, “Representation alignment for generation: Training diffusion transform-
12
ers is easier than you think,” in International Conference on Learning Representations (ICLR), 2025. [23] J. Wang, Z. Wang, H. Pan, Y. Liu, D. Yu, C. Wang, and W. Wang, “Mmgen: Unified multi-modal image generation and understanding in one go,” arXiv preprint arXiv:2503.20644, 2025. [24] D. Xi, J. Wang, Y. Liang, X. Qiu, Y. Huo, R. Wang, C. Zhang, and X. Li, “Omnivdiff: Omni controllable video diffusion for generation and understanding,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 13, 2026, pp. 10 915–10 923. [25] D. Xi, J. Wang, Y. Liang, X. Qiu, J. Liu, H. Pan, Y. Huo, R. Wang, H. Huang, C. Zhang et al., “Ctrlvdiff: Controllable video generation via unified multimodal video diffusion,” arXiv preprint arXiv:2511.21129, 2025. [26] X. Wu, J. Teotia, S. Zhao, and E. Cambria, “Edustory: A unified framework for pedagogically-consistent multi-shot stem instructional video generation,” arXiv preprint arXiv:2605.09378, 2026. [27] Y. Liu, P. Wang, C. Lin, X. Long, J. Wang, L. Liu, T. Komura, and W. Wang, “Nero: Neural geometry and brdf reconstruction of reflective objects from multiview images,” ACM Transactions on Graphics (ToG), vol. 42, no. 4, pp. 1–22, 2023. [28] Y. Jia, X. Wu, L. Hao, Z. Qinglin, Y. Hu, S. Zhao, and W. Fan, “Uni-retrieval: A multi-style retrieval framework for stem’s education,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 10 182– 10 197. [29] X. Wu, Y. Jia, L. Xiao, S. Zhao, F. Chiang, and E. Cambria, “From query to explanation: Uni-rag for multi-modal retrieval-augmented learning in stem,” arXiv preprint arXiv:2507.03868, 2025. [30] X. Wu, Y. Jia, Q. Zhang, Y. Qin, L. Xiao, and S. Zhao, “Towards affective evaluation of stem education: Leveraging mllms in projectbased learning,” IEEE Transactions on Affective Computing, 2026. [31] Y. Jia, J. Xie, S. Jivaganesh, L. Hao, X. Wu, and M. Zhang, “Seeing sound, hearing sight: Uncovering modality bias and conflict of ai models in sound localization,” Advances in neural information processing systems, vol. 38, pp. 148 468–148 499, 2026. [32] E. Cambria, R. Mao, X. Zhang, L. Xiao, T. Shen, and A. Anand, “Senticnet 9: Generative commonsense for emotion ai via conceptual primitive discovery and time shift mechanism,” IEEE Transactions on Computational Social Systems, 2026. [33] Y. Jia, X. Wu, J. Teotia, J. Dong, S. Zhao, P. Koniusz, and E. Cambria, “Towards spatial reasoning and understanding via modeling modality conflict, bias and alignment,” 2026. [34] N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. MeierHellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui, “GLaM: Efficient scaling of language models with mixture-ofexperts,” in Proceedings of the 39th International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol. 162, 2022, pp. 5547–5569. [35] B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus, “ST-MoE: Designing stable and transferable sparse expert models,” 2022. [36] D. Dai, C. Deng, C. Zhao, R. X. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, Z. Xie, Y. K. Li, P. Huang, F. Luo, C. Ruan, Z. Sui, and W. Liang, “DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024, pp. 1280–1297. [37] Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Z. Chen, Q. V. Le, and J. Laudon, “Mixture-of-experts with expert choice routing,” in Advances in Neural Information Processing Systems (NeurIPS), 2022. [38] DeepSeek-AI, “DeepSeek-V3 technical report,” 2024. [39] D. Raposo, S. Ritter, B. Richards, T. Lillicrap, P. C. Humphreys, and A. Santoro, “Mixture-of-depths: Dynamically allocating compute in transformer-based language models,” 2024. [40] C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby, “Scaling vision with sparse mixture of experts,” in Advances in Neural Information Processing Systems (NeurIPS), 2021. [41] Kimi Team, G. Chen, Y. Zhang, J. Su, W. Xu, S. Pan et al., “Attention residuals,” 2026. [42] D. Kukreja, K. Prasad, A. Anand, Z. Wang, E. Cambria, T. Liu, A. B. Ng, S. See, and B. Chatterjee, “Forge: Fused on-register gradient elimination for memory-efficient llm training,” arXiv preprint arXiv:2606.22932, 2026.
IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, MONTH YEAR
[43] M. Shi, Z. Yuan, H. Yang, X. Wang, M. Zheng, X. Tao, W. Zhao, W. Zheng, J. Zhou, J. Lu, P. Wan, D. Zhang, and K. Gai, “DiffMoE: Dynamic token selection for scalable diffusion transformers,” 2025. [44] J. Shao and X. Li, “Ai flow at the network edge,” IEEE Network, 2025. [45] H. An, W. Hu, S. Huang, S. Huang, R. Li, Y. Liang, J. Shao, Y. Song, Z. Wang, C. Yuan et al., “Ai flow: Perspectives, scenarios, and approaches,” Vicinagearth, vol. 3, no. 1, p. 1, 2026. [46] X. Chen, J. Luo, Y. Fan, H. Huang, C. Zhang, and X. Li, “Generative transmission: Rethinking computation, bandwidth, and memory in communication,” 2026. [47] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009, pp. 248–255. [48] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANs trained by a two time-scale update rule converge to a local nash equilibrium,” in Advances in Neural Information Processing Systems (NeurIPS), 2017. [49] J. Ho and T. Salimans, “Classifier-free diffusion guidance,” 2022. [50] W. Zhao, Y. Han, J. Tang, K. Wang, Y. Song, G. Huang, F. Wang, and Y. You, “Dynamic diffusion transformer,” in International Conference on Learning Representations, vol. 2025, 2025, pp. 65 520–65 552. [51] T. Li and K. He, “Back to basics: Let denoising generative models denoise,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 36 115–36 125.
13