Preprint. Under review.
U NIFYING D ISTRIBUTIONAL T RAINING FOR O NE -S TEP V ISUAL G ENERATION Chi Zhang1∗ , Shi Haoyang1,2∗ , Yueyi Liu1,3∗ , Ruichuan An4 , Junkang Zhou5,6 , Chang Li7 Xiuyuan Lu2,8 , Yichi Zhang1 , Bo Wang1 , Yuhang Wu1 , Sen Cui9 , Miao Liu1† 1
arXiv:2609.35763v1 [cs.LG] 28 Sep 2026
5
Tsinghua University Zhejiang University
2 6
Fudan University 3 Xi’an Jiaotong University 4 Peking University DeepSeek-AI 7 ByteDance Seed 8 University of California, Berkeley
9
BAAI
Figure 1: One-step text-to-image samples. FLUX.2 [klein] 4B post-trained with MGFlow.
A BSTRACT Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unified theoretical framework that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates MGFlow, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet 256×256, MGFlow substantially surpasses the FD-Loss baseline, achieving stateof-the-art results with 1.45 FDr6 on pMF-H and 1.64 on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/. † Corresponding author. Contact: [email protected], [email protected]
* Equal contribution.
1
Preprint. Under review.
1
I NTRODUCTION
One-step visual generation aims to synthesize high-quality images with a single network evaluation, avoiding iterative refinement at inference. Existing approaches include learning consistent trajectory endpoints (Song et al., 2023; Song and Dhariwal, 2024; Kim et al., 2024; Luo et al., 2023; Geng et al., 2025b; Lu and Song, 2025), average velocities (Frans et al., 2025; Geng et al., 2025a; 2026; Lu et al., 2026), or distilling pretrained diffusion models (Salimans and Ho, 2022; Yin et al., 2024b; Sauer et al., 2024; Zhou et al., 2024; Yin et al., 2024a). A prominent recent direction is distributional training in frozen representation spaces (Yang et al., 2026a; Deng et al., 2026; Feng et al., 2026). Rather than assigning a fixed target image to each output, these methods compare collections of real and generated features and obtain collective supervision from their distributional mismatch. Each feature’s gradient depends on shared statistics or interactions with other features, rather than only on an individual reconstruction target. This approach raises two fundamental questions: how should feature distributions be modeled, and how should their mismatch guide learning? We introduce a unified theoretical framework for distributional training that separates the distribution model from the matching discrepancy. Sampling and frozen encoders determine the feature populations being compared; the distribution model and matching discrepancy determine what distributional information is recorded and how it is aligned. Wasserstein gradient flow (WGF) (Jordan et al., 1998; Peyré and Cuturi, 2019) acts as the theoretical bridge that relates the distributional discrepancy to a pointwise velocity. This establishes a common formulation for global moment losses and local sample interactions, with existing methods recovered as specific model–discrepancy choices. For example, a Gaussian distribution model with the W2 optimal transport (OT) distance recovers FD-Loss (Yang et al., 2026a). Let pG and qG be Gaussians with the respective means and covariances of the real and generated features. FD-Loss optimizes the Fréchet loss: 1/2 1/2 DFD (qG , pG ) = W22 (qG , pG ) = ∥µq − µp ∥22 + Tr Σq + Σp − 2 Σ1/2 Σ Σ . (1) p q q Thus, the Fréchet discrepancy between global feature statistics is a Gaussian transport cost (Dowson and Landau, 1982; Peyré and Cuturi, 2019). At the population level, differentiating through the moments moves features along the associated affine transport field. The Gaussian model determines which statistics are retained, while the OT discrepancy determines how they guide the movement of generated features. A sample-based density model with Kullback–Leibler (KL) divergence (Kullback and Leibler, 1951) instead recovers Gaussian-kernel Drifting. Drifting (Deng et al., 2026) is specified through kernelweighted attraction and repulsion. For real and generated feature distributions p and qθ , a positive kernel k, and r ∈ {p, qθ }, its mean-shift and update fields are defined respectively as ar (z) =
Ey∼r [k(z, y)(y − z)] , Ey∼r [k(z, y)]
vdrift (z) = ap (z) − aqθ (z).
(2)
The generator regresses samples toward detached targets shifted by this field vdrift . For a Gaussian kernel of bandwidth h, each mean shift satisfies ar (z) = h2 ∇z log rh (z), where rh is the Gaussian kernel density estimate (KDE) of r (Lai et al., 2026). Their difference is therefore the density-level KL velocity evaluated at those KDEs, up to a bandwidth-dependent scale (Cao et al., 2026). Built on top of this theoretical framework, we propose Mixture Gradient Flow (MGFlow). To refine the modeling granularity while capturing global correlations, MGFlow uses Gaussian mixtures (GMs), which retain componentwise means and covariances at an adjustable resolution between global moment matching and sample-centered KDE. We develop both OT-based and score-based matching for this representation. Crucially, we note that an expressive modeling does not itself ensure mode coverage or correct transportation (Wenliang and Kanagawa, 2020): naive posterior assignment and transport may suffer from weight mismatch and mode collapse. MGFlow therefore couples mass-constrained sample assignment with paired component matching, using reference weights to control the mass of each generated component and a shared component correspondence to guide features toward the matched reference components. We evaluate MGFlow under the same sampling and feature encoder settings as FD-Loss. On ImageNet (Russakovsky et al., 2015) 256 × 256, MGFlow achieves a state-of-the-art 2
Preprint. Under review.
Batch Sampling 𝓢 Real Generated …
FD-Loss …
MGFlow
Drifting
[Yang et al., 2026]
Ours
[Deng et al., 2026]
Single Gaussian
Gaussian Mixture
Kernel Density
ℳ = Single Gaussian
ℳ = Gaussian Mixture
ℳ = KDE of Samples
Representation Encoder 𝓔 Real Features 𝑧~𝑝 Generated Features 𝑧~𝑞!
𝒫
𝒬
Distribution Model 𝓜 Discrepancy 𝓓 𝒟(𝒬, 𝒫) → WGF velocity 𝛿ℱ 𝑣𝒟 𝑧 = −𝛻" 𝑧 𝛿𝑄
𝒟=
W22
𝒟=
W22 or KL
𝒟 = KL
Global moments only
Component-wise moments
Local kernel estimates
Unimodal approximation
Adjustable granularity
Isotropic kernel geometry
Figure 2: A unified view of distributional training. Left: sampling S, encoding E, distribution modeling M, and matching discrepancy D define the paradigm. Right: FD-Loss, MGFlow, and Gaussian-kernel Drifting use Gaussian, GM, and KDE representations, respectively. Their matching objectives induce feature-space velocity fields through the Wasserstein gradient flow of their respective distributional objectives.
FDr6 (Yang et al., 2026a) score of 1.45 on pMF-H (Lu et al., 2026) and 1.64 on JiT-H (Li and He, 2026), improving over the FD-Loss baseline with 23% and 38% margins, respectively. For text-toimage generation, MGFlow post-trains FLUX.2 [klein] 4B (Black Forest Labs, 2026) for one-step inference, reaching 0.900 GenEval (Ghosh et al., 2023) and 21.98 PickScore (Kirstain et al., 2023), outperforming all previous methods. Together with comparisons across model–discrepancy combinations, these results show how our novel formulation leads to effective alternatives to existing training prescriptions.
2
A U NIFIED V IEW OF D ISTRIBUTIONAL T RAINING
A collective feature-space distributional training paradigm. We formulate distributional training by (S, E, M, D) (Figure 2): the sampling scheme S supplies sample collections that provide collective supervision; the frozen encoder E maps samples to representation spaces where they are compared. For real and generated feature distributions p and qθ , the distribution model M constructs parametric approximations P and Q, and the discrepancy D measures the mismatch. This construction provides distributional supervision: each feature update depends on the entire sampled population through the distribution model, rather than only on an individual target.
Table 1: Distributional training with distribution models and discrepancies. All methods train pMF-B with Inception features for 10 epochs. All methods use the aligned recipe in Appendix E.1. Model
Method
FDr6 ↓
OT-based (W2 )
Gaussian FD-Loss (Yang et al., 2026a) 13.05 GM K = 4 MGFlow-W2 11.33 Sample-based W-Flow (Han et al., 2026) 12.58 Score-based (KL)
Gaussian GM K = 16
Gaussian KL MGFlow-KL Gaussian-kernel Drifting Sample-based (Deng et al., 2026)
12.73 11.32
FD-Loss and Drifting provide collective supervi12.29 sion through shared moments or sample interactions rather than independent reconstruction targets (Yang et al., 2026a; Deng et al., 2026). FD-Loss is trained in the feature space of three independent encoders, while Drifting uses a latent MAE encoder. By default, we adopt FD-Loss’s sampling scheme S and encoders E, and focus on the distribution model and discrepancy (M, D). A global discrepancy induces a pointwise descent field. To compare loss-based and field-based training prescriptions, we need to connect a scalar distributional mismatch to an update direction for each generated feature. The Wasserstein gradient flow (Jordan et al., 1998; Peyré and Cuturi, 3
Preprint. Under review.
2019) provides this connection by describing how feature locations should move to decrease the chosen discrepancy. Fix P and let F(Q) = D(Q, P ), where Qt is the evolving generated distribution through training time t. Let δF /δQ denote the first variation of the objective, describing its sensitivity to infinitesimal density changes. The corresponding 2-Wasserstein gradient flow induces a descent velocity field vD,t . Under suitable regularity, the velocity and energy dissipation satisfy: δF dzt d vD,t (z) = −∇z , (z) = vD,t (zt ), F(Qt ) = −Ez∼Qt ∥vD,t (z)∥22 ≤ 0. (3) δQ dt dt Q=Qt Thus moving points along the induced field decreases global discrepancy. Details are given in Appendix B.5. FD-Loss and Drifting are special cases of distributional training. We now instantiate this construction with specific distribution models and discrepancies to recover the feature-update fields of FD-Loss and Gaussian-kernel Drifting. For M, let pG = N (µp , Σp ) and qG = N (µq , Σq ) be moment-matched Gaussian models. Alternatively, let ph , qh be Gaussian kernel density estimates (KDEs) of p, qθ using sample-centered kernels with covariance h2 I (Parzen, 1962). For D, we consider an optimal transport (OT) approach using the 2-Wasserstein distance W2 and a scorebased approach with the Kullback-Leibler divergence DKL ; these choices recover the FD-Loss and Gaussian-kernel Drifting fields (Lai et al., 2026; Cao et al., 2026): WGF 2 1 −−→ TqG →pG (z) − z = vFD (z), 2 W2 (qG , pG ) − WGF DKL (qh ∥ph ) −−−→ ∇z log ph (z) − ∇z log qh (z) = h−2 vdrift (z),
(4) where TqG →pG is the closed-form optimal transport map between Gaussians (Peyré and Cuturi, 2019). The Fréchet distance equals W22 , and differentiating it through feature moments moves samples along the first descent field. Gaussian-kernel Drifting instead evaluates the KL velocity on the smoothed densities. The two methods occupy global-moment and sample-local ends of the modeling spectrum, with different discrepancies. Appendices B.1–B.2 give proofs and derivations. Other methods could also fit within this framework. For instance, calculating an OT-based discrepancy on empirical measures gives W-Flow (Han et al., 2026), which constructs optimal transport between batches with Sinkhorn (Cuturi, 2013; Feydy et al., 2019) approximation. Moreover, the framework can inspire new distributional training algorithms. Keeping the single Gaussian surrogate but replacing W2 with KL divergence gives the following closed-form velocity field: −1 vKL (z) = Σ−1 (5) p (µp − z) − Σq (µq − z) which we denote as Gaussian KL. We conduct experiments with all methods mentioned above under the same settings, using the same sampling scheme as FD-Loss and a frozen Inception (Szegedy et al., 2016) encoder. Results are in Table 1. All alternatives outperform FD-Loss on FDr6 , even without direct optimization of the Fréchet distance, demonstrating the versatility of our formulation.
3
MGF LOW
As shown in Section 2, FD-Loss and Gaussian-kernel Drifting take two extremes on the distribution model spectrum: A single Gaussian models only global first and second moments, whereas Gaussian KDE uses sample-centered components and overlooks global covariance structures. In order to improve modeling granularity while capturing global correlations in feature spaces, we introduce Mixture Gradient Flow (MGFlow), a distributional training algorithm that adopts an intermediate distribution model M based on a Gaussian mixture (GM). We first introduce Gaussian-mixture representations and direct OT- and KL-based matching constructions. We then show why representing multiple modes alone is insufficient and develop a coupled allocation-and-update procedure that preserves correspondence with the reference components while still bounding the global discrepancies. 3.1
F ROM A S INGLE G AUSSIAN TO G AUSSIAN M IXTURES
For the distribution models P, Q, consider using Gaussian Mixtures with K Gaussian components: P (z) =
K X
πk pk (z),
Q(z) =
k=1
K X k=1
4
ωk qk (z),
(6)
Preprint. Under review.
where pk = N (µp,k , Σp,k ) and qk = N (µq,k , Σq,k ). P is fit offline by the EM algorithm (Dempster et al., 1977) and remains fixed, while Q summarizes the changing generated feature distribution, which requires updating through training. The case K = 1 recovers a single Gaussian; samplecentered components with a common isotropic covariance (K = B) recover Gaussian KDE at the representation level. Table 1 summarizes the model–discrepancy combinations. Both discrepancies, W2 and KL, extend to this representation. Though Wasserstein Distances between GMs are intractable, restricting transport to Gaussian component pairs gives a tractable upper bound, namely the Mixed Wasserstein distance (Delon and Desolneux, 2020): X W22 (Q, P ) ≤ M W22 (Q, P ) := min Γij W22 (qi , pj ). (7) Γ∈U (ω,π)
i,j
Here U (ω, π) is the set of component couplings with marginals ω and π, and each pair cost is the Gaussian FD. For KL, the global velocity field is the difference of the score functions: X X vglobal (z) = ∇z log P (z) − ∇z log Q(z) = γP,k (z)sp,k (z) − γQ,k (z)sq,k (z), (8) k
k
where sr,k (z) = ∇z log rk (z) = Σ−1 r,k (µr,k − z) for r ∈ {p, q}, γP,k = πk pk /P , and γQ,k = ωk qk /Q. The score of each mixture is a posterior-responsibility-weighted sum of its component scores. However, representing multiple modes does not ensure correct mode proportions. The global P KL field can provide weak inter-mode updates when modes are well separated. Consider P = k πk rk P and Q = k ωk rk with the same separated components rk but different positive weights. Correcting this weight mismatch requires movement between modes, not merely local refinement within each mode. However, in a region dominated by component k, the density ratio is approximately the constant ωk /πk ; the score difference can be small even when the component weights differ substantially: log
Q(z) ωk ≈ log , P (z) πk
vglobal (z) = −∇z log
Q(z) ≈0 P (z)
(9)
Thus, a mismatch can remain visible in the density ratio while producing only a weak feature-update signal in the high-density regions. This is analogous to the score-blindness phenomenon studied byWenliang and Kanagawa (2020), and does not contradict the descent interpretation in Section 2: descent does not guarantee effective transport between modes when the score field is weak. In the top-right state of Figure 3, the two Q-components q1 and q2 carry 85.1% and 14.9% of the mass, respectively; the global velocity field alone provides too little inter-mode transport to correct this imbalance. The same issue can arise for Gaussian-KDE fields when smoothing leaves modes well separated. For score-based matching, mixture expressivity does not by itself provide an effective mechanism for correcting mode proportions. We therefore connect mass allocation explicitly to the feature updates. 3.2
M ASS -C ONSTRAINED A LLOCATION AND PAIRED T RANSPORT
LP-based component allocation. The refer- Table 2: Ablation on component assignment and ence mixture P is fitted offline and kept fixed, transportation. JiT-B trains in Inception feature space whereas Q must track the changing genera- for 10 epochs with K = 4. Timing uses 8 H200 GPUs. tor throughout training. For a single GausKL W2 sian, each generated batch contributes to one FDr6 ↓ s/step FDr6 ↓ s/step set of global moments. A GM instead requires 173.56 0.197 26.78 0.699 component-specific moments, so updating Q Posterior 146.92 0.264 23.73 0.825 requires allocating new features to components. LP-global LP-paired 23.27 0.263 23.42 0.392 A natural choice is posterior soft assignment: each generated sample zn contributes to component k with weight γQ,k (zn ) = ωk qk (zn )/Q(zn ), its posterior responsibility under the current Q. However, due to dynamical reasons, posterior soft assignment cannot transport the correct amount of mass to each component, as shown in Figure 3. 5
Preprint. Under review.
Reference P
t=0
Initial Q0
Component q1
t = 0.2
Component q2
t = 0.8
t=2
Posterior 85.1% / 14.9%
LP-global 85.6% / 14.4%
LP-paired (ours) 50.0% / 50.0%
Figure 3: Toy experiment with component assignment and transport. Score matching with K = 2 moves 2048 particles. Dashed contours show the reference P . Shading and contour counts indicate weights. LP fixes component weights, but global velocity collapses modes. Only LP-paired achieves correct transport.
We instead assign generated features to the fixed reference components pk while constraining the mass allocated to component k to Bπk for a batch of B features. To achieve this, we solve the capacity-constrained linear program (LP) for batch-to-component allocation: X R∗ ∈ arg min Rnk [− log(πk pk (zn ))], subject to R1K = 1B , RT 1B = Bπ. (10) R≥0
n,k
Here R ∈ RB×K , where Rnk denotes the assignment from sample n to component k, which may be fractional. The likelihood cost favors compatible reference components, while the capacity constraints enforce their prescribed masses. Using detached R∗ , we compute the weighted mean and raw second moment for each generated component, maintain these statistics across training batches with EMA as in FD-Loss (Yang et al., 2026a), and recover the component covariances. We set ω = π. Paired OT matching. With matched weights, Γ = diag(π) is feasible in Eq. (7). Empirically, in a pMF-H training run with this objective, the optimal component transport matrix was diagonal at every recorded plan refresh (Appendix B.4). Motivated by this observation, we fix this correspondence and calculate only the diagonal terms in the Mixed Wasserstein distance, optimizing: Lpair-W2 =
K X
πk W22 (qk , pk ) ≥ M W22 (Q, P ) ≥ W22 (Q, P ).
(11)
k=1
This provides an upper bound for W22 . A sufficient condition for diagonal optimality is given in Appendix B.4. Fixed pairing reduces the number of Gaussian pair costs from K 2 to K. We differentiate this objective through the current batch’s contribution to the EMA statistics of the generated components, keeping historical statistics and the LP assignments R∗ detached. Paired score matching. For the KL discrepancy, pairing is more than a matter of efficiency, as constraining the component statistics is not sufficient to constrain the feature updates. The global field in Eq. (8) still weights reference and generated scores by separate mixture responsibilities, rather than the LP assignments. As demonstrated in the second row of Figure 3, the global velocity cannot distinguish the target component for each sample, resulting in mode collapse. To preserve the assignment in the update direction, we instead match corresponding components using X Fpair = πk DKL (qk ∥pk ) ≥ DKL (Q∥P ) (12) k
which also upper-bounds the marginal KL. Appendix B.4 gives the proof and further analysis. Calculating closed-form KL between Gaussians requires heavy computation, so instead we use the WGF-induced velocity. Reusing R∗ for these fields gives the regularized update vpair,n =
K X
∗ Rnk [sλp,k (zn ) − sλq,k (zn )],
sλr,k (z) = (Σr,k + λI)−1 (µr,k − z).
(13)
k=1
Here λ > 0 stabilizes covariance inversion. The same assignment now weights both scores, so the update retains the correspondence used to estimate qk . We evaluate the field using detached, 6
Preprint. Under review.
pre-update statistics and train the generator by detached-target regression: B
Lpair-KL =
1 X 2 ∥zn − sg[zn + ηvpair,n ]∥2 , 2B n=1
η > 0.
(14)
The entire target is detached, and the EMA state is committed after the generator step. The third row in Figure 3 shows that this formulation correctly transports samples from the initial distribution to the reference distribution. An ablation study in Table 2 proves that LP-based component allocation uniformly achieves better results than posterior assignments. Paired transport is crucial especially when training multi-step JiT models with score-based discrepancies. This is possibly because the large gap between the initial one-step generated distribution and the target distribution induces abnormal component allocation. Detailed training and evaluation settings are given in Appendix H. 3.3
T RAINING MGF LOW.
As discussed above, we train two variants of MGFlow: MGFlow-W2 with the paired OT-based objective in Equation 11, and MGFlow-KL with paired score-based updates in Equation 14. For multiple encoders, we use fixed, discrepancy-specific normalizers for the encoder losses: LMGFlow-W2 =
X e
Lepair-W2
LMGFlow-KL =
, W22 (Re , Ve )
X e
Lepair-KL . DKL (Re ∥Ve )
(15)
Here Lepair-W2 and Lepair-KL sum the respective branch losses for encoder e, while Re and Ve are single-Gaussian fits to its real training and validation features. For KL calibration, the same score ridge is added to both covariances. Unlike FD-Loss’s real–generated (R, G) normalization, these fixed (R, V ) scales avoid downweighting harder-to-match feature spaces merely because their current discrepancies are large (Appendix C.2). For text-to-image generation, we concatenate image features with frozen SigLIP2 text features and match their joint distributions (Feng et al., 2026) (Appendices B.6 and C.3.2).
4
E XPERIMENTS
4.1
I NCREASING G AUSSIAN C OMPONENTS
We first examine how distributional granularity affects ImageNet post-training. We compare individual Gaussian mixture models with K ∈ {1, 4, 16} components and their combinations. Beyond selecting a single GM, we consider hierarchical resolution: for a set of component counts K, we sum the corresponding training losses with equal weights, X LK = LK , (16) K∈K
Table 3: Ablation on component counts. Using pMFB. All models train in Inception feature space for 10 epochs. FDr5 evaluates five held-out encoders. Components
FDr6 ↓
FDr5 ↓
FID ↓
K=1 K=4 K = 16
12.73 12.28 11.32
15.04 14.54 13.41
1.95 1.65 1.46
K =1+4 K = 1 + 4 + 16
11.51 11.05
13.59 13.08
1.89 1.51
where LK denotes the MGFlow loss using Kcomponent distribution models. Thus, K = 1 + 4 + 16 combines global moment matching with progressively finer componentwise matching, forming a coarse-to-fine hierarchy of objectives. We train pMF-B in Inception feature space for 10 epochs with different component counts. Among the same 256 × 256 resolution configurations, increasing K from 1 to 16 reduces FDr6 from 12.73 to 11.32 and FID from 1.95 to 1.46, demonstrating that distributional models with finer granularity enhance generation. Hierarchical resolution further improves generalization, lowering FDr6 . These results suggest that coarse and fine distributional constraints provide complementary supervision, motivating multi-resolution matching in the subsequent experiments. By default we use K = 1 + 4 + 16 for MGFlow-KL and K = 1 + 4 for MGFlow-W2 . Larger mixtures are limited by the cost of Gaussian Wasserstein matching for MGFlow-W2 and the samples available to accurately estimate each component’s mean and covariance for MGFlow-KL (see Appendix I.2). 7
Preprint. Under review.
Table 4: Post-training JiT and pMF on ImageNet 256 × 256. All models train in the SIM encoder spaces for 100 epochs. “Model” denotes feature-distribution modeling. Baseline results are from FD-Loss (Yang et al., 2026a), AdvFD (Gao et al., 2026), and AMFD (Liu et al., 2026a). † denotes the full-CFG NFE upper bound for interval CFG; dashes denote unavailable results. Best and second-best results per backbone are bold and underlined, respectively. More baselines are in Table 10. #Params
FDr6 ↓
FDr3 ↓
FID ↓
IS ↑
Pixel-space models, multi-step backbones JiT-B – – 50×2×2† + FD-Loss Gaussian W2 1 + AdvFD Gaussian W2 1 + AMFD-C Gaussian Amortized W2 1 + AMFD-U Gaussian Amortized W2 1 + MGFlow-W2 GM W2 1 + MGFlow-KL GM KL 1 JiT-L – – 50×2×2† + FD-Loss Gaussian W2 1 + AdvFD Gaussian W2 1 + AMFD-C Gaussian Amortized W2 1 + AMFD-U Gaussian Amortized W2 1 + MGFlow-W2 GM W2 1 + MGFlow-KL GM KL 1 JiT-H – – 50×2×2† + FD-Loss Gaussian W2 1 + AdvFD Gaussian W2 1 + AMFD-C Gaussian Amortized W2 1 + AMFD-U Gaussian Amortized W2 1 + MGFlow-W2 GM W2 1 + MGFlow-KL GM KL 1
131M 131M 131M 131M 131M 131M 131M 459M 459M 459M 459M 459M 459M 459M 953M 953M 953M 953M 953M 953M 953M
15.65 5.53 3.92 4.75 3.91 4.43 3.70 10.73 3.24 2.01 2.67 2.02 2.68 1.92 7.66 2.65 1.80 2.15 1.79 2.55 1.64
15.06 8.45 6.03 6.71 6.13 7.36 5.70 10.27 5.46 3.20 3.86 3.12 4.54 2.96 9.07 4.44 2.93 3.10 2.78 4.42 2.57
3.71 1.00 0.79 0.95 0.95 1.45 1.28 2.59 0.77 0.73 0.87 0.85 1.07 1.00 1.97 0.75 0.72 0.85 0.83 0.99 0.94
269.0 344.6 – 319.4 325.2 313.9 308.5 288.5 317.3 – 325.1 319.8 314.9 304.7 296.0 313.0 – 328.1 312.2 307.8 301.9
Pixel-space models, one-step pMF-B – – + FD-Loss Gaussian W2 + AdvFD Gaussian W2 + AMFD-C Gaussian Amortized W2 + AMFD-U Gaussian Amortized W2 + MGFlow-W2 GM W2 + MGFlow-KL GM KL pMF-L – – + FD-Loss Gaussian W2 + AdvFD Gaussian W2 + AMFD-C Gaussian Amortized W2 + AMFD-U Gaussian Amortized W2 + MGFlow-W2 GM W2 + MGFlow-KL GM KL pMF-H – – + FD-Loss Gaussian W2 + AdvFD Gaussian W2 + AMFD-C Gaussian Amortized W2 + AMFD-U Gaussian Amortized W2 + MGFlow-W2 GM W2 + MGFlow-KL GM KL
118M 118M 118M 118M 118M 118M 118M 410M 410M 410M 410M 410M 410M 410M 935M 935M 935M 935M 935M 935M 935M
13.70 3.50 3.32 3.94 3.43 3.03 3.09 9.09 2.09 1.89 2.25 2.01 1.80 1.74 6.87 1.89 1.74 1.93 1.75 1.50 1.45
11.82 4.49 4.22 4.66 4.20 4.02 3.76 7.62 2.72 2.57 2.75 2.57 2.47 2.24 6.09 2.69 2.50 2.43 2.30 2.05 1.88
3.31 0.85 0.81 0.95 0.92 1.65 1.56 2.72 0.78 0.77 0.88 0.86 1.20 1.16 2.29 0.77 0.74 0.86 0.85 1.09 1.07
254.6 331.4 – 315.9 310.1 332.1 323.0 261.7 309.2 – 321.8 306.7 322.4 314.8 267.2 310.1 – 323.0 307.3 316.1 311.0
Method
4.2
Model
Discrepancy
NFE
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
C LASS -C ONDITIONED I MAGE N ET G ENERATION
We post-train the B, L, and H variants of JiT (Li and He, 2026) and pMF (Lu et al., 2026) from their official pretrained weights on ImageNet (Russakovsky et al., 2015) 256 × 256. All use frozen SigLIP (Tschannen et al., 2025), Inception (Szegedy et al., 2016), and MAE (He et al., 2022) encoders (SIM) and a global batch of 1,024, and are trained for 100 epochs. We evaluate post-trained models with 50,000 one-step generated samples. FD-Loss (Yang et al., 2026a) motivates FDr6 as 8
Preprint. Under review.
a more robust metric for ImageNet generation than FID alone, reducing the blind spots of a single representation. We report FDr6 across six representation spaces and FDr3 across the three encoders not used for training, together with FID (Heusel et al., 2017) and Inception Score (IS) (Salimans et al., 2016). Optimization, sampling, and reference-fitting details are given in Appendices C and F; metric definitions are in Appendix E.2. Table 4 compares MGFlow with FD-Loss (Yang et al., 2026a), AdvFD (Gao et al., 2026), and AMFD (Liu et al., 2026a). Both MGFlow-W2 and MGFlow-KL surpass the FD-Loss baseline on every model size. MGFlow-KL achieves state-of-the-art 1.45 FDr6 on pMF-H and 1.64 on JiT-H, improving over the FD-Loss baseline with 23% and 38% margins, respectively. In particular, its gains on the three held-out encoders indicate better generalization beyond the representations used for training, demonstrating that MGFlow can truly match distributions in feature spaces rather than just overfit to the first and second moments. Additionally, MGFlow-KL uses KL-based matching without any FD loss, yet outperforms all the FD-based baselines on both FDr6 and held-out FDr3 , showing that its gains do not rely on directly optimizing these evaluation metrics. See Appendix J for per-encoder results. Qualitative results are demonstrated in Appendix K. 4.3
T EXT- TO -I MAGE G ENERATION
We initialize from the distilled FLUX.2 [klein] 4B model (Black Forest Labs, 2026) and train MGFlow-KL with K = 1+4 for 1,000 steps with a batch size of 1,024. Following the exact protocol adopted in prior works (Liu et al., 2026a; Feng et al., 2026), we use reconstructed COCO (Lin et al., 2014) reference images augmented with GenEval images. We use two versions of MGFlow. For the joint version, we concatenate image features with frozen SigLIP2 (Tschannen et al., 2025) text features (Feng et al., 2026), matching the joint distribution p(x, c) (and thus, by Bayes’ theorem, the conditional distribution) instead of the marginal distribution p(x) for better text-image alignment. By contrast, the image-only variant uses image features alone, matching the marginal distribution. We use the MGFlow-KL variant with K = 1 + 4 hierarchical resolution. The model is trained to generate 512 × 512 images in one step. We report GenEval (Ghosh et al., 2023) and PickScore on Pick-a-Pic (Kirstain et al., 2023). Reference construction and training details are in Appendix C.3.2.
Table 5: Text-to-image generation with FLUX.2 [klein] 4B (Black Forest Labs, 2026) at 512 × 512. We report GenEval (Ghosh et al., 2023) and PickScore (Kirstain et al., 2023); PickScore is evaluated on 499 Pick-a-Pic prompts. Higher is better. Bold and underlining mark the best and second-best scores, respectively. MGFlow-KL uses K = 1 + 4 and 1,000 training steps. Both MGFlow variants use reconstructed COCO references supplemented with images generated from GenEval prompts. Joint matching concatenates image and SigLIP2 text features. Baselines are from AMFD (Liu et al., 2026a); dashes denote unavailable results. Reference construction is in Appendix C.3.2. Method
NFE Single Two Count Colors Position Color attr. Overall PickScore
FLUX.2 [klein] 4B DMD2 (Yin et al., 2024a) FD-SIM (Yang et al., 2026a) iRDM (Feng et al., 2026) AMFD-U-SIM (Liu et al., 2026a) AMFD-C-SIM (Liu et al., 2026a) AMFD-C-10 enc. (Liu et al., 2026a)
4 1 1 1 1 1 1
0.994 0.904 0.791 0.997 0.894 0.806 1.000 0.944 0.716 0.994 0.924 0.756 0.994 0.929 0.741 0.997 0.955 0.791 1.000 0.957 0.778
0.880 0.864 0.878 0.923 0.902 0.923 0.920
0.575 0.603 0.618 0.650 0.638 0.670 0.680
0.623 0.660 0.658 0.708 0.678 0.740 0.733
0.794 0.804 0.802 0.826 0.813 0.846 0.845
21.85 – 21.62 21.82 21.77 21.85 21.82
MGFlow (image-only) MGFlow (joint)
1 1
0.997 0.967 0.844 0.926 1.000 0.982 0.891 0.926
0.633 0.783
0.770 0.818
0.856 0.900
21.86 21.98
With the COCO reference, joint matching reaches a GenEval score of 0.900 in Table 5, compared with 0.794 for the four-step backbone, 0.826 for iRDM, and 0.846 for AMFD-C-SIM. It also achieves the highest PickScore of 21.98 in the table. iRDM (Feng et al., 2026) uses a batch size of 10,240 for 180 steps, whereas AMFD (Liu et al., 2026a) uses 1,024 for 1,500 steps. MGFlow achieves higher GenEval and PickScore scores with 44% and 33% fewer generated training samples, respectively. Qualitative text-to-image examples are provided in Appendix L. 9
Preprint. Under review.
5
R ELATED W ORK
One-step generation can be learned through trajectory consistency, adversarial training, average velocities, or diffusion distillation (Song et al., 2023; Zhang et al., 2025; Geng et al., 2025a; ?). Recent methods instead supervise generated populations through global feature-space discrepancies (Yang et al., 2026a; Feng et al., 2026) or construct updates from attraction–repulsion or transport between batches (Deng et al., 2026; Han et al., 2026). We propose the unified theoretical framework of distributional training that recovers all these works. MGFlow uses a novel Gaussian Mixture Model for distribution approximation that differs from all the above methods. For a detailed discussion of related work, see Appendix A.
6
C ONCLUSION
We presented a unified view of distributional training that separates sampling, encoding, distribution modeling, and matching discrepancy. Wasserstein gradient flow connects global objectives to feature updates, recovering FD-Loss and Gaussian-kernel Drifting as specific choices. MGFlow combines Gaussian mixtures with mass-constrained allocation and paired OT or KL updates. Experiments on ImageNet and text-to-image generation show improvements across training and held-out representations, with KL-based matching improving Fréchet metrics without directly optimizing them. Both modeling granularity and component correspondence matter for one-step generation. More broadly, our findings point toward richer distribution modeling and principled distribution matching as promising directions for advancing one-step generative models.
R EFERENCES Black Forest Labs. FLUX.2 [klein]: Towards Interactive Visual Intelligence. https://bfl.ai/ blog/flux2-klein-towards-interactive-visual-intelligence, 2026. Jiarui Cao, Zixuan Wei, and Yuxin Liu. Gradient flow drifting: Generative modeling via wasserstein gradient flows of KDE-approximated divergences. arXiv preprint arXiv:2603.10592, 2026. Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In Advances in Neural Information Processing Systems, volume 26, 2013. URL https://papers.nips. cc/paper/2013/hash/af21d0c97db2e27e13572cbf59eb343d-Abstract. html. Julie Delon and Agnès Desolneux. A Wasserstein-type distance in the space of Gaussian mixture models. SIAM Journal on Imaging Sciences, 13(2):936–970, 2020. doi: 10.1137/19M1301047. URL https://arxiv.org/abs/1907.05254. Arthur P. Dempster, Nan M. Laird, and Donald B. Rubin. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1):1–22, 1977. doi: 10.1111/j.2517-6161.1977.tb01600.x. Mingyang Deng, He Li, Tianhong Li, Yilun Du, and Kaiming He. Generative modeling via drifting. arXiv preprint arXiv:2602.04770, 2026. D. C. Dowson and B. V. Landau. The fréchet distance between multivariate normal distributions. Journal of Multivariate Analysis, 12(3):450–455, 1982. doi: 10.1016/0047-259X(82)90077-X. Lan Feng, Wuyang Li, Eloi Zablocki, Matthieu Cord, and Alexandre Alahi. Representation distribution matching for one-step visual generation. arXiv preprint arXiv:2607.02375, 2026. URL https://arxiv.org/abs/2607.02375. Jean Feydy, Thibault Séjourné, François-Xavier Vialard, Shun-ichi Amari, Alain Trouvé, and Gabriel Peyré. Interpolating between optimal transport and MMD using Sinkhorn divergences. In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, volume 89 of Proceedings of Machine Learning Research, pages 2681–2690, 2019. URL https://proceedings.mlr.press/v89/feydy19a.html. Kevin Frans, Danijar Hafner, Sergey Levine, and Pieter Abbeel. One step diffusion via shortcut models. In International Conference on Learning Representations, 2025. URL https: //openreview.net/forum?id=OlzB6LnXcS. 10
Preprint. Under review.
Leonard T. Franz, Sebastian Hoffmann, Tim Weiland, Bernhard Schölkopf, and Georg Martius. Drifting fields are not conservative. arXiv preprint arXiv:2604.06333, 2026. Luyu Gao, Yunyi Zhang, Jiawei Han, and Jamie Callan. Scaling deep contrastive learning batch size under memory limited setup. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021), pages 316–321. Association for Computational Linguistics, 2021. URL https://aclanthology.org/2021.repl4nlp-1.31/. Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, and Hao Tang. AdvFD: Boosting visual generation via adversarial fréchet distance loss. arXiv preprint arXiv:2608.11205, 2026. URL https://arxiv.org/abs/2608.11205. Zhengyang Geng, Mingyang Deng, Xingjian Bai, Zico Kolter, and Kaiming He. Mean flows for one-step generative modeling. In Advances in Neural Information Processing Systems, volume 38, pages 75460–75482, 2025a. doi: 10.52202/085713-2534. URL https://proceedings.neurips.cc/paper_files/paper/2025/file/ 6d13e085b79d454da5910e4ca82a3d9d-Paper-Conference.pdf. Zhengyang Geng, Ashwini Pokle, Weijian Luo, Justin Lin, and J. Zico Kolter. Consistency models made easy. In International Conference on Learning Representations, 2025b. URL https: //arxiv.org/abs/2406.14548. Zhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman, J. Zico Kolter, and Kaiming He. Improved mean flows: On the challenges of fastforward generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026. URL https: //arxiv.org/abs/2512.02012. Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. GenEval: An object-focused framework for evaluating text-to-image alignment. In Advances in Neural Information Processing Systems, 2023. URL https://arxiv.org/abs/2310.11513. Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola. A kernel two-sample test. Journal of Machine Learning Research, 13(25):723–773, 2012. URL https://www.jmlr.org/papers/v13/gretton12a.html. Jiaqi Han, Puheng Li, Qiushan Guo, Renyuan Xu, Stefano Ermon, and Emmanuel J. Candès. Onestep generative modeling via Wasserstein gradient flows. arXiv preprint arXiv:2605.11755, 2026. URL https://arxiv.org/abs/2605.11755. Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022. URL https://arxiv. org/abs/2111.06377. Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems, volume 30, 2017. URL https://arxiv.org/ abs/1706.08500. Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. URL https://arxiv.org/abs/2207.12598. Richard Jordan, David Kinderlehrer, and Felix Otto. The variational formulation of the fokker– planck equation. SIAM Journal on Mathematical Analysis, 29(1):1–17, 1998. doi: 10.1137/ S0036141096303359. Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Yutong He, Yuki Mitsufuji, and Stefano Ermon. Consistency trajectory models: Learning probability flow ODE trajectory of diffusion. In International Conference on Learning Representations, 2024. URL https://arxiv.org/abs/2310.02279. Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Picka-pic: An open dataset of user preferences for text-to-image generation. In Advances in Neural Information Processing Systems, 2023. URL https://arxiv.org/abs/2305.01569. S. Kullback and R. A. Leibler. On information and sufficiency. The Annals of Mathematical Statistics, 22(1):79–86, 1951. doi: 10.1214/aoms/1177729694. 11
Preprint. Under review.
Tuomas Kynkäänniemi, Miika Aittala, Tero Karras, Samuli Laine, Timo Aila, and Jaakko Lehtinen. Applying guidance in a limited interval improves sample and distribution quality in diffusion models. In Advances in Neural Information Processing Systems, volume 37, pages 122458–122483, 2024. doi: 10.52202/079017-3892. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ dd540e1c8d26687d56d296e64d35949f-Paper-Conference.pdf. Chieh-Hsin Lai, Bac Nguyen, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Yuki Mitsufuji, Stefano Ermon, and Molei Tao. A unified view of score-based and drifting models. arXiv preprint arXiv:2603.07514, 2026. Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie, and Liang Zheng. REPA-E: Unlocking VAE for end-to-end tuning of latent diffusion transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18262–18272, October 2025. URL https://arxiv.org/abs/2504.10483. Tianhong Li and Kaiming He. Back to basics: Let denoising generative models denoise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026. URL https://arxiv.org/abs/2511.13720. Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He. Autoregressive image generation without vector quantization. In Advances in Neural Information Processing Systems, 2024. URL https://arxiv.org/abs/2406.11838. Yujia Li, Kevin Swersky, and Richard Zemel. Generative moment matching networks. In Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1718–1727, 2015. URL https://proceedings.mlr. press/v37/li15.html. Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014. Wenze Liu, Xintao Wang, Pengfei Wan, and Xiangyu Yue. Amortized moment matching for visual generation. arXiv preprint arXiv:2607.26860, 2026a. URL https://arxiv.org/abs/ 2607.26860. Yueyi Liu, Chi Zhang, Sen Cui, and Miao Liu. Elasticttt: Prior-preserving test-time tuning for video editing. arXiv preprint arXiv:2607.21529, 2026b. Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. URL https://arxiv.org/abs/1711.05101. Cheng Lu and Yang Song. Simplifying, stabilizing and scaling continuous-time consistency models. In International Conference on Learning Representations, 2025. URL https: //openreview.net/forum?id=LyJi5ugyJx. Yiyang Lu, Susie Lu, Qiao Sun, Hanhong Zhao, Zhicheng Jiang, Xianbang Wang, Tianhong Li, Zhengyang Geng, and Kaiming He. One-step latent-free image generation with pixel mean flows. arXiv preprint arXiv:2601.22158, 2026. URL https://arxiv.org/abs/2601.22158. Simian Luo, Yiqin Tan, Longbo Huang, Jian Li, and Hang Zhao. Latent consistency models: Synthesizing high-resolution images with few-step inference. arXiv preprint arXiv:2310.04378, 2023. URL https://arxiv.org/abs/2310.04378. Nanye Ma, Mark Goldstein, Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden, and Saining Xie. SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers. In European Conference on Computer Vision, 2024. URL https://arxiv. org/abs/2401.08740. Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jégou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. DINOv2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. URL https://arxiv.org/abs/2304. 07193. 12
Preprint. Under review.
Emanuel Parzen. On estimation of a probability density function and mode. The Annals of Mathematical Statistics, 33(3):1065–1076, 1962. doi: 10.1214/aoms/1177704472. Gabriel Peyré and Marco Cuturi. Computational optimal transport with applications to data sciences. Foundations and Trends in Machine Learning, 11(5–6):355–607, 2019. doi: 10.1561/ 2200000073. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763, 2021. URL https://proceedings.mlr. press/v139/radford21a.html. Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan Yuille, and Liang-Chieh Chen. FlowAR: Scale-wise autoregressive image generation meets flow matching. In International Conference on Machine Learning, 2025. URL https://arxiv.org/abs/2412.15205. Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li FeiFei. ImageNet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. doi: 10.1007/s11263-015-0816-y. URL https://arxiv.org/ abs/1409.0575. Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, 2022. URL https://arxiv.org/ abs/2202.00512. Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training GANs. In Advances in Neural Information Processing Systems, volume 29, 2016. URL https://arxiv.org/abs/1606.03498. Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast high-resolution image synthesis with latent adversarial diffusion distillation. arXiv preprint arXiv:2403.12015, 2024. URL https://arxiv.org/abs/2403.12015. Yang Song and Prafulla Dhariwal. Improved techniques for training consistency models. In International Conference on Learning Representations, 2024. URL https://proceedings.iclr.cc/paper_files/paper/2024/hash/ 41bd71e7bf7f9fe68f1c936940fd06bd-Abstract-Conference.html. Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 32211–32252, 2023. URL https://proceedings.mlr. press/v202/song23a.html. Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2818–2826, 2016. URL https: //arxiv.org/abs/1512.00567. Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang. Visual autoregressive modeling: Scalable image generation via next-scale prediction. In Advances in Neural Information Processing Systems, 2024. URL https://arxiv.org/abs/2404.02905. Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hénaff, Jeremiah Harmsen, Andreas Steiner, and Xiaohua Zhai. SigLIP 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features. arXiv preprint arXiv:2502.14786, 2025. URL https://arxiv.org/abs/2502.14786. Erkan Turan, Nicolas Dufour, and Maks Ovsjanikov. Generative drifting is secretly score matching: A spectral and variational perspective. arXiv preprint arXiv:2603.09936, 2026. Shuai Wang, Ziteng Gao, Chenhui Zhu, Weilin Huang, and Limin Wang. PixNerd: Pixel neural field diffusion. In International Conference on Learning Representations, 2026a. URL https: //arxiv.org/abs/2507.23268. 13
Preprint. Under review.
Shuai Wang, Zhi Tian, Weilin Huang, and Limin Wang. DDT: Decoupled diffusion transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026b. URL https://arxiv.org/abs/2504.05741. Li K. Wenliang and Heishiro Kanagawa. Blindness of score-based methods to isolated components and mixing proportions. arXiv preprint arXiv:2008.10087, 2020. URL https://arxiv. org/abs/2008.10087. Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie. ConvNeXt V2: Co-designing and scaling ConvNets with masked autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. URL https://arxiv.org/abs/2301.00808. Ge Wu, Shen Zhang, Ruijing Shi, Shanghua Gao, Zhenyuan Chen, Lei Wang, Zhaowei Chen, Hongcheng Gao, Yao Tang, Jian Yang, Ming-Ming Cheng, and Xiang Li. Representation entanglement for generation: Training diffusion transformers is much easier than you think. In Advances in Neural Information Processing Systems, 2025. URL https://arxiv.org/abs/ 2507.01467. Jiawei Yang, Zhengyang Geng, Xuan Ju, Yonglong Tian, and Yue Wang. Representation fréchet loss for visual generation. arXiv preprint arXiv:2604.28190, 2026a. Jiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian, and Yue Wang. Latent denoising makes good tokenizers. In International Conference on Learning Representations, 2026b. URL https: //arxiv.org/abs/2507.15856. Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. URL https://arxiv.org/abs/2501.01423. Tianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang, Eli Shechtman, Fredo Durand, and William T. Freeman. Improved distribution matching distillation for fast image synthesis. In Advances in Neural Information Processing Systems, 2024a. URL https://arxiv.org/ abs/2405.14867. Tianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman, Frédo Durand, William T. Freeman, and Taesung Park. One-step diffusion with distribution matching distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024b. URL https: //arxiv.org/abs/2311.18828. Qihang Yu, Qihao Liu, Ju He, Xinyang Zhang, Yang Liu, Liang-Chieh Chen, and Xi Chen. Autoregressive image generation with masked bit modeling. arXiv preprint arXiv:2602.09024, 2026. URL https://arxiv.org/abs/2602.09024. Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. In International Conference on Learning Representations, 2025. URL https: //arxiv.org/abs/2410.06940. Chi Zhang, Zehua Chen, Kaiwen Zheng, and Jun Zhu. Voicebridge: Designing latent bridge models for general speech restoration at scale. arXiv e-prints, pages arXiv–2509, 2025. Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, and Miao Liu. From scores to samples: Elastic forcing for autoregressive video generation. arXiv e-prints, pages arXiv–2609, 2026a. Chi Zhang, Haoyang Shi, Yueyi Liu, Zhaokun Yan, Yishu Yin, Yuhang Wu, and Miao Liu. Interacvid: Building a real interactive audio-visual response dataset from live-chat videos. arXiv preprint arXiv:2608.01157, 2026b. Boyang Zheng, Nanye Ma, Shengbang Tong, and Saining Xie. Diffusion transformers with representation autoencoders. In International Conference on Learning Representations, 2026. URL https://arxiv.org/abs/2510.11690. Mingyuan Zhou, Huangjie Zheng, Zhendong Wang, Mingzhang Yin, and Hai Huang. Score identity distillation: Exponentially fast distillation of pretrained diffusion models for one-step generation. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 62307–62331, 2024. URL https://proceedings.mlr.press/v235/zhou24x.html. 14
Preprint. Under review.
A PPENDIX C ONTENTS Page
A
B
C
Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
16
A.1
One-step generation and distribution matching . . . . . . . . . . . . . . . . . . . . . .
16
A.2
Drifting and its gradient-flow interpretations . . . . . . . . . . . . . . . . . . . . . . .
16
A.3
Representation-based distributional post-training . . . . . . . . . . . . . . . . . . . . .
17
A.4
Mixture transport and componentwise scores . . . . . . . . . . . . . . . . . . . . . . .
17
Derivations and Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
17
B.1
Derivation of the KDE–KL interpretation of Drifting . . . . . . . . . . . . . . . . . . .
17
B.2
Derivation of the Gaussian OT interpretation of FD-Loss . . . . . . . . . . . . . . . .
18
B.3
Global KL Velocity for Gaussian Mixtures . . . . . . . . . . . . . . . . . . . . . . . .
20
B.4
Component Matching and Paired Updates . . . . . . . . . . . . . . . . . . . . . . . . .
20
B.5
Energy descent and detached-target regression . . . . . . . . . . . . . . . . . . . . . .
21
B.6
Conditional image scores in a joint Gaussian . . . . . . . . . . . . . . . . . . . . . . .
22
Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
23
C.1
Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
23
C.2
Weighting multiple representation losses . . . . . . . . . . . . . . . . . . . . . . . . .
26
C.3
Training configurations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
27
C.4
Statistics EMA schedule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
29
D
Additional ImageNet Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
29
E
Evaluation and Comparison Protocols . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
30
E.1
The 10-epoch distribution-model comparison . . . . . . . . . . . . . . . . . . . . . . .
30
E.2
Evaluation in training and held-out representations . . . . . . . . . . . . . . . . . . . .
30
F
Reference GMs and Component Structure . . . . . . . . . . . . . . . . . . . . . . . . . . .
31
G
Score Ridge Settings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
33
H
Component Matching . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
34
I
Computation and Training Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
34
I.1
Training-time comparison and protocol . . . . . . . . . . . . . . . . . . . . . . . . . .
34
I.2
Scaling the number of components . . . . . . . . . . . . . . . . . . . . . . . . . . . .
35
J
Results in Individual Evaluation Representations . . . . . . . . . . . . . . . . . . . . . . .
37
K
Qualitative ImageNet Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
38
L
Additional Text-to-Image Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
42
15
Preprint. Under review.
A
R ELATED W ORK
A.1
O NE - STEP GENERATION AND DISTRIBUTION MATCHING
One-step generation can be learned by approximating a sampling trajectory or by directly matching the output distribution. The technique is widely useful in multimodal generation (Black Forest Labs, 2026), editing (Liu et al., 2026b), and interaction (Zhang et al., 2026b). Consistency models (Song et al., 2023) map points on the same probability-flow trajectory to a common endpoint and support both distillation and training without a teacher. MeanFlow (Geng et al., 2025a) instead learns an average velocity over a time interval, using its relation to the instantaneous velocity as a training target. Improved Mean Flows (Geng et al., 2026) and pixel Mean Flows (Lu et al., 2026) further develop this approach. These methods determine how a generator produces an image in one step. Distributional post-training optimizes the distribution of those images and can be applied to an already trained one-step model. Our experiments use both pMF (Lu et al., 2026) and a one-step initialization from JiT (Li and He, 2026). Distribution Matching Distillation (DMD) (Yin et al., 2024b) expresses a reverse-KL gradient through the difference between real and generated scores at noisy image distributions. A pretrained diffusion model provides the real score, while a separate diffusion model estimates the generated score. Its original formulation also uses a regression loss on teacher-generated pairs. DMD2 (Yin et al., 2024a) removes this regression requirement, uses two-timescale updates to improve the generated-score estimate, and incorporates an adversarial loss on real images. MGFlow also uses a difference of scores for KL matching, but computes these scores from explicit Gaussian mixtures in frozen representation spaces. It does not train diffusion score networks for the matching loss. Kernel distribution matching predates recent one-step diffusion models. Generative Moment Matching Networks (Li et al., 2015) train a generator with maximum mean discrepancy (MMD), a kernel two-sample criterion (Gretton et al., 2012). A characteristic kernel can distinguish distributions beyond their first two moments, whereas a Gaussian Fréchet objective depends only on means and covariances. Our GM representation retains componentwise moments and assignments, allowing distributional training to distinguish modes that a single Gaussian cannot separate. A.2
D RIFTING AND ITS GRADIENT- FLOW INTERPRETATIONS
Drifting (Deng et al., 2026) constructs a field from attraction to real features and repulsion from generated features. The generator is trained to regress its features toward detached field-shifted targets. The iterative distribution update takes place during training; inference uses a single generator evaluation. This separates training-time movement of a distribution from the denoising trajectory used by a diffusion sampler. Several works analyze the relation between these fields and scores. Lai et al. (2026) connect Gaussian-kernel mean shifts to scores of smoothed densities and analyze the residual for more general radial kernels. Turan et al. (2026) develop a spectral and variational view, relating Gaussiankernel Drifting to score differences and studying the effect of bandwidth. Cao et al. (2026) formulate Drifting through Wasserstein gradient flows of KDE-approximated divergences, including extensions beyond KL. The choice of kernel and normalization matters: Franz et al. (2026) show that general normalized Drifting fields need not be conservative, with the Gaussian kernel providing an exception. Our KDE–KL connection uses this Gaussian-kernel setting; it does not identify every Drifting variant with the same KL gradient flow. Wasserstein gradient flows describe steepest descent of a distributional energy under transport geometry (Jordan et al., 1998; Peyré and Cuturi, 2019). W-Flow (Han et al., 2026) applies this perspective to one-step generation using the Sinkhorn divergence between empirical measures. Its field subtracts the self-transport barycentric projection from the cross-distribution projection. W-Flow and Gaussian-kernel Drifting thus use different discrepancies even though both construct fields from sample interactions. Table 1 groups them as sample-based methods, with empirical measures for W-Flow and KDE for Gaussian-kernel Drifting. MGFlow instead estimates a finite collection of component statistics and evaluates transport or score fields from them. 16
Preprint. Under review.
A.3
R EPRESENTATION - BASED DISTRIBUTIONAL POST- TRAINING
FD-Loss (Yang et al., 2026a) directly optimizes the Fréchet distance between real and generated feature statistics. The Gaussian distance has a closed form (Dowson and Landau, 1982), so training requires no learned critic. Real statistics can be precomputed, while queues or exponential moving averages provide generated statistics beyond a single batch. Gradients pass through the current batch’s contribution. Using several frozen encoders extends the objective beyond a single representation. Our Gaussian OT case recovers this moment-based objective. Replacing OT by KL gives a different field with the same Gaussian representation; using a GM changes the representation of the distribution itself. Representation Distribution Matching (RDM) (Feng et al., 2026) organizes visual generation around discrepancies between feature pushforward distributions. Its iRDM implementation uses a Nyström MMD estimator, fixed reference statistics, and joint image–text matching for conditional generation. Elastic Forcing (Zhang et al., 2026a) further expand this approach to videos, presenting a hybrid estimator balancing the computational-statistical tradeoff. These work makes clear that the representation, reference, estimator, and discrepancy all affect training. Our study focuses on the distribution model and its update: we compare Gaussian, GM, and sample-based constructions, then address component assignment and correspondence for GM training. For text-to-image generation, we follow iRDM in concatenating image and text features for joint matching. AdvFD (Gao et al., 2026) augments frozen representations with a learned representation that maximizes the Fréchet discrepancy while the generator minimizes it. Whitening real features constrains the learned representation and prevents a trivial increase through feature scaling. Its adversarial update adapts the representation to the current generator. MGFlow keeps the encoders fixed and instead refines the distribution model within each representation. Amortized Moment Matching (AMFD) (Liu et al., 2026a) represents moment matching through affine denoising operators and amortizes these operators with neural networks. Its formulation supports representation spaces and native generative spaces without explicitly maintaining all covariance matrices in the loss. MGFlow uses explicit full-covariance components and closed-form Gaussian scores or pair costs. The main additional problem is then to assign new generated samples to components while keeping them matched to the reference mixture. Frozen visual representations are also used inside generative models. REPA (Yu et al., 2025) aligns diffusion-transformer hidden states with pretrained features, while representation autoencoders (Zheng et al., 2026) use pretrained representations as a generative latent space. These uses differ from matching the distribution of generated outputs in several feature spaces. The latter lets us change the training objective without changing the generator’s sampling architecture. A.4
M IXTURE TRANSPORT AND COMPONENTWISE SCORES
The Wasserstein distance between two Gaussians has a closed form, but the same is not true for arbitrary Gaussian mixtures. Delon and Desolneux (2020) define M W2 by restricting the coupling to Gaussian component pairs. The squared cost upper-bounds W22 between the mixtures. MGFlow builds on this discrepancy, but its LP acts at the sample level: it allocates generated features to components with prescribed reference masses. This differs from transporting mass between two already fitted mixtures. For KL, mixture scores can be insensitive to mixing proportions when components are well separated (Wenliang and Kanagawa, 2020). MGFlow combines explicit mass constraints with paired component scores, rather than relying on independently weighted marginal scores. Appendix B.4 relates this update to labelled KL.
B
D ERIVATIONS AND P ROOFS
B.1
D ERIVATION OF THE KDE–KL INTERPRETATION OF D RIFTING
We derive the mean-shift–score identity and its KL interpretation in Eq. (4). 17
Preprint. Under review.
From the mean-shift field to the KDE score. For a fixed bandwidth h > 0, let kh (z, y) = N (z; y, h2 I) and define Z rh (z) = kh (z, y) dr(y), r ∈ {p, qθ }. (17) Differentiating with respect to the query z, while holding r fixed, gives y−z kh (z, y), 2 Rh ∇z kh (z, y) dr(y) ∇z log rh (z) = R kh (z, y) dr(y) 1 Ey∼r [kh (z, y)(y − z)] ar (z) = 2 . = h Ey∼r [kh (z, y)] h2 ∇z kh (z, y) =
(18)
The Gaussian kernel and its first derivatives are bounded, so the derivative can pass through the integral. The Gaussian-kernel drifting field vdrift = ap − aqθ therefore satisfies vdrift (z) = h2 ∇z log ph (z) − ∇z log qh (z) (19) (Lai et al., 2026; Turan et al., 2026; Franz et al., 2026). From KL to the score difference. For smooth positive densities ρ and fixed ph , write E(ρ) = R R ρ log(ρ/ph ) dz. A mass-preserving perturbation ρ + εη, with η dz = 0, gives Z d ρ(z) E(ρ + εη) = log + 1 η(z) dz, dε ph (z) ε=0 δE (z) = log ρ(z) − log ph (z) + 1. (20) δρ Taking the negative spatial gradient gives the Wasserstein velocity vρ (z) = −∇z
δE (z) = ∇z log ph (z) − ∇z log ρ(z) δρ
(21)
(Jordan et al., 1998). Evaluating at ρ = qh proves vdrift = h2 vKL (Cao et al., 2026). This is the KL field evaluated at the smoothed densities. For comparison, varying the composite functional G(q) = DKL (kh ∗ q∥ph ) with respect to q gives Z δG qh (z) (x) = kh (z, x) log + 1 dz. (22) δq ph (z) R Here δqh (z) = kh (z, x) δq(x) dx introduces an additional smoothing operation; its negative gradient is not generally the field in Eq. (21). Detached-target regression.
Drifting uses the regression loss (Deng et al., 2026)
Ldrift =
1 X 2 ∥zi − sg[zi + vdrift (zi )]∥2 . B i
(23)
The general calculation in Appendix B.5 gives ∇zi Ldrift = −2vdrift (zi )/B. B.2
D ERIVATION OF THE G AUSSIAN OT INTERPRETATION OF FD-L OSS
We derive the Gaussian transport cost, its feature gradient, and the finite-sample factors underlying the FD field in Eq. (4). 18
Preprint. Under review.
The Gaussian optimal transport cost and map. Let qG = N (µq , Σq ) and pG = N (µp , Σp ), with positive-definite covariances. All matrix square roots below are symmetric positive definite. The optimal Gaussian transport map is 1/2 1/2 −1/2 T (z) = µp + A(z − µq ), A = Σ−1/2 (Σ1/2 Σq . (24) q q Σp Σq ) Indeed, AΣq A = Σp , so T sends qG to pG . Since A is positive definite, T is the gradient of a convex quadratic and is optimal for quadratic transport (Peyré and Cuturi, 2019). Writing z − T (z) = (µq − µp ) + (I − A)(z − µq ), we obtain
W22 (qG , pG ) = EqG ∥z − T (z)∥22 = ∥µq − µp ∥22 + tr (I − A)Σq (I − A) = ∥µq − µp ∥22 + tr(Σq + Σp − 2AΣq ) 1/2 1/2 = ∥µq − µp ∥22 + tr Σq + Σp − 2(Σ1/2 = DFD . Σ Σ ) p q q
(25)
The third line uses AΣq A = Σp ; the last uses cyclicity of the trace. This is the Gaussian Fréchet formula (Dowson and Landau, 1982). Differentiating the mean and covariance. Set F = 21 DFD , keeping the reference moments fixed. The mean term gives ∇µq F = µq − µp . For the covariance term, use the equivalent cross term 1/2 1/2 tr(M 1/2 ), where M = Σp Σq Σp . The two cross-term matrices are BB T and B T B for B = 1/2 1/2 Σq Σp , so they have the same eigenvalues. To differentiate its trace, write S = M 1/2 and differentiate S 2 = M : S dS + dS S = dM =⇒ d tr(M 1/2 ) = 12 tr(M −1/2 dM ). (26) The implication follows by multiplying by S −1 and taking the trace; it does not require M and dM to commute. Therefore 1/2 dΣq F = 21 tr(dΣq ) − 12 tr(M −1/2 Σ1/2 p dΣq Σp ) (27) = 12 tr (I − A) dΣq . 1/2
1/2
Here Σp M −1/2 Σp = A: both are positive-definite solutions of XΣq X = Σp , whose unique 1/2 1/2 solution follows by squaring Σq XΣq . Thus ∇µq F = µq − µp , ∇Σq F = 12 (I − A). (28) From moment derivatives to feature movement. Let q be any generated distribution with finite second moments and positive-definite covariance. Perturb its features by zε = z +εu(z), and denote their distribution by qε . Differentiating the mean and Σq = Eq [zz T ] − µq µT q at ε = 0 gives µ̇q = Eq [u(z)], T Σ̇q = Eq [u(z)z T + zu(z)T ] − µ̇q µT q − µq µ̇q
= Eq [u(z)(z − µq )T + (z − µq )u(z)T ]. Combining these with Eq. (28) yields d F (µqε , Σqε ) = (µq − µp )T µ̇q + 12 tr((I − A)Σ̇q ) dε 0 h i T = Eq µq − µp + (I − A)(z − µq ) u(z)
(29)
= Eq [(z − T (z))T u(z)]. (30) The two covariance terms combine because I − A is symmetric. Hence the moment objective has descent field T (z) − z. When q = qG , this is the Wasserstein velocity of 12 W22 (q, pG ): vFD,t (z) = TqG,t →pG (z) − z, ∂t qG,t + ∇ · (qG,t vFD,t ) = 0. (31) The affine field preserves Gaussianity. For non-Gaussian q, the same feature gradient differentiates the moment objective, but T need only match the reference mean and covariance, not the full distribution. Setting v = vFD and η = 1 in the regression formula of Appendix B.5 gives the target sg[T (zi )] and gradient (zi − T (zi ))/B. This recovers the same field; using DFD = 2F instead of F multiplies its gradient by two. 19
Preprint. Under review.
Finite-sample and EMA factors. For a feature zi contributing weight ci to the estimated mean and raw second moment, while all other contributions are fixed, dµq = ci dzi , dΣq = ci dzi (zi − µq )T + (zi − µq )dziT . (32) Substitution into Eq. (28) gives ∇zi F = ci µq − µp + (I − A)(zi − µq ) = ci (zi − T (zi )). (33) For moments averaged over N features, ci = 1/N ; detached queue entries affect the moments but receive no gradient. For EMA updates of the mean and raw second moment with decay β and current batch size B, ci = (1 − β)/B, with T evaluated at the updated moments (Yang et al., 2026a). P If the empirical covariance instead uses Σq = (N − 1)−1 j (zj − µq )(zj − µq )T , its derivative has coefficient 1/(N − 1), while the mean retains 1/N : µq − µp (I − A)(zi − µq ) ∇z i F = + . (34) N N −1 P Here the terms from differentiating the centering cancel because j (zj − µq ) = 0. B.3
G LOBAL KL V ELOCITY FOR G AUSSIAN M IXTURES
Let P and Q have positive mixture weights and positive-definite component covariances. Applying Eq. (21) to these densities gives vglobal = ∇ log P − ∇ log Q. To evaluate each mixture score, P differentiate Q(z) = k ωk qk (z): P ωk ∇z qk (z) ∇z log Q(z) = k Q(z) X ωk qk (z) X = ∇z log qk (z) = γQ,k (z)sq,k (z). (35) Q(z) k
k
For qk = N (µq,k , Σq,k ), −1 sq,k (z) = − 12 ∇z (z − µq,k )T Σ−1 q,k (z − µq,k ) = Σq,k (µq,k − z).
(36)
Applying the same calculation to P yields Eq. (8). The two scores use their respective posterior responsibilities γQ,k = ωk qk /Q and γP,k = πk pk /P . B.4
C OMPONENT M ATCHING AND PAIRED U PDATES
Mixture transport and its paired upper bound. For each component pair, let γij Pbe an optimal coupling of qi and pj , with cost Cij = W22 (qi , pj ). For any Γ ∈ U (ω, π), γ = i,j Γij γij has P P P marginals i ωi qi = Q and j πj pj = P . Its transport cost is i,j Γij Cij , so minimizing over Γ gives W22 (Q, P ) ≤ M W22 (Q, P ) (Delon and Desolneux, 2020). When ω = π, the feasible coupling diag(π) gives X W22 (Q, P ) ≤ M W22 (Q, P ) ≤ πk Ckk , (37) k
which proves the two bounds in Eq. (11). When is diagonal component transport optimal? Suppose Cii ≤ Cij for every i, j and both mixture weights are π. For any feasible Γ, its row sums give X X X Γij Cij − πi Cii = Γij (Cij − Cii ) ≥ 0. (38) i,j
i
i,j
The diagonal coupling attains equality and is therefore optimal. It is the unique optimum if Cii < Cij for all j ̸= i. Matched weights alone do not imply this condition: swapping the labels of two equally weighted, distinct Gaussians makes off-diagonal pairing optimal. Empirically, We inspected a pMF-H run using SIM encoders, K = 1 + 4, and a global batch of 1,024. For each encoder, the training code forms all 16 Gaussian costs and solves the component transport LP every 100 updates. The logs cover 34.8 epochs and contain 435 distinct plan refreshes per encoder. Every recorded plan has diagonal mass 1 and exactly four nonzero entries, giving Γ = diag(π) in all three feature spaces. 20
Preprint. Under review.
Paired KL and the marginal KL bound. For fixed positive weights πk summing to one, introe z) = πk qk (z) and Pe(k, z) = πk pk (z). Since the weights cancel duce the labelled distributions Q(k, inside the ratio, XZ X πk qk (z) e Pe) = πk qk (z) log DKL (Q∥ dz = πk DKL (qk ∥pk ) = Fpair . (39) πk pk (z) k
k
e z) = Q(z)γQ,k (z) and Pe(k, z) = P (z)γP,k (z). Then To relate this to marginal KL, factor Q(k, XZ Q(z) γQ,k (z) Fpair = Q(z)γQ,k (z) log + log dz P (z) γP,k (z) k # " X γQ,k (z) = DKL (Q∥P ) + Ez∼Q γQ,k (z) log γP,k (z) k
= DKL (Q∥P ) + Ez∼Q [DKL (γQ (· | z)∥γP (· | z))] ≥ DKL (Q∥P ). (40) P The second line uses k γQ,k = 1. Equality holds exactly when the posterior label distributions agree Q-almost everywhere. From the paired energy to component fields. Let uk move component qk , and set gk = ∇z log(qk /pk ). Using the KL variation in Eq. (20) and ∂t qk = −∇ · (qk uk ) gives X d Fpair = πk Eqk [gkT uk ], (41) dt k P with vanishing boundary terms. product Wasserstein distance k πk W22 (qk , qk′ ) corP The squared responds to squared speed k πk Eqk ∥uk ∥22 . Its steepest-descent velocity minimizes X 1X πk Eqk ∥uk + gk ∥22 − ∥gk ∥22 . (42) πk Eqk gkT uk + 12 ∥uk ∥22 = 2 k
k
The minimum is attained at vk (z) = −gk (z) = ∇z log pk (z) − ∇z log qk (z).
(43)
Thus the common weights in the energy and metric do not introduce an extra factor πk into the component velocity (Jordan et al., 1998; Peyré and Cuturi, 2019). The regularized scores in Eq. (13) are those of rkλ = N (µr,k , Σr,k + λI). Training combines their ∗ differences with the detached LP assignments Rnk and applies the regression gradient derived in ∗ Appendix B.5. Unlike the marginal KL velocity in Eq. (8), this update uses Rnk for both scores rather than separate mixture posteriors. B.5
E NERGY DESCENT AND DETACHED - TARGET REGRESSION
Continuous distributional descent. Let ϕt = δF /δQ|Q=Qt , where t denotes continuous flow time. For a smooth density satisfying ∂t Qt + ∇ · (Qt vt ) = 0, assume sufficient decay or no-flux boundary conditions so that the boundary term below vanishes. The chain rule and integration by parts give (Jordan et al., 1998; Peyré and Cuturi, 2019) Z d F(Qt ) = ϕt (z) ∂t Qt (z) dz dt Z = − ϕt (z)∇ · (Qt (z)vt (z)) dz Z = Qt (z)∇ϕt (z)T vt (z) dz. (44) Substituting vt = −∇ϕt proves Eq. (3): d F(Qt ) = −Ez∼Qt ∥vt (z)∥22 ≤ 0. dt 21
(45)
Preprint. Under review.
Applying a field toPgenerator training. For zn = E(Gθ (ξn , cn ), cn ), define z̄n = sg[zn + ηvn ] and Lv = (2B)−1 n ∥zn − z̄n ∥22 . The entire target is constant during differentiation, so ∇zn Lv =
zn − z̄n η = − vn , B B
∇θ Lv = −
η X T J vn , B n n
Jn =
∂zn . ∂θ
(46)
No derivative of the estimated field enters this gradient. The regression transfers the field through the encoder and generator Jacobians; the continuous-flow energy identity is not a finite-step guarantee for the parameter optimizer. B.6
C ONDITIONAL IMAGE SCORES IN A JOINT G AUSSIAN
Let p(z, u) be a joint Gaussian over image features z and text features u, with µ Σzz Σzu µ= z , Σ= ≻ 0. µu Σuz Σuu
(47)
Write a = z − µz , b = u − µu , and [ξz ; ξu ] = Σ−1 [a; b]. The Gaussian image score is −ξz . Instead of forming the full inverse, solve the block equations: Σzz ξz + Σzu ξu = a, Σuz ξz + Σuu ξu = b.
(48)
The second equation gives ξu = Σ−1 uu (b − Σuz ξz ). Substituting into the first yields (Σzz − Σzu Σ−1 Σuz ) ξz = a − Σzu Σ−1 uu b. {z uu } |
(49)
Σz|u
Define µz|u = µz + Σzu Σ−1 uu (u − µu ). We obtain ∇z log p(z, u) = −ξz = Σ−1 z|u (µz|u − z) = ∇z log p(z | u).
(50)
The last equality also follows from log p(z, u) = log p(z | u) + log p(u), since u is fixed. For paired GM matching, we apply this calculation to each joint Gaussian component and combine the ∗ image-score differences with Rnk . Image–text cross-covariance therefore affects the image update even though text features receive no gradient.
22
Preprint. Under review.
C
I MPLEMENTATION
C.1
A LGORITHMS
We fit the real-data distributions once, initialize generated statistics from the pretrained generator, and then update the generator and its statistics jointly. Each encoder and component count has a separate reference and moment state. A K = 1 + 4 configuration therefore contains one Gaussian branch and one four-component branch. Algorithms 1 and 2 are run once for each encoder–branch pair; Algorithm 3 is repeated at each training step. We write E(x, c) for the features being matched: image features for image-only training, or concatenated image and text features for joint matching. Reference fitting. Algorithm 1 gives the full-data EM update (Dempster et al., 1977). Features are cached and processed in blocks, but parameters are updated only after a complete pass. The saved covariance is the raw centered second moment. Algorithm 1 Full-covariance reference GM fitting Require: Cached real features Y = {yn }N n=1 , component count K, initial parameters (π, µ, Σ), maximum EM iterations Tmax PK Ensure: Frozen reference P = k=1 πk N (µk , Σk ) 1: if K = 1 then P P 2: µ1 ← N −1 n yn ; Σ1 ← N −1 n yn ynT − µ1 µT1 3: return P = N (µ1 , Σ1 ), with π1 = 1 4: end if 5: for t = 1, . . . , Tmax do 6: (π old , µold , Σold ) ← (π, µ, Σ) 7: Set Nk = 0, bk = 0, Sk = 0 for all k ▷ Reset full-data accumulators 8: for each feature block with indices I do 9: For all n ∈ I and k, compute the E-step responsibilities: old π old N (yn ; µold k , Σk ) rnk ← P k old old old j πj N (yn ; µj , ΣjP ) 10: For all k, accumulate Nk ← Nk + n∈I rnk P P 11: bk ← bk + n∈I rnk yn ; Sk ← Sk + n∈I rnk yn ynT for all k 12: end for 13: For all k, perform the M-step after processing all N features: πk ← Nk /N ; µk ← bk /Nk ; Σk ← Sk /Nk − µk µTk 14: Compute parameter changes relative to (π old , µold , Σold ) 15: if the convergence test below passes then 16: break 17: end if 18: end for P 19: return frozen reference P = k πk N (µk , Σk )
The final full-data fits start from a hard-assignment initialization. We check maximum absolute weight change and relative RMS changes in means and covariances, using thresholds 0.002, 0.006, and 0.020, respectively. The test must pass on two consecutive updates without an increase in these changes. These checks concern the fitted reference; the training-time LP does not refit or alter P .
23
Preprint. Under review.
P Generated statistics. For assignments Rnk , define the batch statistics b ak = B −1 n Rnk , bbk = P P B −1 n Rnk zn , and Sbk = B −1 n Rnk zn znT . The stored state is H = {ak , bk , Sk }k , from which µq,k = bk /ak and Σq,k = Sk /ak − µq,k µT q,k . Algorithm 2 initializes this state without a generator update. Subsequent updates use an EMA, following FD-Loss (Yang et al., 2026a). For the LP assignments R∗ in Eq. (10), the capacity constraint fixes ak = b ak = πk . The conditional batch mean and raw second moment are therefore X 1 X ∗ ck = 1 µ bk = Rnk zn , M R∗ zn znT . (51) Bπk n Bπk n nk Writing Mq,k = Sk /ak , the EMA with decay β gives µ+ µk , q,k = βµq,k +(1−β)b
+ ck , Mq,k = βMq,k +(1−β)M
+ + + T Σ+ q,k = Mq,k −µq,k (µq,k ) . (52)
Algorithm 2 Warm-starting generated component statistics P Require: Pretrained generator Gθ0 , frozen encoder E, reference P = k πk pk , sample budget N0 , assignment batch size B0 Ensure: Initial moment state H = {ak , bk , Sk }K k=1 ; no generator update 1: Disable gradient recording; set processed count m = 0 2: Set totals Ak = 0, Uk = 0, Vk = 0 for all k 3: while m < N0 do 4: B ← min(B0 , N0 − m) ▷ Include the final partial batch 5: Sample {(ξn , cn )}B n=1 using the training sampling scheme 6: Compute zn = E(Gθ0 (ξn , cn ), cn ) for all n 7: if K = 1 then 8: Set Rn1 = 1 for every n 9: else 10: Cnk ← − log(πk pk (zn )) Pfor all n, k 11: Solve R ∈ arg minR≥0 n,k Rnk Cnk subject to R1K = 1B and RT 1B = Bπ 12: end if P P 13: For all k, accumulate Ak ← Ak + n Rnk ; Uk ← Uk + n Rnk zn P 14: Vk ← Vk + n Rnk zn znT for all k; m ← m + B 15: end while 16: For all k, set (ak , bk , Sk ) ← (Ak , Uk , Vk )/N0 17: return H = {ak , bk , Sk }k ▷ ak = πk by the LP constraints
We use B0 = 1,024 and N0 = 50,000. Encoder microbatches are collected into these assignment batches before solving the LP. Initialization averages the accumulated moments.
24
Preprint. Under review.
LP-paired training. Algorithm 3 distinguishes the two gradient paths. OT differentiates through the current batch in the candidate EMA state. KL evaluates the field using the stored state before its update and detaches the regression target. In both cases the assignment is fixed during backpropagation, and the EMA is committed once per global batch. The Gaussian OT cost uses raw covariances, clipping numerically negative eigenvalues to zero when evaluating matrix square roots. Here sg stops gradients without changing values. Encoder parameters remain fixed, but gradients pass through their image inputs to the generator. Within each encoder–branch iteration below, we omit the indices (e, K) on component parameters and moment entries. Algorithm 3 One generator update with LP-paired OT or KL Require: Generator Gθ and its optimizer; frozen encoders Ee and references Pe,K Require: Stored states He,K , branch counts K ∈ K, weights we,K , global batch size B, EMA decay βt Require: Matching objective (OT or KL); for KL, score ridges λe,K and field scale η Ensure: Updated generator parameters and one EMA update per encoder–branch pair 1: Clear generator gradients; set L ← 0 2: Sample {(ξn , cn )}B n=1 ; generate xn = Gθ (ξn , cn ) once 3: For each encoder e, compute ze,n = Ee (xn , cn ) for all n 4: for each encoder e and branch K do P 5: Set zn = ze,n and read the fixed reference Pe,K = k πk pk 6: if K = 1 then 7: Set Rn1 = 1 for every n 8: else 9: Compute Cnk = − log(πk pk (sg[zn ])) and solve Eq. (10) 10: end if 11: R ← sg[R] ▷ Reuse this assignment for moments and paired fields P P 12: b ak ← B −1 n Rnk ; b bk ← B −1 n Rnk zn for all k P 13: Sbk ← B −1 n Rnk zn znT for all k + b 14: He,K ← βt sg[He,K ] + (1 − βt )(b a, b b, S) Keep He,K unchanged until after the optimizer step. 15: if OT matching then + + + + + + + T 16: µ+ q,k ← bk /ak ; Σq,k ← Sk /ak − µq,k (µq,k ) for all k P + + 2 17: ℓe,K ← k πk W2 (N (µq,k , Σq,k ), pk ) 18: else ▷ KL matching 19: From stored He,K , set µq,k = bk /ak and Σq,k = Sk /ak − µq,k µTq,k 20: Without gradients, solve (Σr,k + λe,K I)sr,k,n = µr,k − zn forP r ∈ {p, q}, all k, n, using Cholesky factors shared across samples 21: vn ← k Rnk (sp,k,n − sq,k,n ); z̄n ← sg[zn + ηvn ] P 22: ℓe,K ← (2B)−1 n ∥zn − z̄n ∥22 23: end if 24: L ← L + we,K ℓe,K 25: end for 26: Backpropagate L to θ and take one optimizer step + 27: Commit He,K ← sg[He,K ] for every encoder and branch
Global-batch gradients with limited memory. When a global batch spans several devices or accumulation passes, we first collect detached features and solve the LP on the complete batch. We compute the loss gradient with respect to these features, then replay each generator microbatch with the same noise, conditions, and random state. Sequential encoder vector–Jacobian products propagate its slice of the feature gradient to the generator (Gao et al., 2021). The feature loss is already normalized by the global batch size; parameter gradients are summed without a second batch-size normalization. The assignment and candidate EMA state remain fixed across replay passes.
25
Preprint. Under review.
C.2
W EIGHTING MULTIPLE REPRESENTATION LOSSES
The discrepancies have different scales across encoders. FD-Loss (Yang et al., 2026a) divides each encoder’s FD by its detached current value plus a constant, giving a weight [sg{FD(Re , Ge )}+c]−1 . Here Re and Ge denote real and generated feature statistics. This weight changes with the generator. We instead use the discrepancy between real training features Re and real validation features Ve to set a fixed scale for each encoder. For OT, we use we = 1/FD(Re , Ve ), with the same encoder weight for every component-count branch. In SigLIP/Inception/MAE order, the denominators are (0.62469, 1.67956, 0.04212). Table 6 compares the published FD-Loss results with our aligned single-Gaussian FD runs. The aligned recipe uses this fixed normalization and the statistics schedule in Appendix C.4, without a GM branch. It reduces FDr6 on pMF-B. Table 6: Single-Gaussian FD with fixed reference normalization. FDr6 ↓ after 100 epochs with SIM encoders. FD-Loss results are from Yang et al. (2026a); the aligned runs use our statisticsupdate recipe. Training recipe
Encoder normalization
pMF-B
FD-Loss Aligned Gaussian FD
sg{FD(R, G)} + c FD(R, V )
3.50 3.17
For KL, we similarly use we = 1/DKL (Re ∥Ve ), where Re and Ve are single-Gaussian fits. These calibration distributions are distinct from the training objective DKL (Q∥P ). The resulting SigLIP/Inception/MAE weights are (0.327, 0.110, 0.276) for the ImageNet reference. Each weight is reused across the encoder’s Gaussian and mixture branches; we do not estimate a separate marginal GM KL to normalize each branch. For multiple branches, the final loss weight is we,K = we αK . All multi-branch configurations use αK = 1 for every included branch. Text-to-image weights are recalculated for the corresponding image-only or joint reference.
26
Preprint. Under review.
C.3 C.3.1
T RAINING CONFIGURATIONS I MAGE N ET
Table 7 summarizes the common ImageNet settings for the six JiT/pMF backbones. We use frozen SigLIP2 (Tschannen et al., 2025), Inception-v3 (Szegedy et al., 2016), and MAE (He et al., 2022) encoders, abbreviated as SIM. All six models use AdamW (Loshchilov and Hutter, 2019) and a global batch size of 1,024, and are trained for 100 epochs. One epoch corresponds to 1,250 optimizer updates. Learning rates are 10−5 for JiT and 10−6 for pMF, with five warmup epochs followed by cosine decay. We use no gradient clipping or dropout. Table 7: ImageNet post-training configurations. Slash-separated pMF sampling values follow B/L/H order. Configuration Model and representations Model sizes Image resolution Initialization Frozen encoders Training Training epochs Global batch size Optimizer Learning rate Learning-rate schedule Warmup epochs Weight decay Precision Statistics estimation Estimator Warm-start samples EMA decay Sampling and evaluation NFE CFG Noise scale Evaluation samples
JiT (Li and He, 2026)
pMF (Lu et al., 2026)
B, L, H
B, L, H 256 × 256 Official pretrained weights SigLIP2, Inception-v3, MAE 100 1,024 AdamW, (β1 , β2 ) = (0.9, 0.95)
10−5
10−6 Cosine decay 5 0 BF16 mixed precision
EMA 50,000 from the pretrained generator 0.995 → 0.999, linear over epochs 8–32 1 1.0 1.0
8.5 / 7.0 / 7.0 1.0 / 1.0 / 2.0 50,000
The guidance intervals (Ho and Salimans, 2022; Kynkäänniemi et al., 2024) for pMF-B/L/H are [0.1, 0.7], [0.2, 0.7], and [0.2, 0.6]; JiT uses [0.1, 1.0]. Evaluation uses 50,000 one-step samples. Each component count in a configuration has its own fitted reference and generated statistics. Encoder and branch weights are specified in Appendix C.2.
27
Preprint. Under review.
C.3.2
T EXT- TO - IMAGE GENERATION
Table 5 uses two checkpoints initialized from the official distilled FLUX.2 [klein] 4B model (Black Forest Labs, 2026). All use MGFlow-KL with equally weighted K = 1 and K = 4 branches and frozen SigLIP, MAE, and Inception image encoders. The component-count choice is discussed in Appendix I.2. Following iRDM (Feng et al., 2026), the joint variant concatenates image features with scaled, normalized SigLIP2 (Tschannen et al., 2025) text features and retains the full image– text cross-covariance. For each caption c, we normalize its text embedding as b t(c) = t(c)/∥t(c)∥2 and form the joint feature [fe (x); βe b t(c)], without additional normalization of the image feature fe (x). We set βe = 0.25 me /mt , where me and mt are the square roots of the median pairwise squared distances of image and normalized text features on the same 2,000 reference samples (seed 3407), excluding self-pairs. In SigLIP/Inception/MAE order, βe ≈ (3.76114, 4.11487, 0.51504). These scales remain fixed during reference fitting and training and are shared by the K = 1 and K = 4 branches. The image-only variant omits text from the matching objective, not from the generator conditioning. Reference data. We follow the reference-data construction protocol of iRDM (Feng et al., 2026) and AMFD (Liu et al., 2026a). For each of 82,783 COCO (Lin et al., 2014) train2014 captions, we retain the three highest-PickScore images from 24 teacher-generated candidates, yielding 248,349 reconstructed images. We add 53,357 accepted teacher images from 553 GenEval prompts, with at most 100 accepted images per prompt. Both variants use this shared reference set of 301,706 images, sampled uniformly by image. Training and evaluation. Table 8 lists the optimizer, statistics, and sampling settings. Inverse-KL encoder weights are computed for each variant. Evaluation uses one Euler step. GenEval (Ghosh et al., 2023) averages its six category scores equally, using the same prompt set as the GenEval reference subset. PickScore (Kirstain et al., 2023) is evaluated on 499 Pick-a-Pic prompts. All evaluation protocols are aligned with AMFD (Liu et al., 2026a). Table 8: Text-to-image post-training configurations. Both variants use reconstructed COCO and GenEval reference images. Configuration Model and reference Initialization Image resolution Image encoders Text features Reference images Distribution branches Branch weights
Joint
Image-only
Distilled FLUX.2 [klein] 4B 512 × 512 SigLIP2, MAE, Inception-v3 Frozen SigLIP2 – 301,706, including 53,357 GenEval images K =1+4 1:1
Optimization and statistics Training steps Global batch size Optimizer Learning rate Learning-rate schedule Statistics EMA decay
1,000 1,024 AdamW (Loshchilov and Hutter, 2019) 5 × 10−6 150-step warmup, then constant 0.99
Sampling and evaluation NFE / CFG Text context length GenEval samples PickScore prompts
1/1 512 tokens 2,212 (four per prompt) 499 Pick-a-Pic prompts
28
Preprint. Under review.
C.4
S TATISTICS EMA SCHEDULE
An EMA averages out batch noise but also delays the response to a changing generator. We use decay 0.995 for the first eight ImageNet epochs, increase it linearly to 0.999 over epochs 8–32, and keep it fixed thereafter. The smaller initial decay lets the component statistics follow early distribution changes; the larger final decay smooths the estimates later in training. This EMA updates feature statistics, not generator parameters. Table 9 compares the schedule with a constant 0.999 decay for JiT-B in the aligned single-Gaussian FD implementation. After 50 epochs, the schedule reduces FDr6 from 4.99 to 4.87.
Table 9: Statistics EMA decay. JiT-B trained for 50 epochs with aligned Gaussian-FD, fixed FD(R, V ) normalization, and SIM encoders. We report FDr6 ↓.
D
Statistics decay
FDr6 ↓
Constant 0.999 0.995 → 0.999
4.99 4.87
A DDITIONAL I MAGE N ET R ESULTS
Table 10 supplements the JiT/pMF comparison in Table 4 with other discrete-, latent-, and pixelspace generators.
Table 10: Additional ImageNet 256×256 baselines. Published results from FD-Loss (Yang et al., 2026a), AdvFD (Gao et al., 2026), and AMFD (Liu et al., 2026a). FDr6 averages normalized Fréchet distances over six encoders; FDr3 excludes SigLIP, Inception, and MAE. † denotes the fullCFG NFE upper bound for interval CFG; dashes denote unavailable results. JiT/pMF results are in Table 4. #Params FDr6 ↓ FDr3 ↓ FID ↓
IS ↑
Method (K )
NFE
Space
Reference (real images) 50k validation images
N/A
N/A
N/A
1.00
1.00
1.68
232.2
10×2 256×2×4
discrete discrete
2B 1.1B
6.70 3.57
6.77 3.35
1.97 1.01
304.6 281.9
250×2 256×2×100 50×2† 256×2×100 256×2×100
latent latent latent latent latent
675M 478M 1.9B 942M 478M
8.44 6.68 6.13 5.61 5.49
9.20 7.32 6.16 6.31 6.05
2.12 1.80 1.68 1.56 1.39
256.7 293.4 274.1 299.5 306.2
250×2† 250×2† 250×2 250×2 250×2† 50×2†
latent latent latent latent latent latent
685M 675M 675M 675M 676M 839M
4.64 5.45 4.57 5.70 3.04 3.26
5.15 6.05 5.02 6.38 3.33 3.92
1.54 1.42 1.42 1.26 1.17 1.16
302.9 306.1 294.3 309.3 298.3 261.0
1 1 2
latent latent latent
463M 610M 610M
10.92 8.39 7.48
11.32 8.72 7.79
1.53 1.82 1.61
257.2 278.9 289.1
100×2
pixel pixel
1.0B 465M
5.01 10.51
– 11.18
2.10 1.43
318.8 305.8
Discrete-space models VAR-d30 (Tian et al., 2024) BAR-L (Yu et al., 2026) Latent-space models, multi-step without semantic distillation SiT-XL/2 (Ma et al., 2024) MAR-L (Li et al., 2024) FlowAR-H (Ren et al., 2025) MAR-H (Li et al., 2024) MAR-L, DeTok (Yang et al., 2026b) with semantic distillation REG (Wu et al., 2025) SiT-XL/2-REPA (Yu et al., 2025) LightningDiT (Yao et al., 2025) DDT-XL (Wang et al., 2026b) REPA-E (Leng et al., 2025) RAE-XL (Zheng et al., 2026) Latent-space models, one-step Drift-L (latent) (Deng et al., 2026) iMF-XL (Geng et al., 2026) iMF-XL (Geng et al., 2026) Pixel-space models PixNerd-XL (Wang et al., 2026a) Drift-L (pixel) (Deng et al., 2026)
1
29
Preprint. Under review.
E
E VALUATION AND C OMPARISON P ROTOCOLS
E.1
T HE 10- EPOCH DISTRIBUTION - MODEL COMPARISON
Shared setup. All six methods in Table 1 start from the same pretrained pMF-B and use frozen Inception features on ImageNet 256 × 256. We train for 10 epochs with a global batch size of 1,024 and a learning rate of 10−6 . Sampling uses one step, CFG 8.5, a guidance interval of [0.1, 0.7], and a noise scale of 1. Class labels are drawn uniformly, and each feature objective matches the global class-marginal distribution. We evaluate 50,000 generated images using FDr6 under the protocol in Appendix E.2. The Gaussian and GM methods fit full-covariance reference distributions to real features and keep them fixed during training. Generated-distribution statistics are initialized from 50,000 samples of the pretrained generator and updated by EMA of first and second moments. GM methods additionally use the capacity-constrained LP in Eq. (10) to assign each batch to components with the reference weights. FD-Loss. The aligned FD-Loss (Yang et al., 2026a) baseline uses a single Gaussian and minimizes its squared Wasserstein distance to the reference. It retains the Gaussian reference, statistics estimator, and fixed real-training/validation normalization of our OT implementation, with no mixture term. Gradients pass through the current batch’s contribution to the EMA statistics, while historical statistics are detached. Gaussian KL. This baseline uses K = 1 and the Gaussian score difference vKL (z) = sp (z) − sq (z). We train with detached-target regression in Eq. (14), using η = 1, and update the generated statistics after the generator step. MGFlow-W2 . We use a single K = 4 branch with LP assignments and minimize the weighted sum of paired Gaussian costs in Eq. (11). The loss uses the same fixed normalization as the aligned FD-Loss baseline. We differentiate through the current batch statistics, keeping the LP assignments and historical statistics detached. MGFlow-KL. We use a single K = 16 branch. The LP assignments weight the paired component score differences in Eq. (13). Training uses the same detached-target regression and statistics-update order as Gaussian KL. Neither GM method includes an additional Gaussian branch or another component count. W-Flow. W-Flow (Han et al., 2026) uses the released quadratic-cost Sinkhorn objective with ϵ = 0.05, ten iterations, and its feature and force normalization. Each update uses 1,024 gradientcarrying queries, an independent batch of 1,024 generated support samples, and 1,024 real featurebank samples. Gaussian-kernel Drifting. Gaussian-kernel Drifting (Deng et al., 2026; Turan et al., 2026) uses the current generated batch as detached negative support and 1,024 real feature-bank samples. Its fixed bandwidths satisfy 2σ 2 ∈ {0.02, 0.05, 0.2} in normalized feature coordinates, with the selfinteraction diagonal masked and per-scale force normalization. Both sample-based methods gather the complete global batch before computing their fields. Their sample interactions replace the moment estimator, so they require no generated statistics initialization or EMA. E.2
E VALUATION IN TRAINING AND HELD - OUT REPRESENTATIONS
Representation encoders. Table 11 lists the six frozen encoders used for ImageNet evaluation. SIM uses SigLIP2, Inception-v3, and MAE for training; ConvNeXt-v2, DINOv2, and CLIP are held out. We use the pooled or CLS features without the classification or contrastive projection head. The five timm encoders use bicubic resizing and their pretrained input normalization; Inception-v3 uses TensorFlow-compatible FID preprocessing. 30
Preprint. Under review.
Table 11: ImageNet representation encoders. Input denotes the resized image side length before patch padding. The first three encoders form SIM; the last three define FDr3 . Encoder
Variant
SigLIP2 (Tschannen et al., 2025) ViT-SO400M/16 Inception-v3 (Szegedy et al., 2016) FID weights MAE (He et al., 2022) ViT-L/16 ConvNeXt-v2 (Woo et al., 2023) DINOv2 (Oquab et al., 2023) CLIP (Radford et al., 2021)
Base, IN-22K → IN-1K ViT-L/14 ViT-L/14, OpenAI
Input
Dimension
Feature
224 299 224
1,152 2,048 1,024
Attention pool Global average pool CLS token
224 256 256
1,024 1,024 1,024
Global average pool CLS token CLS token
Metric aggregation. We evaluate 50,000 generated images using FDr6 (Yang et al., 2026a). For an evaluation encoder e, let Fe be the Fréchet distance between generated and real features, and let be be the corresponding real-validation reference value. For Inception features, the unnormalized distance Fe is FID (FD-Inception) (Heusel et al., 2017). The normalized score is FDre = Fe /be . We use the arithmetic mean across the six evaluation spaces, 1 X FDr6 = FDre , H6 = {Inception, ConvNeXt, DINOv2, MAE, SigLIP, CLIP}. 6 e∈H6 (53) For a single training encoder etrain , we additionally report X 1 FDre . (54) FDr5 = 5 e∈H6 \{etrain }
For SIM training, the corresponding held-out score is 1 (FDrConvNeXt + FDrDINOv2 + FDrCLIP ) . (55) 3 IS (Salimans et al., 2016) is the exponentiated mean KL divergence between each generated image’s Inception-v3 class probabilities and their marginal over generated images. FDr3 =
F
R EFERENCE GM S AND C OMPONENT S TRUCTURE
Reference fitting is described in Algorithm 1. This section examines the structure of the fitted components. We measure class–component association on 200 randomly sampled training images per ImageNet class, for 200,000 images in total. For class c, let 1 X rck = γk (yn ), qc = max rck . (56) k Nc n:y ∈c n
Here qc is the average mass assigned to the dominant component of class c. In Table 12, the mean image-level peak remains above 99%, while class concentration decreases at K = 16. Individual images have sharp assignments, but images within the same class are distributed across several components. Figure 4 shows associations beyond animals and vehicles. For SigLIP, furniture and prepared food share a dominant component at K = 4 but peak in distinct components at K = 16. The grouping also depends on the encoder: at K = 4, Inception places dogs and road vehicles in the same dominant component, whereas MAE and SigLIP separate them.
31
Preprint. Under review.
Table 12: Class concentration versus image-level assignment. We use 200 images per ImageNet class. Class concentration qc is the mean responsibility of class c for its dominant component; “Image peak” averages the maximum responsibility of each image. Encoder
K
Mean qc
Classes with qc ≥ 0.9
Image peak
MAE SigLIP Inception
4 4 4
81.24% 91.12% 82.23%
398 744 461
99.88% 99.96% 99.97%
MAE SigLIP Inception
16 16 16
64.79% 71.88% 77.21%
113 229 305
99.79% 99.94% 99.96%
MAE | K = 4
SigLIP | K = 4
Inception | K = 4
Dogs (118) Domestic cats (5) Birds (59) Fish (16) Reptiles (35) Insects (27) Primates (20) Road vehicles (20) Furniture (20) Musical instruments (26) Clothing (32) Fruit & vegetables (21) Prepared food (18) 1
2
3
4
1
2
3
4
1
2
3
4
Component
Component
Component
MAE | K = 16
SigLIP | K = 16
Inception | K = 16
Dogs (118) Domestic cats (5) Birds (59) Fish (16) Reptiles (35) Insects (27) Primates (20) Road vehicles (20) Furniture (20) Musical instruments (26) Clothing (32) Fruit & vegetables (21) Prepared food (18) 1
4
8
12
16
1
4
Component
8
12
16
1
4
8
Component 0.00
0.25
0.50
12
16
Component 0.75
1.00
Mean component responsibility
Figure 4: Semantic groups across independently fitted mixtures. Rows show mean posterior responsibilities for 13 label-defined groups covering 417 ImageNet classes, using 200 images per class. We average rck equally over the classes in each group; parentheses give the number of classes. Component indices are local to each GM, not aligned across panels.
32
Preprint. Under review.
G
S CORE R IDGE S ETTINGS
We vary the score ridge λ with K = 1, training pMF-B for 10 epochs with MAE, SigLIP, or Inception alone. This preliminary sweep initializes the generated statistics with 5,000 samples, retains a prior of the same size, and uses initial gradient-norm calibration. The aligned comparison in Table 1 instead uses 50,000 initialization samples, no prior, and fixed loss scaling. Table 13 reports the sweep results; their absolute scores are not directly comparable to the aligned runs. The lowest FDr6 is obtained at λ = 0.001 for MAE and λ = 0.03 for both SigLIP and Inception, reaching 5.769, 5.742, and 10.720, respectively. A larger ridge value does not consistently improve the score. The setting with the lowest discrepancy in the training representation can also differ from the one with the lowest FDr6 : for Inception, FID is lowest at λ = 0.01, whereas FDr6 is lowest at 0.03. Table 13: Single-encoder ridge selection. pMF-B trained with MGFlow-KL, K = 1, for 10 epochs. All metrics are lower-is-better. Bold marks the lowest FDr6 for each training encoder. λ
FDr6
FID
FDMAE
FDSigLIP
MAE
−5
10 10−4 3 × 10−4 5 × 10−4 10−3 2 × 10−3 3 × 10−3 10−2
5.867 5.821 5.800 5.837 5.769 5.881 5.902 6.312
2.559 2.749 2.867 2.966 3.207 3.504 3.876 5.288
0.091 0.089 0.086 0.077 0.076 0.072 0.067 0.064
8.470 8.326 8.260 8.353 8.107 8.274 8.380 8.780
SigLIP
10−5 10−4 10−3 3 × 10−3 5 × 10−3 10−2 3 × 10−2 5 × 10−2
6.497 6.110 6.010 6.103 5.975 5.970 5.742 5.831
6.116 6.535 6.487 7.311 7.215 7.000 7.338 7.605
0.525 0.414 0.410 0.416 0.400 0.408 0.366 0.368
3.918 3.834 3.730 3.707 3.588 3.576 3.380 3.406
Inception
10−5 10−4 3 × 10−4 10−3 2 × 10−3 3 × 10−3 5 × 10−3 10−2 3 × 10−2 5 × 10−2
11.109 10.966 10.957 11.064 10.874 10.828 10.995 10.898 10.720 10.836
1.153 1.072 1.087 1.139 1.087 1.096 1.105 1.057 1.104 1.105
0.353 0.366 0.355 0.351 0.351 0.352 0.360 0.359 0.367 0.354
16.464 16.070 16.118 16.496 16.047 15.793 16.211 15.727 15.528 15.731
Encoder
Ridge configuration. For ImageNet, the K = 1 branches use λ1 = (0.03, 0.03, 0.001) in SigLIP/ Inception/MAE order. These values lie at the (56.34%, 62.26%, 73.24%) percentiles of the corresponding ImageNet K = 1 reference covariance spectra. For text-to-image generation, we select the eigenvalues at these same percentiles in the T2I K = 1 reference spectra and use nearby rounded values for λ1 (Table 14). Joint matching uses the full image–text covariance; image-only matching uses the image covariance. We use λ4 = 3λ1 in both tasks and λ16 = 9λ1 for ImageNet, with the same ridge added to the reference and generated covariances.
33
Preprint. Under review.
Table 14: T2I covariance eigenvalue quantiles and score ridges. The K = 1 references use 301,706 COCO and GenEval images. Matched denotes the eigenvalue at the encoder-specific ImageNet percentile; λ1 is the value used in training. 25th
50th
75th
Matched
λ1
Joint image–text SigLIP 1.10 × 10−4 Inception 2.16 × 10−3 MAE 1.06 × 10−5
1.90 × 10−3 5.60 × 10−3 5.36 × 10−5
1.71 × 10−2 2.10 × 10−2 2.14 × 10−4
3.22 × 10−3 9.44 × 10−3 1.91 × 10−4
0.003 0.01 0.0002
Image-only SigLIP Inception MAE
1.33 × 10−2 1.00 × 10−2 1.51 × 10−4
6.88 × 10−2 4.30 × 10−2 5.42 × 10−4
1.97 × 10−2 1.87 × 10−2 4.83 × 10−4
0.02 0.02 0.0005
Encoder
H
6.90 × 10−4 4.44 × 10−3 5.35 × 10−5
C OMPONENT M ATCHING
Assignment-ablation settings. All configurations in Table 2 use K = 4, without a K = 1 term. We train for 10 epochs and evaluate with 50k images. Training time is measured on eight H200 GPUs using the protocol in Appendix I.1. The Posterior variant uses the generated mixture’s posterior responsibilities to update component statistics. LP-global instead uses the capacity-constrained assignment in Eq. (10), while retaining the global KL field or full component transport for W2 . LP-paired uses the same LP assignment but replaces these updates with paired component scores or diagonal component transport, respectively.
I
C OMPUTATION AND T RAINING E FFICIENCY
I.1
T RAINING - TIME COMPARISON AND PROTOCOL
Table 15 compares training times for the six JiT/pMF backbones. MGFlow-KL runs faster per training step than the official FD-Loss implementation on all L/H backbones, but slightly slower on the two B backbones under the benchmark settings below. Table 15: Training time per step. Seconds per optimizer update on eight H200 GPUs with a global batch of 1,024. MGFlow-W2 uses K = 1 + 4; MGFlow-KL uses K = 1 + 4 + 16. Training states and precision settings are specified below. Lower is better. Backbone
FD-Loss (Yang et al., 2026a)
MGFlow-W2
MGFlow-KL
JiT-B JiT-L JiT-H
0.799 1.131 2.192
0.969 1.137 1.368
0.846 1.009 1.247
pMF-B pMF-L pMF-H
0.829 1.660 2.572
0.874 1.103 1.703
0.882 1.112 1.882
The training times in Tables 2 and 15 are measured on the same node with eight H200 GPUs and a global batch of 1,024. We use 128 samples per GPU unless memory requires gradient accumulation, as listed in Table 16. All assignment/objective ablations use a = 1. We synchronize CUDA and record the slowest rank for each complete optimizer step, including generation, all training encoders, the objective, backward passes, communication, the optimizer, and native metric reduction. Compilation, statistics initialization, checkpointing, evaluation, and benchmark-log I/O are excluded. After at least ten warmup steps, we report the first twenty-step mean with both coefficient of variation and relative drift between its two halves at most 5%. GPUs are not shared with other jobs. 34
Preprint. Under review.
Training state and precision. The assignment/objective ablations, MGFlow-W2 (K = 1+4), and FD-Loss are timed from their pretrained-base initial training states. The MGFlow-KL (K = 1 + 4 + 16) column instead reports plateau measurements from 60-epoch checkpoints, restoring the model, optimizer, and all nine encoder–branch statistics queues. MGFlow uses BF16 neural forwards with its original covariance precision. FD-Loss uses the official implementation1 , retaining FP32 training, enabled TF32, and the original FD-statistics precision, without added autocast. These are end-to-end implementation timings, with the training states and precision settings specified above. Table 16: Gradient accumulation in the timing benchmark. Entries are a; the per-GPU microbatch is 128/a and the global batch remains 1,024. Values above one follow a recorded CUDA OOM at the preceding larger microbatch. Backbone
MGFlow-W2
MGFlow-KL
Official FD-Loss
JiT-B JiT-L JiT-H pMF-B pMF-L pMF-H
1 1 1 1 1 2
1 1 1 1 1 2
1 1 2 1 2 4
Full-batch FD-Loss with accumulation. At a = 1, we use the upstream training step. For a > 1, we evaluate one FD objective over the full global batch, then replay microbatches to backpropagate its feature gradients, as in Appendix C.1. Each global batch performs one optimizer update and one statistics update; replay overhead is included in the measured time. Why use plateau KL measurements? LP solve time can change during training. On JiT-L, step time falls from 1.805 seconds at initialization to 1.009 seconds after 60 epochs, with CPU LP time falling from 1.031 to 0.241 seconds. The corresponding pMF-B/L times change little: 0.890/1.115 seconds initially and 0.882/1.112 seconds after 60 epochs. I.2
S CALING THE NUMBER OF COMPONENTS
Computational cost. Consider one encoder with feature dimension d, batch size B, and K fullcovariance components. Both variants form the Gaussian assignment costs and update component moments in O(BKd2 ) time. The component moment state requires O(Kd2 ) storage per encoder. The sample LP has BK variables and B + K − 1 independent equality constraints; denote its solve time by TLP (B, K). MGFlow-W2 evaluates K paired Gaussian costs. Each requires a dense spectral decomposition for the matrix-square-root trace, giving O(Kd3 ) work. Evaluating all component pairs instead would require O(K 2 d3 ) work before solving a K ×K component transport LP (Delon and Desolneux, 2020). Fixed pairing removes this component-level LP, but retains the sample-assignment LP. MGFlow-KL uses Cholesky factorizations in O(Kd3 ) time and evaluates the Gaussian scores in O(BKd2 ) time. Global and paired KL scores have the same leading dense cost: pairing reuses R∗ instead of computing two mixture posteriors, but still evaluates K component scores per feature. Excluding the shared neural-network computation, both paired variants therefore have cost O(BKd2 + Kd3 ) + TLP (B, K).
(57)
For joint resolutions, these costs sum over K ∈ K, while generated images and encoder features are shared. Why KL is cheaper in practice. The cubic terms involve different matrix operations. For KL, we factor the covariance used in score evaluation as Cq,k = Lq,k L⊤ q,k and compute each score by two triangular solves: Lq,k u = µq,k − z, L⊤ (58) q,k sq,k (z) = u. 1
https://github.com/Jiawei-Yang/FD-Loss, commit 5c03b8112fec.
35
Preprint. Under review.
Each generated-component factor is reused across the batch, and reference factors are cached. Cholesky factorization has a smaller computational constant than dense spectral decomposition; the subsequent solves need neither eigenvectors nor an explicit covariance inverse. 1/2
1/2
For W2 , even with the reference square root cached, each update forms Ck = Σp,k Σq,k Σp,k through 1/2 two dense matrix multiplications and computes tr(Ck ) from its eigenvalues. Gradients then pass
through this spectral term, the matrix products, and the current batch’s contribution to the covariance. KL instead computes the entire field with gradients stopped. Its backward pass differentiates only the squared regression loss with respect to the features, followed by the shared encoder and generator backward passes. KL therefore saves the spectral computation, its backward pass, and the memory needed to differentiate the current-batch covariance estimate. The practical advantage lies in these cheaper matrix operations and the shorter backward path, rather than a lower asymptotic order. Larger Wasserstein objectives. We extend the timing protocol in Appendix I.1 to MGFlow-W2 with K = 16 and K = 1 + 4 + 16 on another eight-H200 node, using SIM features and a global batch size of 1,024. Runs start from pretrained JiT-H and pMF-H weights. JiT-H uses a per-GPU batch of 128 without accumulation; pMF-H uses 64 with two accumulation steps after an OOM at 128, matching the accepted geometry in Table 16. Table 17 compares these measurements with Table 15. Relative to K = 1 + 4, the two larger Wasserstein objectives take 1.68 × /1.95× as long per step on JiT-H and 1.26 × /1.40× on pMF-H. This cost motivates our default K = 1 + 4 for MGFlow-W2 . MGFlow-KL uses K = 1 + 4 + 16 at 1.247 and 1.882 seconds per step on the two backbones in the plateau benchmark above. Table 17: Training cost of larger mixtures for MGFlow-W2 . Seconds per optimizer update on eight H200 GPUs with SIM features and global batch size 1,024. All runs start from pretrained weights. The K = 1 + 4 result is from Table 15. Backbone
K =1+4
K = 16
K = 1 + 4 + 16
JiT-H pMF-H
1.368 1.703
2.294 2.146
2.661 2.384
Sample support for component statistics. Increasing K also divides the statistics budget among more components. Under the LP constraints, component k receives total assignment mass Bπk per batch and N0 πk over the N0 initialization samples. We fit K = 64 full-covariance GMs to N = 1,281,167 real features for each encoder with seeds 3407–3410. All twelve EM runs reach the parameter-change stopping criterion. No Inception or MAE fit satisfies πmax /πmin ≤ 10, whereas all four SigLIP fits do (Table 18). Inception is the most restrictive: its smallest component has only 978–1,358 units of reference responsibility mass in d = 2,048 dimensions. With N0 = 50,000 and B = 1,024, the corresponding LP budgets are only 38–53 at initialization and 0.78–1.09 per batch. These small components provide weak support for estimating a full covariance. EMA accumulates information over time, but extending the averaging window also slows adaptation to the changing generator. This samplesupport limitation motivates retaining K = 16 as the largest branch in MGFlow-KL rather than using K = 64 across the three encoders. Table 18: Component balance and sample support at K = 64. Ranges across seeds 3407–3410. N πmin is the smallest component’s reference responsibility mass; N0 πmin and Bπmin are its LP allocation budgets for 50,000-sample initialization and a batch of 1,024. Passed denotes the number of fits with πmax /πmin ≤ 10. Encoder
πmax /πmin
N πmin
N0 πmin
Bπmin
Passed
Inception MAE SigLIP
50.58–75.54 15.44–16.66 3.02–5.64
978–1,358 2,766–3,112 7,767–10,146
38.2–53.0 107.9–121.5 303.1–396.0
0.78–1.09 2.21–2.49 6.21–8.11
0/4 0/4 4/4
36
Preprint. Under review.
Component count for text-to-image generation. Our text-to-image reference set contains only 301,706 images. At K = 16, the Inception fits have weight ratios πmax /πmin = 79.61 for joint matching and 72.31 for image-only matching. Their smallest components have reference responsibility masses N πmin of only 699 and 683 in 3,200- and 2,048-dimensional feature spaces, respectively. These small components provide insufficient sample support for stable full-covariance estimation. At K = 4, both weight ratios are below 6.2, and the smallest components each have over 25,000 units of reference responsibility mass. We therefore use K = 1 + 4 for both text-to-image variants.
J
R ESULTS IN I NDIVIDUAL E VALUATION R EPRESENTATIONS
Table 19 separates the training and held-out terms of FDr6 for the configurations in Table 4. Each value is the encoder’s Fréchet distance divided by its real-validation baseline. Table 19: Evaluation in training and held-out representations. Normalized Fréchet distances after 100 epochs with SIM encoders; lower is better. MGFlow-W2 uses K = 1 + 4, and MGFlowKL uses K = 1+4+16. FD-Loss results (Yang et al., 2026a) are supplemented with 50,000-sample evaluations of the released SIM checkpoints for JiT-B/L and pMF-B. Training encoders Backbone
Method
JiT-B
FD-Loss MGFlow-W2 MGFlow-KL FD-Loss MGFlow-W2 MGFlow-KL FD-Loss MGFlow-W2 MGFlow-KL FD-Loss MGFlow-W2 MGFlow-KL FD-Loss MGFlow-W2 MGFlow-KL FD-Loss MGFlow-W2 MGFlow-KL
JiT-L
JiT-H
pMF-B
pMF-L
pMF-H
Held-out encoders
Inception
MAE
SigLIP
ConvNeXt
DINOv2
CLIP
0.60 0.87 0.76 0.46 0.63 0.59 0.45 0.59 0.56 0.51 0.98 0.93 0.47 0.71 0.69 0.46 0.65 0.64
4.22 1.63 1.75 0.66 0.58 0.53 0.43 0.40 0.33 1.86 1.78 1.99 0.56 0.67 0.64 0.35 0.45 0.40
3.34 2.05 2.60 1.98 1.24 1.50 1.68 1.08 1.25 5.37 3.36 4.35 3.03 2.00 2.40 2.46 1.72 2.03
1.33 1.44 1.22 0.87 1.00 0.97 0.86 0.80 0.77 0.77 1.19 1.09 0.57 0.81 0.65 0.57 0.78 0.68
5.13 4.97 4.62 2.63 2.63 2.35 2.10 2.05 1.85 4.10 3.64 3.59 2.21 2.03 1.99 1.74 1.68 1.61
18.78 15.65 11.26 12.91 10.00 5.55 10.37 10.39 5.09 8.51 7.24 6.58 5.68 4.58 4.08 5.77 3.70 3.35
For pMF-L/H, MGFlow-KL improves DINOv2 and CLIP over FD-Loss, while FD-Loss retains lower ConvNeXt distance.
37
Preprint. Under review.
K
Q UALITATIVE I MAGE N ET R ESULTS JiT-H + FD-Loss
JiT-H + MGFlow-W2
JiT-H + MGFlow-KL
class 0107: jellyfish
class 0145: king penguin
class 0001: goldfish
class 0319: dragonfly
class 0309: bee
class 0291: lion
Figure 5: Uncurated ImageNet 256 × 256 samples from JiT-H. FD-Loss (Yang et al., 2026a) (left), MGFlow-W2 (middle), and MGFlow-KL (right), all with one NFE and identical input noise at corresponding positions.
38
Preprint. Under review.
JiT-H + FD-Loss
JiT-H + MGFlow-W2
JiT-H + MGFlow-KL
class 0105: koala
class 0396: lionfish
class 0084: peacock
class 0293: cheetah
class 0323: monarch
class 0089: sulphur-crested cockatoo
Figure 6: Additional uncurated ImageNet 256 × 256 samples from JiT-H. FD-Loss (left), MGFlow-W2 (middle), and MGFlow-KL (right), all with one NFE. Corresponding images use identical input noise.
39
Preprint. Under review.
pMF-H + FD-Loss
pMF-H + MGFlow-W2
pMF-H + MGFlow-KL
class 0937: broccoli
class 0954: banana
class 0852: tennis ball
class 0483: castle
class 0404: airliner
class 0402: acoustic guitar
Figure 7: Uncurated ImageNet 256 × 256 samples from pMF-H. FD-Loss (Yang et al., 2026a) (left), MGFlow-W2 (middle), and MGFlow-KL (right), all with one NFE and identical input noise at corresponding positions.
40
Preprint. Under review.
pMF-H + FD-Loss
pMF-H + MGFlow-W2
pMF-H + MGFlow-KL
class 0945: bell pepper
class 0953: pineapple
class 0820: steam locomotive
class 0967: espresso
class 0933: cheeseburger
class 0779: school bus
Figure 8: Additional uncurated ImageNet 256 × 256 samples from pMF-H. FD-Loss (left), MGFlow-W2 (middle), and MGFlow-KL (right), all with one NFE. Corresponding images use identical input noise.
41
Preprint. Under review.
L
A DDITIONAL T EXT- TO -I MAGE E XAMPLES
Prompt. Giant glowing letters spelling “Nina Lu” rising above a neon-lit futuristic city skyline at night, low-angle wide shot, rain-slicked streets reflecting pink and blue light, cinematic sci-fi digital art.
Prompt. A beautiful young woman with long dark hair stands in a sunlit garden, soft golden light on her face, delicate flowers around her, shallow depth of field, elegant portrait photography.
Prompt. A fluffy cream-colored alpaca standing in a sunny green meadow, soft morning light, gentle expression, shallow depth of field, pastel color palette, whimsical children’s book illustration style.
Prompt. A red sports car drifts through a rain-soaked stadium beneath a blazing and neat MGFlow neon sign; flying spray, wet reflections, low-angle sports photograph
Prompt. Claymation fashion portrait, woman in purple beret and orange ruffled collar posing beside a solid-colored awning, terracotta storefront, morning light casting soft shadows, medium close-up, stop-motion clay texture.
Prompt. A cute plush puppy with two large floppy ears, sitting on a clean table, seen from a side angle, soft daylight, shallow depth of field, cozy still-life photograph.
Prompt. A hidden valley revealed through a rocky archway, lush green terraces, mist drifting between cliffs, a still turquoise pool reflecting soft dawn light, wide landscape view, painterly digital art.
Prompt. Fashion portrait in watercolor: a Somali model in an indigo headwrap and ochre linen jacket poses against a sunlit mud-brick wall, three-quarter view, soft dry-brush washes, warm dusty palette, loose paper texture.
Prompt. Wide cinematic still of a towering sculpture of weathered copper and blue resin standing in a misty deserted plaza at soft twilight, camera pulled far back to show the full statue small within the empty square, calm deep blue tones.
Figure 9: One-step text-to-image samples. Selected 512 × 512 images from FLUX.2 [klein] 4B post-trained with MGFlow for 1,000 steps.
42