λ-C ONTROLLED GRPO: T URNING F LOW-M ATCHING R ATIO I NSTABILITY INTO A B UDGETED R ESOURCE Parivesh Priye ∗ Rivian and Volkswagen Group Technologies
Yufeng Wang Stony Brook University
arXiv:2609.22041v1 [cs.LG] 18 Sep 2026
Meeshawn Marathe Rivian and Volkswagen Group Technologies
Ramit Pahwa Rivian and Volkswagen Group Technologies
A BSTRACT Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable in a way specific to multi-step denoising: the policy update changes systematically across denoising steps, with importance ratios drifting below one, becoming increasingly dispersed, clipping at different rates, and leaving fewer usable samples late in training. Prior work treats these effects as separate failure modes and addresses each with a handtuned stabilizer. We show instead that they arise from a single per-step quantity, which we call path variance. This quantity is determined exactly by the sampler’s Gaussian transition kernel and can be estimated cheaply during training. This reframes instability as a resource that can be measured and budgeted rather than a collection of symptoms to repair. Our method, λ-Controlled GRPO, calibrates importance-ratio behavior from this predicted law rather than from noisy empirical statistics, and allocates gradient effort across denoising steps according to their predicted cost. The two scales governing the update are fixed by standard policy choices rather than introduced as free tuning parameters. On a text-toimage model under two reward settings, rendering difficult target text scored by optical character recognition and matching human preferences scored by a preference model, λ-Controlled GRPO improves both text accuracy and preference reward over the strongest empirical stabilizer. It also keeps late-step path variance within its intended budget, precisely where the baseline systematically overshoots. The result is a Flow-GRPO update calibrated by its own transition law rather than stabilized after instability appears.
1
I NTRODUCTION
Reinforcement learning has become a standard tool for aligning generative models with reward signals, and has recently extended from language models to the flow-matching and rectified-flow models used for image generation. The key step, Flow-GRPO, converts a deterministic flow-matching sampler into a stochastic policy whose denoising trajectory has a tractable per-step likelihood, enabling policy-gradient updates from the group relative policy optimization (GRPO) family. This makes online reward optimization possible for image generators, but introduces a stability problem specific to multi-step denoising: the policy importance ratio, which measures how much more or less likely the new policy is to take a sampled step than the old one, behaves very differently across denoising timesteps. Recent work reports that ratios shift below one at some steps, become highly dispersed at others, clip asymmetrically in the surrogate objective, and eventually leave few effective samples late in training. These effects have largely been treated as separate empirical symptoms and addressed with hand-designed stabilizers, such as timestep-wise normalization of observed ratios and manually chosen gradient reweighting. What is missing is a common explanation for why these pathologies emerge and which quantity an update should control. We identify such a quantity ∗
Corresponding Author.
1
directly from the finite-grid Gaussian transition kernel used by the sampler. A single per-step scalar, which we call the path variance, 2
λk = σk−1 (bθ − bold )(Xk , tk ) ∆tk , determines the conditional mean and variance of the step log-ratio, and therefore predicts its drift, dispersion, left shift, clipping imbalance, and loss of usable samples within one law. The same scalar is not only diagnostic. It defines the natural cost of a policy update at each denoising step and therefore provides a principled quantity for allocating gradient effort across the trajectory. A practical subtlety arises because real implementations typically store a mean-reduced logprobability rather than the full coordinate sum. This reduction splits the path-variance effect into two related quantities: a centering term that governs the mean shift of the log-ratio and a variance term that governs its spread. Keeping these roles separate is essential for the theoretical prediction to match the statistics measured during training. This paper makes four contributions. First, we derive the finite-grid importance-ratio law and its mean-reduced corollary, obtaining an exact per-step prediction for the log-ratio statistics from the path variance. Second, we give a cheap online estimator of this quantity using values already produced by the sampler. Third, we turn the prediction into an algorithm, λ-Controlled GRPO, which replaces empirical ratio normalization with analytic calibration and allocates gradient effort according to predicted path-variance cost, with both governing scales determined by standard policy choices rather than introduced as free hyperparameters. Fourth, across two reward regimes for a textto-image model, we show that the analytic update improves task quality over the strongest empirical stabilizer while keeping late-step path-variance spend near its intended budget. Throughout, we first verify that the estimated scalar predicts the measured log-ratio statistics before modifying the optimizer, so the resulting intervention is grounded in a confirmed transition-law prediction rather than an assumed training heuristic.
2
R ELATED WORK
Reinforcement learning for flow-matching models. Reinforcement learning has become a standard tool for aligning generative models with rewards, beginning in language with reward models learned from human preferences (Christiano et al., 2017; Ouyang et al., 2022) and policy or preference optimization methods including trust-region and proximal policy optimization (Schulman et al., 2015; 2017), group relative policy optimization (GRPO) (Shao et al., 2024; Guo et al., 2025), and direct preference optimization (Rafailov et al., 2023). The same principle was extended to image diffusion models by treating denoising as a sequential decision process optimized with policy gradients (Black et al., 2024; Fan et al., 2023), differentiating through the sampler under differentiable rewards (Clark et al., 2024; Prabhudesai et al., 2023; Eyring et al., 2024), and adapting preference optimization to diffusion models (Wallace et al., 2024). Flow-GRPO (Liu et al., 2025) and DanceGRPO (Xue et al., 2025) bring GRPO to flow-matching image generators by converting the deterministic sampler into a stochastic policy through an ODE-to-SDE construction with Gaussian transition likelihoods and online updates, making sampled denoising trajectories admit tractable per-step importance ratios. This is the policy class studied here, and it raises a central question: how much policy movement can each denoising step absorb before its likelihood ratio becomes unstable? The closest empirical antecedent is GRPO-Guard (Wang et al., 2025a), which reports the same ratio pathologies studied here, including implicit over-optimization, ratios shifted below one, timestep inconsistency, and clipping imbalance, and stabilizes training through empirical ratio normalization and regulated clipping. That work establishes the failure mode, but its correction is not derived from the transition kernel. It recenters measured ratios after instability appears rather than identifying a quantity that predicts their mean shift and variance before the update, or determines how much gradient a timestep should receive per unit cost. Our finite-grid analysis (Section 3.2) supplies that quantity: conditioned on the current state, the step log-ratio is Gaussian with mean −λk /2 and variance λk . A single λk therefore determines the typical left shift, ratio dispersion, timestep inconsistency, and clipping imbalance together. Empirical normalization and gradient reweighting can thus be viewed as symptom-level corrections to a quantity that is predictable from the transition law itself. Timestep, sampling, and surrogate-ratio variants. A rapidly growing literature improves FlowGRPO from complementary directions. Several methods modify where, when, or how denoising 2
trajectories are sampled or weighted. TempFlow-GRPO (He et al., 2025) emphasizes timestep allocation, DenseGRPO (Deng et al., 2026) densifies reward feedback across denoising steps, SmartGRPO (Yu et al., 2025) searches over noise perturbations, BranchGRPO (Li et al., 2025b) introduces structured branching, MixGRPO (Li et al., 2025a) changes the ODE-to-SDE mixture, CPS (Wang & Yu, 2025) studies coefficient-preserving sampling, Pref-GRPO (Wang et al., 2025b) replaces pointwise rewards with pairwise preferences to improve optimization stability, OP-GRPO (Zhang et al., 2026) targets replay and off-policy efficiency, and Flow-Factory (Ping et al., 2026) organizes the broader design space. These works demonstrate that sampling and credit assignment strongly affect optimization, but none derives a per-step likelihood-ratio resource from the Gaussian transition kernel and uses that quantity to calibrate both ratio statistics and gradient allocation. A λ budget is therefore largely complementary to these approaches. FPO (McAllister et al., 2025) takes a different route by replacing exact likelihood computation with a flow-matching surrogate ratio, which is useful when exact transition likelihoods are unavailable. In Flow-GRPO, however, the Euler– Maruyama Gaussian transition is explicit, allowing us to analyze the exact finite-grid ratio and use its path-variance scalar as a predictive optimization budget. Prior work therefore provides the policy class, empirical stabilizers, sampling strategies, and surrogate objectives; our contribution is to connect these phenomena through a single analytically tractable quantity that both predicts per-step ratio behavior and determines how gradient effort should be allocated.
3
P RELIMINARIES AND METHODS
3.1
P RELIMINARIES
Flow matching. A rectified-flow model (Liu et al., 2023; Lipman et al., 2023; Albergo et al., 2025) transports data to noise along the linear interpolant xt = (1 − t)x0 + tx1 ,
x0 ∼ pdata ,
x1 ∼ N (0, I),
t ∈ [0, 1],
(3.1)
and learns a velocity field vθ (x, t) that predicts x1 − x0 . Deterministic generation then solves the ordinary differential equation dXt = vθ (Xt , t) dt. Rectified flow and closely related diffusion models (Ho et al., 2020; Song et al., 2021) underlie modern text-to-image generators such as SD3 (Esser et al., 2024) and FLUX (Labs et al., 2025), whose transformer backbones follow the diffusiontransformer line (Peebles & Xie, 2023; Ma et al., 2024). Stochastic policy. GRPO requires stochastic trajectories with tractable transition densities. FlowGRPO (Liu et al., 2025) therefore replaces the deterministic ODE with a marginal-preserving stochastic differential equation, dXt = bθ (Xt , t) dt + σt dWt , whose drift bθ is fixed through Fokker–Planck matching to the flow-matching score and whose diffusion σt follows a prescribed schedule, both derived in Section C. Training discretizes this process on a grid 0 = t0 < · · · < tK = 1 with step sizes ∆tk , yielding the Euler–Maruyama Gaussian transition kernel Kθ (xk+1 | xk ) = N xk + bθ (xk , tk )∆tk , σk σk⊤ ∆tk , (3.2) which is the central object of our analysis. GRPO objective. For a prompt c, a group of G trajectories is sampled from the old policy and bi (Shao et al., 2024), with evaluated by a reward model, producing group-relative advantages A rewards standardized within each group. Each sampled transition xik → xik+1 carries the per-step importance ratio Kθ (xik+1 | xik , c) rki (θ) = , (3.3) Kold (xik+1 | xik , c) bi , clip(ri , 1 − ϵ, 1 + ϵ)A bi and the clipped surrogate (Schulman et al., 2017) optimizes min rki A k with clipping radius ϵ. The next subsection characterizes how rk varies across denoising timesteps. 3.2
T HE FINITE - GRID PATH - RATIO LAW
The behavior of rk is determined exactly by the transition kernel under one assumption satisfied by the fixed-diffusion schedules used in Flow-GRPO. Assumption 3.1 (Fixed diffusion during the policy update). At each timestep, the new and old transition kernels share the same diffusion σk σk⊤ and differ only in their drift. 3
A direct Gaussian calculation (Section D) then yields the one-step log-ratio in closed form. Proposition 3.2 (One-step Gaussian ratio identity). Let hk = σk−1 (bθ − bold )(Xk , tk ) and ∆Wk ∼ N (0, ∆tk I). Then ξk := log
Kθ (Xk+1 | Xk ) 2 = ⟨hk , ∆Wk ⟩ − 12 ∥hk ∥ ∆tk . Kold (Xk+1 | Xk )
(3.4)
The squared coefficient defines the quantity that drives the paper, the per-step path variance 2
λk := ∥hk ∥ ∆tk = σk−1 (bθ − bold )(Xk , tk )
2
∆tk ,
(3.5)
which measures the squared policy displacement at timestep k in units of the transition noise, equivalently the local Mahalanobis distance between the new and old drifts. This single scalar determines the complete conditional ratio law. Proposition 3.3 (Conditional lognormal ratio law). Conditioned on Ftk , the one-step log-ratio is Gaussian, ξk | Ftk ∼ N − λ2k , λk , (3.6) so, with rk = exp(ξk ), E[rk | Ftk ] = 1,
median(rk | Ftk ) = e−λk /2 ,
Var(rk | Ftk ) = eλk − 1.
(3.7)
A single λk therefore controls the negative log-ratio drift, the left-shifted typical ratio, the ratio variance, the imbalance between the two PPO clipping tails, and the loss of effective samples. Closedform clipping-tail expressions are given in Section D. Empirical stabilizers from prior work can therefore be viewed as correcting several symptoms of one analytically predictable quantity. One implementation detail is essential. Practical Flow-GRPO code stores a mean-reduced logprobability rather than the full coordinate sum. Let ak,m = (µθ,k,m − µold,k,m )/sk,m denote the normalized mean shift for coordinate m. The mean-reduced log-ratio is D 1 X ak,m ϵk,m − 12 a2k,m , ξ¯k = D m=1
iid
ϵk,m ∼ N (0, 1),
(3.8)
and the reduction separates the path-variance effect into a centering scale and a variance scale, 1 X 2 1 X 2 ak,m , λvar a . (3.9) λcenter = k = k D m D2 m k,m Proposition 3.4 (Reduced-log-prob ratio law). For the mean-reduced log-ratio in Equation (3.8), , Var(ξ¯k ) = λvar conditioned on Ftk , E[ξ¯k ] = − 12 λcenter k k . The two scales differ by the latent dimension. This explains why the predicted mean-shift law remains pronounced while the raw variance curve can appear nearly flat, and why centering and variance must be treated separately in the optimizer. 3.3
λ-C ONTROLLED GRPO
Our method estimates the path variance online and treats it as a resource to budget across denoising timesteps. It has two analytic components, and the constants governing both are determined by standard policy choices rather than tuned as free hyperparameters. Full derivations are given in Sections C and E. Estimating the scalar. The transition means µθ,k = Xk + bθ ∆tk and µold,k are already available during sampling, so the reduction-matched estimators add negligible cost: D i D i X X µθ,k,m − µiold,k,m 2 µθ,k,m − µiold,k,m 2 bcenter = 1 bvar = 1 λ , λ , (3.10) i,k i,k D m=1 sk,m D2 m=1 sk,m √ where sk,m = σk,m ∆tk . Section C gives the endpoint scaling of λk under the Flow-GRPO schedule.
4
LambdaNorm-T: analytic ratio calibration. Because exponentiating the mean-reduced log-ratio does not preserve the mean-one property of the full likelihood ratio, we calibrate it from its predicted law instead of reusing it directly. We center and standardize yi,k = ξ¯i,k using the predicted moments and then rescale to a target variance ν: bcenter yi,k + 1 λ √ i,k zi,k = q 2 (3.11) , rei,k = exp ν zi,k − 12 ν . bvar + ελ λ i,k If zi,k is standard normal, then E[e ri,k ] ≈ 1, so rei,k becomes a calibrated, mean-one surrogate ratio. The scale ν is not introduced as an independent tuning parameter. PPO already specifies a clipping radius ϵ, so we set ν = ϵ2 . Damp-only λ-gradient weighting. We next allocate gradient effort according to predicted pathvariance cost. A timestep is damped only when its predicted center exceeds a budget τ : q h i bcenter + ελ ) , bi , clip(e bi , wλ = min 1, τ /(λ L = − Ei,k wλ min rei,k A ri,k , 1 − ϵ, 1 + ϵ)A k
k
k
(3.12) bcenter is the batch mean of the per-sample estimates in Equation (3.10), and all estimates where λ k are detached so timesteps already within budget are never amplified. The budget is determined by a target retained effective-sample fraction q (Elvira et al., 2018) over the late timesteps Klate . Under log q (Section E). We report realized spend using the the lognormal path-weight heuristic, τ = − |K late | mean late-step path variance X 1 bcenter , Λlate = λ (3.13) k |Klate | k∈Klate
which the budget τ is intended to control. The complete intervention is therefore specified by two interpretable policy choices, the PPO clipping radius ϵ and the retained fraction q. We use ϵ = 0.2 and q = 0.95 for OCR, and the tighter setting ϵ = 0.1 and q = 0.99 for the denser PickScore reward.
4
E XPERIMENTAL FINDINGS
The experiments answer three questions in sequence. The tiny-SD3 diagnostics first test whether the finite-grid law predicts the log-ratio statistics observed during training. The primal-dual experiments then test whether path variance is not only descriptive but also controllable. Finally, the SD3.5 experiments evaluate whether the analytic method, LambdaNorm-T with damp-only λgradient weighting, improves task quality over the empirical stabilizer under matched budgets. The primary held-out comparison is reported in Table 2. 4.1
T INY-SD3 LAW AUDIT
The tiny-SD3 audit tests whether the finite-grid path-variance law appears in a working Flow-GRPO implementation. We train a small public SD3 debug pipeline with a compressibility reward and average every statistic over three independent runs from fresh random initializations (configuration in Section B). Here, log rk denotes the mean-reduced log-ratio ξ¯k from Equation (3.8), which is the bcenter closely quantity stored by the implementation. In the baseline run, the drift-based estimator λ k center b b tracks the negative log-ratio drift, as predicted by the reduced law λk ≈ −2 E[log rk ]. At the bcenter = 0.0054 predicts a drift of −0.0027, exactly matching the final timestep, for example, λ 7 bvar also lies on the correct scale for the observed variance measured −0.0027. The corresponding λ k (Figure 1). This agreement provides the first direct evidence that the derived law predicts statistics measurable during training. We next add λ-z clipping to test whether controlling the predicted quantity also controls the meabcenter sured pathology. Averaged over three seeds and the final 50 updates, it reduces the late-step λ 7 −4 −5 from 0.0054 to 1.0 × 10 and the log-ratio drift from −0.0027 to −4 × 10 , while leaving reward essentially unchanged (Figure 2; full statistics in Table 1). Because this experiment uses a small debug model whose effective sample size remains near one, it is a law audit rather than a performance bcenter and λ bvar act as accurate online predictors of the reduced result. Its purpose is to establish that λ k k log-ratio center and variance, and that directly reducing λ suppresses the corresponding pathology. 5
Figure 1: Baseline diagnostics from the tiny-SD3 law audit. The key comparisons are between bvar and Var(log bcenter and −2E[log b d rk ) under the reduced-log-prob convention. rk ], and between λ λ k k
Figure 2: Baseline versus λ-z clipping in the tiny-SD3 law audit. λ-z clipping sharply reduces the bcenter , λ bvar , log-ratio mean drift, reduced log-ratio variance, and typical-ratio depression. late-step λ k k
6
Metric at timestep 7
Baseline
λ-z clip
bcenter λ 7 bvar λ 7 b E[log r7 ] d Var(log r7 ) Median ratio Per-step ESS / n Path ESS / n Clip fraction Reward average
0.005425 3.31×10−7 −0.002693 7.32×10−7 0.997264 0.9999993 0.999996 0.1169 −0.008114
0.000100 6.12×10−9 −0.000041 8.83×10−9 0.999949 0.99999996 1.000000 0.0730 −0.008238
Table 1: Tiny-SD3 law audit at the final timestep. Statistics are averaged over three seeds and the final 50 updates. λ-z clipping reduces predicted path variance, log-ratio drift and variance, and typical-ratio depression while leaving reward essentially unchanged. 4.2
T INY-SD3 STRESS RESULT: λ IS CONTROLLABLE
We next test whether λ is actionable rather than merely descriptive. As a control ablation, we impose a primal-dual λ budget (Section G). Across three paired tiny-SD3 stress seeds, the controller bcenter improves the compressibility reward on every seed while reducing the average final-step λ 7 from 0.998 to 0.166. This result establishes that the derived scalar can steer optimization. The main method below removes the dual machinery and instead uses the analytic prediction directly through LambdaNorm-T and damp-only weighting. 4.3
SD3.5 HARD -OCR RESULTS
Our main real-model experiment fine-tunes SD3.5 Medium with low-rank adaptation (LoRA) on the hard-OCR task. We use the same OCR dataset and reward definition as the public Flow-GRPO and GRPO-Guard configurations, but with smaller update statistics, namely one GPU, smaller groups, fewer batches per epoch, and shorter training. Checkpoints are saved at fixed training-step intervals and selected using a held-out protocol. A random 256 OCR test prompts are used only for checkpoint selection, while the remaining 762 prompts are reserved for final evaluation. Alongside the FlowGRPO-style OCR reward, we report character error rate (CER), defined as the fraction of characters that must be edited to recover the target, together with exact-match and substring-match rates. The latter require the complete target or a contiguous portion of it to appear in the generated image. PickScore (Kirstain et al., 2023) and CLIP score (Hessel et al., 2021) serve as optional image-quality guardrails. The primal-dual controller from Section G is included only as a baseline, denoted “Dual,” alongside empirical RatioNorm. Under this protocol, empirical RatioNorm selects checkpoint 40, before the late-training pathology emerges, while the analytic method and primal-dual baseline select checkpoint 80. On the held-out 762-prompt test split, LambdaNorm-T with damp-only λ weighting outperforms both baselines on OCR reward, CER, exact match, and substring match (Table 2). Method RatioNorm Dual λ-Controlled GRPO
Ckpt.
OCR ↑
CER ↓
Exact ↑
Substr. ↑
Λlate ↓
40 80 80
0.5632 0.5501 0.5831
0.4368 0.4499 0.4169
12.99% 12.07% 14.44%
19.29% 16.80% 22.83%
0.0 0.000217 0.000332
Table 2: SD3.5 hard-OCR held-out results on 762 test prompts. The analytic LambdaNorm-T plus damp-only λ-gradient method improves over the early-stopped empirical RatioNorm checkpoint by 0.0199 OCR reward, 1.44 exact-match points, and 3.54 substring-match points, while keeping realized late-step spend Λlate far below the target budget τ = − log(0.95)/5 ≈ 0.01026. On the 256-prompt validation split used for checkpoint selection, the same method, LambdaNorm-T with damp-only weighting at ϵ = 0.2 and q = 0.95, improves the stricter transcript-success metrics much more strongly than the mean OCR reward (Table 3). 7
Method RatioNorm Dual λ-Controlled GRPO
Ckpt.
OCR ↑
CER ↓
Exact ↑
Substr. ↑
Λlate ↓
40 80 80
0.5744 0.5643 0.5657
0.4256 0.4357 0.4343
10.94% 8.59% 15.23%
14.84% 14.06% 23.83%
0.0 0.000217 0.000332
Table 3: Random 256-prompt SD3.5 hard-OCR validation comparison. The analytic λ-controlled method substantially improves exact-match and substring-match rates over empirical RatioNorm and the primal-dual ablation while keeping predicted late-step path-variance spend near zero. Relative to RatioNorm, exact match improves by approximately 39% and substring match by approximately 61%. The SD3.5 hard-OCR experiment provides the strongest evidence for the main algorithmic claim. Replacing empirical RatioNorm with the finite-grid Gaussian calibration of LambdaNorm-T and shaping gradient allocation using the predicted λ cost in Equation (3.12) improves the metrics that directly test whether the requested text appears in the image. On the held-out 762-prompt test split, λ-Controlled GRPO gains +0.0199 OCR reward, +1.44 exact-match points, and +3.54 substringmatch points over RatioNorm. On the 256-prompt validation split, the strict-metric gains are larger still, approximately 39% for exact match and 61% for substring match. The law audit already bcenter predicts reduced log-ratio drift. Here, using that same prediction to calibrate established that λ k ratios and allocate gradient effort translates into improved task performance, while the primal-dual runs serve only as a controllability ablation. 4.4
SD3.5 P ICK S CORE RESULTS
To test whether the benefit is specific to OCR, we repeat the protocol with PickScore (Kirstain et al., 2023), a human-preference reward model in the same broad family as ImageReward (Xu et al., 2023) and HPSv2 (Wu et al., 2023). Because this reward is denser, we use the tighter operating point ϵ = 0.1 and q = 0.99, giving ν = 0.01 and τ ≈ 2.0 × 10−3 . These values are fixed before training. We compare against empirical RatioNorm on the identical training pipeline. Every saved checkpoint at steps 40, 60, and 80 is evaluated on a random 256-prompt validation split. Each method’s best checkpoint is selected by mean PickScore and then evaluated once on the disjoint 762-prompt test complement. Method RatioNorm λ-Controlled GRPO
Ckpt.
Mean PickScore ↑
Std.
Prompt-level comparison
40 80
0.8283 0.8358
0.057 0.056
— wins 467/762 (61.3%)
Paired ∆ (Ours − RatioNorm): +0.0075, 95% bootstrap CI [+0.0059, +0.0092], t ≈ 9.1
Table 4: Disjoint 762-prompt SD3.5 PickScore test split. Representative checkpoints are selected using Table 6 and evaluated once on the held-out complement. The paired 95% bootstrap confidence interval excludes zero, and λ-Controlled GRPO beats RatioNorm on 61.3% of prompts with no ties. RatioNorm selects checkpoint 40 on validation; its later checkpoints perform worse and are therefore not evaluated on the test split. The PickScore experiment reproduces the same optimization mechanism observed in OCR. Empiribcenter rises silently to approxical RatioNorm has no explicit path-variance budget, so its late-step λ mately 3.4τ by step 80, while mean PickScore falls from 0.8325 at step 40 to 0.8126 at step 80. By contrast, λ-Controlled GRPO keeps late-step path variance near τ and improves monotonically over training. On the held-out test split, the paired bootstrap separates the +0.0075 gain from zero across 762 prompts, and the analytic method wins on a clear majority of prompt-level comparisons. 4.5
G ENERALIZATION TO A SECOND BACKBONE : FLUX
To test whether the analytic method transfers to another flow-matching backbone, we additionally use a broader campaign which trains both SD3.5 and FLUX.1-dev on three tasks, PickScore, hardOCR, and the GenEval compositional-generation benchmark (Ghosh et al., 2023), and compares 8
λ-Controlled GRPO against empirical RatioNorm and an additional untuned Flow-GRPO baseline. The held-out FLUX results are summarized in Table 5, and matched held-out GenEval image comparisons for the SD3.5 runs of this campaign are shown in Section J. Primary metric ↑
Training clip fraction ↓
FLUX task
Flow-GRPO
RatioNorm
λ-Ctrl
Flow-GRPO
RatioNorm
λ-Ctrl
PickScore OCR reward GenEval
0.8430 0.6434 0.6545
0.8383 0.6453 0.6569
0.8429 0.6606 0.6538
0.961 0.965 0.990
0.888 0.964 0.915
0.089 0.060 0.100
Table 5: Held-out FLUX.1-dev results from the broader single-seed 8-GPU A6000 reproduction campaign. We compare λ-Controlled GRPO with untuned Flow-GRPO and empirical RatioNorm on PickScore (762 paired prompts), hard-OCR (762 prompts), and GenEval (2,148 images). The primary metric is mean PickScore, mean OCR reward, or mean GenEval composite score. Training clip fraction is averaged over the 50 updates preceding the selected checkpoint. The analytic method improves over RatioNorm on PickScore and OCR and reduces clipping by approximately an order of magnitude on all three tasks; on GenEval, the primary metrics are statistically indistinguishable. The FLUX experiments extend the SD3.5 findings to a second backbone, although the tasklevel improvements are not uniform. On FLUX PickScore, λ-Controlled GRPO essentially ties the untuned Flow-GRPO baseline, with a difference of −0.0001 and a 95% confidence interval of [−0.0013, +0.0011], while outperforming empirical RatioNorm by +0.0046, with interval [+0.0036, +0.0057] and a 63.4% prompt-level win rate. On FLUX hard-OCR, the analytic method achieves the best observed reward and strict metrics, but the confidence intervals cross zero under this single-seed evaluation, so the improvement remains directional rather than resolved. On GenEval, the three methods are statistically indistinguishable. The most stable finding across backbones is mechanistic. On all three FLUX tasks, the analytic approach lowers the required training clip fraction by about one order of magnitude, mirroring the stabilization pattern previously seen on SD3.5. This value represents the calibrated surrogate ratio that our method optimizes, not a directly comparable raw ratio computed identically across algorithms. Accordingly, we view FLUX as directional support for cross-backbone generalization, most clearly for the predicted clipping behavior rather than as proof of consistent task-level superiority.
5
C ONCLUSION AND LIMITATIONS
In this paper we asked whether the timestep-dependent ratio instability observed in Flow-GRPO has a single underlying cause that can be controlled rather than merely patched. It does. The finite-grid transition kernel shows that negative log-ratio drift, timestep-varying variance, clipping imbalance, and the loss of usable samples are all governed by one per-step scalar, the path variance λk , which fixes the conditional mean and variance of the step log-ratio exactly. Empirical stabilizers from prior work can therefore be interpreted as symptom-level corrections to this quantity. The more direct intervention is to control the cause itself: calibrate ratios from the predicted law and allocate gradient effort according to predicted path-variance cost, with the governing scales set by the PPO clipping radius and a target sample-retention level rather than introduced as free tuning parameters. The central insight is that path variance behaves as a resource already measured by the sampler. It can be estimated online, budgeted across timesteps, and spent through the optimizer, turning FlowGRPO from an empirically stabilized procedure into one calibrated by its own transition law. These conclusions hold within a defined operating range. The theory is cleanest when the compared policies share the same diffusion coefficient; policy-dependent diffusion, reversed-time conventions, or modified diffusion schedules require a corresponding finite-grid analysis. The Gaussian theory also predicts the ideal transition law, but does not guarantee that its estimated path variance will exactly match realized statistics once approximate scores, finite discretization, and implementation details enter. This is why we verify the predicted law empirically before using it to control the optimizer. The evidence is similarly bounded. The primary SD3.5 comparisons use a single-GPU pilot protocol, while a broader single-seed campaign shows that task-level gains are not uniform: they are strongest for human-preference optimization, favorable but unresolved for text rendering, 9
and neutral on GenEval. These boundaries define the regime in which the present conclusions should be read. Within that regime, path variance is directly measured and explicitly budgeted, providing a concrete account of when and why the resulting Flow-GRPO updates remain controlled. R EPRODUCIBILITY STATEMENT The theoretical claims are stated with their assumptions in Section 3.2 and proved in full in Section D. The estimator of the path-variance scalar and the two policy-derived scales that fix the method are specified in Section 3.3, and the exact experimental configuration, including the models, timestep counts, optimization schedules, reward definitions, and the per-regime operating points, is collected in Section B. Code to reproduce the training and evaluation is provided as supplementary material. E THICS STATEMENT This work studies the optimization stability of reinforcement learning for text-to-image models and does not involve human subjects or the release of new data. The image generators and reward models used are existing public artifacts, and the reward-optimization procedure inherits their known limitations, including the potential to amplify biases present in the underlying models. We see no additional ethical concerns specific to the analysis presented here. AI USE STATEMENT A large language model assisted the authors in writing and polishing the language of the manuscript; the authors verified that all claims, proofs, mathematical formulations, and reported values are valid, checked them against the implementation and results, and take full responsibility for the manuscript.
R EFERENCES Michael S. Albergo, Nicholas M. Boffi, and Eric Vanden-Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. Journal of Machine Learning Research (JMLR), 26, 2025. arXiv:2303.08797. Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine. Training diffusion models with reinforcement learning. In International Conference on Learning Representations (ICLR), 2024. Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems (NeurIPS), 2017. Kevin Clark, Paul Vicol, Kevin Swersky, and David J. Fleet. Directly fine-tuning diffusion models on differentiable rewards. In International Conference on Learning Representations (ICLR), 2024. Haoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang, Yu Liu, Changxin Gao, and Nong Sang. DenseGRPO: From sparse to dense reward for flow matching model alignment. arXiv preprint arXiv:2601.20218, 2026. Víctor Elvira, Luca Martino, and Christian P. Robert. Rethinking the effective sample size. International Statistical Review, 90(3):525–550, 2018. doi: 10.1111/insr.12500. arXiv:1809.04129. Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In International Conference on Machine Learning (ICML), 2024. Luca Eyring, Shyamgopal Karthik, Karsten Roth, Alexey Dosovitskiy, and Zeynep Akata. ReNO: Enhancing one-step text-to-image models through reward-based noise optimization. In Advances in Neural Information Processing Systems (NeurIPS), 2024. 10
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, et al. DPOK: Reinforcement learning for fine-tuning text-to-image diffusion models. In Advances in Neural Information Processing Systems (NeurIPS), 2023. Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text-to-image alignment. Advances in Neural Information Processing Systems, 36: 52132–52152, 2023. Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, et al. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. doi: 10.1038/s41586-025-09422-z. Xiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li, Jian Yang, Dacheng Yin, Fengyun Rao, and Bo Zhang. TempFlow-GRPO: When timing matters for GRPO in flow models. arXiv preprint arXiv:2508.04324, 2025. Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: A reference-free evaluation metric for image captioning. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021. doi: 10.18653/v1/2021.emnlp-main.595. Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), 2020. Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Picka-pic: An open dataset of user preferences for text-to-image generation. In Advances in Neural Information Processing Systems (NeurIPS), 2023. Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, Sumith Kulal, Kyle Lacey, Yam Levi, Cheng Li, Dominik Lorenz, Jonas Müller, Dustin Podell, Robin Rombach, Harry Saini, Axel Sauer, and Luke Smith. Flux.1 kontext: Flow matching for in-context image generation and editing in latent space, 2025. URL https://arxiv.org/abs/2506.15742. Junzhe Li, Yutao Cui, Tao Huang, Yinping Ma, Chun Fan, Yiming Cheng, Miles Yang, Zhao Zhong, and Liefeng Bo. MixGRPO: Unlocking flow-based GRPO efficiency with mixed ODE-SDE. arXiv preprint arXiv:2507.21802, 2025a. Yuming Li, Yikai Wang, Yuying Zhu, Zhongyu Zhao, Ming Lu, Qi She, and Shanghang Zhang. BranchGRPO: Stable and efficient GRPO with structured branching in diffusion models. arXiv preprint arXiv:2509.06040, 2025b. Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. In International Conference on Learning Representations (ICLR), 2023. Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-GRPO: Training flow matching models via online RL. arXiv preprint arXiv:2505.05470, 2025. Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In International Conference on Learning Representations (ICLR), 2023. Nanye Ma, Mark Goldstein, Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden, and Saining Xie. SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers. In European Conference on Computer Vision (ECCV), 2024. David McAllister, Songwei Ge, Brent Yi, Chung Min Kim, Ethan Weber, Hongsuk Choi, Haiwen Feng, and Angjoo Kanazawa. Flow matching policy gradients. arXiv preprint arXiv:2507.21053, 2025. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, et al. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS), 2022. doi: 10. 52202/068431-2011. 11
William Peebles and Saining Xie. Scalable diffusion models with transformers. In IEEE/CVF International Conference on Computer Vision (ICCV), 2023. doi: 10.1109/ICCV51070.2023. 00387. Bowen Ping, Chengyou Jia, Minnan Luo, Hangwei Qian, and Ivor Tsang. A unified framework for reinforcement learning in flow-matching models. arXiv:2602.12529, 2026.
Flow-Factory: arXiv preprint
Mihir Prabhudesai, Anirudh Goyal, Deepak Pathak, and Katerina Fragkiadaki. Aligning text-toimage diffusion models with reward backpropagation. arXiv preprint arXiv:2310.03739, 2023. Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems (NeurIPS), 2023. doi: 10.52202/ 075280-2338. John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel. Trust region policy optimization. In International Conference on Machine Learning (ICML), 2015. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR), 2021. Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, et al. Diffusion model alignment using direct preference optimization. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. doi: 10.1109/CVPR52733.2024.00786. Feng Wang and Zihao Yu. Coefficients-preserving sampling for reinforcement learning with flow matching. arXiv preprint arXiv:2509.05952, 2025. Jing Wang, Jiajun Liang, Jie Liu, Henglin Liu, Gongye Liu, Jun Zheng, Wanyuan Pang, Ao Ma, Zhenyu Xie, Xintao Wang, Meng Wang, Pengfei Wan, and Xiaodan Liang. GRPO-Guard: Mitigating implicit over-optimization in flow matching via regulated clipping. arXiv preprint arXiv:2510.22319, 2025a. Yibin Wang, Zhimin Li, Yuhang Zang, Yujie Zhou, Jiazi Bu, Chunyu Wang, Qinglin Lu, Cheng Jin, and Jiaqi Wang. Pref-GRPO: Pairwise preference reward-based GRPO for stable text-to-image reinforcement learning. arXiv preprint arXiv:2508.20751, 2025b. Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhao, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-toimage synthesis. arXiv preprint arXiv:2306.09341, 2023. Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. ImageReward: Learning and evaluating human preferences for text-to-image generation. In Advances in Neural Information Processing Systems (NeurIPS), 2023. Zeyue Xue, Jie Wu, Yu Gao, Fangyuan Kong, Lingting Zhu, Mengzhao Chen, Zhiheng Liu, Wei Liu, Qiushan Guo, Weilin Huang, and Ping Luo. DanceGRPO: Unleashing GRPO on visual generation. arXiv preprint arXiv:2505.07818, 2025. Benjamin Yu, Jackie Liu, and Justin Cui. Smart-GRPO: Smartly sampling noise for efficient RL of flow-matching models. arXiv preprint arXiv:2510.02654, 2025. Liyu Zhang, Kehan Li, Tingrui Han, Tao Zhao, Yuxuan Sheng, Shibo He, and Chao Li. OP-GRPO: Efficient off-policy GRPO for flow-matching models. arXiv preprint arXiv:2604.04142, 2026.
12
A
OVERVIEW OF THE APPENDICES
The appendices are grouped by purpose. The first group covers methodology and theory: the experimental configuration, the flow-matching preliminaries and drift, the derivations and proofs behind the finite-grid law, the full derivation of the analytic ratio calibration, and the diagnostic and alternative estimators of the path-variance scalar. The next section gives the primal-dual baseline and the diagnostic clipping rule used in the law audit, which appear in the main text only as comparisons. The last three are experimental evidence, presenting the qualitative held-out comparisons for the OCR, PickScore, and GenEval reward regimes.
B
E XPERIMENTAL CONFIGURATION
The tiny SD3 law audit was deliberately small. It used a tiny public SD3 debug pipeline with a JPEG-compressibility reward and K = 8 denoising timesteps, trained for 50 epochs with 32 samples per epoch, 5 inner epochs, and a batch size of 8. The two matched configurations were a baseline diagnostic run and a λ-z-clipping run, and every reported statistic is the average of the last 50 updates over 3 seeds, where a seed is an independent training run from a fresh random initialization. The SD3.5 experiments fine-tuned SD3.5 Medium with low-rank adaptation on a single GPU, using the same OCR and PickScore datasets and reward definitions as the public Flow-GRPO and GRPOGuard configurations, but with smaller groups, fewer batches per epoch, and shorter training than the original multi-GPU regime. The two policy choices that define the method were fixed before training in each reward regime: the PPO step radius ϵ and the effective-sample-size retention target q, from which ν = ϵ2 and τ = − log(q)/|Klate | follow. The OCR runs used ϵ = 0.2 and q = 0.95; the denser PickScore reward used the tighter ϵ = 0.1 and q = 0.99.
C
F LOW- MATCHING PRELIMINARIES AND DRIFT
Section 3 states the marginal-preserving SDE and its Gaussian transition kernel. This appendix gives the drift, the diffusion schedule, and the endpoint behavior of the path-variance scalar. The deterministic flow-matching sampler solves dXt = vθ (Xt , t) dt.
(C.1)
GRPO needs stochastic trajectories with a Gaussian transition density, so Flow-GRPO turns Equation (C.1) into the marginal-preserving SDE dXt = bθ dt + σt dWt . If pt is the ODE marginal density, Fokker–Planck matching gives bθ (x, t) = vθ (x, t) +
σt2 ∇ log pt (x). 2
(C.2)
For the linear interpolant Equation (3.1), the score identity is ∇ log pt (x) = −
x + (1 − t)vt (x) , t
(C.3)
so substituting the estimated velocity yields the working drift bθ (x, t) = vθ (x, t) −
σt2 x + (1 − t)vθ (x, t) . 2t
(C.4)
The minus sign comes from the negative score in Equation (C.3). Flow-GRPO commonly uses the schedule q t t σt = a 1−t , σt2 = a2 1−t , (C.5) which is small near the data endpoint t = 0 and large near the noise endpoint t = 1. 13
C.1
E NDPOINT SCALING OF THE PATH VARIANCE 2
Let ∆(tk ) = E[∥bθ (Xk , tk ) − bold (Xk , tk )∥ ]. For scalar σk , the mean path variance is ∆(tk ) ∆tk , σk2
(C.6)
∆(tk )(1 − tk ) ∆tk . a 2 tk
(C.7)
λ̄k = and with the schedule Equation (C.5), λ̄k =
The low-noise endpoint is therefore problematic when ∆(t)/t does not vanish: if ∆(t) is bounded below near t = 0 then λ̄k diverges as tk → 0, whereas ∆(t) = O(t) keeps the endpoint contribution finite. The governing diagnostic is the variance-normalized drift scale ∆(t)/σt2 .
D
D ERIVATIONS AND PROOFS
This appendix collects the derivations and proofs for the finite-grid path-ratio law of Section 3.2. Proof of Proposition 3.2. The kernels are Gaussian with the same covariance Σk ∆tk = σk σk⊤ ∆tk and means xk + bθ ∆tk and xk + bold ∆tk . The determinant terms in the Gaussian densities cancel. Expanding the difference of the two quadratic forms and using Xk+1 − Xk − bold ∆tk = σk ∆Wk gives 1 −1 2 σ δbk ∆tk , (D.1) ξk = σk−1 δbk , ∆Wk − 2 k which is Equation (3.4). D.1
W HY THE PATH - VARIANCE SCALAR TAKES THIS FORM
Start with the scalar case. Let the old transition be pold = N (µold , s2 ) and the new transition be pθ = N (µθ , s2 ) with the same variance. Write the mean shift in units of the standard deviation as a=
µθ − µold . s
(D.2)
If X ∼ pold , then X = µold + sZ with Z ∼ N (0, 1), and direct substitution into the Gaussian densities gives pθ (X) 1 log = aZ − a2 . (D.3) pold (X) 2 Thus the same number a2 controls both statistics of the log-ratio: pθ (X) 1 2 pθ (X) Eold log =− a , Varold log = a2 . pold (X) 2 pold (X)
(D.4)
This equality is the reason for defining a path-variance scalar: it is the squared policy move measured in noise units. The multivariate case is the same calculation after whitening the noise. Suppose pold = N (µold , Σ) and pθ = N (µθ , Σ) share covariance Σ. Let a = Σ−1/2 (µθ − µold ).
(D.5)
Under pold , write X = µold + Σ1/2 Z with Z ∼ N (0, I). Expanding the two quadratic forms gives log so that
pθ (X) 1 2 = ⟨a, Z⟩ − ∥a∥ , pold (X) 2
pθ (X) 1 2 Eold log = − ∥a∥ , pold (X) 2
Varold 14
pθ (X) log pold (X)
(D.6)
2
= ∥a∥ .
(D.7)
The scalar that appears from the likelihood ratio is therefore the Mahalanobis distance between the two means, 2 ∥a∥ = (µθ − µold )⊤ Σ−1 (µθ − µold ). (D.8) For one Euler–Maruyama Flow-GRPO step, the old and new kernels share covariance Σk ∆tk = σk σk⊤ ∆tk and differ only in mean by µθ,k − µold,k = (bθ − bold )(Xk , tk )∆tk . Substituting these into Equation (D.8) gives 2 (µθ,k − µold,k )⊤ (Σk ∆tk )−1 (µθ,k − µold,k ) = σk−1 bθ − bold (Xk , tk ) ∆tk = λk , (D.9) so the one-step log-ratio has mean −λk /2 and variance λk , and λk is the scalar the Gaussian likelihood ratio itself uses for its drift and variance. Proof of Proposition 3.3. Conditioned on Ftk , hk is fixed and ∆Wk ∼ N (0, ∆tk I). Hence 2 ⟨hk , ∆Wk ⟩ ∼ N (0, ∥hk ∥ ∆tk ) = N (0, λk ). Subtracting λk /2 gives Equation (3.6). Since 2 rk = exp(ξk ), rk is lognormal. If Z ∼ N (µ, s2 ), then E[eZ ] = eµ+s /2 , median(eZ ) = eµ , 2 2 and Var(eZ ) = e2µ+s (es − 1). Plugging µ = −λk /2 and s2 = λk gives Equation (3.7). Proof of Proposition 3.4. Taking the conditional expectation of Equation (3.8) removes the noise P term and leaves −(2D)−1 m a2k,m = −λcenter /2. Since the ϵk,m are independent with unit varik ance, D 1 X 2 a = λvar (D.10) Var(ξ¯k | Ftk ) = 2 k . D m=1 k,m Corollary D.1 (Total variance for reduced log-ratios). For mean-reduced latent log-probabilities, the batch variance combines conditional transition noise with heterogeneity in the conditional cen1 center ters, Var(ξ¯k ) = E[λvar ). k ] + 4 Var(λk Proof of Corollary D.1. Apply the law of total variance to Proposition 3.4: Var(ξ¯k ) = E[Var(ξ¯k | Ftk )] + Var(E[ξ¯k | Ftk ])
(D.11)
center = E[λvar /2). k ] + Var(−λk
(D.12)
The same argument gives the unconditional moments of the full log-ratio, which are useful when comparing against batch statistics that mix many conditional states. Corollary D.2 (Unconditional first and second moments of log-ratios). Let λ̄k = E[λk ]. Then 1 E[ξk ] = − λ̄k , 2
1 Var(ξk ) = λ̄k + Var(λk ) ≥ λ̄k . 4
(D.13)
Proof. The expectation follows by the tower property. For the variance, use the law of total variance, Var(ξk ) = E[Var(ξk | Ftk )] + Var(E[ξk | Ftk ]) = E[λk ] + Var(−λk /2).
E
(D.14)
A NALYTIC RATIO CALIBRATION
This section gives the longer derivation of LambdaNorm-T, the analytic ratio calibration used by the main method in Section 3.3. The question is whether the reduced-log-ratio law can replace empirical RatioNorm itself. We call the construction LambdaNorm-T, where “T” denotes a target log-ratio variance. Centering alone is not enough. If yi,k := log ri,k , 15
(E.1)
then the reduced law gives 1 E[yi,k | Ftk ] = − λcenter , Var(yi,k | Ftk ) = λvar (E.2) i,k . 2 i,k The reduced ratio exp(yi,k ) is therefore a surrogate ratio rather than the full path Radon–Nikodym derivative, and it generally does not preserve the mean-one likelihood-ratio convention: 1 center 1 var E[exp(yi,k ) | Ftk ] = exp − λi,k + λi,k . (E.3) 2 2 The analytically standardized coordinate is therefore zi,k =
yi,k + 12 λcenter i,k q . var λi,k + ε
(E.4)
Under the Gaussian transition approximation, zi,k is approximately standard normal. Using exp(zi,k ) directly would impose an arbitrary and often too-large surrogate ratio scale. LambdaNorm-T instead chooses a target log-ratio variance νk > 0 and reconstructs a calibrated surrogate log-ratio: √ 1 ℓei,k = νk zi,k − νk , rei,k = exp(ℓei,k ). (E.5) 2 If zi,k ∼ N (0, 1), then ℓei,k ∼ N (−νk /2, νk ) and E[e ri,k ] ≈ 1. The subtraction −νk /2 is important: it preserves the mean-one ratio convention after choosing the PPO-scale variance. The target variance νk is tied to the PPO clipping scale rather than tuned as an unconstrained normalization constant. If the clipped objective uses a ratio window [1 − ϵ, 1 + ϵ], the corresponding positive log-ratio boundary is log(1 + ϵ). We therefore use νk = log(1 + ϵ)2 ≈ ϵ2
(E.6)
as the default scale, with small sensitivity sweeps around it. For the standard PPO-scale choice ϵ = 0.2, this gives ν ≈ 0.04, i.e. a target surrogate log-ratio standard deviation of about 0.2. Equivalently, define the centered reduced log-ratio 1 . ℓci,k = yi,k + λcenter 2 i,k Then LambdaNorm-T can be written as 1 2 var ℓei,k = Ti,k ℓci,k − Ti,k λi,k , 2
(E.7)
s Ti,k =
νk . var λi,k + ε
(E.8)
This form makes clear that empirical RatioNorm is being replaced by analytic centering plus analytic variance calibration. In implementation we stop-gradient through λcenter and λvar inside the normalizer, so the model cannot win by gaming the coordinate system. The budget penalty still uses differentiable λcenter , because that term must shape the update. The PPO objective uses rei,k in place of the empirical RatioNorm ratio: LLambdaNormT = −Ei,k [min (e ri,k Ai , clip(e ri,k , 1 − ϵ, 1 + ϵ)Ai )] . PPO
(E.9)
The main method combines this calibrated surrogate ratio with the damp-only λ-gradient weight from Equation (3.12). The earlier primal-dual ablation instead adds the exponential-moving-average (EMA) smoothed budget penalty from Section G.1: X 1 Ei [λcenter ]. (E.10) L = LLambdaNormT + βλ,t ϕα (Λlate − τ ) , Λlate = i,k PPO |Klate | k∈Klate
There is one further scale issue. For a Gaussian transition Xk+1 | Xk ∼ N (µθ,k , s2k I), the score with respect to the mean scales like ∇µ log p(Xk+1 | Xk ) =
Xk+1 − µθ,k , s2k 16
∥∇µ log p∥ scales like s−1 k .
(E.11)
Thus low-noise timesteps can receive disproportionately large gradients. A simple analytic compensation is an optional per-step loss weight p sk wkscore = clip (E.12) , wmin , wmax , sk = E[σk ∆tk ], sref which is applied to the per-step PPO loss. This is an ablation, not part of the minimal LambdaNormT definition. The initial LambdaNorm-T ablation grid varies around the clipping-derived central value: √ ν ∈ {0.01, 0.025, 0.04, 0.09}, ν ∈ {0.10, 0.158, 0.20, 0.30}.
(E.13)
The central setting is ν = 0.04, with tighter and looser alternatives included to test sensitivity. We also include a score-scale-weighted variant at the same target variance. LambdaNorm-T answers the sharper analytic question: can the finite-grid likelihood law replace most or all of empirical RatioNorm? The primary comparison reports reward together with pathvariance spend: RLambdaNormT+dual ≳ RRatioNorm
Λlate,LambdaNormT+dual < Λlate,RatioNorm . (E.14) This comparison tests whether analytic centering and variance calibration recover the stabilizing effect of empirical RatioNorm while preserving the same λ-budgeted control variable.
F
and
D IAGNOSTIC AND ALTERNATIVE ESTIMATORS OF THE PATH - VARIANCE SCALAR
The main method uses the transition mean-shift estimator in Section 3.3. The estimators below are useful for diagnostics, sanity checks, and implementation variants, but they are not used to train the main SD3.5 hard-OCR result. F.1
V ELOCITY- ONLY SIMPLIFICATION OF THE DRIFT
Often the model naturally exposes vθ , not bθ . If both policies use the same scalar schedule σt and the marginal-preserving drift Equation (C.4), then σ 2 (1 − t) bθ (x, t) − bold (x, t) = 1 − t (vθ (x, t) − vold (x, t)) . (F.1) 2t Therefore
2 2 2 vθ (xik , tk , ci ) − vold (xik , tk , ci ) bvel = 1 − σk (1 − tk ) λ ∆tk . i,k 2tk σk2
For the schedule σt2 = a2 t/(1 − t), the bracket simplifies to 1 − a2 /2, so 2 2 vθ (xik , tk , ci ) − vold (xik , tk , ci ) a2 vel b λi,k = 1 − ∆tk . 2 a2 tk /(1 − tk )
(F.2)
(F.3)
This identity is useful only when applied with the same time convention and same SDE drift used by the sampler. F.2
R ATIO - BASED VALIDATION
The theorem also predicts bk /2, E[log rk ] ≈ −λ
bk . Var(log rk ) ≈ λ
(F.4)
bmean = −2E b i [log ri,k ]. λ k
(F.5)
Thus one can estimate λk from empirical log-ratios: bvar = Var d i (log ri,k ), λ k 17
For mean-reduced latent log-probabilities, use the two-component validation from Proposition 3.4 instead: ¯r ]≈λ bcenter , ¯r )≈λ bvar . b i [log d i (log −2E Var (F.6) i,k i,k k k These ratio-based estimators are the convention used in the law-audit plots in Section 4. They are noisier than the transition mean-shift estimator, but they are a stringent test of the theory: if the finite-grid law is explanatory, then the mean-shift estimate and the observed log-ratio statistics should align up to sampling noise.
G
P RIMAL - DUAL BASELINE AND DIAGNOSTIC CLIPPING
This appendix formalizes the two constructions used outside the main method: the primal-dual ablation reported as a baseline in Tables 2 and 3, and the diagnostic z-space normalization used in the Tiny-SD3 law-audit intervention in Section 4. Both are referenced from the main text; neither is part of the λ-Controlled GRPO algorithm. G.1
E XPONENTIALLY SMOOTHED PRIMAL - DUAL BUDGET
The primal-dual ablation constrains the late-step path-variance mean X 1 bcenter , Λlate = λ k |Klate |
(G.1)
k∈Klate
through max J(θ) θ
subject to
Λlate ≤ τ.
With empirical RatioNorm as J(θ), the ablation uses the smooth hinge 2 softplus(αu) ϕα (u) = α
(G.2)
(G.3)
and primal loss Lt (θ) = LGRPO/RatioNorm + βλ,t ϕα (Λlate (θ) − τ ) . The multiplier is updated outside autograd from a detached EMA:
(G.4)
λ̄late,t = (1 − γ)λ̄late,t−1 + γ Λlate (θt )sg , (G.5) (G.6) βλ,t+1 = Π[0,βmax ] βλ,t + ηβ λ̄late,t − τ − δ . This ablation established that λ can actively control optimization, but the main method replaces the multiplier dynamics with direct damp-only λ-gradient weighting. G.2
D IAGNOSTIC NORMALIZATION AND CLIPPING
The reduced law gives the analytic diagnostic zi,k =
bcenter /2 log ri,k + λ i,k q . bvar + ε λ i,k
(G.7)
One can either log zi,k as a law audit or clip it, z̃i,k = clip(zi,k , −c, c),
(G.8)
and reconstruct
q gr = − 1 λ bcenter + z̃i,k λ bvar + ε. log i,k i,k 2 i,k This implies the adaptive interval q q 1 bcenter 1 bcenter var var b b log ri,k ∈ − λi,k − c λi,k , − λi,k + c λi,k . 2 2
18
(G.9)
(G.10)
H
Q UALITATIVE HELD - OUT OCR COMPARISONS
The aggregate held-out metrics in Table 2 are supported by a qualitative sweep on the same 762prompt test complement. The examples below use matched prompt indices and matched random seeds across methods: empirical RatioNorm at step 40, the primal-dual ablation at step 80, and LambdaNorm-T with damp-only λ weights at step 80. All examples are drawn from the disjoint 762-prompt held-out complement and were not used for checkpoint selection. The first block shows clear transcript-level wins for λ-Controlled GRPO. The second block deliberately includes mixed and failure cases, because the method improves the distribution of outcomes but does not solve text rendering uniformly. Ours
RatioNorm
Dual
Prompt. A futuristic time machine with a sleek, metallic finish, its dial set to “Best Day Ever Reload”. The machine is surrounded by a glowing aura, with digital readouts and holographic interfaces displaying vibrant colors, set against a backdrop of a twilight sky.
Figure 3: Qualitative held-out test comparison for prompt 319. Target text: Best Day Ever Reload. λ-Controlled GRPO renders the complete target in the central display, while RatioNorm and the primal-dual ablation produce the object and lighting but fail to place readable target text. Ours
RatioNorm
Dual
Prompt. A bustling amusement park with a vibrant entrance arch prominently displaying “Height Limit 48in”, surrounded by excited children and their parents, with colorful banners and playful music in the background.
Figure 4: Qualitative held-out test comparison for prompt 107. Target text: Height Limit 48in. The analytic method renders the full target on the amusement-park arch. RatioNorm and the primal-dual ablation retain the colorful park scene, but their sign text corrupts the numerical suffix or middle word.
19
Ours
RatioNorm
Dual
Prompt. A charming bakery window with a vintage wooden frame, adorned with a decal that reads “Fresh Daily” in elegant cursive. Sunlight streams through, casting a warm glow on the display of freshly baked bread and pastries.
Figure 5: Qualitative held-out test comparison for prompt 145. Target text: Fresh Daily. The analytic method gives the cleanest bakery window composition and a readable decal, with only a small character-level artifact in “Fresh.” RatioNorm repeats the text, and the primal-dual output is more typographically fragmented.
Ours
RatioNorm
Dual
Prompt. A bustling school cafeteria on a Friday, with a large, colorful sign displaying the “Pizza Friday Special” menu. Students in vibrant uniforms gather excitedly, pointing at the mouth-watering pizzas arranged on the serving counter.
Figure 6: Qualitative held-out test comparison for prompt 366. Target text: Pizza Friday Special. λ-Controlled GRPO writes the full cafeteria banner cleanly. RatioNorm and the primaldual ablation preserve the school-lunch scene, but their banner text is partial or corrupted.
20
Ours
RatioNorm
Dual
Prompt. An astronaut sits at a desk inside Moon Base Alpha, writing in a journal. The base’s futuristic interior is illuminated by soft blue lights, and a large window behind the astronaut showcases the barren lunar landscape and Earth rising above the horizon. “Moon Base Alpha” is prominently displayed on a plaque nearby.
Figure 7: Qualitative held-out test comparison for prompt 747. Target text: Moon Base Alpha. λ-Controlled GRPO keeps the full phrase readable on the base sign. The baselines generate coherent lunar-base scenes, but their signage repeats or corrupts the target phrase.
Ours
RatioNorm
Dual
Prompt. A gritty urban street at night, with a dive bar neon sign glowing brightly, reading “Cold Beer Here”, casting a warm, inviting glow through the misty air, reflecting off the wet pavement.
Figure 8: Qualitative held-out test comparison for prompt 168. Target text: Cold Beer Here. The analytic method produces the most legible target phrase and strongest neon-street atmosphere. The remaining error is spatial: “Here” is readable but spills onto the adjacent facade rather than remaining entirely within the main sign.
21
Ours
RatioNorm
Dual
Prompt. A serene desert scene with a weathered signpost pointing towards “Hidden Water Spring”, surrounded by tall palm trees and golden sand dunes, under a clear blue sky.
Figure 9: Qualitative held-out test comparison for prompt 176. Target text: Hidden Water Spring. The analytic method gives an exact, centered, readable sign in the intended desert scene. The comparison highlights the kind of sample counted by the stricter exact and substring metrics rather than only by mean OCR reward. Ours
RatioNorm
Dual
Prompt. A close-up photograph of an engraved silver ring with the inscription “Forever Yours” delicately etched into its surface, set against a soft, blurred background of romantic, warm tones.
Figure 10: Qualitative held-out test comparison for prompt 002. Target text: Forever Yours. Both λ-Controlled GRPO and RatioNorm render this short phrase cleanly; the primal-dual ablation introduces a character-level error. Ours
RatioNorm
Dual
Prompt. A neon bike rental sign glowing “Ride the City” stands out against a dark urban backdrop, its vibrant colors reflecting off wet pavements in a bustling night scene.
Figure 11: Qualitative held-out test comparison for prompt 047. Target text: Ride the City. λ-Controlled GRPO preserves the complete phrase, while the baselines tend to omit or repeat words.
22
Ours
RatioNorm
Dual
Prompt. A glowing Magic 8 Ball floats in a dimly lit room, its triangular window displaying the answer “Ask Again Later” in shimmering, ethereal blue text, surrounded by a soft, mystical aura.
Figure 12: Qualitative held-out test comparison for prompt 130. Target text: Ask Again Later. The proposed method keeps the three-word answer legible inside the triangular display; the baselines introduce spelling noise. Ours
RatioNorm
Dual
Prompt. A realistic photograph of a vintage vending machine with a prominent sign that reads “Exact Change Only”, set against a slightly worn brick wall, with a few coins scattered at its base.
Figure 13: Qualitative held-out test comparison for prompt 321. Target text: Exact Change Only. λ-Controlled GRPO places the phrase on the vending machine sign, while the baselines produce less recognizable signage. Ours
RatioNorm
Dual
Prompt. A futuristic space hotel lobby with a sleek, glowing sign that reads “ZeroG Suites”, surrounded by minimalist decor and large windows showcasing the vastness of space outside.
Figure 14: Qualitative held-out test comparison for prompt 456. Target text: ZeroG Suites. The proposed method renders the hotel sign cleanly; the baselines are close but introduce extra or incorrect letters.
23
Ours
RatioNorm
Dual
Prompt. A vast, icy landscape with a stark, metal sign reading “Titanic Memorial Site” standing firmly on a snow-covered rock near a towering iceberg, the cold waters of the North Atlantic stretching endlessly into the horizon.
Figure 15: Qualitative held-out test comparison for prompt 654. Target text: Titanic Memorial Site. λ-Controlled GRPO preserves the full memorial sign text in the icy scene, while the baselines retain only fragments.
Ours
RatioNorm
Dual
Prompt. A realistic photograph of a handwritten note, with the words “Eat Healthy Today” clearly visible, taped to the door of a modern refrigerator in a well-lit kitchen.
Figure 16: Qualitative held-out test comparison for prompt 706. Target text: Eat Healthy Today. The proposed method keeps the note text readable; the baselines drop or corrupt parts of the phrase.
24
H.1
M IXED AND FAILURE CASES
The examples above show where the analytic intervention is most visually clear. For calibration, we also scan the disjoint held-out test complement for prompts where the proposed method is comparable to RatioNorm, worse than RatioNorm, or where all methods remain far from the requested text. These cases are useful scientifically: they separate an average improvement in transcript success from a claim of uniform dominance. Ours
RatioNorm
Dual
Prompt. A vast, empty city square with a large, blank billboard standing prominently, awaiting the words “Your Ad Here” to be filled in, under a clear blue sky with a few fluffy clouds.
Figure 17: Mixed held-out test comparison for prompt 718. Target text: Your Ad Here. This is a clear negative case for the proposed method: λ-Controlled GRPO leaves the billboard effectively blank, while both empirical RatioNorm and the primal-dual ablation place the intended phrase cleanly in the center. Ours
RatioNorm
Dual
Prompt. A detailed amusement park map with colorful pathways and attractions, prominently marking “You Are Here” with a large, red arrow. The map is held by a friendly park employee, standing in front of a vibrant, bustling entrance.
Figure 18: Mixed held-out test comparison for prompt 617. Target text: You Are Here. The proposed method renders the main phrase but adds extra corrupted text, which hurts the transcript metric. RatioNorm and the primal-dual ablation give cleaner target text in this particular sample.
25
Ours
RatioNorm
Dual
Prompt. A high-tech laboratory setting with a test tube labeled “Sample XZ42” on a sleek, illuminated stand, surrounded by advanced scientific equipment and glowing monitors displaying complex data. The scene is modern and sterile, with a scientist in the background observing through a protective visor.
Figure 19: Failure held-out test comparison for prompt 026. Target text: Sample XZ42. All three methods produce plausible laboratory imagery, but none renders the alphanumeric label correctly. This illustrates a residual failure mode for short technical strings and small curved surfaces.
Ours
RatioNorm
Dual
Prompt. Retro diner scene with a red and white checkered placemat featuring the text “Todays Special Atomic Burger” in bold, vintage font. The placemat is slightly worn, with a classic 1950s diner background.
Figure 20: Mixed held-out test comparison for prompt 159. Target text: Todays Special Atomic Burger. This is a case where empirical RatioNorm is better: it keeps the main phrase nearly intact, while λ-Controlled GRPO produces a plausible diner placemat but scrambles the target into extra and reordered words.
26
Ours
RatioNorm
Dual
Prompt. “Skate or Die” slogan prominently displayed on a vibrant, colorful skateboard deck, set against a backdrop of a bustling urban skate park at sunset, with skaters in motion and graffiti-covered walls, capturing the rebellious spirit and dynamic energy of skate culture.
Figure 21: Mixed held-out test comparison for prompt 054. Target text: Skate or Die. All three samples are visually plausible skateboard-deck renderings, but this is not a clean win for the proposed method: the OCR system fails on the stylized version, while the RatioNorm and primaldual samples are closer to the normalized transcript.
Ours
RatioNorm
Dual
Prompt. A weathered treasure map laid out on an old wooden table, with “X Marks the Spot” clearly visible in the center, surrounded by intricate illustrations of mountains, forests, and a distant coastline, all under the warm glow of a vintage lamp.
Figure 22: Failure held-out test comparison for prompt 040. Target text: X Marks the Spot. All methods capture the map-and-lamp composition, but none renders the complete target phrase. This suggests that some failures are driven by scene/text placement difficulty rather than only by ratio instability.
27
Table 6 reports the per-checkpoint PickScore validation trajectory that underlies the winner selection in Table 4. Method RatioNorm RatioNorm RatioNorm λ-Controlled GRPO λ-Controlled GRPO λ-Controlled GRPO
Ckpt.
Mean PickScore ↑
Std.
Λlate ↓
40 60 80 40 60 80
0.8325 0.8266 0.8126 0.8365 0.8373 0.8380
0.057 0.057 0.059 0.057 0.055 0.057
3.8×10−5 6.4×10−5 6.80×10−3 4.65×10−3 1.9×10−4 3.34×10−3
Table 6: Random 256-prompt SD3.5 PickScore validation comparison. Every λ-Controlled GRPO checkpoint beats every RatioNorm checkpoint. RatioNorm peaks at step 40 and then regresses as its late-step λ spend jumps past τ ≈ 2.0×10−3 by step 80, an overspend of about 3.4τ . The analytic method improves PickScore monotonically across 40 → 60 → 80 and keeps Λlate at most comparable to τ .
I
Q UALITATIVE HELD - OUT P ICK S CORE COMPARISONS
The PickScore aggregate result in Table 4 is illustrated by a qualitative sweep on the same 762prompt held-out test complement. Comparisons are matched by prompt index and random seed across methods: empirical RatioNorm at step 40 (the validation-favored early stop) and λ-Controlled GRPO at step 80 (the validation-favored late stop). Every example below is from the disjoint 762prompt held-out complement and was not used for checkpoint selection. Captions report the prompt and the raw PickScore values verbatim, so the reader can judge each pair. The first block shows PickScore wins for the analytic method, the second block shows near-ties or PickScore losses, and the third block shows cases where both methods score low on the reward model. I.1
P ICK S CORE HITS ( ANALYTIC METHOD WINS ) λ-Controlled GRPO
RatioNorm
Prompt. anime girl in the space with a sign saying ‘sdxl’ as text,4k
Figure 23: Held-out test prompt 274. PickScore: Ours = 0.9137, RatioNorm = 0.8235, ∆ = +0.0902.
28
λ-Controlled GRPO
RatioNorm
Prompt. medieval painting of a man in a blue gown using a cellphone
Figure 24: Held-out test prompt 521. PickScore: Ours = 0.9014, RatioNorm = 0.8180, ∆ = +0.0834.
λ-Controlled GRPO
RatioNorm
Prompt. portrait of a giraffe in a fiery thunderstorm, digital art, hyper detailed
Figure 25: Held-out test prompt 853. PickScore: Ours = 0.9170, RatioNorm = 0.8512, ∆ = +0.0658.
29
λ-Controlled GRPO
RatioNorm
Prompt. Working in a diamond mine, Midjourney v5 style, insanely detailed, photorealistic, 8k, volumetric lighting
Figure 26: Held-out test prompt 223. PickScore: Ours = 0.8623, RatioNorm = 0.7737, ∆ = +0.0886.
λ-Controlled GRPO
RatioNorm
Prompt. Cyberpunk, Ancient India style, sci fi, silver on a black background, bas-relief, cyborgs, neon lighting, contrasting shadows, three-dimensional sculpture, high resolution, 8k detail, baroque, clear edges, technology, mechanisms
Figure 27: Held-out test prompt 682. PickScore: Ours = 0.8182, RatioNorm = 0.7329, ∆ = +0.0853.
30
λ-Controlled GRPO
RatioNorm
Prompt. Synesthesia, musical notes flying away from a music flower, charcoal hyperrealistic colourful art print by robert longo and alicexz
Figure 28: Held-out test prompt 626. PickScore: Ours = 0.8725, RatioNorm = 0.7951, ∆ = +0.0774.
λ-Controlled GRPO
RatioNorm
Prompt. A set of museum-quality emerald bracelets and beads in green in a display box at the auction, 32k, highest resolution, hyper realistic
Figure 29: Held-out test prompt 27. PickScore: Ours = 0.8773, RatioNorm = 0.8050, ∆ = +0.0723.
31
λ-Controlled GRPO
RatioNorm
Prompt. film still of Neytiri from Avatar
Figure 30: Held-out test prompt 118. PickScore: Ours = 0.8441, RatioNorm = 0.7749, ∆ = +0.0692.
λ-Controlled GRPO
RatioNorm
Prompt. rutger hauer from blade runner standing in the rain, green light, vhs quality, film grain, wet, very sad and reluctant expression, wearing a biomechanical suit, scifi, digital painting, artstation, concept art
Figure 31: Held-out test prompt 835. PickScore: Ours = 0.8748, RatioNorm = 0.8071, ∆ = +0.0677.
32
λ-Controlled GRPO
RatioNorm
Prompt. a portal to snowy mountain, standing in a warm summer field
Figure 32: Held-out test prompt 1003. PickScore: Ours = 0.9609, RatioNorm = 0.8936, ∆ = +0.0673.
λ-Controlled GRPO
RatioNorm
Prompt. art print of a cute fire elemental pokemon by league of legends. finally evolutionary stage
Figure 33: Held-out test prompt 514. PickScore: Ours = 0.9264, RatioNorm = 0.8594, ∆ = +0.0670.
33
λ-Controlled GRPO
RatioNorm
Prompt. Nissan GT-R in a parking lot, raining, film grain, moody
Figure 34: Held-out test prompt 149. PickScore: Ours = 0.8943, RatioNorm = 0.8288, ∆ = +0.0655.
λ-Controlled GRPO
RatioNorm
Prompt. a manga drawing of naruto fighting goku
Figure 35: Held-out test prompt 130. PickScore: Ours = 0.8538, RatioNorm = 0.7916, ∆ = +0.0622.
34
λ-Controlled GRPO
RatioNorm
Prompt. Movie still of star wars young luke skywalker working as mechanic in a garage, extremely detailed, intricate, high resolution, hdr, trending on artstation
Figure 36: Held-out test prompt 852. PickScore: Ours = 0.8465, RatioNorm = 0.7915, ∆ = +0.0550.
λ-Controlled GRPO
RatioNorm
Prompt. photo of king kong lifting a landrover defender in the jungle river, misty mud rocks, headlights chrome detailing
Figure 37: Held-out test prompt 307. PickScore: Ours = 0.8618, RatioNorm = 0.8076, ∆ = +0.0542.
35
I.2
P ICK S CORE NEAR - TIES AND R ATIO N ORM WINS λ-Controlled GRPO
RatioNorm
Prompt. analog style picture of a lizard dressed as a knight in armour
Figure 38: Held-out test prompt 942. PickScore: Ours = 0.8598, RatioNorm = 0.9018, ∆ = −0.0420. λ-Controlled GRPO
RatioNorm
Prompt. A fortnite map inspired by Star Wars
Figure 39: Held-out test prompt 597. PickScore: Ours = 0.8291, RatioNorm = 0.8680, ∆ = −0.0389.
36
I.3
P ICK S CORE JOINT FAILURES ( BOTH METHODS SCORE LOW ) λ-Controlled GRPO
RatioNorm
Prompt. Young Victoria coren-mitchell, insanely detailed, photorealistic, 8k, ultra high resolution, volumetric lighting, taken with canon eos,
Figure 40: Held-out test prompt 391. PickScore: Ours = 0.6642, RatioNorm = 0.6773, ∆ = −0.0131. Both methods land well below the aggregate mean; this prompt class (photorealistic named-individual likeness) is a known hard case for SD3.5-LoRA regardless of the ratio controller. λ-Controlled GRPO
RatioNorm
Prompt. an epic view of a demonic Rose-ringed parakeet cyborg inside an ironmaiden robot, wearing a noble robe, large view, a surrealist painting, aralan bean and Philippe Druillet, hiromu arakawa, volumetric lighting, detailed shadows
Figure 41: Held-out test prompt 10. PickScore: Ours = 0.6661, RatioNorm = 0.6675, ∆ = −0.0015. Heavy compositional nesting with named artist styles; both methods score near the bottom of the distribution.
37
J
Q UALITATIVE HELD - OUT G EN E VAL COMPARISONS
This appendix illustrates the compositional GenEval regime with matched held-out image comparisons between the analytic λ-Controlled GRPO method and empirical RatioNorm on SD3.5. Each pair uses the same prompt and the same random seed at each method’s validation-selected checkpoint, and every example is drawn from the disjoint held-out test complement that was not used for checkpoint selection. Captions report the prompt, its GenEval category tag, and the raw per-prompt GenEval and strict-accuracy scores, so the reader can judge each pair. The first block shows GenEval wins for the analytic method, the second shows joint successes where both methods pass, and the third shows cases where RatioNorm wins. J.1
G EN E VAL HITS ( ANALYTIC METHOD WINS ) λ-Controlled GRPO
RatioNorm
Prompt. a photo of a toothbrush below a pizza
(position)
Figure 42: Held-out test prompt 117. GenEval: Ours = 1.00, RatioNorm = 0.00. Strict accuracy: Ours = 1, RatioNorm = 0.
38
λ-Controlled GRPO
RatioNorm
Prompt. a photo of a green couch and an orange umbrella
(color_attr)
Figure 43: Held-out test prompt 124. GenEval: Ours = 1.00, RatioNorm = 0.00. Strict accuracy: Ours = 1, RatioNorm = 0.
λ-Controlled GRPO
RatioNorm
Prompt. a photo of a toaster below a traffic light
(position)
Figure 44: Held-out test prompt 234. GenEval: Ours = 1.00, RatioNorm = −0.33. Strict accuracy: Ours = 1, RatioNorm = 0.
39
λ-Controlled GRPO
RatioNorm
Prompt. a photo of three sinks
(counting)
Figure 45: Held-out test prompt 423. GenEval: Ours = 1.00, RatioNorm = 0.33. Strict accuracy: Ours = 1, RatioNorm = 0.
40
J.2
G EN E VAL JOINT SUCCESSES ( BOTH METHODS PASS ) λ-Controlled GRPO
RatioNorm
Prompt. a photo of a cat
(single_object)
Figure 46: Held-out test prompt 5. GenEval: Ours = 1.00, RatioNorm = 1.00. Both methods produce a strict pass; included to show that the matched-seed protocol is not exploiting failure modes of one side. λ-Controlled GRPO
RatioNorm
Prompt. a photo of two trains
(counting)
Figure 47: Held-out test prompt 8. GenEval: Ours = 1.00, RatioNorm = 1.00.
41
λ-Controlled GRPO
RatioNorm
Prompt. a photo of a yellow bird and a black motorcycle
(color_attr)
Figure 48: Held-out test prompt 9. GenEval: Ours = 1.00, RatioNorm = 1.00.
λ-Controlled GRPO
RatioNorm
Prompt. a photo of a zebra left of an elephant
(position)
Figure 49: Held-out test prompt 25. GenEval: Ours = 1.00, RatioNorm = 1.00. A rare position joint success on a category that is otherwise near-floor for both methods.
42
J.3
G EN E VAL LOSSES (R ATIO N ORM WINS ) λ-Controlled GRPO
RatioNorm
Prompt. a photo of a kite above a toothbrush
(position)
Figure 50: Held-out test prompt 39. GenEval: Ours = 0.33, RatioNorm = 1.00. Strict accuracy: Ours = 0, RatioNorm = 1. λ-Controlled GRPO
RatioNorm
Prompt. a photo of three baseball bats
(counting)
Figure 51: Held-out test prompt 62. GenEval: Ours = 0.00, RatioNorm = 1.00. Strict accuracy: Ours = 0, RatioNorm = 1.
43
λ-Controlled GRPO
RatioNorm
Prompt. a photo of a green bus
(colors)
Figure 52: Held-out test prompt 105. GenEval: Ours = 0.00, RatioNorm = 1.00. Strict accuracy: Ours = 0, RatioNorm = 1. The bus is recognized in both images, but only the RatioNorm rendering is judged strictly green by the GenEval color classifier.
λ-Controlled GRPO
RatioNorm
Prompt. a photo of a pizza and a bench
(two_object)
Figure 53: Held-out test prompt 136. GenEval: Ours = 0.00, RatioNorm = 1.00. Strict accuracy: Ours = 0, RatioNorm = 1.
44