Preprint
B EYOND T OKEN -L OCAL I MITATION : R EWARD -C OMPATIBLE T EMPORAL C REDIT A SSIGNMENT FOR O N -P OLICY D ISTILLATION Shiqi Liu1,2 , Zeyu He1,2 , Letian Tao1,2 , Guojian Zhan1,2 , Jiaxin Gao1 , Feihong Zhang1 , Jingliang Duan1,2 , Wei Xiong2 , Kehua Sheng2 , Bo Zhang2 , Yang Guan1,B , Shengbo Eben Li1,B
arXiv:2609.16937v1 [cs.LG] 15 Sep 2026
1
School of Vehicle and Mobility & College of AI, Tsinghua University 2 Didi Voyager Labs, DiDi Autonomous Driving
A BSTRACT On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost of horizon-dependent variance. We establish a unified temporal-credit view of these formulations, showing that practical token-level OPD can be interpreted as a temporal approximation to the sequence-level reverse-KL gradient. Building on this connection, we propose γOPD, which uses discounted temporal credit assignment to balance long-horizon supervision and optimization stability, while admitting a horizon-independent variance bound. We further develop a rewardcompatible bounded mixing (RBM) mechanism for γOPD that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization. Experiments on mathematical and code reasoning demonstrate consistent improvements over existing OPD methods across vanilla, size-mismatched, and multi-teacher distillation settings.
(a)
(b)
Figure 1: Core Idea. (a) Vanilla OPD (left) focuses only on the immediate next step and closely follows the teacher’s local guidance. We argue that OPD should instead account for the influence of future reasoning steps and should not merely imitate the teacher token by token (right). (b) By incorporating future outcomes through temporal credit assignment and balancing teacher supervision with verifiable rewards, γOPD achieves strong overall performance across mathematical and code reasoning tasks in the multi-teacher distillation setting, as detailed in Section 4.
1
I NTRODUCTION
Recent large language models (LLMs), including Kimi K3 (Kimi Team et al., 2026), GLM-5 (GLM5 Team, 2026), DeepSeek-V4 (DeepSeek-AI et al., 2026), and Nemotron-Cascade 2 (Yang et al., 2026b), have demonstrated strong reasoning capabilities across mathematics, coding, and science. B
Corresponding author: S. E. Li and Y. Guan; email: [email protected].
1
Preprint
To further consolidate and integrate such capabilities, on-policy distillation (OPD) (Lu & Thinking Machines Lab, 2025) has emerged as a key approach in LLM post-training. It provides dense teacher supervision on trajectories sampled from the current student policy, promoting high-quality reasoning while mitigating exposure bias. To avoid costly full-vocabulary reverse-KL computation, existing OPD methods (Yang et al., 2026a; Jin et al., 2026; Oh et al., 2026) reformulate the objective in a policy-gradient form and approximate it using a one-sample token-level Monte Carlo estimator (Li et al., 2026). While computationally efficient, this practical estimator introduces non-negligible bias relative to the original sequence-level objective, effectively altering the optimization problem (Yang et al., 2026a). Other methods (Fu et al., 2026) instead adopt sequence-level Monte Carlo estimators to more faithfully preserve the original objective. However, accumulating future credit over the remaining response leads to increasing variance as the sequence grows, resulting in unstable optimization for long reasoning trajectories. In this work, we present a unified analysis of token-level and sequence-level OPD, deriving their underlying objectives and clarifying the relationship between their policy-gradient formulations. Building on this insight, we propose γOPD, which introduces discounted temporal credit assignment to interpolate between token-local and sequence-level supervision while admitting a horizonindependent variance bound. To complement teacher supervision with outcome-level guidance, we further develop a reward-compatible bounded mixing mechanism that incorporates verifiable outcome feedback without sacrificing the stability benefits of temporal discounting. Overall, our main contributions are summarized as follows: • We establish a unified temporal-credit view of OPD, showing that practical token-level OPD can be interpreted as a temporally truncated approximation to the sequence-level reverse-KL gradient. • Building on this connection, we propose γOPD, a discounted temporal-credit surrogate that interpolates between token-level and sequence-level credit assignment, while admitting a horizon-independent variance bound. • We further develop Reward-Compatible Bounded Mixing (RBM) for γOPD, which combines discounted teacher-derived credit with verifiable outcome rewards while preventing the teacher signal from dominating task-level supervision.
2
P RELIMINARIES
2.1
N OTATION
We consider autoregressive reasoning tasks over a discrete vocabulary V, where we aim to optimize a student policy πθ parameterized by θ under the guidance of a fixed teacher policy π ∗ . Let x ∼ D denote a prompt sequence sampled from the reasoning dataset. Given x, the policy generates a complete response trajectory y = (y1 , y2 , . . . , yT ) ∈ V T , where T = |y| is the number of tokens in the trajectory, and yt ∈ V is the token generated at step t. At decoding step t, the policy conditions on the historical context prefix ht := (x, y<t ), which consists of the prompt and all previously generated reasoning tokens. We use DKL (·∥·) to denote the KL divergence. For notational simplicity, we define the token-level and sequence-level log-ratios as ∆t ≜ log πθ (yt | ht ) − log π ∗ (yt | ht ), ∆(y) ≜ log πθ (y | x) − log π ∗ (y | x) =
T X
(1a) ∆t .
(1b)
t=1
2.2
O N -P OLICY D ISTILLATION (OPD)
On-policy distillation (OPD) (Agarwal et al., 2024; DeepSeek-AI et al., 2026) trains the student by minimizing the reverse KL divergence from the student policy πθ to the teacher policy π ∗ , evaluated on trajectories induced by the current student itself: JOPD (θ) = Ex∼D [DKL (πθ (y | x) ∥ π ∗ (y | x))] . 2
(2)
Preprint
By training on trajectories sampled from the current student policy, on-policy learning reduces exposure bias and is well suited to long chain-of-thought reasoning tasks such as mathematical problem solving (GLM-5 Team, 2026). Following recent practice (Lu & Thinking Machines Lab, 2025; Li et al., 2026), the empirical implementation of OPD typically converts the reverse-KL minimization in equation 2 into an RL-style surrogate gradient: " T # X OPD (3) ∇θ JOPD (θ) ≈ −Ex∼D At ∇θ log πθ (yt | ht ) , t=1
where AOPD ≜ −∆t = log π ∗ (yt | ht ) − log πθ (yt | ht ) is the token-level OPD advantage. t 2.3
RL WITH V ERIFIABLE R EWARDS (RLVR)
Reinforcement learning with verifiable rewards (RLVR) (Shao et al., 2024; Guo et al., 2025; Yu et al., 2025) has also become a key approach for improving the reasoning ability of LLMs. For a prompt x ∼ D, the policy πθ generates an output sequence y = (y1 , . . . , yT ) ∼ πθ (· | x). The generated output is evaluated by an external verifier, such as a code compiler or a mathematical rule checker, which provides a sparse sequence-level reward scalar R(x, y) ∈ {−1, 1}. The RLVR objective is JRLVR (θ) = Ex∼D, y∼πθ (·|x) [R(x, y)] . Compared with OPD, RLVR provides a sparse but verifier-grounded sequence-level correctness signal, whereas OPD supplies dense token-level supervision through teacher guidance, which may nevertheless be limited by the teacher’s suboptimal behavior.
3
M ETHODOLOGY
3.1
T OKEN - LEVEL OPD AS A T EMPORAL A PPROXIMATION
Although the practical OPD objective in equation 3 is widely used, it should be understood as a temporal approximation rather than the exact policy gradient of the OPD objective in equation 2. To clarify this distinction, we formally define token-level OPD and sequence-level OPD as follows: seq (θ) ≜ Ex∼D [DKL (πθ (y | x) ∥ π ∗ (y | x))] , JOPD " T # X token ∗ JOPD (θ) ≜ Ex∼D Eht ∼dθ̄ (·|x) [DKL (πθ (yt | ht ) ∥ π (yt | ht ))] ,
(4a) (4b)
t=1 seq where JOPD denotes the sequence-level reverse-KL objective, i.e., the original OPD objective defined in equation 2, which compares the student and teacher distributions over complete response token trajectories. In contrast, JOPD denotes the token-level objective, which compares their next-token distributions at sampled prefixes. The distribution dθ̄ (ht | x) is the prefix distribution induced by the current student policy, where ht = (x, y<t ). The notation θ̄ indicates stop-gradient, namely this prefix distribution is treated as fixed when differentiating with respect to θ.
The two objectives are equivalent at the level of function values when the stop-gradient reference parameter is evaluated at the current policy parameter. Specifically, if θ̄ = θ, then the sampled prefix token distribution in JOPD matches the autoregressive prefix distribution induced by πθ , and we have seq token JOPD (θ) = JOPD (θ) θ̄=θ .
However, this value-level equivalence does not imply gradient-level equivalence. The reason is seq token that JOPD differentiates through the autoregressive distribution over future prefixes, whereas JOPD stops this dependence through dθ̄ (ht | x). This distinction leads to different temporal credit assignments, as shown below. Proposition 3.1 (Token-level OPD gradient). Under the stop-gradient treatment of the prefix distritoken bution dθ̄ (ht | x), the gradient of JOPD (θ) in equation 4b is given by " T # X token ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t ∇θ log πθ (yt | ht ) . (5) t=1
3
Preprint
Token-level OPD
low Sequence-level OPD
Teacher
Verifier
Temporal-credit OPD ( OPD)
high (a) Temporal credit Assignment
(b) Reward-compatible bounded mixing
Figure 2: Overview of γOPD. (a) Token-level OPD considers only the log-ratio of the current token, whereas sequence-level OPD accumulates the log-ratios of all future tokens. γOPD introduces a discount factor γ to interpolate between these two extremes. (b) The γOPD advantage is first normalized and then combined with the verifiable task reward, balancing dense teacher guidance with task-level verification signals. The proof is provided in Appendix A. token (θ) recovers the practical OPD update in equation 3. MeanConsequently, the gradient of JOPD seq while, the gradient of the sequence-level OPD objective JOPD (θ) takes the following form: Proposition 3.2 (Sequence-level OPD gradient). The gradient of the sequence-level OPD objective seq (θ) in equation 4a is given by JOPD " T ! # T X X seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t′ ∇θ log πθ (yt | ht ) . (6) t′ =t
t=1
The proof is provided in Appendix B. Propositions 3.1 and 3.2 show that practical OPD is a token-level approximation to sequence-level OPD that neglects the influence of subsequent tokens. As illustrated in Figure 2, under the sequencePT level objective, the score term at step t is weighted by the future log-ratio t′ =t ∆t′ , since yt affects all future histories ht+1 , . . . , hT . In contrast, practical OPD keeps only the local term ∆t . This gap stems from the stop-gradient treatment of dθ̄ (ht | x) in equation 4b. 3.2
T EMPORAL -C REDIT O N -P OLICY D ISTILLATION (γOPD)
Compared with the token-level gradient in Proposition 3.1, the sequence-level gradient in equation 6 accounts for the indirect effect of each token on future autoregressive prefixes, and is therefore unbiased with respect to the sequence-level objective in equation 2. However, for long reasoning trajecPT tories, the accumulated future log-ratio t′ =t ∆t′ can have large variance, which may destabilize training. To balance bias and variance in temporal credit assignment, we propose Temporal-Credit On-Policy Distillation (γOPD). The key idea is to introduce a discount factor γ ∈ [0, 1] to control how much future OPD credit is assigned to the current token. Specifically, we define the temporal-credit surrogate gradient as " T # X (γ) gγ (θ) ≜ −Ex∼D, y∼πθ (·|x) At ∇θ log πθ (yt | ht ) , t=1
where the discounted γOPD advantage is defined as (γ)
At
≜−
T X t′ =t
4
′
γ t −t ∆t′ .
(7)
Preprint
(1)
The discount factor γ controls the temporal horizon of OPD credit assignment. When γ = 1, At becomes the sequence-level OPD return-to-go, recovering equation 6. When γ = 0, it reduces to (0) the token-level OPD advantage At = −∆t , recovering equation 5. Thus, for 0 < γ < 1, gγ (θ) is deliberately introduced as a biased surrogate that interpolates between these two endpoint gradients, allowing its variance to be controlled through γ, as formalized in the following theorem. Theorem 3.3 (Variance Stability of γOPD). Assume that the token-level OPD advantage has a (0) 2 bounded second moment, i.e., E[(At )2 ] ≤ σ∆ for all t. Then, the sequence-level OPD credit admits a horizon-dependent variance bound, whereas the γOPD credit admits a horizon-independent variance bound: (1) 2 Var At ≤ (T − t + 1)2 σ∆ , 2 σ∆ (γ) Var At ≤ , ∀γ ∈ [0, 1). (1 − γ)2 The proof is provided in Appendix C. Theorem 3.3 shows that the variance of sequence-level OPD credit can grow quadratically with the remaining sequence horizon in the worst case, whereas γOPD admits a horizon-independent variance bound for any fixed γ < 1. Adjusting γ therefore provides a principled trade-off between training stability and long-horizon credit propagation. 3.3
γOPD WITH R EWARD -C OMPATIBLE B OUNDED M IXING (RBM)
Although γOPD improves temporal credit assignment, its optimization signal is still fundamentally derived from the teacher distribution. Consequently, the student may remain constrained by the teacher policy and lack an explicit task-level optimization signal for solving verifiable reasoning problems. To move beyond this teacher-imitation ceiling, we incorporate the verifiable task reward R(x, y) ∈ {−1, 1} into γOPD. A naive combination, however, can be problematic because the magnitude of (γ) the teacher-derived advantage At may vary substantially across responses and potentially dominate the bounded task reward. We therefore introduce a reward-compatible bounded mixing (RBM) mechanism: the γOPD credit is first calibrated by its response-wise mean absolute magnitude, then bounded through a softsign transformation, and finally combined with the verifiable reward. These operations can be written compactly as (γ)
b(γ) A ≜ t
At 1 T
(γ)
PT
k=1 Ak
(γ)
+ R(x, y).
(8)
+ At
(γ)
When the denominator vanishes, equivalently when Ak = 0 for all k, we define the normalized teacher term as zero. The first term is equivalent to applying softsign after response-wise mean(γ) absolute normalization. It is monotonic in At and bounded within (−1, 1), thereby retaining the relative token-level OPD credit while preventing the teacher-derived signal from overwhelming the b(γ) task reward. Since R(x, y) ∈ {−1, 1}, it follows that sign(A t ) = R(x, y). Thus, RBM is rewardcompatible by construction: the verifiable reward determines whether the sampled response is reinforced or suppressed, while the bounded γOPD credit modulates the token-level update strength. The resulting reward-compatible γOPD surrogate gradient is " T # X (γ) b ĝγ (θ) ≜ −Ex∼D, y∼π (·|x) At ∇θ log πθ (yt | ht ) . θ
(9)
t=1
Therefore, RBM combines sparse verifiable rewards for task-level learning beyond teacher imitation with bounded γOPD for dense token-level credit assignment. Unless otherwise specified, γOPD refers to its use within RBM throughout the remainder of this paper. 5
Preprint
Table 1: Mathematical reasoning results under vanilla and size-mismatch distillation. We report accuracy and pass rate across four mathematical reasoning benchmarks. AIME24
Method
Acc.
Pass
AIME25 Acc.
Pass
AMC23 Acc.
Pass
MATH500
Average
Acc.
Pass
Acc.
Pass
Qwen3-4B-Math → Qwen3-4B Student Teacher
22.60 57.40
60.00 76.67
20.83 51.67
33.33 66.67
60.16 93.52
92.50 97.50
67.40 69.80
72.60 78.40
42.75 68.10
64.61 79.81
JustRL STAPO
53.54 55.31
83.33 83.33
42.50 50.00
56.67 66.67
88.75 90.16
97.50 97.50
68.00 68.35
70.40 69.80
63.20 65.95
76.97 79.33
OPD ExOPD REOPOLD AOPD TOPD γOPD
56.46 57.19 56.98 56.56 54.79 60.94
80.00 83.33 83.33 86.67 80.00 86.67
50.83 52.50 51.67 52.50 51.67 54.17
63.33 63.33 63.33 66.67 63.33 66.67
92.81 93.13 93.59 93.28 93.44 92.89
95.00 97.50 97.50 97.50 97.50 97.50
68.05 68.90 68.70 69.35 69.95 69.95
72.20 74.20 72.60 73.00 72.40 78.80
67.04 67.93 67.74 67.92 67.46 69.49
77.63 79.59 79.19 80.96 78.31 82.41
Qwen3-4B-Math → Qwen3-1.7B Student Teacher
11.77 57.40
46.67 76.67
8.33 51.67
20.00 66.67
39.84 93.52
90.00 97.50
59.20 69.80
72.80 78.40
29.79 68.10
57.37 79.81
JustRL STAPO
37.60 35.52
66.67 63.33
36.67 30.83
50.00 46.67
78.67 77.73
95.00 95.00
65.10 64.95
68.80 68.80
54.51 52.26
70.12 68.45
OPD ExOPD REOPOLD AOPD TOPD γOPD
37.71 38.54 37.50 37.71 39.38 42.08
66.67 66.67 66.67 70.00 70.00 70.00
31.25 31.67 30.83 32.50 30.83 35.00
40.00 43.33 43.33 43.33 36.67 46.67
76.02 76.72 77.73 79.38 79.14 80.86
92.50 92.50 92.50 95.00 95.00 95.00
65.40 65.05 66.70 65.85 65.60 68.10
69.80 68.40 69.00 70.80 70.00 73.20
52.60 53.00 53.19 53.86 53.74 56.51
67.24 67.73 67.88 69.78 67.92 71.22
4
E XPERIMENTS
4.1
S ETTINGS
Benchmarks. Our experiments evaluate both mathematical and code reasoning abilities. For mathematical reasoning, we train on DeepMath (He et al., 2025), retaining problems with difficulty level at least 6, and evaluate on AIME24 (Li et al., 2024), AIME25 (OpenCompass, 2025), AMC23 (Li et al., 2024), and MATH500 (Hendrycks et al., 2021). For code reasoning, we train on the 25Ksample Eurus-RL-Code dataset (Cui et al., 2026) and evaluate on HumanEval+, MBPP+ (Liu et al., 2023), and the v6 split of LiveCodeBench (Jain et al., 2025), covering problems from February 2025 to May 2025. Detailed training and evaluation configurations are provided in Appendix E. Models. We conduct experiments under three distillation settings: (1) vanilla distillation, where we distill a Qwen3-4B-Math (Yang et al., 2026a) teacher into a Qwen3-4B (Yang et al., 2025) student for mathematical reasoning; (2) size-mismatch distillation, where we distill the same Qwen3-4BMath teacher into a smaller Qwen3-1.7B student for mathematical reasoning; and (3) multi-teacher distillation, where we distill Qwen3-4B-Math and Qwen3-4B-Code (Yang et al., 2026a) teachers into a single Qwen3-4B student for both math and code reasoning. Baselines. We compare γOPD against representative OPD-based baselines, including vanilla OPD (Lu & Thinking Machines Lab, 2025), ExOPD (Yang et al., 2026a), REOPOLD (Ko et al., 2026), AOPD (Jia et al., 2026), and TOPD (Zhang et al., 2026), as well as RL-based methods, including JustRL (He et al., 2026) and STAPO (Liu et al., 2026). For γOPD, we use γ = 0.99 6
Preprint
0.4 0.2 0.0
0
10
20 30 Step
40
10
10
0
10
−2
10
−4
Grad Norm
0.6
REOPOLD
0
10
20 30 Step
10
AOPD
1
−1
40
0
10
20 30 Step
γOPD
TOPD Response Length
ExOPD Absolute Advantage
Verifiable Reward
OPD
10000
5000
40
0
10
20 30 Step
40
Figure 3: Training dynamics under vanilla distillation. Verifiable reward, absolute advantage, gradient norm, and response length throughout training are reported. Due to the truncation mechanism, TOPD’s verifiable reward remains mostly below zero throughout training. Table 2: Accuracy on mathematical and code reasoning benchmarks under multi-teacher distillation. TotalAvg is the equal-weighted average of the mean mathematical score and the mean code score. Method Student Teacher OPD ExOPD REOPOLD AOPD TOPD γOPD
AIME24 22.60 57.40 55.21 57.50 58.13 55.52 57.81 59.27
Mathematical Reasoning AIME25 AMC23 MATH500 20.83 51.67 52.50 52.50 52.08 55.83 51.67 55.00
60.16 93.52 93.13 92.34 91.88 93.05 92.34 94.38
67.40 69.80 68.65 68.25 68.55 68.70 69.50 69.00
Code Reasoning HumanEval+ MBPP+ 75.61 82.32 84.15 85.37 82.32 84.76 84.15 89.02
61.90 70.63 67.99 71.16 69.31 69.05 69.58 70.63
LCB 17.14 25.71 26.00 26.71 28.00 26.71 27.86 28.14
TotalAvg 47.15 63.82 63.37 64.36 63.77 64.22 64.18 66.00
by default unless otherwise specified. All methods are implemented based on veRL (Sheng et al., 2025). 4.2
M AIN R ESULTS
Vanilla Distillation. As shown in Table 1, γOPD consistently outperforms existing OPD baselines in the vanilla distillation setting. Compared with the strongest baseline for each metric, γOPD achieves relative improvements of 2.30% in AvgAcc and 1.80% in AvgPass. Notably, it also surpasses the teacher on both averaged metrics. These results demonstrate the effectiveness of γOPD and suggest that the proposed RBM can better leverage dense teacher supervision while enabling improvement beyond direct imitation. Size-Mismatch Distillation. The size-mismatch setting (Qwen3-4B-Math → Qwen3-1.7B) is more challenging because the student has substantially lower capacity than the teacher. Consequently, as shown in Table 1, all student methods remain below the teacher. Nevertheless, γOPD substantially narrows the performance gap and achieves the best overall performance. Compared with the strongest baseline for each metric, γOPD achieves relative improvements of 4.92% in AvgAcc and 2.06% in AvgPass. These results highlight the benefit of temporal credit assignment when distilling a stronger teacher into a lower-capacity student. Multi-Teacher Distillation. We next evaluate whether γOPD remains effective in the multi-teacher distillation setting. As shown in Table 2, it consistently performs strongly across mathematical reasoning benchmarks and achieves the best results on HumanEval+ and LCB among the code reasoning tasks. Overall, γOPD improves TotalAvg by 1.87% over the strongest baseline, AOPD. Notably, it also outperforms the teacher on average in both mathematical and code reasoning. These results demonstrate that γOPD remains effective when jointly training a shared student with domainspecialized teachers. Training Dynamics. We further visualize the training dynamics of different methods under the vanilla distillation setting in Figure 3. Throughout training, γOPD consistently achieves the highest verifiable reward, indicating the largest proportion of correctly solved problems. It also maintains a stable absolute OPD advantage and exhibits the smallest fluctuations in gradient norm among all 7
Preprint
Table 3: Ablation study of the proposed components on AIME benchmarks. γ, Mix, and Norm correspond to temporal discounting, reward mixing, and bounded normalization, respectively. ∆ Avg. denotes the absolute improvement in average accuracy over the vanilla OPD baseline. Components γ ✓ ✓ – ✓
Mix
Norm
Vanilla OPD – – ✓ – ✓ ✓ ✓ ✓
Accuracy (%)
∆ Avg.
AIME24
AIME25
Avg.
56.46 58.29 58.81 57.55 60.94
50.83 53.08 52.67 53.75 54.17
53.65 55.69 55.74 55.65 57.56
– +2.04 +2.09 +2.00 +3.91
methods, suggesting more consistent temporal credit assignment and optimization. Moreover, the response length of γOPD converges to a shorter and more stable range than those of most baselines, with the exception of TOPD, which explicitly truncates the distillation signal. Together, these results suggest that γOPD enables more stable optimization while achieving stronger mathematical reasoning performance. 4.3
A BLATION S TUDIES
We isolate the effects of temporal discounting, reward mixing, and bounded normalization on the AIME benchmarks in Table 3. Temporal discounting alone improves the average accuracy by 2.04 points over vanilla OPD. Removing temporal discounting from the full method while retaining reward mixing and normalization reduces the gain from 3.91 to 2.00 points, confirming its complementary contribution. Moreover, adding naive reward mixing to temporal discounting brings only a marginal 0.05-point improvement, whereas the full combination including bounded normalization yields the strongest overall performance. These results suggest that temporal credit assignment and reward-compatible normalization provide complementary benefits, with their combination yielding the strongest performance. We further analyze the computational and memory overhead of the individual components of γOPD. Compared with standard OPD, γOPD introduces only about 0.1% additional computation time with negligible memory overhead; detailed profiling results are provided in Appendix F.1. 4.4
D ETAILED A NALYSIS
Sensitivity to γ. We further investigate the sensitivity of temporal credit assignment to the discount factor γ without RBM. Specifically, we distill a Qwen3-4B-Math teacher into a Qwen3-1.7B student using different values of γ, with the results shown in Figure 4. Here, γ = 0 reduces to local OPD, whereas γ = 1 corresponds to the undiscounted sequence-level return-to-go. Due to the long response horizon, the sequence-level variant produces substantially larger gradient norms and suffers from rapid entropy collapse, resulting in inferior reasoning performance. In contrast, local OPD and the discounted variants steadily improve during training. Among them, γ = 0.99 achieves the best performance, outperforming smaller values such as γ = 0.9, while exhibiting distinct gradient-norm and entropy dynamics. These results suggest that a large but sub-unity discount factor provides the best balance between long-range temporal credit assignment and optimization stability. Visualization of Token Advantages. We further visualize the normalized token-level advantages (0) (0.99) (0) At and At in Figure 5. In the correct response, At provides sparse local supervision, assigning a noticeable negative signal to only a few tokens while exerting little influence on most of (0.99) the reasoning process. In contrast, At propagates information from subsequent steps, producing a smoother signal and assigning stronger positive credit to the key mathematical derivation while (0) mildly penalizing redundant text. For the incorrect response, At strongly penalizes several tokens (γ) that are only weakly related to the actual reasoning error, whereas At concentrates stronger negative credit on the erroneous formula derivation without excessively penalizing the final token. These examples illustrate that temporal credit assignment can redistribute supervision toward reasoning 8
Preprint
γ = 0.9
γ = 0.99
γ=1 0.4
0.3 0.2
10
1
10
0
Entropy
Grad Norm
AIME24 Acc (avg@32)
γ=0 0.4
0.3
0.2
0.1 0
10
20
30
40 Step
50
60
70
0
10
20
30
40 Step
50
60
70
0
10
20
30
40 Step
50
60
70
Figure 4: Training dynamics under different γ. OPD distills a Qwen3-4B-Math teacher into a Qwen3-1.7B student, with validation accuracy, gradient norm, and policy entropy reported. (a) Correct answer
(b) Incorrect answer $$ \sum_{k=1}^{6} \tan^2
We are given a function $ f(x, y, z) = xyz $ and a region $W$ At(0)
\frac{(7-1)(7-2)}{3} defined by the
inequalities $ 0 ......
\frac{(7-1)(7-2)}{3}
-1.0
\right) =
= \frac{6 \cdot 5}{3} = \frac{30}{3}
$$ \sum_{k=1}^{6} \tan^2
We are given a function $ f(x, y, z) = xyz $ and a region $W$
inequalities
\frac{k\pi}{7}
= 10 $$ --- ### Final Answer: $$ \boxed{10} $$<|im_end|>
At(0.99) defined by the
\left(
$ 0 ......
\left(
\frac{k\pi}{7}
\right) =
= \frac{6 \cdot 5}{3} = \frac{30}{3}
= 10 $$ --- ### Final Answer: $$ \boxed{10} $$<|im_end|>
-0.5
0.0
0.5
1.0
At
Figure 5: Token-level advantage visualization. (a) Correct response; (b) incorrect response. The (0) (0.99) top row shows normalized At , and the bottom row shows normalized At . Blue and red denote negative and positive advantages, respectively, with color intensity indicating magnitude. Complete response examples are provided in Appendix F.2.
steps that are more relevant to the final outcome, yielding smoother and more semantically aligned token-level signals. Complete response-level visualizations are provided in Appendix F.2.
5
R ELATED W ORK
Credit Assignment for OPD. Recent studies have explored credit assignment in OPD, motivated by the potentially unreliable reasoning trajectories generated by relatively weak student policies(Liu et al., 2026; Yu et al., 2026). Truncation-based methods (Zhou et al., 2026; Zhang et al., 2026) terminate rollouts early or mask the learning signals of tokens beyond selected positions, thereby reducing the influence of unreliable continuations. Entropy-aware OPD (Jin et al., 2026) augments reverse KL with forward KL on tokens where the teacher has high entropy, while ExOPD (Yang et al., 2026a) introduces a reference model to calibrate the OPD advantage. Fu et al. (2026) analyze the bias–variance trade-off between token-level and sequence-level reverse-KL estimators and propose teacher top-K local-support matching, yet a principled balance between sequence-level objective fidelity and token-level optimization stability remains unresolved. Reward-Guided OPD. Recent studies incorporate verifiable rewards into OPD to complement teacher-derived supervision. REOPOLD (Ko et al., 2026) combines reward clipping, entropy-based sampling, and exploration-to-refinement scheduling, while SCOPE (Zheng et al., 2026) routes incorrect trajectories to teacher-perplexity-weighted KL distillation and correct trajectories to studentperplexity-weighted MLE. AOPD (Jia et al., 2026) preserves positive reinforcement while replacing non-positive token updates with localized teacher imitation, whereas RWOPD (Zou et al., 2026) weights teacher KL gradients using verifier rewards. Other works (Ding et al., 2026; Zhan et al., 2026) further stabilize optimization through advantage compression and warmup-then-anneal scheduling. Nevertheless, these methods often introduce additional complexity through group-based 9
Preprint
advantage estimation, log-ratio correction, or verifier-weighted teacher KL objectives, potentially incurring substantial computational overhead.
6
C ONCLUSION
In this work, we unified token-level and sequence-level OPD from a temporal credit assignment perspective and proposed γOPD to balance long-horizon supervision and optimization stability. We further introduced RBM to combine discounted teacher guidance with verifiable outcome rewards. Empirical results on mathematical and code reasoning demonstrate that γOPD remains effective across different teacher–student configurations and distillation scenarios. Due to computational constraints, our experiments are currently limited to models with fewer than 10B parameters. Evaluating γOPD at larger scales and exploring more flexible temporal credit assignment schemes, such as adaptive or entropy-aware discounting, are promising directions for future work.
R EFERENCES Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from selfgenerated mistakes. In International Conference on Learning Representations, 2024. Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Yuchen Zhang, Jiacheng Chen, et al. Process reinforcement through implicit rewards. Transactions on Machine Learning Research, 2026. URL https://arxiv.org/abs/2502.01456. DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, et al. DeepSeek-V4: Towards highly efficient million-token context intelligence, 2026. URL https://arxiv.org/abs/ 2606.19348. Yifan Ding, Xincheng Wei, Yoshua Y. Li, Ziheng Li, Yuquan Lu, Siyu Zhang, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, and Yun Chen. SAF-OPD: Stable advantage fusion for OnPolicy Distillation. arXiv preprint arXiv:2607.29209, 2026. URL https://arxiv.org/ abs/2607.29209. Yuqian Fu, Haohuan Huang, Kaiwen Jiang, Jiacai Liu, Zhuo Jiang, Yuanheng Zhu, and Dongbin Zhao. Revisiting On-Policy Distillation: Empirical failure modes and simple fixes. arXiv preprint arXiv:2603.25562, 2026. URL https://arxiv.org/abs/2603.25562. GLM-5 Team. GLM-5: from Vibe Coding to Agentic Engineering, 2026. URL https://arxiv. org/abs/2602.15763. Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, et al. DeepSeekR1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645(8081):633–638, 2025. doi: 10.1038/s41586-025-09422-z. Bingxiang He, Zekai Qu, Zeyuan Liu, Yinghao Chen, Yuxin Zuo, Cheng Qian, Kaiyan Zhang, Weize Chen, Chaojun Xiao, Ganqu Cui, Ning Ding, and Zhiyuan Liu. JustRL: Scaling a 1.5B LLM with a Simple RL Recipe. In ICLR Blogposts 2026, 2026. Zhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu, Xingyu Chen, Yue Wang, Linfeng Song, Dian Yu, Zhenwen Liang, Wenxuan Wang, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu. DeepMath-103K: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning. arXiv preprint arXiv:2504.11456, 2025. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. 10
Preprint
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. LiveCodeBench: Holistic and contamination free evaluation of large language models for code. In International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=chfJJYC3iL. Nan Jia, Haojin Yang, Xing Ma, Jiesong Lian, Shuailiang Zhang, Weipeng Zhang, Ke Zeng, Xunliang Cai, and Zequn Sun. Asymmetric On-Policy Distillation: Bridging exploitation and imitation at the token level. arXiv preprint arXiv:2605.06387, 2026. doi: 10.48550/arXiv.2605.06387. URL https://arxiv.org/abs/2605.06387. Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, and Kimin Lee. Entropy-aware On-Policy Distillation of language models. In Proceedings of the 43rd International Conference on Machine Learning, volume 306. PMLR, 2026. doi: 10.48550/arXiv.2603.07079. URL https://arxiv.org/abs/2603.07079. Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, et al. Kimi K3: Open Frontier Intelligence, 2026. URL https://arxiv.org/abs/2607.24653. Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen, and Pashmina Cameron. Scaling reasoning efficiently via relaxed on-policy distillation, 2026. URL https://arxiv.org/abs/2603. 11137. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023. Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, et al. NuminaMath: The largest public dataset in AI4Maths with 860k pairs of competition math problems and solutions. Hugging Face dataset, 2024. URL https://huggingface.co/datasets/AI-MO/NuminaMath-CoT. Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huanang Gao, Wenkai Yang, Zhiyuan Liu, and Ning Ding. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe, 2026. URL https://arxiv. org/abs/2604.13016. Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Advances in Neural Information Processing Systems, volume 36, pp. 21558–21572, 2023. URL https://openreview.net/forum?id=1qvx610Cu7. Shiqi Liu, Zeyu He, Guojian Zhan, Letian Tao, Zhilong Zheng, Jiang Wu, Yinuo Wang, Yang Guan, Kehua Sheng, Bo Zhang, Keqiang Li, Jingliang Duan, and Shengbo Eben Li. STAPO: Stabilizing reinforcement learning for LLMs by silencing rare spurious tokens. arXiv preprint arXiv:2602.15620, 2026. Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. URL https://thinkingmachines.ai/blog/ on-policy-distillation. Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi, and Yohan Jo. KL for a KL: On-Policy Distillation with control variate baseline. arXiv preprint arXiv:2605.07865, 2026. doi: 10.48550/ arXiv.2605.07865. URL https://arxiv.org/abs/2605.07865. Presented as a poster at the AI for Math Workshop, ICML 2026. OpenCompass. AIME2025 dataset. https://huggingface.co/datasets/ opencompass/AIME2025, 2025. Accessed: 2026-08-09. Zhixin Ren, Yau Lyu, Congrong Li, Liping Zhang, and Shengbo Eben Li. Momentum as residualdriven multiplier correction for deep learning optimization, 2026. URL https://arxiv. org/abs/2608.12925. 11
Preprint
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. URL https://arxiv.org/abs/2402.03300. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. HybridFlow: A flexible and efficient RLHF framework. In Proceedings of the Twentieth European Conference on Computer Systems, pp. 1279–1297, 2025. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. URL https://arxiv.org/abs/2505.09388. Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, and Yankai Lin. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation. arXiv preprint arXiv:2602.12125, 2026a. URL https://arxiv.org/abs/2602.12125. Zhuolin Yang, Zihan Liu, Yang Chen, Wenliang Dai, Boxin Wang, Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He, et al. Nemotron-Cascade 2: Post-training LLMs with cascade RL and multi-domain on-policy distillation. arXiv preprint arXiv:2603.19220, 2026b. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, et al. DAPO: An open-source LLM reinforcement learning system at scale. In Advances in Neural Information Processing Systems, volume 38, pp. 125532–125554, 2025. doi: 10.52202/085713-3775. Zhouyang Yu, Guojian Zhan, Yang Guan, Jingliang Duan, Letian Tao, and Shengbo Eben Li. Taming aleatoric impulse in off-policy reinforcement learning. In Proceedings of the 43rd International Conference on Machine Learning, 2026. Guojian Zhan, Xiangteng Zhang, Feihong Zhang, Letian Tao, and Shengbo Eben Li. Bicriteria policy optimization for high-accuracy reinforcement learning. IEEE Transactions on Neural Networks and Learning Systems, 37(1):312–326, 2026. doi: 10.1109/TNNLS.2025.3605362. Yaocheng Zhang, Jiajun Chai, Yuqian Fu, Songjun Tu, Xiaohan Wang, Wei Lin, Guojun Yin, Qichao Zhang, Yuanheng Zhu, and Dongbin Zhao. Are full rollouts necessary for On-Policy Distillation? arXiv preprint arXiv:2605.31490, 2026. doi: 10.48550/arXiv.2605.31490. URL https:// arxiv.org/abs/2605.31490. Binbin Zheng, Xing Ma, Yiheng Liang, Jingqing Ruan, Xiaoliang Fu, Kepeng Lin, Benchang Zhu, Ke Zeng, and Xunliang Cai. SCOPE: Signal-calibrated On-Policy Distillation enhancement with dual-path adaptive weighting. arXiv preprint arXiv:2604.10688, 2026. doi: 10.48550/arXiv.2604. 10688. URL https://arxiv.org/abs/2604.10688. Ziheng Zhou, Jiaqi Li, Huacong Tang, Ying Nian Wu, and Demetri Terzopoulos. Less is more: Early stopping rollout for On-Policy Distillation. arXiv preprint arXiv:2605.27028, 2026. doi: 10.48550/arXiv.2605.27028. URL https://arxiv.org/abs/2605.27028. Qingyun Zou, Yingze Li, Tianen Liu, Bingsheng He, and Weng-Fai Wong. Reward-weighted OnPolicy Distillation with an open property-equivalence verifier for NL-to-SVA generation. arXiv preprint arXiv:2605.13501, 2026. doi: 10.48550/arXiv.2605.13501. URL https://arxiv. org/abs/2605.13501.
12
Preprint
A PPENDIX A
P ROOF OF P ROPOSITION 3.1
Starting from the token-level OPD objective in equation 4b, we have " T # X token ∗ JOPD (θ) = Ex∼D Eht ∼dθ̄ (·|x) [DKL (πθ (· | ht ) ∥ π (· | ht ))] . t=1
Since the prefix distribution dθ̄ (ht | x) is treated with stop-gradient, it is fixed when differentiating with respect to θ. Therefore, the gradient can be moved inside the expectation over prefixes: " T " ## X X token ∗ ∇θ JOPD (θ) = Ex∼D Eht ∼dθ̄ (·|x) ∇θ πθ (a | ht ) (log πθ (a | ht ) − log π (a | ht )) . t=1
a∈V
(10) For a fixed prefix ht , define ∆(a; ht ) ≜ log πθ (a | ht ) − log π ∗ (a | ht ). This is the full-vocabulary counterpart of the sampled token-level log-ratio ∆t in equation 1a. Then the gradient of the inner next-token KL term is X ∇θ πθ (a | ht )∆(a; ht ) a∈V
=
X
∇θ πθ (a | ht )∆(a; ht ) +
a∈V
X
(11) πθ (a | ht )∇θ log πθ (a | ht ).
a∈V
The second term in equation 11 vanishes because X X X πθ (a | ht )∇θ log πθ (a | ht ) = ∇θ πθ (a | ht ) = ∇θ πθ (a | ht ) = 0. a∈V
a∈V
a∈V
Using the score-function identity ∇θ πθ (a | ht ) = πθ (a | ht )∇θ log πθ (a | ht ), we obtain ∇θ DKL (πθ (· | ht ) ∥ π ∗ (· | ht )) =
X
πθ (a | ht )∆(a; ht )∇θ log πθ (a | ht ) (12)
a∈V
= Eyt ∼πθ (·|ht ) [∆t ∇θ log πθ (yt | ht )] , where ∆t = ∆(yt ; ht ) follows from equation 1a. Substituting equation 12 into equation 10 gives " T # X token ∇θ JOPD (θ) = Ex∼D Eht ∼dθ̄ (·|x), yt ∼πθ (·|ht ) [∆t ∇θ log πθ (yt | ht )] .
(13)
t=1
When the stop-gradient prefix distribution is evaluated at the current policy, the prefix ht together with the next token yt can be generated by an on-policy trajectory y ∼ πθ (· | x), while gradients are not propagated through the prefix-sampling distribution. Hence, equation 13 can be written as " T # X token ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t ∇θ log πθ (yt | ht ) , t=1
which is exactly the token-level OPD gradient in equation 5. This proves Proposition 3.1. 13
Preprint
B
P ROOF OF P ROPOSITION 3.2
Starting from the sequence-level OPD objective in equation 4a, we expand the reverse KL over complete response trajectories: X seq JOPD (θ) = Ex∼D πθ (y | x)∆(y) , y∈V T
where ∆(y) is the sequence-level log-ratio defined in equation 1b. Taking the gradient with respect to θ gives X X seq ∇θ JOPD (θ) = Ex∼D ∇θ πθ (y | x)∆(y) + πθ (y | x)∇θ log πθ (y | x) . (14) y∈V T
y∈V T
The second term in equation 14 vanishes because X X πθ (y | x)∇θ log πθ (y | x) = ∇θ πθ (y | x) = ∇θ 1 = 0. y∈V T
y∈V T
Using the score-function identity, ∇θ πθ (y | x) = πθ (y | x)∇θ log πθ (y | x), we obtain seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) [∆(y)∇θ log πθ (y | x)] .
(15)
By the autoregressive factorization and the log-ratio decomposition in equation 1, we have ∆(y) =
T X
∆ t′ ,
∇θ log πθ (y | x) =
t′ =1
T X
∇θ log πθ (yt | ht ).
(16)
t=1
Substituting equation 16 into equation 15 yields seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x)
" T T XX
# ∆t′ ∇θ log πθ (yt | ht ) .
(17)
t=1 t′ =1
It remains to remove the terms with t′ < t. For t′ < t, ∆t′ is determined by earlier tokens and is therefore measurable with respect to ht . Hence, conditioning on ht , Eyt ∼πθ (·|ht ) [∆t′ ∇θ log πθ (yt | ht )] = ∆t′ Eyt ∼πθ (·|ht ) [∇θ log πθ (yt | ht )] = 0, where the last equality follows from the token-level score identity X Eyt ∼πθ (·|ht ) [∇θ log πθ (yt | ht )] = ∇θ πθ (yt | ht ) = 0. yt ∈V
Therefore, all terms with t′ < t in equation 17 vanish in expectation. Keeping only the terms with t′ ≥ t, we obtain " T ! # T X X seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t′ ∇θ log πθ (yt | ht ) , t=1
t′ =t
which is exactly the sequence-level gradient in equation 6. This proves Proposition 3.2. 14
Preprint
C
P ROOF OF T HEOREM 3.3
Proof. Let Nt = T − t + 1 and define h i (0) (0) e(0) A . τ = Aτ − E Aτ Since e(0) A τ e(0) we have A τ
2
2 2
2 (0) 2 ≤ E A ≤ σ∆ , = Var A(0) τ τ
≤ σ∆ .
For the sequence-level OPD credit, the triangle inequality in L2 gives r
T X (1) e(0) Var At = A τ τ =t
≤
T X
2
e(0) A τ
τ =t
2
≤ Nt σ∆ .
Therefore, (1) 2 2 Var At ≤ Nt2 σ∆ = (T − t + 1)2 σ∆ . (0)
This quadratic dependence is attainable in the worst case. In particular, if Aτ 2 , then {t, . . . , T }, where E[Z] = 0 and Var(Z) = σ∆
= Z for all τ ∈
(1) 2 Var At = Var(Nt Z) = Nt2 σ∆ . For the γOPD credit, similarly, r
T X (γ) e(0) Var At = γ τ −t A τ τ =t
≤ σ∆
2
T X
γ τ −t
τ =t
= σ∆
σ∆ 1 − γ Nt ≤ , 1−γ 1−γ
∀γ ∈ [0, 1).
Squaring both sides yields (γ) Var At ≤
2 σ∆ , (1 − γ)2
∀γ ∈ [0, 1).
Thus, sequence-level OPD can exhibit variance that grows quadratically with the remaining horizon, whereas γOPD admits a horizon-independent variance bound for every fixed γ < 1.
D
A LGORITHM
The complete reward-enhanced γOPD procedure is summarized in Algorithm 1. 15
Preprint
Table 4: Training hyperparameters for the OPD-based methods. The RLVR baselines use a rollout group size of 8. Hyperparameter Value Training framework Rollout backend Math reward function Code reward function Train batch size Responses per prompt PPO mini batch size PPO micro batch size/GPU Max prompt length Max response length Optimizer Weight decay Learning rate Training temperature / top-p Validation temperature / top-p Rollout tensor parallel size
veRL vLLM DAPO boxed verifier Execution-based verifier 1024 1 1024 1 2048 16384 RADAR 0.01 1 × 10−5 1.0 / 1.0 1.0 / 1.0 4
Algorithm 1 Reward-Enhanced Temporal-Credit On-Policy Distillation (γOPD) Require: Dataset D, student policy πθ , teacher policy π ∗ , discount factor γ, batch size B 1: for each training iteration do B 2: Sample prompts {xi }B i=1 ∼ D and responses {yi }i=1 ∼ πθ (· | xi ) 3: Evaluate the verifier to obtain {Ri }B , where R ← R(xi , yi ) i i=1 4: for each mini-batch do 5: for each response (x, y, R) in the mini-batch do 6: for t = 1, . . . , T do 7: ∆t ← log πθ (yt | ht ) − log π ∗ (yt | ht ) 8: end for (γ) 9: Compute {Ât }Tt=1 according to equation 7 and equation 8 10: end for 11: Update θ using equation 9 12: end for 13: end for 14: return final student policy πθ
E
E XPERIMENT D ETAILS
All methods, including γOPD, are implemented with veRL (Sheng et al., 2025), using vLLM (Kwon et al., 2023) for rollouts and FSDP for actor training. For mathematical reasoning, we train on DeepMath-103K (He et al., 2025), retaining examples with difficulty level at least 6. Each prompt is formatted in chat style and appended with “Please output the final answer within \boxed{}.” For code reasoning, we train on Eurus-RL-Code (Cui et al., 2026). For both domains, verifier outcomes are mapped to binary rewards in {+1, −1}. Math responses are evaluated with a DAPO-style boxed-answer verifier (Yu et al., 2025), while code responses are verified through execution-based unit tests. In the multi-teacher setting, each sample is routed to its domain-specific teacher and verifier while sharing the same student policy. OPD uses reverse-KL token advantages without an additional KL reward penalty. Unless otherwise specified, all models use RADAR (Ren et al., 2026) optimizer, a constant learning rate on 32 NVIDIA H20 GPUs, with each training run taking approximately 1–3 days; full hyperparameters are provided in Table 4. 16
Preprint
Table 5: Computational and memory overhead of the additional γOPD operations. Operation Time (ms/step) Memory (MB) Discounted temporal credit Bounded advantage shaping Verifier-reward mixing
609.85 0.42 0.04
0.063 0.125 0.188
Total additional overhead
610.32
0.375
573,669.00
81,254.40
0.1064%
0.00046%
Vanilla OPD update (reference) Relative overhead
We use temperature 1.0 and top-p 1.0 for all evaluations. For mathematical reasoning, we evaluate on AIME24 (Li et al., 2024), AIME25 (OpenCompass, 2025), AMC23 (Li et al., 2024), and MATH500 (Hendrycks et al., 2021), with a maximum prompt length of 2048 and response length of 16384. Following the repeated-sampling convention, AIME24, AIME25, and AMC23 are evaluated with 32 samples per problem, while MATH500, MinervaMath, and OlympiadBench use 4 samples per problem; we report both accuracy and pass rate. All predictions are scored using the same boxed-answer verifier as in training. For code reasoning, we evaluate HumanEval+ and MBPP+ (Liu et al., 2023) with EvalPlus, and the v6 split of LiveCodeBench (Jain et al., 2025). HumanEval+ and MBPP+ use one sample per problem, whereas LiveCodeBench uses four samples with a maximum generation length of 16384 tokens.
F
S UPPLEMENTARY E XPERIMENTAL R ESULTS
F.1
C OMPUTATIONAL T IME AND M EMORY A NALYSIS
We profile the additional computation introduced by γOPD using the same 4B student/teacher models and 32-GPU setup as the main experiments. Statistics are averaged over training steps 2–100 after warm-up. The extra operations include discounted temporal-credit computation, bounded advantage shaping, and verifier-reward mixing. Overall, γOPD adds only 0.1064% computation time and 0.00046% memory relative to a full OPD update, indicating negligible training overhead. F.2
C OMPLETE T OKEN -L EVEL A DVANTAGE V ISUALIZATION
This appendix provides the complete token-level visualizations behind Figure 5. Each response is shown as a sequence of consecutive full-width panels rather than as a collection of small subfigures, so that individual tokens and advantage magnitudes remain readable. The upper row in every panel (0) shows the local advantage At , while the lower row shows the normalized temporally propagated (0.99) advantage At . (0)
(0.99)
As shown in Figure 6, for the correct response, At gives a sparse local signal, whereas At propagates information from later reasoning steps and assigns stronger positive credit to the key (0) mathematical derivation. As shown in Figure 7, for the incorrect response, At penalizes tokens (0.99) largely unrelated to the actual error, while At assigns stronger negative credit to the erroneous formula derivation without excessively penalizing the final end-of-sequence token.
17
Preprint
At(0) At(0.99)
are are
given given
a a
defined defined
by by
the the
We We
need need
to to
find find
average average by: by:
value value
over over
y, y,
of of
$ W $ W
W W
$. $.
$ $
x x
\le \le
1 1
(since (since
to to
set set
set set
up up
This This
order order
for for
as: as:
$ $ $ $
x x
z z
can can $$ $$
\int_{z=0}^y \int_{z=0}^y
W W
$ $
$ $
y y
range range
$ $
$ $
z z
0 0
to to
xyz xyz xyz xyz
$ $
\, \,
z z
$, $,
y y
dz dz
\, \,
$ $
given given
is is
the the
y y
$, $,
$ $
\le \le
\le \le
in in
order order
is: is:
z z
Let's Let's
terms terms
$, $,
we we
from from
0 0
to to
$ $
y y
$. $.
So, So,
dV dV
= =
the the
can can
fix fix
$ $
$, $,
and and
the the
dx dx
$$ $$
of of
z z
y y
\le \le octant octant
1 1
the the
$. $.
So So
we we
limits. limits.
Since Since
the the
can can
choose choose
an an
z z
$, $,
$ $
y y
$, $,
$ $
determine determine order order x x
integral integral
of of
$, $,
for for
\int_{x=0}^1 \int_{x=0}^1 \, \,
$ $
here. here.
$ $
of of
x x
f f
integration integration
we we
$. $.
$ $
and and x x
$, $,
W} W} the the
first first
\le \le
1 1
} }
of of
0 0
the the
x x
18
$$ $$
express express
\le \le
The The
bounds bounds
appropriate appropriate
x x
We We
compute compute
with with
\le \le
dy dy
to to
important important
\le \le
$. $.
of of
the the
y y
\le \le
is is
think think
\, \,
$ W $ $ W $
integral integral
in in
possible possible
x x
$ $
1 1
region. region.
non-negative), non-negative),
y y
range range
from from
$ W $ $ W $
\le \le
this this
by: by:
region region
can can
\le \le
\le \le
x x
need need
defined defined
a a
let's let's
can can
\iiint_W \iiint_W
is is
A A
But But
\le \le
Determine Determine
integration integration
$ $
Since Since $, $,
triple triple
as as
variables. variables.
easier. easier.
variables. variables. each each
the the
we we
Alternatively, Alternatively, is is
compute compute
integral integral
$ $
region region
a a
region region
a a
are are
of of
the the
over over
we we
triple triple
by by
$ $
over over
1: 1:
integral, integral,
bounded bounded
$ $
f f
\le \le
first, first,
ordered ordered
the the
is is
$ $
f f
defines defines
are are
region region
of of
and and y y
z z
So, So,
variables variables
order order
z z
then then
$ $
the the
$. $.
$$ $$
region region
The The
which which
$ $
dV dV
The The
order. order.
x x
\, \,
Step Step
up up
\le \le
$ $
\frac{1}{\text{Volume \frac{1}{\text{Volume
### ###
inequalities inequalities
At(0) At(0.99)
$ $
-----
all all
0 0
xyz xyz
= =
= =
and and
$$ $$
z) z)
value value
function function
z) z)
$, $,
integration integration
To To
a a
y, y, $ $
average average of of
f(x, f(x,
volume volume
need need
f(x, f(x,
\text{Average} \text{Average}
\iiint_W \iiint_W
\le \le
$ $
inequalities inequalities the the
$$ $$
At(0) At(0.99)
function function
then then each each
can can
the the
be be
for for
$ $ set set
y y
$, $, up up
\int_{y=0}^x \int_{y=0}^x Alternatively, Alternatively,
we we
can can
Preprint
At(0) At(0.99)
reverse reverse
the the
order order
order order
for for
now. now.
First, First,
integration, integration,
of of -----
### ###
compute compute
the the xyz xyz
\int_{z=0}^y \int_{z=0}^y
over over
$ $
y y
$: $:
dy dy
$$ $$ is is
the the
The The
volume volume
function function
$ $
\iiint_W \iiint_W
$$ $$ 1 1
\, \,
\, \, dz dz
\left[ \left[
y, y,
dV dV
= =
dy dy
\, \,
dx dx
y y
$$ $$
= =
\frac{1}{3} \frac{1}{3}
= =
$$ $$
\frac{1}{6} \frac{1}{6}
$. $.
-----
$$ $$
Step Step
19
4: 4:
over over
$ $
So So
$ $
Compute Compute
the the
integral integral
of of
triple triple $ $
xyz xyz
= =
W W
$, $,
which which
\int_{y=0}^x \int_{y=0}^x inner inner
integral: integral:
\int_{y=0}^x \int_{y=0}^x
y y $$ $$
dy dy
\frac{1}{2} \frac{1}{2}
\cdot \cdot
= =
\frac{1}{2} \frac{1}{2}
\cdot \cdot
the the
volume volume
Compute Compute
the the
of of
$ $
average average
W W
= =
Then: Then:
= =
\right]_0^1 \right]_0^1
### ###
3: 3:
= =
z) z)
\frac{x^2}{2} \frac{x^2}{2}
= = dx dx
$$ $$
Step Step
1 1
y, y,
f(x, f(x,
Compute Compute
Then: Then:
\frac{1}{6} \frac{1}{6}
\frac{x^6}{6} \frac{x^6}{6}
$ $
the the = =
= =
\int_{x=0}^1 \int_{x=0}^1
$$ $$
\int_{x=0}^1 \int_{x=0}^1
\frac{x^2}{2} \frac{x^2}{2}
\frac{x^3}{3} \frac{x^3}{3}
is is
integral integral
\frac{1}{6} \frac{1}{6}
### ###
z) z)
\right]_0^x \right]_0^x
\int_{x=0}^1 \int_{x=0}^1
$$ $$
\, \,
\frac{y^2}{2} \frac{y^2}{2}
\left[ \left[
$ W $ $ W $
f(x, f(x,
1 1
dz dz 1 1
of of
$: $:
of of
-----
= =
\frac{x}{2} \frac{x}{2}
\left[ \left[
integral integral $. $.
\left[ \left[
\frac{y^4}{4} \frac{y^4}{4}
\cdot \cdot
\frac{1}{48} \frac{1}{48}
$ W $ $ W $
x x
\cdot \cdot
triple triple
$$ $$
\frac{x^4}{4} \frac{x^4}{4} $ $
\frac{1}{8} \frac{1}{8}
So So
\int_{z=0}^y \int_{z=0}^y
$$ $$
\frac{1}{8} \frac{1}{8}
= =
xy xy
= =
\left[ \left[
over over
$: $:
the the
\cdot \cdot
integrate integrate
Now, Now,
$ $
\int_{z=0}^y \int_{z=0}^y
\cdot \cdot
\frac{x}{2} \frac{x}{2}
= =
constant constant
be: be:
\frac{x}{2} \frac{x}{2}
= =
= =
z z
\frac{y^2}{2} \frac{y^2}{2}
\cdot \cdot
y^3}{2} y^3}{2}
dy dy
this this integral integral
$ $
dz dz
\frac{x \frac{x
\frac{1}{48} \frac{1}{48}
should should
xy xy
= =
\, \,
\int_{y=0}^x \int_{y=0}^x
\right]_0^1 \right]_0^1
At(0) At(0.99)
z z
$$ $$
dx dx
the the
\int_0^y \int_0^y
into into
\frac{x^5}{8} \frac{x^5}{8}
of of
xy xy
= =
this this
$$ $$
volume volume
dz dz
with with
over over
substitute substitute
\frac{x^5}{8} \frac{x^5}{8}
$ $
triple triple
Now, Now,
y^3 y^3
W W
the the
$$ $$
= =
$ $
Compute Compute
y^3}{2} y^3}{2}
\right]_0^x \right]_0^x
over over
stick stick
integral integral
\right]_0^y \right]_0^y
\int_{0}^x \int_{0}^x
At(0) At(0.99)
2: 2:
let's let's
inner inner
\, \,
\frac{z^2}{2} \frac{z^2}{2} \frac{x \frac{x
Step Step
but but
$ $
is is value value
$ $
Preprint
At(0) At(0.99)
average average
Now, Now,
value value
\text{Triple \text{Triple = =
6 6
$ $
is: is:
$$ $$
is is
\frac{1}{1/6} \frac{1}{1/6}
= =
\frac{1}{48} \frac{1}{48}
let let
xyz xyz
\frac{1}{\text{Volume}} \frac{1}{\text{Volume}}
$$ $$
integral} integral}
\cdot \cdot
Wait, Wait,
is: is:
me me $ $
check check
1/48 1/48
that that
$, $,
the the
\frac{1/48}{1/6} \frac{1/48}{1/6}
\frac{6}{48} \frac{6}{48} Final Final
\frac{6}{48} \frac{6}{48}
again: again:
Wait, Wait, is is
\frac{1}{8} \frac{1}{8}
$$ $$
the the 1/6 1/6
$, $,
so so
-0.5
of of
average average
the the \cdot \cdot
Yes, Yes,
$$ $$
integral integral
triple triple
that's that's
$ $
value value
\frac{6}{1} \frac{6}{1}
= =
-----
### ###
correct. correct.
\boxed{\dfrac{1}{8}} \boxed{\dfrac{1}{8}}
$$ $$
-1.0
\frac{1}{8} \frac{1}{8}
= =
\frac{1}{48} \frac{1}{48}
= =
= =
Answer: Answer:
$ $
\frac{1}{48} \frac{1}{48}
\cdot \cdot
= =
volume volume
\times \times
$$<|im_end|> $$<|im_end|>
0.0
0.5
1.0
At
Figure 6. Complete token-level advantage visualization for the correct response. The response is (0.99) (0) in the lower shown in token order across seven consecutive panels, with At in the upper row and At row. Blue and red indicate negative and positive advantages, respectively. At(0) At(0.99)
We We
are are
asked asked
to to
evaluate evaluate
the the
\frac{k\pi}{7} \frac{k\pi}{7}
\left( \left(
\right) \right)
trigonometric trigonometric the the
spaced spaced
angles: angles:
sum sum
of of
small small
$ $
n n
squares squares
for for
-----
### ###
$. $.
### ### is is
Step Step
a a
tangent tangent
of of
integer integer 2: 2:
1: 1:
Use Use
$ $
n n
> >
2 2
Apply Apply
the the
a a
known known
identity identity
functions functions
equally equally
at at
\tan^2 \tan^2
\left( \left(
\frac{(n-1)(n-2)}{3} \frac{(n-1)(n-2)}{3}
= =
Step Step
\tan^2 \tan^2
useful useful
\sum_{k=1}^{n-1} \sum_{k=1}^{n-1}
\right) \right) works works
-----
There There
$$ $$
\frac{k\pi}{n} \frac{k\pi}{n}
\sum_{k=1}^{6} \sum_{k=1}^{6}
$$ $$
$$ $$
identity identity
involving involving
identity identity
sum: sum:
$. $.
$$ $$ verify verify
Let's Let's identity identity
In In
our our
This This it it
for for
case, case,
At(0)
n
=
7
$,
so
the
sum
becomes:
$$
\sum_{k=1}^{6}
\tan^2
\left(
At(0.99)
n
=
7
$,
so
the
sum
becomes:
$$
\sum_{k=1}^{6}
\tan^2
\left(
\frac{k\pi}{7}
\right)
=
\frac{(7-1)(7-2)}{3}
=
\frac{6
\frac{k\pi}{7}
\right)
=
\frac{(7-1)(7-2)}{3}
=
\frac{6
$ $
\cdot
5}{3}
=
\frac{30}{3}
=
10
$$
---
###
Final
Answer:
$$
\cdot
5}{3}
=
\frac{30}{3}
=
10
$$
---
###
Final
Answer:
$$
-1.0
\boxed{10}
$$<|im_end|>
\boxed{10}
$$<|im_end|> -0.5
0.0
0.5
1.0
At
Figure 7. Complete token-level advantage visualization for the incorrect response. The two consec(0) (0.99) utive panels show how At can penalize tokens unrelated to the actual error, whereas At assigns stronger negative credit to the erroneous formula derivation.
20