Conceptio › Archive › arXiv CS
arXiv CSopen access

Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Preprint

B EYOND T OKEN -L OCAL I MITATION : R EWARD -C OMPATIBLE T EMPORAL C REDIT A SSIGNMENT FOR O N -P OLICY D ISTILLATION Shiqi Liu1,2 , Zeyu He1,2 , Letian Tao1,2 , Guojian Zhan1,2 , Jiaxin Gao1 , Feihong Zhang1 , Jingliang Duan1,2 , Wei Xiong2 , Kehua Sheng2 , Bo Zhang2 , Yang Guan1,B , Shengbo Eben Li1,B

arXiv:2609.16937v1 [cs.LG] 15 Sep 2026

1

School of Vehicle and Mobility & College of AI, Tsinghua University 2 Didi Voyager Labs, DiDi Autonomous Driving

A BSTRACT On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost of horizon-dependent variance. We establish a unified temporal-credit view of these formulations, showing that practical token-level OPD can be interpreted as a temporal approximation to the sequence-level reverse-KL gradient. Building on this connection, we propose γOPD, which uses discounted temporal credit assignment to balance long-horizon supervision and optimization stability, while admitting a horizon-independent variance bound. We further develop a rewardcompatible bounded mixing (RBM) mechanism for γOPD that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization. Experiments on mathematical and code reasoning demonstrate consistent improvements over existing OPD methods across vanilla, size-mismatched, and multi-teacher distillation settings.

(a)

(b)

Figure 1: Core Idea. (a) Vanilla OPD (left) focuses only on the immediate next step and closely follows the teacher’s local guidance. We argue that OPD should instead account for the influence of future reasoning steps and should not merely imitate the teacher token by token (right). (b) By incorporating future outcomes through temporal credit assignment and balancing teacher supervision with verifiable rewards, γOPD achieves strong overall performance across mathematical and code reasoning tasks in the multi-teacher distillation setting, as detailed in Section 4.

1

I NTRODUCTION

Recent large language models (LLMs), including Kimi K3 (Kimi Team et al., 2026), GLM-5 (GLM5 Team, 2026), DeepSeek-V4 (DeepSeek-AI et al., 2026), and Nemotron-Cascade 2 (Yang et al., 2026b), have demonstrated strong reasoning capabilities across mathematics, coding, and science. B

Corresponding author: S. E. Li and Y. Guan; email: [email protected].

1

Preprint

To further consolidate and integrate such capabilities, on-policy distillation (OPD) (Lu & Thinking Machines Lab, 2025) has emerged as a key approach in LLM post-training. It provides dense teacher supervision on trajectories sampled from the current student policy, promoting high-quality reasoning while mitigating exposure bias. To avoid costly full-vocabulary reverse-KL computation, existing OPD methods (Yang et al., 2026a; Jin et al., 2026; Oh et al., 2026) reformulate the objective in a policy-gradient form and approximate it using a one-sample token-level Monte Carlo estimator (Li et al., 2026). While computationally efficient, this practical estimator introduces non-negligible bias relative to the original sequence-level objective, effectively altering the optimization problem (Yang et al., 2026a). Other methods (Fu et al., 2026) instead adopt sequence-level Monte Carlo estimators to more faithfully preserve the original objective. However, accumulating future credit over the remaining response leads to increasing variance as the sequence grows, resulting in unstable optimization for long reasoning trajectories. In this work, we present a unified analysis of token-level and sequence-level OPD, deriving their underlying objectives and clarifying the relationship between their policy-gradient formulations. Building on this insight, we propose γOPD, which introduces discounted temporal credit assignment to interpolate between token-local and sequence-level supervision while admitting a horizonindependent variance bound. To complement teacher supervision with outcome-level guidance, we further develop a reward-compatible bounded mixing mechanism that incorporates verifiable outcome feedback without sacrificing the stability benefits of temporal discounting. Overall, our main contributions are summarized as follows: • We establish a unified temporal-credit view of OPD, showing that practical token-level OPD can be interpreted as a temporally truncated approximation to the sequence-level reverse-KL gradient. • Building on this connection, we propose γOPD, a discounted temporal-credit surrogate that interpolates between token-level and sequence-level credit assignment, while admitting a horizon-independent variance bound. • We further develop Reward-Compatible Bounded Mixing (RBM) for γOPD, which combines discounted teacher-derived credit with verifiable outcome rewards while preventing the teacher signal from dominating task-level supervision.

2

P RELIMINARIES

2.1

N OTATION

We consider autoregressive reasoning tasks over a discrete vocabulary V, where we aim to optimize a student policy πθ parameterized by θ under the guidance of a fixed teacher policy π ∗ . Let x ∼ D denote a prompt sequence sampled from the reasoning dataset. Given x, the policy generates a complete response trajectory y = (y1 , y2 , . . . , yT ) ∈ V T , where T = |y| is the number of tokens in the trajectory, and yt ∈ V is the token generated at step t. At decoding step t, the policy conditions on the historical context prefix ht := (x, y<t ), which consists of the prompt and all previously generated reasoning tokens. We use DKL (·∥·) to denote the KL divergence. For notational simplicity, we define the token-level and sequence-level log-ratios as ∆t ≜ log πθ (yt | ht ) − log π ∗ (yt | ht ), ∆(y) ≜ log πθ (y | x) − log π ∗ (y | x) =

T X

(1a) ∆t .

(1b)

t=1

2.2

O N -P OLICY D ISTILLATION (OPD)

On-policy distillation (OPD) (Agarwal et al., 2024; DeepSeek-AI et al., 2026) trains the student by minimizing the reverse KL divergence from the student policy πθ to the teacher policy π ∗ , evaluated on trajectories induced by the current student itself: JOPD (θ) = Ex∼D [DKL (πθ (y | x) ∥ π ∗ (y | x))] . 2

(2)

Preprint

By training on trajectories sampled from the current student policy, on-policy learning reduces exposure bias and is well suited to long chain-of-thought reasoning tasks such as mathematical problem solving (GLM-5 Team, 2026). Following recent practice (Lu & Thinking Machines Lab, 2025; Li et al., 2026), the empirical implementation of OPD typically converts the reverse-KL minimization in equation 2 into an RL-style surrogate gradient: " T # X OPD (3) ∇θ JOPD (θ) ≈ −Ex∼D At ∇θ log πθ (yt | ht ) , t=1

where AOPD ≜ −∆t = log π ∗ (yt | ht ) − log πθ (yt | ht ) is the token-level OPD advantage. t 2.3

RL WITH V ERIFIABLE R EWARDS (RLVR)

Reinforcement learning with verifiable rewards (RLVR) (Shao et al., 2024; Guo et al., 2025; Yu et al., 2025) has also become a key approach for improving the reasoning ability of LLMs. For a prompt x ∼ D, the policy πθ generates an output sequence y = (y1 , . . . , yT ) ∼ πθ (· | x). The generated output is evaluated by an external verifier, such as a code compiler or a mathematical rule checker, which provides a sparse sequence-level reward scalar R(x, y) ∈ {−1, 1}. The RLVR objective is JRLVR (θ) = Ex∼D, y∼πθ (·|x) [R(x, y)] . Compared with OPD, RLVR provides a sparse but verifier-grounded sequence-level correctness signal, whereas OPD supplies dense token-level supervision through teacher guidance, which may nevertheless be limited by the teacher’s suboptimal behavior.

3

M ETHODOLOGY

3.1

T OKEN - LEVEL OPD AS A T EMPORAL A PPROXIMATION

Although the practical OPD objective in equation 3 is widely used, it should be understood as a temporal approximation rather than the exact policy gradient of the OPD objective in equation 2. To clarify this distinction, we formally define token-level OPD and sequence-level OPD as follows: seq (θ) ≜ Ex∼D [DKL (πθ (y | x) ∥ π ∗ (y | x))] , JOPD " T # X token ∗ JOPD (θ) ≜ Ex∼D Eht ∼dθ̄ (·|x) [DKL (πθ (yt | ht ) ∥ π (yt | ht ))] ,

(4a) (4b)

t=1 seq where JOPD denotes the sequence-level reverse-KL objective, i.e., the original OPD objective defined in equation 2, which compares the student and teacher distributions over complete response token trajectories. In contrast, JOPD denotes the token-level objective, which compares their next-token distributions at sampled prefixes. The distribution dθ̄ (ht | x) is the prefix distribution induced by the current student policy, where ht = (x, y<t ). The notation θ̄ indicates stop-gradient, namely this prefix distribution is treated as fixed when differentiating with respect to θ.

The two objectives are equivalent at the level of function values when the stop-gradient reference parameter is evaluated at the current policy parameter. Specifically, if θ̄ = θ, then the sampled prefix token distribution in JOPD matches the autoregressive prefix distribution induced by πθ , and we have seq token JOPD (θ) = JOPD (θ) θ̄=θ .

However, this value-level equivalence does not imply gradient-level equivalence. The reason is seq token that JOPD differentiates through the autoregressive distribution over future prefixes, whereas JOPD stops this dependence through dθ̄ (ht | x). This distinction leads to different temporal credit assignments, as shown below. Proposition 3.1 (Token-level OPD gradient). Under the stop-gradient treatment of the prefix distritoken bution dθ̄ (ht | x), the gradient of JOPD (θ) in equation 4b is given by " T # X token ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t ∇θ log πθ (yt | ht ) . (5) t=1

3

Preprint

Token-level OPD

low Sequence-level OPD

Teacher

Verifier

Temporal-credit OPD ( OPD)

high (a) Temporal credit Assignment

(b) Reward-compatible bounded mixing

Figure 2: Overview of γOPD. (a) Token-level OPD considers only the log-ratio of the current token, whereas sequence-level OPD accumulates the log-ratios of all future tokens. γOPD introduces a discount factor γ to interpolate between these two extremes. (b) The γOPD advantage is first normalized and then combined with the verifiable task reward, balancing dense teacher guidance with task-level verification signals. The proof is provided in Appendix A. token (θ) recovers the practical OPD update in equation 3. MeanConsequently, the gradient of JOPD seq while, the gradient of the sequence-level OPD objective JOPD (θ) takes the following form: Proposition 3.2 (Sequence-level OPD gradient). The gradient of the sequence-level OPD objective seq (θ) in equation 4a is given by JOPD " T ! # T X X seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t′ ∇θ log πθ (yt | ht ) . (6) t′ =t

t=1

The proof is provided in Appendix B. Propositions 3.1 and 3.2 show that practical OPD is a token-level approximation to sequence-level OPD that neglects the influence of subsequent tokens. As illustrated in Figure 2, under the sequencePT level objective, the score term at step t is weighted by the future log-ratio t′ =t ∆t′ , since yt affects all future histories ht+1 , . . . , hT . In contrast, practical OPD keeps only the local term ∆t . This gap stems from the stop-gradient treatment of dθ̄ (ht | x) in equation 4b. 3.2

T EMPORAL -C REDIT O N -P OLICY D ISTILLATION (γOPD)

Compared with the token-level gradient in Proposition 3.1, the sequence-level gradient in equation 6 accounts for the indirect effect of each token on future autoregressive prefixes, and is therefore unbiased with respect to the sequence-level objective in equation 2. However, for long reasoning trajecPT tories, the accumulated future log-ratio t′ =t ∆t′ can have large variance, which may destabilize training. To balance bias and variance in temporal credit assignment, we propose Temporal-Credit On-Policy Distillation (γOPD). The key idea is to introduce a discount factor γ ∈ [0, 1] to control how much future OPD credit is assigned to the current token. Specifically, we define the temporal-credit surrogate gradient as " T # X (γ) gγ (θ) ≜ −Ex∼D, y∼πθ (·|x) At ∇θ log πθ (yt | ht ) , t=1

where the discounted γOPD advantage is defined as (γ)

At

≜−

T X t′ =t

4

′

γ t −t ∆t′ .

(7)

Preprint

(1)

The discount factor γ controls the temporal horizon of OPD credit assignment. When γ = 1, At becomes the sequence-level OPD return-to-go, recovering equation 6. When γ = 0, it reduces to (0) the token-level OPD advantage At = −∆t , recovering equation 5. Thus, for 0 < γ < 1, gγ (θ) is deliberately introduced as a biased surrogate that interpolates between these two endpoint gradients, allowing its variance to be controlled through γ, as formalized in the following theorem. Theorem 3.3 (Variance Stability of γOPD). Assume that the token-level OPD advantage has a (0) 2 bounded second moment, i.e., E[(At )2 ] ≤ σ∆ for all t. Then, the sequence-level OPD credit admits a horizon-dependent variance bound, whereas the γOPD credit admits a horizon-independent variance bound:   (1) 2 Var At ≤ (T − t + 1)2 σ∆ ,   2 σ∆ (γ) Var At ≤ , ∀γ ∈ [0, 1). (1 − γ)2 The proof is provided in Appendix C. Theorem 3.3 shows that the variance of sequence-level OPD credit can grow quadratically with the remaining sequence horizon in the worst case, whereas γOPD admits a horizon-independent variance bound for any fixed γ < 1. Adjusting γ therefore provides a principled trade-off between training stability and long-horizon credit propagation. 3.3

γOPD WITH R EWARD -C OMPATIBLE B OUNDED M IXING (RBM)

Although γOPD improves temporal credit assignment, its optimization signal is still fundamentally derived from the teacher distribution. Consequently, the student may remain constrained by the teacher policy and lack an explicit task-level optimization signal for solving verifiable reasoning problems. To move beyond this teacher-imitation ceiling, we incorporate the verifiable task reward R(x, y) ∈ {−1, 1} into γOPD. A naive combination, however, can be problematic because the magnitude of (γ) the teacher-derived advantage At may vary substantially across responses and potentially dominate the bounded task reward. We therefore introduce a reward-compatible bounded mixing (RBM) mechanism: the γOPD credit is first calibrated by its response-wise mean absolute magnitude, then bounded through a softsign transformation, and finally combined with the verifiable reward. These operations can be written compactly as (γ)

b(γ) A ≜ t

At 1 T

(γ)

PT

k=1 Ak

(γ)

+ R(x, y).

(8)

+ At

(γ)

When the denominator vanishes, equivalently when Ak = 0 for all k, we define the normalized teacher term as zero. The first term is equivalent to applying softsign after response-wise mean(γ) absolute normalization. It is monotonic in At and bounded within (−1, 1), thereby retaining the relative token-level OPD credit while preventing the teacher-derived signal from overwhelming the b(γ) task reward. Since R(x, y) ∈ {−1, 1}, it follows that sign(A t ) = R(x, y). Thus, RBM is rewardcompatible by construction: the verifiable reward determines whether the sampled response is reinforced or suppressed, while the bounded γOPD credit modulates the token-level update strength. The resulting reward-compatible γOPD surrogate gradient is " T # X (γ) b ĝγ (θ) ≜ −Ex∼D, y∼π (·|x) At ∇θ log πθ (yt | ht ) . θ

(9)

t=1

Therefore, RBM combines sparse verifiable rewards for task-level learning beyond teacher imitation with bounded γOPD for dense token-level credit assignment. Unless otherwise specified, γOPD refers to its use within RBM throughout the remainder of this paper. 5

Preprint

Table 1: Mathematical reasoning results under vanilla and size-mismatch distillation. We report accuracy and pass rate across four mathematical reasoning benchmarks. AIME24

Method

Acc.

Pass

AIME25 Acc.

Pass

AMC23 Acc.

Pass

MATH500

Average

Acc.

Pass

Acc.

Pass

Qwen3-4B-Math → Qwen3-4B Student Teacher

22.60 57.40

60.00 76.67

20.83 51.67

33.33 66.67

60.16 93.52

92.50 97.50

67.40 69.80

72.60 78.40

42.75 68.10

64.61 79.81

JustRL STAPO

53.54 55.31

83.33 83.33

42.50 50.00

56.67 66.67

88.75 90.16

97.50 97.50

68.00 68.35

70.40 69.80

63.20 65.95

76.97 79.33

OPD ExOPD REOPOLD AOPD TOPD γOPD

56.46 57.19 56.98 56.56 54.79 60.94

80.00 83.33 83.33 86.67 80.00 86.67

50.83 52.50 51.67 52.50 51.67 54.17

63.33 63.33 63.33 66.67 63.33 66.67

92.81 93.13 93.59 93.28 93.44 92.89

95.00 97.50 97.50 97.50 97.50 97.50

68.05 68.90 68.70 69.35 69.95 69.95

72.20 74.20 72.60 73.00 72.40 78.80

67.04 67.93 67.74 67.92 67.46 69.49

77.63 79.59 79.19 80.96 78.31 82.41

Qwen3-4B-Math → Qwen3-1.7B Student Teacher

11.77 57.40

46.67 76.67

8.33 51.67

20.00 66.67

39.84 93.52

90.00 97.50

59.20 69.80

72.80 78.40

29.79 68.10

57.37 79.81

JustRL STAPO

37.60 35.52

66.67 63.33

36.67 30.83

50.00 46.67

78.67 77.73

95.00 95.00

65.10 64.95

68.80 68.80

54.51 52.26

70.12 68.45

OPD ExOPD REOPOLD AOPD TOPD γOPD

37.71 38.54 37.50 37.71 39.38 42.08

66.67 66.67 66.67 70.00 70.00 70.00

31.25 31.67 30.83 32.50 30.83 35.00

40.00 43.33 43.33 43.33 36.67 46.67

76.02 76.72 77.73 79.38 79.14 80.86

92.50 92.50 92.50 95.00 95.00 95.00

65.40 65.05 66.70 65.85 65.60 68.10

69.80 68.40 69.00 70.80 70.00 73.20

52.60 53.00 53.19 53.86 53.74 56.51

67.24 67.73 67.88 69.78 67.92 71.22

4

E XPERIMENTS

4.1

S ETTINGS

Benchmarks. Our experiments evaluate both mathematical and code reasoning abilities. For mathematical reasoning, we train on DeepMath (He et al., 2025), retaining problems with difficulty level at least 6, and evaluate on AIME24 (Li et al., 2024), AIME25 (OpenCompass, 2025), AMC23 (Li et al., 2024), and MATH500 (Hendrycks et al., 2021). For code reasoning, we train on the 25Ksample Eurus-RL-Code dataset (Cui et al., 2026) and evaluate on HumanEval+, MBPP+ (Liu et al., 2023), and the v6 split of LiveCodeBench (Jain et al., 2025), covering problems from February 2025 to May 2025. Detailed training and evaluation configurations are provided in Appendix E. Models. We conduct experiments under three distillation settings: (1) vanilla distillation, where we distill a Qwen3-4B-Math (Yang et al., 2026a) teacher into a Qwen3-4B (Yang et al., 2025) student for mathematical reasoning; (2) size-mismatch distillation, where we distill the same Qwen3-4BMath teacher into a smaller Qwen3-1.7B student for mathematical reasoning; and (3) multi-teacher distillation, where we distill Qwen3-4B-Math and Qwen3-4B-Code (Yang et al., 2026a) teachers into a single Qwen3-4B student for both math and code reasoning. Baselines. We compare γOPD against representative OPD-based baselines, including vanilla OPD (Lu & Thinking Machines Lab, 2025), ExOPD (Yang et al., 2026a), REOPOLD (Ko et al., 2026), AOPD (Jia et al., 2026), and TOPD (Zhang et al., 2026), as well as RL-based methods, including JustRL (He et al., 2026) and STAPO (Liu et al., 2026). For γOPD, we use γ = 0.99 6

Preprint

0.4 0.2 0.0

0

10

20 30 Step

40

10

10

0

10

−2

10

−4

Grad Norm

0.6

REOPOLD

0

10

20 30 Step

10

AOPD

1

−1

40

0

10

20 30 Step

γOPD

TOPD Response Length

ExOPD Absolute Advantage

Verifiable Reward

OPD

10000

5000

40

0

10

20 30 Step

40

Figure 3: Training dynamics under vanilla distillation. Verifiable reward, absolute advantage, gradient norm, and response length throughout training are reported. Due to the truncation mechanism, TOPD’s verifiable reward remains mostly below zero throughout training. Table 2: Accuracy on mathematical and code reasoning benchmarks under multi-teacher distillation. TotalAvg is the equal-weighted average of the mean mathematical score and the mean code score. Method Student Teacher OPD ExOPD REOPOLD AOPD TOPD γOPD

AIME24 22.60 57.40 55.21 57.50 58.13 55.52 57.81 59.27

Mathematical Reasoning AIME25 AMC23 MATH500 20.83 51.67 52.50 52.50 52.08 55.83 51.67 55.00

60.16 93.52 93.13 92.34 91.88 93.05 92.34 94.38

67.40 69.80 68.65 68.25 68.55 68.70 69.50 69.00

Code Reasoning HumanEval+ MBPP+ 75.61 82.32 84.15 85.37 82.32 84.76 84.15 89.02

61.90 70.63 67.99 71.16 69.31 69.05 69.58 70.63

LCB 17.14 25.71 26.00 26.71 28.00 26.71 27.86 28.14

TotalAvg 47.15 63.82 63.37 64.36 63.77 64.22 64.18 66.00

by default unless otherwise specified. All methods are implemented based on veRL (Sheng et al., 2025). 4.2

M AIN R ESULTS

Vanilla Distillation. As shown in Table 1, γOPD consistently outperforms existing OPD baselines in the vanilla distillation setting. Compared with the strongest baseline for each metric, γOPD achieves relative improvements of 2.30% in AvgAcc and 1.80% in AvgPass. Notably, it also surpasses the teacher on both averaged metrics. These results demonstrate the effectiveness of γOPD and suggest that the proposed RBM can better leverage dense teacher supervision while enabling improvement beyond direct imitation. Size-Mismatch Distillation. The size-mismatch setting (Qwen3-4B-Math → Qwen3-1.7B) is more challenging because the student has substantially lower capacity than the teacher. Consequently, as shown in Table 1, all student methods remain below the teacher. Nevertheless, γOPD substantially narrows the performance gap and achieves the best overall performance. Compared with the strongest baseline for each metric, γOPD achieves relative improvements of 4.92% in AvgAcc and 2.06% in AvgPass. These results highlight the benefit of temporal credit assignment when distilling a stronger teacher into a lower-capacity student. Multi-Teacher Distillation. We next evaluate whether γOPD remains effective in the multi-teacher distillation setting. As shown in Table 2, it consistently performs strongly across mathematical reasoning benchmarks and achieves the best results on HumanEval+ and LCB among the code reasoning tasks. Overall, γOPD improves TotalAvg by 1.87% over the strongest baseline, AOPD. Notably, it also outperforms the teacher on average in both mathematical and code reasoning. These results demonstrate that γOPD remains effective when jointly training a shared student with domainspecialized teachers. Training Dynamics. We further visualize the training dynamics of different methods under the vanilla distillation setting in Figure 3. Throughout training, γOPD consistently achieves the highest verifiable reward, indicating the largest proportion of correctly solved problems. It also maintains a stable absolute OPD advantage and exhibits the smallest fluctuations in gradient norm among all 7

Preprint

Table 3: Ablation study of the proposed components on AIME benchmarks. γ, Mix, and Norm correspond to temporal discounting, reward mixing, and bounded normalization, respectively. ∆ Avg. denotes the absolute improvement in average accuracy over the vanilla OPD baseline. Components γ ✓ ✓ – ✓

Mix

Norm

Vanilla OPD – – ✓ – ✓ ✓ ✓ ✓

Accuracy (%)

∆ Avg.

AIME24

AIME25

Avg.

56.46 58.29 58.81 57.55 60.94

50.83 53.08 52.67 53.75 54.17

53.65 55.69 55.74 55.65 57.56

– +2.04 +2.09 +2.00 +3.91

methods, suggesting more consistent temporal credit assignment and optimization. Moreover, the response length of γOPD converges to a shorter and more stable range than those of most baselines, with the exception of TOPD, which explicitly truncates the distillation signal. Together, these results suggest that γOPD enables more stable optimization while achieving stronger mathematical reasoning performance. 4.3

A BLATION S TUDIES

We isolate the effects of temporal discounting, reward mixing, and bounded normalization on the AIME benchmarks in Table 3. Temporal discounting alone improves the average accuracy by 2.04 points over vanilla OPD. Removing temporal discounting from the full method while retaining reward mixing and normalization reduces the gain from 3.91 to 2.00 points, confirming its complementary contribution. Moreover, adding naive reward mixing to temporal discounting brings only a marginal 0.05-point improvement, whereas the full combination including bounded normalization yields the strongest overall performance. These results suggest that temporal credit assignment and reward-compatible normalization provide complementary benefits, with their combination yielding the strongest performance. We further analyze the computational and memory overhead of the individual components of γOPD. Compared with standard OPD, γOPD introduces only about 0.1% additional computation time with negligible memory overhead; detailed profiling results are provided in Appendix F.1. 4.4

D ETAILED A NALYSIS

Sensitivity to γ. We further investigate the sensitivity of temporal credit assignment to the discount factor γ without RBM. Specifically, we distill a Qwen3-4B-Math teacher into a Qwen3-1.7B student using different values of γ, with the results shown in Figure 4. Here, γ = 0 reduces to local OPD, whereas γ = 1 corresponds to the undiscounted sequence-level return-to-go. Due to the long response horizon, the sequence-level variant produces substantially larger gradient norms and suffers from rapid entropy collapse, resulting in inferior reasoning performance. In contrast, local OPD and the discounted variants steadily improve during training. Among them, γ = 0.99 achieves the best performance, outperforming smaller values such as γ = 0.9, while exhibiting distinct gradient-norm and entropy dynamics. These results suggest that a large but sub-unity discount factor provides the best balance between long-range temporal credit assignment and optimization stability. Visualization of Token Advantages. We further visualize the normalized token-level advantages (0) (0.99) (0) At and At in Figure 5. In the correct response, At provides sparse local supervision, assigning a noticeable negative signal to only a few tokens while exerting little influence on most of (0.99) the reasoning process. In contrast, At propagates information from subsequent steps, producing a smoother signal and assigning stronger positive credit to the key mathematical derivation while (0) mildly penalizing redundant text. For the incorrect response, At strongly penalizes several tokens (γ) that are only weakly related to the actual reasoning error, whereas At concentrates stronger negative credit on the erroneous formula derivation without excessively penalizing the final token. These examples illustrate that temporal credit assignment can redistribute supervision toward reasoning 8

Preprint

γ = 0.9

γ = 0.99

γ=1 0.4

0.3 0.2

10

1

10

0

Entropy

Grad Norm

AIME24 Acc (avg@32)

γ=0 0.4

0.3

0.2

0.1 0

10

20

30

40 Step

50

60

70

0

10

20

30

40 Step

50

60

70

0

10

20

30

40 Step

50

60

70

Figure 4: Training dynamics under different γ. OPD distills a Qwen3-4B-Math teacher into a Qwen3-1.7B student, with validation accuracy, gradient norm, and policy entropy reported. (a) Correct answer

(b) Incorrect answer $$ \sum_{k=1}^{6} \tan^2

We are given a function $ f(x, y, z) = xyz $ and a region $W$ At(0)

\frac{(7-1)(7-2)}{3} defined by the

inequalities $ 0 ......

\frac{(7-1)(7-2)}{3}

-1.0

\right) =

= \frac{6 \cdot 5}{3} = \frac{30}{3}

$$ \sum_{k=1}^{6} \tan^2

We are given a function $ f(x, y, z) = xyz $ and a region $W$

inequalities

\frac{k\pi}{7}

= 10 $$ --- ### Final Answer: $$ \boxed{10} $$<|im_end|>

At(0.99) defined by the

\left(

$ 0 ......

\left(

\frac{k\pi}{7}

\right) =

= \frac{6 \cdot 5}{3} = \frac{30}{3}

= 10 $$ --- ### Final Answer: $$ \boxed{10} $$<|im_end|>

-0.5

0.0

0.5

1.0

At

Figure 5: Token-level advantage visualization. (a) Correct response; (b) incorrect response. The (0) (0.99) top row shows normalized At , and the bottom row shows normalized At . Blue and red denote negative and positive advantages, respectively, with color intensity indicating magnitude. Complete response examples are provided in Appendix F.2.

steps that are more relevant to the final outcome, yielding smoother and more semantically aligned token-level signals. Complete response-level visualizations are provided in Appendix F.2.

5

R ELATED W ORK

Credit Assignment for OPD. Recent studies have explored credit assignment in OPD, motivated by the potentially unreliable reasoning trajectories generated by relatively weak student policies(Liu et al., 2026; Yu et al., 2026). Truncation-based methods (Zhou et al., 2026; Zhang et al., 2026) terminate rollouts early or mask the learning signals of tokens beyond selected positions, thereby reducing the influence of unreliable continuations. Entropy-aware OPD (Jin et al., 2026) augments reverse KL with forward KL on tokens where the teacher has high entropy, while ExOPD (Yang et al., 2026a) introduces a reference model to calibrate the OPD advantage. Fu et al. (2026) analyze the bias–variance trade-off between token-level and sequence-level reverse-KL estimators and propose teacher top-K local-support matching, yet a principled balance between sequence-level objective fidelity and token-level optimization stability remains unresolved. Reward-Guided OPD. Recent studies incorporate verifiable rewards into OPD to complement teacher-derived supervision. REOPOLD (Ko et al., 2026) combines reward clipping, entropy-based sampling, and exploration-to-refinement scheduling, while SCOPE (Zheng et al., 2026) routes incorrect trajectories to teacher-perplexity-weighted KL distillation and correct trajectories to studentperplexity-weighted MLE. AOPD (Jia et al., 2026) preserves positive reinforcement while replacing non-positive token updates with localized teacher imitation, whereas RWOPD (Zou et al., 2026) weights teacher KL gradients using verifier rewards. Other works (Ding et al., 2026; Zhan et al., 2026) further stabilize optimization through advantage compression and warmup-then-anneal scheduling. Nevertheless, these methods often introduce additional complexity through group-based 9

Preprint

advantage estimation, log-ratio correction, or verifier-weighted teacher KL objectives, potentially incurring substantial computational overhead.

6

C ONCLUSION

In this work, we unified token-level and sequence-level OPD from a temporal credit assignment perspective and proposed γOPD to balance long-horizon supervision and optimization stability. We further introduced RBM to combine discounted teacher guidance with verifiable outcome rewards. Empirical results on mathematical and code reasoning demonstrate that γOPD remains effective across different teacher–student configurations and distillation scenarios. Due to computational constraints, our experiments are currently limited to models with fewer than 10B parameters. Evaluating γOPD at larger scales and exploring more flexible temporal credit assignment schemes, such as adaptive or entropy-aware discounting, are promising directions for future work.

R EFERENCES Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from selfgenerated mistakes. In International Conference on Learning Representations, 2024. Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Yuchen Zhang, Jiacheng Chen, et al. Process reinforcement through implicit rewards. Transactions on Machine Learning Research, 2026. URL https://arxiv.org/abs/2502.01456. DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, et al. DeepSeek-V4: Towards highly efficient million-token context intelligence, 2026. URL https://arxiv.org/abs/ 2606.19348. Yifan Ding, Xincheng Wei, Yoshua Y. Li, Ziheng Li, Yuquan Lu, Siyu Zhang, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, and Yun Chen. SAF-OPD: Stable advantage fusion for OnPolicy Distillation. arXiv preprint arXiv:2607.29209, 2026. URL https://arxiv.org/ abs/2607.29209. Yuqian Fu, Haohuan Huang, Kaiwen Jiang, Jiacai Liu, Zhuo Jiang, Yuanheng Zhu, and Dongbin Zhao. Revisiting On-Policy Distillation: Empirical failure modes and simple fixes. arXiv preprint arXiv:2603.25562, 2026. URL https://arxiv.org/abs/2603.25562. GLM-5 Team. GLM-5: from Vibe Coding to Agentic Engineering, 2026. URL https://arxiv. org/abs/2602.15763. Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, et al. DeepSeekR1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645(8081):633–638, 2025. doi: 10.1038/s41586-025-09422-z. Bingxiang He, Zekai Qu, Zeyuan Liu, Yinghao Chen, Yuxin Zuo, Cheng Qian, Kaiyan Zhang, Weize Chen, Chaojun Xiao, Ganqu Cui, Ning Ding, and Zhiyuan Liu. JustRL: Scaling a 1.5B LLM with a Simple RL Recipe. In ICLR Blogposts 2026, 2026. Zhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu, Xingyu Chen, Yue Wang, Linfeng Song, Dian Yu, Zhenwen Liang, Wenxuan Wang, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu. DeepMath-103K: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning. arXiv preprint arXiv:2504.11456, 2025. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. 10

Preprint

Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. LiveCodeBench: Holistic and contamination free evaluation of large language models for code. In International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=chfJJYC3iL. Nan Jia, Haojin Yang, Xing Ma, Jiesong Lian, Shuailiang Zhang, Weipeng Zhang, Ke Zeng, Xunliang Cai, and Zequn Sun. Asymmetric On-Policy Distillation: Bridging exploitation and imitation at the token level. arXiv preprint arXiv:2605.06387, 2026. doi: 10.48550/arXiv.2605.06387. URL https://arxiv.org/abs/2605.06387. Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, and Kimin Lee. Entropy-aware On-Policy Distillation of language models. In Proceedings of the 43rd International Conference on Machine Learning, volume 306. PMLR, 2026. doi: 10.48550/arXiv.2603.07079. URL https://arxiv.org/abs/2603.07079. Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, et al. Kimi K3: Open Frontier Intelligence, 2026. URL https://arxiv.org/abs/2607.24653. Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen, and Pashmina Cameron. Scaling reasoning efficiently via relaxed on-policy distillation, 2026. URL https://arxiv.org/abs/2603. 11137. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023. Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, et al. NuminaMath: The largest public dataset in AI4Maths with 860k pairs of competition math problems and solutions. Hugging Face dataset, 2024. URL https://huggingface.co/datasets/AI-MO/NuminaMath-CoT. Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huanang Gao, Wenkai Yang, Zhiyuan Liu, and Ning Ding. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe, 2026. URL https://arxiv. org/abs/2604.13016. Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Advances in Neural Information Processing Systems, volume 36, pp. 21558–21572, 2023. URL https://openreview.net/forum?id=1qvx610Cu7. Shiqi Liu, Zeyu He, Guojian Zhan, Letian Tao, Zhilong Zheng, Jiang Wu, Yinuo Wang, Yang Guan, Kehua Sheng, Bo Zhang, Keqiang Li, Jingliang Duan, and Shengbo Eben Li. STAPO: Stabilizing reinforcement learning for LLMs by silencing rare spurious tokens. arXiv preprint arXiv:2602.15620, 2026. Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. URL https://thinkingmachines.ai/blog/ on-policy-distillation. Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi, and Yohan Jo. KL for a KL: On-Policy Distillation with control variate baseline. arXiv preprint arXiv:2605.07865, 2026. doi: 10.48550/ arXiv.2605.07865. URL https://arxiv.org/abs/2605.07865. Presented as a poster at the AI for Math Workshop, ICML 2026. OpenCompass. AIME2025 dataset. https://huggingface.co/datasets/ opencompass/AIME2025, 2025. Accessed: 2026-08-09. Zhixin Ren, Yau Lyu, Congrong Li, Liping Zhang, and Shengbo Eben Li. Momentum as residualdriven multiplier correction for deep learning optimization, 2026. URL https://arxiv. org/abs/2608.12925. 11

Preprint

Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. URL https://arxiv.org/abs/2402.03300. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. HybridFlow: A flexible and efficient RLHF framework. In Proceedings of the Twentieth European Conference on Computer Systems, pp. 1279–1297, 2025. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. URL https://arxiv.org/abs/2505.09388. Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, and Yankai Lin. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation. arXiv preprint arXiv:2602.12125, 2026a. URL https://arxiv.org/abs/2602.12125. Zhuolin Yang, Zihan Liu, Yang Chen, Wenliang Dai, Boxin Wang, Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He, et al. Nemotron-Cascade 2: Post-training LLMs with cascade RL and multi-domain on-policy distillation. arXiv preprint arXiv:2603.19220, 2026b. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, et al. DAPO: An open-source LLM reinforcement learning system at scale. In Advances in Neural Information Processing Systems, volume 38, pp. 125532–125554, 2025. doi: 10.52202/085713-3775. Zhouyang Yu, Guojian Zhan, Yang Guan, Jingliang Duan, Letian Tao, and Shengbo Eben Li. Taming aleatoric impulse in off-policy reinforcement learning. In Proceedings of the 43rd International Conference on Machine Learning, 2026. Guojian Zhan, Xiangteng Zhang, Feihong Zhang, Letian Tao, and Shengbo Eben Li. Bicriteria policy optimization for high-accuracy reinforcement learning. IEEE Transactions on Neural Networks and Learning Systems, 37(1):312–326, 2026. doi: 10.1109/TNNLS.2025.3605362. Yaocheng Zhang, Jiajun Chai, Yuqian Fu, Songjun Tu, Xiaohan Wang, Wei Lin, Guojun Yin, Qichao Zhang, Yuanheng Zhu, and Dongbin Zhao. Are full rollouts necessary for On-Policy Distillation? arXiv preprint arXiv:2605.31490, 2026. doi: 10.48550/arXiv.2605.31490. URL https:// arxiv.org/abs/2605.31490. Binbin Zheng, Xing Ma, Yiheng Liang, Jingqing Ruan, Xiaoliang Fu, Kepeng Lin, Benchang Zhu, Ke Zeng, and Xunliang Cai. SCOPE: Signal-calibrated On-Policy Distillation enhancement with dual-path adaptive weighting. arXiv preprint arXiv:2604.10688, 2026. doi: 10.48550/arXiv.2604. 10688. URL https://arxiv.org/abs/2604.10688. Ziheng Zhou, Jiaqi Li, Huacong Tang, Ying Nian Wu, and Demetri Terzopoulos. Less is more: Early stopping rollout for On-Policy Distillation. arXiv preprint arXiv:2605.27028, 2026. doi: 10.48550/arXiv.2605.27028. URL https://arxiv.org/abs/2605.27028. Qingyun Zou, Yingze Li, Tianen Liu, Bingsheng He, and Weng-Fai Wong. Reward-weighted OnPolicy Distillation with an open property-equivalence verifier for NL-to-SVA generation. arXiv preprint arXiv:2605.13501, 2026. doi: 10.48550/arXiv.2605.13501. URL https://arxiv. org/abs/2605.13501.

12

Preprint

A PPENDIX A

P ROOF OF P ROPOSITION 3.1

Starting from the token-level OPD objective in equation 4b, we have " T # X token ∗ JOPD (θ) = Ex∼D Eht ∼dθ̄ (·|x) [DKL (πθ (· | ht ) ∥ π (· | ht ))] . t=1

Since the prefix distribution dθ̄ (ht | x) is treated with stop-gradient, it is fixed when differentiating with respect to θ. Therefore, the gradient can be moved inside the expectation over prefixes: " T " ## X X token ∗ ∇θ JOPD (θ) = Ex∼D Eht ∼dθ̄ (·|x) ∇θ πθ (a | ht ) (log πθ (a | ht ) − log π (a | ht )) . t=1

a∈V

(10) For a fixed prefix ht , define ∆(a; ht ) ≜ log πθ (a | ht ) − log π ∗ (a | ht ). This is the full-vocabulary counterpart of the sampled token-level log-ratio ∆t in equation 1a. Then the gradient of the inner next-token KL term is X ∇θ πθ (a | ht )∆(a; ht ) a∈V

=

X

∇θ πθ (a | ht )∆(a; ht ) +

a∈V

X

(11) πθ (a | ht )∇θ log πθ (a | ht ).

a∈V

The second term in equation 11 vanishes because X X X πθ (a | ht )∇θ log πθ (a | ht ) = ∇θ πθ (a | ht ) = ∇θ πθ (a | ht ) = 0. a∈V

a∈V

a∈V

Using the score-function identity ∇θ πθ (a | ht ) = πθ (a | ht )∇θ log πθ (a | ht ), we obtain ∇θ DKL (πθ (· | ht ) ∥ π ∗ (· | ht )) =

X

πθ (a | ht )∆(a; ht )∇θ log πθ (a | ht ) (12)

a∈V

= Eyt ∼πθ (·|ht ) [∆t ∇θ log πθ (yt | ht )] , where ∆t = ∆(yt ; ht ) follows from equation 1a. Substituting equation 12 into equation 10 gives " T # X token ∇θ JOPD (θ) = Ex∼D Eht ∼dθ̄ (·|x), yt ∼πθ (·|ht ) [∆t ∇θ log πθ (yt | ht )] .

(13)

t=1

When the stop-gradient prefix distribution is evaluated at the current policy, the prefix ht together with the next token yt can be generated by an on-policy trajectory y ∼ πθ (· | x), while gradients are not propagated through the prefix-sampling distribution. Hence, equation 13 can be written as " T # X token ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t ∇θ log πθ (yt | ht ) , t=1

which is exactly the token-level OPD gradient in equation 5. This proves Proposition 3.1. 13

Preprint

B

P ROOF OF P ROPOSITION 3.2

Starting from the sequence-level OPD objective in equation 4a, we expand the reverse KL over complete response trajectories:   X seq JOPD (θ) = Ex∼D  πθ (y | x)∆(y) , y∈V T

where ∆(y) is the sequence-level log-ratio defined in equation 1b. Taking the gradient with respect to θ gives   X X seq ∇θ JOPD (θ) = Ex∼D  ∇θ πθ (y | x)∆(y) + πθ (y | x)∇θ log πθ (y | x) . (14) y∈V T

y∈V T

The second term in equation 14 vanishes because X X πθ (y | x)∇θ log πθ (y | x) = ∇θ πθ (y | x) = ∇θ 1 = 0. y∈V T

y∈V T

Using the score-function identity, ∇θ πθ (y | x) = πθ (y | x)∇θ log πθ (y | x), we obtain seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) [∆(y)∇θ log πθ (y | x)] .

(15)

By the autoregressive factorization and the log-ratio decomposition in equation 1, we have ∆(y) =

T X

∆ t′ ,

∇θ log πθ (y | x) =

t′ =1

T X

∇θ log πθ (yt | ht ).

(16)

t=1

Substituting equation 16 into equation 15 yields seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x)

" T T XX

# ∆t′ ∇θ log πθ (yt | ht ) .

(17)

t=1 t′ =1

It remains to remove the terms with t′ < t. For t′ < t, ∆t′ is determined by earlier tokens and is therefore measurable with respect to ht . Hence, conditioning on ht , Eyt ∼πθ (·|ht ) [∆t′ ∇θ log πθ (yt | ht )] = ∆t′ Eyt ∼πθ (·|ht ) [∇θ log πθ (yt | ht )] = 0, where the last equality follows from the token-level score identity X Eyt ∼πθ (·|ht ) [∇θ log πθ (yt | ht )] = ∇θ πθ (yt | ht ) = 0. yt ∈V

Therefore, all terms with t′ < t in equation 17 vanish in expectation. Keeping only the terms with t′ ≥ t, we obtain " T ! # T X X seq ∇θ JOPD (θ) = Ex∼D, y∼πθ (·|x) ∆t′ ∇θ log πθ (yt | ht ) , t=1

t′ =t

which is exactly the sequence-level gradient in equation 6. This proves Proposition 3.2. 14

Preprint

C

P ROOF OF T HEOREM 3.3

Proof. Let Nt = T − t + 1 and define h i (0) (0) e(0) A . τ = Aτ − E Aτ Since e(0) A τ e(0) we have A τ

2

2 2

 2    (0) 2 ≤ E A ≤ σ∆ , = Var A(0) τ τ

≤ σ∆ .

For the sequence-level OPD credit, the triangle inequality in L2 gives r

T   X (1) e(0) Var At = A τ τ =t

≤

T X

2

e(0) A τ

τ =t

2

≤ Nt σ∆ .

Therefore,   (1) 2 2 Var At ≤ Nt2 σ∆ = (T − t + 1)2 σ∆ . (0)

This quadratic dependence is attainable in the worst case. In particular, if Aτ 2 , then {t, . . . , T }, where E[Z] = 0 and Var(Z) = σ∆

= Z for all τ ∈

  (1) 2 Var At = Var(Nt Z) = Nt2 σ∆ . For the γOPD credit, similarly, r

T   X (γ) e(0) Var At = γ τ −t A τ τ =t

≤ σ∆

2

T X

γ τ −t

τ =t

= σ∆

σ∆ 1 − γ Nt ≤ , 1−γ 1−γ

∀γ ∈ [0, 1).

Squaring both sides yields   (γ) Var At ≤

2 σ∆ , (1 − γ)2

∀γ ∈ [0, 1).

Thus, sequence-level OPD can exhibit variance that grows quadratically with the remaining horizon, whereas γOPD admits a horizon-independent variance bound for every fixed γ < 1.

D

A LGORITHM

The complete reward-enhanced γOPD procedure is summarized in Algorithm 1. 15

Preprint

Table 4: Training hyperparameters for the OPD-based methods. The RLVR baselines use a rollout group size of 8. Hyperparameter Value Training framework Rollout backend Math reward function Code reward function Train batch size Responses per prompt PPO mini batch size PPO micro batch size/GPU Max prompt length Max response length Optimizer Weight decay Learning rate Training temperature / top-p Validation temperature / top-p Rollout tensor parallel size

veRL vLLM DAPO boxed verifier Execution-based verifier 1024 1 1024 1 2048 16384 RADAR 0.01 1 × 10−5 1.0 / 1.0 1.0 / 1.0 4

Algorithm 1 Reward-Enhanced Temporal-Credit On-Policy Distillation (γOPD) Require: Dataset D, student policy πθ , teacher policy π ∗ , discount factor γ, batch size B 1: for each training iteration do B 2: Sample prompts {xi }B i=1 ∼ D and responses {yi }i=1 ∼ πθ (· | xi ) 3: Evaluate the verifier to obtain {Ri }B , where R ← R(xi , yi ) i i=1 4: for each mini-batch do 5: for each response (x, y, R) in the mini-batch do 6: for t = 1, . . . , T do 7: ∆t ← log πθ (yt | ht ) − log π ∗ (yt | ht ) 8: end for (γ) 9: Compute {Ât }Tt=1 according to equation 7 and equation 8 10: end for 11: Update θ using equation 9 12: end for 13: end for 14: return final student policy πθ

E

E XPERIMENT D ETAILS

All methods, including γOPD, are implemented with veRL (Sheng et al., 2025), using vLLM (Kwon et al., 2023) for rollouts and FSDP for actor training. For mathematical reasoning, we train on DeepMath-103K (He et al., 2025), retaining examples with difficulty level at least 6. Each prompt is formatted in chat style and appended with “Please output the final answer within \boxed{}.” For code reasoning, we train on Eurus-RL-Code (Cui et al., 2026). For both domains, verifier outcomes are mapped to binary rewards in {+1, −1}. Math responses are evaluated with a DAPO-style boxed-answer verifier (Yu et al., 2025), while code responses are verified through execution-based unit tests. In the multi-teacher setting, each sample is routed to its domain-specific teacher and verifier while sharing the same student policy. OPD uses reverse-KL token advantages without an additional KL reward penalty. Unless otherwise specified, all models use RADAR (Ren et al., 2026) optimizer, a constant learning rate on 32 NVIDIA H20 GPUs, with each training run taking approximately 1–3 days; full hyperparameters are provided in Table 4. 16

Preprint

Table 5: Computational and memory overhead of the additional γOPD operations. Operation Time (ms/step) Memory (MB) Discounted temporal credit Bounded advantage shaping Verifier-reward mixing

609.85 0.42 0.04

0.063 0.125 0.188

Total additional overhead

610.32

0.375

573,669.00

81,254.40

0.1064%

0.00046%

Vanilla OPD update (reference) Relative overhead

We use temperature 1.0 and top-p 1.0 for all evaluations. For mathematical reasoning, we evaluate on AIME24 (Li et al., 2024), AIME25 (OpenCompass, 2025), AMC23 (Li et al., 2024), and MATH500 (Hendrycks et al., 2021), with a maximum prompt length of 2048 and response length of 16384. Following the repeated-sampling convention, AIME24, AIME25, and AMC23 are evaluated with 32 samples per problem, while MATH500, MinervaMath, and OlympiadBench use 4 samples per problem; we report both accuracy and pass rate. All predictions are scored using the same boxed-answer verifier as in training. For code reasoning, we evaluate HumanEval+ and MBPP+ (Liu et al., 2023) with EvalPlus, and the v6 split of LiveCodeBench (Jain et al., 2025). HumanEval+ and MBPP+ use one sample per problem, whereas LiveCodeBench uses four samples with a maximum generation length of 16384 tokens.

F

S UPPLEMENTARY E XPERIMENTAL R ESULTS

F.1

C OMPUTATIONAL T IME AND M EMORY A NALYSIS

We profile the additional computation introduced by γOPD using the same 4B student/teacher models and 32-GPU setup as the main experiments. Statistics are averaged over training steps 2–100 after warm-up. The extra operations include discounted temporal-credit computation, bounded advantage shaping, and verifier-reward mixing. Overall, γOPD adds only 0.1064% computation time and 0.00046% memory relative to a full OPD update, indicating negligible training overhead. F.2

C OMPLETE T OKEN -L EVEL A DVANTAGE V ISUALIZATION

This appendix provides the complete token-level visualizations behind Figure 5. Each response is shown as a sequence of consecutive full-width panels rather than as a collection of small subfigures, so that individual tokens and advantage magnitudes remain readable. The upper row in every panel (0) shows the local advantage At , while the lower row shows the normalized temporally propagated (0.99) advantage At . (0)

(0.99)

As shown in Figure 6, for the correct response, At gives a sparse local signal, whereas At propagates information from later reasoning steps and assigns stronger positive credit to the key (0) mathematical derivation. As shown in Figure 7, for the incorrect response, At penalizes tokens (0.99) largely unrelated to the actual error, while At assigns stronger negative credit to the erroneous formula derivation without excessively penalizing the final end-of-sequence token.

17

Preprint

At(0) At(0.99)

are are

given given

a a

defined defined

by by

the the

We We

need need

to to

find find

average average by: by:

value value

over over

y, y,

of of

$ W $ W

W W

$. $.

$ $

x x

\le \le

1 1

(since (since

to to

set set

set set

up up

This This

order order

for for

as: as:

$ $ $ $

x x

z z

can can $$ $$

\int_{z=0}^y \int_{z=0}^y

W W

$ $

$ $

y y

range range

$ $

$ $

z z

0 0

to to

xyz xyz xyz xyz

$ $

\, \,

z z

$, $,

y y

dz dz

\, \,

$ $

given given

is is

the the

y y

$, $,

$ $

\le \le

\le \le

in in

order order

is: is:

z z

Let's Let's

terms terms

$, $,

we we

from from

0 0

to to

$ $

y y

$. $.

So, So,

dV dV

= =

the the

can can

fix fix

$ $

$, $,

and and

the the

dx dx

$$ $$

of of

z z

y y

\le \le octant octant

1 1

the the

$. $.

So So

we we

limits. limits.

Since Since

the the

can can

choose choose

an an

z z

$, $,

$ $

y y

$, $,

$ $

determine determine order order x x

integral integral

of of

$, $,

for for

\int_{x=0}^1 \int_{x=0}^1 \, \,

$ $

here. here.

$ $

of of

x x

f f

integration integration

we we

$. $.

$ $

and and x x

$, $,

W} W} the the

first first

\le \le

1 1

} }

of of

0 0

the the

x x

18

$$ $$

express express

\le \le

The The

bounds bounds

appropriate appropriate

x x

We We

compute compute

with with

\le \le

dy dy

to to

important important

\le \le

$. $.

of of

the the

y y

\le \le

is is

think think

\, \,

$ W $ $ W $

integral integral

in in

possible possible

x x

$ $

1 1

region. region.

non-negative), non-negative),

y y

range range

from from

$ W $ $ W $

\le \le

this this

by: by:

region region

can can

\le \le

\le \le

x x

need need

defined defined

a a

let's let's

can can

\iiint_W \iiint_W

is is

A A

But But

\le \le

Determine Determine

integration integration

$ $

Since Since $, $,

triple triple

as as

variables. variables.

easier. easier.

variables. variables. each each

the the

we we

Alternatively, Alternatively, is is

compute compute

integral integral

$ $

region region

a a

region region

a a

are are

of of

the the

over over

we we

triple triple

by by

$ $

over over

1: 1:

integral, integral,

bounded bounded

$ $

f f

\le \le

first, first,

ordered ordered

the the

is is

$ $

f f

defines defines

are are

region region

of of

and and y y

z z

So, So,

variables variables

order order

z z

then then

$ $

the the

$. $.

$$ $$

region region

The The

which which

$ $

dV dV

The The

order. order.

x x

\, \,

Step Step

up up

\le \le

$ $

\frac{1}{\text{Volume \frac{1}{\text{Volume

### ###

inequalities inequalities

At(0) At(0.99)

$ $

-----

all all

0 0

xyz xyz

= =

= =

and and

$$ $$

z) z)

value value

function function

z) z)

$, $,

integration integration

To To

a a

y, y, $ $

average average of of

f(x, f(x,

volume volume

need need

f(x, f(x,

\text{Average} \text{Average}

\iiint_W \iiint_W

\le \le

$ $

inequalities inequalities the the

$$ $$

At(0) At(0.99)

function function

then then each each

can can

the the

be be

for for

$ $ set set

y y

$, $, up up

\int_{y=0}^x \int_{y=0}^x Alternatively, Alternatively,

we we

can can

Preprint

At(0) At(0.99)

reverse reverse

the the

order order

order order

for for

now. now.

First, First,

integration, integration,

of of -----

### ###

compute compute

the the xyz xyz

\int_{z=0}^y \int_{z=0}^y

over over

$ $

y y

$: $:

dy dy

$$ $$ is is

the the

The The

volume volume

function function

$ $

\iiint_W \iiint_W

$$ $$ 1 1

\, \,

\, \, dz dz

\left[ \left[

y, y,

dV dV

= =

dy dy

\, \,

dx dx

y y

$$ $$

= =

\frac{1}{3} \frac{1}{3}

= =

$$ $$

\frac{1}{6} \frac{1}{6}

$. $.

-----

$$ $$

Step Step

19

4: 4:

over over

$ $

So So

$ $

Compute Compute

the the

integral integral

of of

triple triple $ $

xyz xyz

= =

W W

$, $,

which which

\int_{y=0}^x \int_{y=0}^x inner inner

integral: integral:

\int_{y=0}^x \int_{y=0}^x

y y $$ $$

dy dy

\frac{1}{2} \frac{1}{2}

\cdot \cdot

= =

\frac{1}{2} \frac{1}{2}

\cdot \cdot

the the

volume volume

Compute Compute

the the

of of

$ $

average average

W W

= =

Then: Then:

= =

\right]_0^1 \right]_0^1

### ###

3: 3:

= =

z) z)

\frac{x^2}{2} \frac{x^2}{2}

= = dx dx

$$ $$

Step Step

1 1

y, y,

f(x, f(x,

Compute Compute

Then: Then:

\frac{1}{6} \frac{1}{6}

\frac{x^6}{6} \frac{x^6}{6}

$ $

the the = =

= =

\int_{x=0}^1 \int_{x=0}^1

$$ $$

\int_{x=0}^1 \int_{x=0}^1

\frac{x^2}{2} \frac{x^2}{2}

\frac{x^3}{3} \frac{x^3}{3}

is is

integral integral

\frac{1}{6} \frac{1}{6}

### ###

z) z)

\right]_0^x \right]_0^x

\int_{x=0}^1 \int_{x=0}^1

$$ $$

\, \,

\frac{y^2}{2} \frac{y^2}{2}

\left[ \left[

$ W $ $ W $

f(x, f(x,

1 1

dz dz 1 1

of of

$: $:

of of

-----

= =

\frac{x}{2} \frac{x}{2}

\left[ \left[

integral integral $. $.

\left[ \left[

\frac{y^4}{4} \frac{y^4}{4}

\cdot \cdot

\frac{1}{48} \frac{1}{48}

$ W $ $ W $

x x

\cdot \cdot

triple triple

$$ $$

\frac{x^4}{4} \frac{x^4}{4} $ $

\frac{1}{8} \frac{1}{8}

So So

\int_{z=0}^y \int_{z=0}^y

$$ $$

\frac{1}{8} \frac{1}{8}

= =

xy xy

= =

\left[ \left[

over over

$: $:

the the

\cdot \cdot

integrate integrate

Now, Now,

$ $

\int_{z=0}^y \int_{z=0}^y

\cdot \cdot

\frac{x}{2} \frac{x}{2}

= =

constant constant

be: be:

\frac{x}{2} \frac{x}{2}

= =

= =

z z

\frac{y^2}{2} \frac{y^2}{2}

\cdot \cdot

y^3}{2} y^3}{2}

dy dy

this this integral integral

$ $

dz dz

\frac{x \frac{x

\frac{1}{48} \frac{1}{48}

should should

xy xy

= =

\, \,

\int_{y=0}^x \int_{y=0}^x

\right]_0^1 \right]_0^1

At(0) At(0.99)

z z

$$ $$

dx dx

the the

\int_0^y \int_0^y

into into

\frac{x^5}{8} \frac{x^5}{8}

of of

xy xy

= =

this this

$$ $$

volume volume

dz dz

with with

over over

substitute substitute

\frac{x^5}{8} \frac{x^5}{8}

$ $

triple triple

Now, Now,

y^3 y^3

W W

the the

$$ $$

= =

$ $

Compute Compute

y^3}{2} y^3}{2}

\right]_0^x \right]_0^x

over over

stick stick

integral integral

\right]_0^y \right]_0^y

\int_{0}^x \int_{0}^x

At(0) At(0.99)

2: 2:

let's let's

inner inner

\, \,

\frac{z^2}{2} \frac{z^2}{2} \frac{x \frac{x

Step Step

but but

$ $

is is value value

$ $

Preprint

At(0) At(0.99)

average average

Now, Now,

value value

\text{Triple \text{Triple = =

6 6

$ $

is: is:

$$ $$

is is

\frac{1}{1/6} \frac{1}{1/6}

= =

\frac{1}{48} \frac{1}{48}

let let

xyz xyz

\frac{1}{\text{Volume}} \frac{1}{\text{Volume}}

$$ $$

integral} integral}

\cdot \cdot

Wait, Wait,

is: is:

me me $ $

check check

1/48 1/48

that that

$, $,

the the

\frac{1/48}{1/6} \frac{1/48}{1/6}

\frac{6}{48} \frac{6}{48} Final Final

\frac{6}{48} \frac{6}{48}

again: again:

Wait, Wait, is is

\frac{1}{8} \frac{1}{8}

$$ $$

the the 1/6 1/6

$, $,

so so

-0.5

of of

average average

the the \cdot \cdot

Yes, Yes,

$$ $$

integral integral

triple triple

that's that's

$ $

value value

\frac{6}{1} \frac{6}{1}

= =

-----

### ###

correct. correct.

\boxed{\dfrac{1}{8}} \boxed{\dfrac{1}{8}}

$$ $$

-1.0

\frac{1}{8} \frac{1}{8}

= =

\frac{1}{48} \frac{1}{48}

= =

= =

Answer: Answer:

$ $

\frac{1}{48} \frac{1}{48}

\cdot \cdot

= =

volume volume

\times \times

$$<|im_end|> $$<|im_end|>

0.0

0.5

1.0

At

Figure 6. Complete token-level advantage visualization for the correct response. The response is (0.99) (0) in the lower shown in token order across seven consecutive panels, with At in the upper row and At row. Blue and red indicate negative and positive advantages, respectively. At(0) At(0.99)

We We

are are

asked asked

to to

evaluate evaluate

the the

\frac{k\pi}{7} \frac{k\pi}{7}

\left( \left(

\right) \right)

trigonometric trigonometric the the

spaced spaced

angles: angles:

sum sum

of of

small small

$ $

n n

squares squares

for for

-----

### ###

$. $.

### ### is is

Step Step

a a

tangent tangent

of of

integer integer 2: 2:

1: 1:

Use Use

$ $

n n

> >

2 2

Apply Apply

the the

a a

known known

identity identity

functions functions

equally equally

at at

\tan^2 \tan^2

\left( \left(

\frac{(n-1)(n-2)}{3} \frac{(n-1)(n-2)}{3}

= =

Step Step

\tan^2 \tan^2

useful useful

\sum_{k=1}^{n-1} \sum_{k=1}^{n-1}

\right) \right) works works

-----

There There

$$ $$

\frac{k\pi}{n} \frac{k\pi}{n}

\sum_{k=1}^{6} \sum_{k=1}^{6}

$$ $$

$$ $$

identity identity

involving involving

identity identity

sum: sum:

$. $.

$$ $$ verify verify

Let's Let's identity identity

In In

our our

This This it it

for for

case, case,

At(0)

n

=

7

$,

so

the

sum

becomes:

$$

\sum_{k=1}^{6}

\tan^2

\left(

At(0.99)

n

=

7

$,

so

the

sum

becomes:

$$

\sum_{k=1}^{6}

\tan^2

\left(

\frac{k\pi}{7}

\right)

=

\frac{(7-1)(7-2)}{3}

=

\frac{6

\frac{k\pi}{7}

\right)

=

\frac{(7-1)(7-2)}{3}

=

\frac{6

$ $

\cdot

5}{3}

=

\frac{30}{3}

=

10

$$

---

###

Final

Answer:

$$

\cdot

5}{3}

=

\frac{30}{3}

=

10

$$

---

###

Final

Answer:

$$

-1.0

\boxed{10}

$$<|im_end|>

\boxed{10}

$$<|im_end|> -0.5

0.0

0.5

1.0

At

Figure 7. Complete token-level advantage visualization for the incorrect response. The two consec(0) (0.99) utive panels show how At can penalize tokens unrelated to the actual error, whereas At assigns stronger negative credit to the erroneous formula derivation.

20

Record · ID 919416 · SHA-256 999b7595f7607b1b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.