Conceptio › Archive › arXiv CS
arXiv CSopen access

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Calibrating Teacher–Student Discrepancy for On-Policy Distillation

Qiangqiang He1

arXiv:2609.21619v1 [cs.AI] 18 Sep 2026

1

Jin Li2

MingCai Chen3

State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China 2 College of Software Engineering, Southeast University, Nanjing, China 3 Nanjing University of Posts and Telecommunications, Nanjing, China

[email protected] [email protected] [email protected]

Abstract On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher–student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger teacher-side likelihood shifts, thereby encouraging the student to learn more of the teacher’s own deviation. We introduce Calibrated On-Policy Distillation (Cal-OPD), which estimates the teacher’s self-deviation region through positive and negative privileged interventions and calibrates the original teacher–student discrepancy by retaining only the component that lies beyond this region. Experiments on mathematical reasoning benchmarks show that, while retaining only about 52–65% of the original teacher–student discrepancy as the optimization signal, Cal-OPD consistently outperforms standard OPD and its variants across model scales.

1 Introduction Knowledge distillation (KD) transfers capabilities from a stronger teacher to a weaker student(Hinton et al., 2015), but conventional off-policy distillation suffers from distribution mismatch between fixed distillation data and the student’s evolving policy distribution(Agarwal et al., 2024; Gu et al., 2024). On-policy distillation (OPD) mitigates this mismatch by training directly on student-generated trajectories and using teacher–student token-level likelihood discrepancies as dense supervision(Agarwal et al., 2024; Yang et al., 2025; Jin et al., 2026). Compared with reinforcement learning with verifiable rewards (RLVR), which relies on sparse outcome-level rewards(Wen et al., 2026; Guo et al., 2025; Yu et al., 2026a), OPD provides token-level guidance throughout the trajectory, enabling more direct and efficient reasoning post-training(Yang et al., 2025; Jin et al., 2026). Standard OPD implicitly treats the teacher likelihood assigned to each student token as an equally reliable reference. Yet its stability varies substantially across tokens. Recent studies on reasoning models suggest that reasoning tokens tend to remain relatively stable, whereas stylistic or surface-form tokens, such as discourse markers and formatting choices, are substantially more sensitive to contextual interventions even when both the question and student rollout are held fixed(Pan et al., 2026; He et al., 2026). Related work further shows that the reliability of teacher supervision can vary across tokens and reasoning positions(Liu et al., 2026). We refer to this token-specific variability in teacher likelihood as Teacher Self-Deviation (TSD). Consequently, the observed teacher–student discrepancy reflects not only the underlying capability gap but also deviations arising from the teacher itself, which standard OPD indiscriminately incorporates into the learning signal. 1

Calibrated On-Policy Distillation

Preprint

Teacher Self-Deviation Region

Standard OPD Calibrated OPD

���������� ����ℎ��

����ℎ��

Privileged OPD

�������

Figure 1: Conceptual comparison of standard OPD, privileged OPD, and Cal-OPD. Cal-OPD removes the TSD-explained discrepancy and retains only the residual beyond the TSD region.

Privileged OPD extends standard OPD by conditioning a stronger teacher on additional training-time information, such as reference solutions, final answers, or hints(Ye et al., 2026; Yu et al., 2026b; Kaur et al., 2026). While such privileged context is intended to improve teacher supervision, it also induces further shifts in the teacher distribution. Existing studies show that these shifts can introduce information-asymmetry effects(Yu et al., 2026b), shortcut behavior(Tian et al., 2026), and even degrade performance in thinking models(Kaur et al., 2026). From the perspective of teacher self-deviation, privileged OPD therefore amplifies the teacher-side deviations already embedded in the teacher–student discrepancy, causing the student to learn more context-induced teacher variation, as illustrated in Figure 1. To address this issue, we propose Calibrated On-Policy Distillation (Cal-OPD), which first estimates the teacher’s self-deviation region and then learns only the teacher–student discrepancy that lies beyond it. Specifically, Cal-OPD applies positive and negative privileged interventions, whose contrasting semantics provide complementary probes of the teacher’s contextual variability and enable an approximation of its token-level self-deviation region. Unlike privileged OPD, these interventions are not directly distilled into the student, but are instead used to measure and calibrate the teacher reference. Cal-OPD then removes the portion of the original discrepancy covered by the estimated self-deviation region, retaining only the residual as the learning signal and setting it to zero when the student likelihood falls within this estimated region. Our contributions are summarized as follows: • We identify and empirically characterize Teacher Self-Deviation (TSD), showing that teacher likelihood is not an equally stable reference across tokens and contexts, and that learning these teacher-side deviations, which are further amplified by privileged context, can substantially degrade OPD for long-CoT reasoning. • We introduce Calibrated On-Policy Distillation (Cal-OPD), which uses positive and negative privileged interventions to estimate the teacher’s self-deviation region and retains only the teacher–student discrepancy beyond it. Privileged information is thus used to calibrate the teacher reference rather than directly supervise the student. • Extensive experiments on mathematical reasoning benchmarks show that Cal-OPD retains only about 52–65% of the original teacher–student discrepancy for optimization, yet consistently outperforms standard OPD and its variants across model scales, while alleviating the response-length expansion observed in standard and privileged OPD.

2

Calibrated On-Policy Distillation

Preprint

2 Understanding Teacher Self-Deviation 2.1 Preliminaries On-Policy Distillation. Let 𝜋 𝑆 = 𝜋 𝜃 denote the student policy parameterized by 𝜃, and let 𝜋𝑇 denote a fixed teacher policy. Given a problem 𝑥 ∼ D, the student generates an on-policy rollout 𝑦 = (𝑦 1 , . . . , 𝑦 𝑇 ) ∼ 𝜋 𝑆 (· | 𝑥),

(1)

where 𝑦 <𝑡 = (𝑦 1 , . . . , 𝑦 𝑡 −1 ) denotes the prefix of the student response preceding token 𝑦 𝑡 . For each student-generated token, we evaluate the teacher and student under the same problem 𝑥 and student response prefix 𝑦 <𝑡 , and define their token-level log-likelihoods and discrepancy as ℓ𝑡𝑇 = log 𝜋𝑇 (𝑦 𝑡 | 𝑥, 𝑦 <𝑡 ),

𝐴𝑡OPD = ℓ𝑡𝑇 − ℓ𝑡𝑆 .

ℓ𝑡𝑆 = log 𝜋 𝑆 (𝑦 𝑡 | 𝑥, 𝑦 <𝑡 ),

OPD uses 𝐴𝑡OPD as the token-level advantage for optimizing the student: # " 𝑇 1 ∑︁  OPD  log 𝜋 𝑆 (𝑦 𝑡 | 𝑥, 𝑦 <𝑡 ) , sg 𝐴𝑡 LOPD (𝜃) = −E 𝑥∼D, 𝑦∼ 𝜋𝑆 (· | 𝑥 ) 𝑇 𝑡=1

(2)

(3)

where sg(·) denotes the stop-gradient operator. Accordingly, 𝐴𝑡OPD > 0 increases the student likelihood of 𝑦 𝑡 , whereas 𝐴𝑡OPD < 0 decreases it. This formulation treats the teacher likelihood ℓ𝑡𝑇 as the token-level reference against which the student is optimized. Teacher Self-Deviation. For a fixed problem 𝑥, student rollout 𝑦, and token position 𝑡, we introduce a teacher-side contextual intervention 𝑐 while keeping the evaluated student trajectory unchanged. The intervention is provided only to the teacher, together with the problem 𝑥, and precedes the student rollout prefix 𝑦 <𝑡 in the teacher context. We use 𝑐 0 = ∅ to denote the original setting without additional context. The teacher likelihoods under intervention 𝑐 and the original setting 𝑐 0 , together with the induced likelihood variation, are defined as ℓ𝑡𝑇 (𝑐) = log 𝜋𝑇 (𝑦 𝑡 | 𝑥, 𝑐, 𝑦 <𝑡 ),

ℓ𝑡𝑇 (𝑐 0 ) = log 𝜋𝑇 (𝑦 𝑡 | 𝑥, 𝑦 <𝑡 ) = ℓ𝑡𝑇 ,

Δ𝑇𝑡 (𝑐) = ℓ𝑡𝑇 (𝑐) − ℓ𝑡𝑇 (𝑐 0 ).

(4)

Since 𝑥, 𝑦 <𝑡 , and 𝑦 𝑡 remain fixed throughout the comparison, Δ𝑇𝑡 (𝑐) captures the change in teacher likelihood induced by the additional context 𝑐, rather than any change in the evaluated trajectory. We refer to this context-induced variation in teacher likelihood as teacher self-deviation (TSD). Given an intervention set C, we use the induced likelihood variations to estimate the downward and upward magnitudes of TSD:     𝑑ˆ𝑡↓ = max 0, − min Δ𝑇𝑡 (𝑐) ,

𝑑ˆ𝑡↑ = max 0, max Δ𝑇𝑡 (𝑐) .

𝑐∈ C

𝑐∈ C

(5)

Both 𝑑ˆ𝑡↓ and 𝑑ˆ𝑡↑ are non-negative and measure the maximum downward and upward deviations in teacher loglikelihood, respectively. These quantities yield an empirical estimate of the TSD region: h i R̂ 𝑡𝑇 (C) = ℓ𝑡𝑇 − 𝑑ˆ𝑡↓ , ℓ𝑡𝑇 + 𝑑ˆ𝑡↑ . (6) Accordingly, R̂ 𝑡𝑇 (C) is a finite-intervention approximation to the underlying teacher self-deviation region R 𝑡𝑇 , rather than an exhaustive characterization of all possible contextual variation. TSD characterizes variability in the teacher likelihood, while student-generated rollouts provide the trajectories on which this variability is measured. 3

Calibrated On-Policy Distillation

Preprint

Table 1: Contextual interventions used to probe TSD. The shared problem context 𝑥 is omitted for brevity. For answer- and solution-level interventions, the positive and negative variants share the same template and differ only in the inserted privileged content. Symbol pos

𝑐 inst neg

𝑐 inst pos

𝑐 eval neg

𝑐 eval

pos/neg

𝑐 ans

pos/neg

𝑐 sol

Contextual Intervention Task-Agnostic Instructions Please reason through the problem carefully and thoroughly. Verify intermediate steps and provide a complete, rigorous solution. Please solve the problem quickly and directly. Avoid unnecessary elaboration or extensive verification and reach the final answer efficiently. Evaluative Feedback A gold-standard verifier has judged that the following solution reaches the correct final answer. The reasoning is rigorous, coherent, and mathematically sound. A gold-standard verifier has judged that the following solution does not reach the correct final answer. The reasoning is flawed, incoherent, and mathematically unreliable. Answer-Level Privilege A verified ground-truth answer is provided as a reliable reference: [correct / incorrect answer]. Use it to guide your reasoning while providing a complete and logically coherent solution. Solution-Level Privilege A reference solution is provided as additional guidance: [correct / unrelated solution]. Use it while independently providing a complete and logically coherent solution.

2.2 Teacher Self-Deviation Is Not Reliable Knowledge TSD captures contextual changes in the teacher likelihood assigned to a fixed token. If TSD faithfully reflected task-relevant knowledge, its variation should be systematically tied to the information introduced by the context. In particular, TSD should depend on the presence of task-specific information and respond consistently to the correctness of that information. Otherwise, the observed likelihood shift cannot be reliably interpreted as a knowledge signal. Table 1 summarizes the contextual interventions used in our analysis, spanning task-agnostic instructions, evaluative feedback, and answer- and solution-level privileged information. We conduct this analysis with Qwen3-1.7B as the student and Qwen3-8B(Yang et al., 2025) as the teacher, both in thinking mode, over 6,528 questions sampled from DAPO-17k(Yu et al., 2026a), with one response per question and approximately 60 million response tokens evaluated across all intervention conditions. Details of data collection and preparation for TSD analysis are provided in Appendix A.2. TSD Emerges Without Task Knowledge. We first find that substantial TSD emerges even in the absence of external task-specific knowledge, and that the affected token positions are largely preserved when richer privileged information is subsequently introduced. This indicates that part of the teacher-side variability amplified by privileged context is already present before answer- or solution-level knowledge is provided. To quantify this effect, we group the positive and negative variants of each intervention group as  pos neg C𝑔 = 𝑐 𝑔 , 𝑐 𝑔 , 𝑔 ∈ {inst, eval, ans, sol}. (7) Given a threshold 𝜏, we define the set of tokens exhibiting significant TSD under group C𝑔 as   𝑇 S𝑔 (𝜏) = 𝑡 : max Δ𝑡 (𝑐) > 𝜏 . 𝑐∈ C𝑔

(8)

Figure 2(a) reports the prevalence of significant TSD under each intervention, together with the union prevalence defined by S𝑔 (𝜏). At 𝜏 = 0.01, task-agnostic instructions induce significant TSD on 20.2% and 25.8% of tokens under the positive and negative variants, respectively, with their union covering 29.8% of all tokens. This prevalence 4

Calibrated On-Policy Distillation

Preprint

Positive

Negative

Group Union Sg (τ)

τ = 0.01

τ = 0.01

1.00

39.6

C inst

100.0%

86.5%

86.0%

98.3%

C eval

88.7%

100.0%

87.6%

98.6%

C ans

87.4%

86.8%

100.0%

98.2%

C sol

74.2%

72.6%

72.9%

100.0%

34.4

37.0 29.8 25.8

0.80

29.4

29.1 25.0

25.1

Source C i

Tokens with Significant TSD (%)

40

30

25.2

21.9

20.2

20

10

Evaluation

Answer

0.60

0.40

C inst Instruction

Ret i→j (τ)

Solution

C eval

C ans

C sol

Target C j

(a) Prevalence of Significant TSD

(b) Retention of Significant TSD

Figure 2: TSD across contextual interventions. Significant TSD emerges even under task-agnostic interventions, and the affected token positions are largely retained under richer privileged contexts. Ovl g (τ)

Agr g (τ)

100

Percentage (%)

τ = 0.01

88.7

83.0 80

SDR g (τ)

71.3

70.7

74.7

61.1

80.3 79.1 66.1

60 40 20 0

Evaluation

Answer

Solution

Figure 3: Consistency of TSD under contrasting interventions. TSD occurs at overlapping token positions, remains directionally aligned, and is dominated by a shared common-mode component.

is comparable to the 29.1% observed under evaluative feedback and the 29.4% under answer-level privilege, despite task-agnostic interventions providing neither an answer nor a solution. Solution-level privilege further expands the affected set to 39.6%, showing that richer privileged context broadens TSD rather than creating it from scratch. We measure whether TSD under less informative interventions persists under richer privileged contexts. For groups C𝑖 and C 𝑗 , we define the retention of significant TSD as Ret𝑖→ 𝑗 (𝜏) =

S𝑖 (𝜏) ∩ S 𝑗 (𝜏) . |S𝑖 (𝜏)|

(9)

Figure 2(b) reveals a clear asymmetry. Under solution-level privilege, 98.3%, 98.6%, and 98.2% of tokens exhibiting significant TSD under task-agnostic instructions, evaluative feedback, and answer-level privilege remain significant, respectively, whereas only 74.2%, 72.6%, and 72.9% of solution-level significant-TSD tokens remain significant under these less informative interventions. Thus, richer privileged context largely preserves previously affected positions while extending TSD to additional positions. Together, these results show that task knowledge is not necessary for TSD to emerge, while richer privileged contexts mainly broaden the affected set rather than introducing an entirely new deviation pattern.

5

Calibrated On-Policy Distillation

Preprint pos

Table 2: Token forms with the highest and lowest significant-TSD rates 𝜌 𝜏 (𝑣) under 𝑐 sol at 𝜏 = 0.01, restricted to forms occurring more than 20,000 times. |Δ| denotes the mean TSD magnitude over significant occurrences. Values are rounded for display; rankings are based on the unrounded 𝜌 𝜏 (𝑣). For compactness, altern. abbreviates alternatively. 𝝆𝝉 (%)

|𝚫|

# Token

1 maybe 2 however 3 therefore 4 consider 5 altern. 6 earlier

97.4 97.4 95.9 95.4 94.3 94.2

0.59 0.63 0.46 0.38 0.58 0.41

Highest 𝜌 𝜏 7 seems 93.6 0.41 8 another 92.9 0.50 9 since 91.8 0.44 10 try 91.8 0.40 11 think 91.6 0.40 12 says 90.4 0.25

1 }{ 2 ### 3 0 4 9 5 8 √ 6

2.8 0.25 3.2 0.13 5.0 0.36 6.8 0.38 7.2 0.39 7.2 0.28

# Token

𝝆𝝉 (%)

|𝚫|

# Token

𝝆𝝉 (%)

|𝚫|

13 wait 14 check 15 let 16 here 17 because 18 now

90.0 0.42 89.6 0.28 89.5 0.50 89.4 0.36 89.4 0.40 89.2 0.49

13 frac 14 { 15 4 16 2 17 3 18 _i

8.0 0.22 8.0 0.23 8.8 0.37 8.8 0.32 9.2 0.36 9.2 0.28

Lowest 𝜌 𝜏 7 ◦ 8 _ 9 6 10 7 11 5 12 𝜃

7.3 0.17 7.3 0.29 7.4 0.40 7.7 0.39 7.8 0.38 7.9 0.28

TSD Is Largely Insensitive to Intervention Semantics. We find that TSD is largely insensitive to both semantic polarity and correctness. To quantify this consistency, we compare the positive and negative variants of evaluative, pos answer-level, and solution-level interventions using three metrics. For intervention group 𝑔, let Δ+𝑡 = Δ𝑇𝑡 (𝑐 𝑔 ) neg and Δ𝑡− = Δ𝑇𝑡 (𝑐 𝑔 ), with corresponding significant-TSD sets 𝑆 𝑔+ (𝜏) = {𝑡 : |Δ+𝑡 | > 𝜏} and 𝑆 𝑔− (𝜏) = {𝑡 : |Δ𝑡− | > 𝜏}. Overlap measures whether significant TSD occurs at the same token positions, while Agreement measures whether the two interventions shift teacher likelihood in the same direction at jointly significant positions:  + −  Í |𝑆 𝑔+ (𝜏) ∩ 𝑆 𝑔− (𝜏)| 𝑡 ∈𝑆𝑔+ ( 𝜏 )∩𝑆𝑔− ( 𝜏 ) 1 Δ𝑡 Δ𝑡 > 0 Ovl𝑔 (𝜏) = + , Agr𝑔 (𝜏) = . (10) |𝑆 𝑔 (𝜏) ∪ 𝑆 𝑔− (𝜏)| |𝑆 𝑔+ (𝜏) ∩ 𝑆 𝑔− (𝜏)| We further measure how much paired variation is shared between the two interventions. Let 𝐼𝑔 (𝜏) = 𝑆 𝑔+ (𝜏) ∩ 𝑆 𝑔− (𝜏) denote the set of jointly significant tokens. We decompose the paired deviations into shared and contrastive components and define the Shared Deviation Ratio (SDR) as Í Δ+𝑡 + Δ𝑡− Δ+𝑡 − Δ𝑡− 𝑡 ∈ 𝐼𝑔 ( 𝜏 ) |𝑀𝑡 | 𝑀𝑡 = , 𝐷𝑡 = , SDR𝑔 (𝜏) = Í . (11) 2 2 𝑡 ∈ 𝐼𝑔 ( 𝜏 ) (|𝑀𝑡 | + |𝐷 𝑡 |) A higher shared deviation ratio indicates that, among jointly significant tokens, paired TSD is dominated by variation shared across interventions, with less attributable to their semantic difference. Figure 3 reveals that TSD remains highly structured under contrasting intervention semantics. Reversing answer correctness still yields 88.7% directional agreement and a 74.7% shared deviation ratio, showing that paired likelihood shifts are dominated by a shared response rather than the correctness contrast itself. Solution-level interventions exhibit the highest positional overlap at 80.3%, yet the lowest shared deviation ratio at 66.1%, suggesting that richer context alters how TSD varies more than where it emerges. Across all groups, substantial overlap, high directional agreement, and predominantly shared variation persist. These results show that TSD is largely insensitive to intervention semantics and correctness, further indicating that such teacher likelihood shifts cannot be reliably interpreted as task knowledge. Additional analyses across multiple thresholds and model scales in Appendix A.3 reproduce these findings.

6

Calibrated On-Policy Distillation

Preprint

2.3 TSD Concentrates on Surface-Form Tokens TSD is strongly concentrated on surface-form tokens rather than tokens carrying mathematical content. HighTSD forms are natural-language markers that organize, qualify, or redirect the reasoning text without directly encoding problem-specific mathematical information, whereas low-TSD forms are dominated by numbers, mathematical pos symbols, and notation. Under 𝑐 sol , we merge tokenizer variants corresponding to the same surface form and restrict the analysis to forms  occurring more than 20,000 times. For each token form 𝑣, we define the significant-TSD pos rate as 𝜌 𝜏 (𝑣) = Pr |Δ𝑇𝑡 (𝑐 sol )| > 𝜏 | 𝑦 𝑡 = 𝑣 . Unlike occurrence counts, 𝜌 𝜏 (𝑣) measures how often a token exhibits significant TSD when it appears, thereby controlling for differences in token frequency. Table 2 reveals a clear separation between token forms with the highest and lowest significant-TSD rates at 𝜏 = 0.01. The 18 highest-ranked forms are dominated by natural-language surface expressions such as maybe, however, therefore, consider, and alternatively, with 𝜌 𝜏 exceeding 89% throughout, meaning significant TSD appears in about nine of ten occurrences. In contrast, the 18 lowest-ranked forms consist predominantly √ of digits, mathematical symbols, and notation such as 0, , 𝜃, and frac, all with 𝜌 𝜏 below 9.3%. This nearly order-of-magnitude separation shows that significant TSD occurs far more frequently on surface-form tokens than on tokens directly expressing mathematical content. The large |Δ| values in Table 2 mainly reflect the stronger teacher-side shifts induced by solution-level privilege. Appendix A.4 extends the analysis to the top-24 and bottom-24 token forms across different contextual interventions, showing that weaker interventions substantially reduce deviation magnitudes while preserving the same separation between surface-form tokens and mathematical or symbolic forms. Appendix A.5 provides a trajectory-level case study illustrating how TSD is distributed throughout a complete reasoning trace.

3 Calibrated On-Policy Distillation The preceding analysis shows that teacher likelihood is not an equally reliable pointwise reference: it exhibits substantial TSD that cannot be reliably interpreted as a knowledge signal. We therefore propose Calibrated OnPolicy Distillation (Cal-OPD), which decomposes the observed teacher–student discrepancy into a TSD-explained component and a calibrated residual, and distills only the discrepancy that remains beyond the teacher’s estimated self-deviation region. 𝐴𝑡OPD |{z} Teacher–Student Discrepancy

=

𝐴𝑡Cal |{z} Calibrated Teacher–Student Discrepancy

+

𝐴𝑡TSD |{z}

.

(12)

TSD-Explained Discrepancy

For each token 𝑦 𝑡 , Cal-OPD probes the teacher with two contrasting interventions 𝑐pos and 𝑐neg , producing neg pos Δ𝑡 = Δ𝑇𝑡 (𝑐pos ) and Δ𝑡 = Δ𝑇𝑡 (𝑐neg ). We define their maximum downward and upward deviations as 𝑑ˆ𝑡↓ = neg neg pos pos max(0, −Δ𝑡 , −Δ𝑡 ) and 𝑑ˆ𝑡↑ = max(0, Δ𝑡 , Δ𝑡 ). Since two interventions provide only a finite probe of the underlying TSD, we introduce a relaxation factor 𝜆 ≥ 1 and estimate the TSD region as h i R̂ 𝑡𝑇 = ℓ𝑡𝑇 − 𝜆 𝑑ˆ𝑡↓ , ℓ𝑡𝑇 + 𝜆 𝑑ˆ𝑡↑ = [𝐿 𝑇𝑡 , 𝑈𝑡𝑇 ].

(13)

Cal-OPD then removes the portion of the teacher–student discrepancy covered by the estimated TSD region and retains only the residual beyond its boundary:     𝐴𝑡Cal = 𝐿 𝑇𝑡 − ℓ𝑡𝑆 + − ℓ𝑡𝑆 − 𝑈𝑡𝑇 + , [𝑧] + = max(𝑧, 0). (14) When ℓ𝑡𝑆 ∈ R̂ 𝑡𝑇 , the observed teacher–student discrepancy is fully covered by the estimated TSD region and 𝐴𝑡Cal = 0; otherwise, 𝐴𝑡Cal retains only the discrepancy beyond the nearest boundary. Accordingly, 𝐴𝑡TSD = 𝐴𝑡OPD − 𝐴𝑡Cal . Cal-OPD follows the standard OPD objective in Eq. 3, replacing 𝐴𝑡OPD with 𝐴𝑡Cal as the token-level advantage. 7

Calibrated On-Policy Distillation

Preprint

Table 3: Main results on mathematical reasoning benchmarks. The best result among distillation methods for each teacher–student configuration is shown in bold. Subscripts on Cal-OPD indicate gains over the corresponding student. For compactness, 4B-2507 and 30B-2507 denote Qwen3-4B-Thinking-2507 and Qwen3-30B-A3B-Thinking-2507, respectively. Model / Method

AMC23

AIME24

AIME25

AIME26

HMMT26

MATH500

Avg.

Student Teacher OPD ExOPD EOPD Uni-OPD Privileged-OPD Cal-OPD

80.8 94.8 82.0 81.7 81.6 81.4 83.3 84.5+3.7

4B-2507 → Qwen3-1.7B 40.6 32.5 29.2 56.5 53.1 54.2 39.6 35.0 34.6 43.1 35.4 34.2 44.0 36.7 38.8 42.1 34.4 32.7 39.8 30.4 31.0 45.0+4.4 37.1+4.6 36.3+7.1

21.8 25.2 22.9 25.6 19.7 22.5 21.8 24.8+3.0

90.2 95.4 90.6 90.6 90.6 89.8 89.4 90.8+0.6

49.2 63.2 50.8 51.8 51.9 50.5 49.3 53.1+3.9

Student Teacher OPD ExOPD EOPD Uni-OPD Privileged-OPD Cal-OPD

95.2 97.8 94.2 92.0 95.3 93.4 93.4 96.3+1.1

30B-2507 → Qwen3-4B 64.8 56.3 57.9 74.8 65.4 66.0 62.1 57.5 55.2 66.9 59.4 56.3 62.7 54.6 60.0 67.9 59.8 57.9 58.8 52.1 54.0 67.1+2.3 61.3+5.0 60.2+2.3

30.1 33.9 31.1 33.7 33.5 31.4 31.3 33.3+3.2

95.4 96.9 95.0 95.6 96.7 95.4 94.4 95.8+0.4

66.6 72.5 65.9 67.3 67.1 67.6 64.0 69.0+2.4

4 Experiments 4.1 Experimental Setup Datasets. We use DAPO-17K (Yu et al., 2026a) filtered by Qwen3-235B-A22B-Instruct-2507(Yang et al., 2025) as the training dataset. The filtering and solution-annotation procedure is described in Appendix A.2. For evaluation, we consider six reasoning benchmarks: AMC23, AIME24, AIME25, AIME26, HMMT26, and MATH500 (Lightman et al., 2024). These benchmarks span a range of difficulty, from standard problem solving to high-difficulty competition mathematics. Models and Baselines. We evaluate Cal-OPD on two teacher–student configurations from the Qwen3 family (Yang et al., 2025): Qwen3-4B-Thinking-2507→Qwen3-1.7B and Qwen3-30B-A3B-Thinking-2507→Qwen3-4B, covering two distinct scale regimes. We compare against five representative OPD baselines: standard OPD (Agarwal et al., 2024), which directly optimizes teacher–student discrepancies; ExOPD (Yang et al., 2026a), augmented with reward extrapolation; EOPD (Jin et al., 2026), with entropy-aware forward-KL supervision; Uni-OPD (Hou et al., 2026a), with exploration and outcome-guided calibration; and Privileged-OPD (Ye et al., 2026; Kaur et al., 2026), where the teacher is additionally conditioned on the ground-truth reference solution. This evaluates if Cal-OPD remains effective across varying capacity gaps and diverse modifications of the standard OPD objective. We report Avg@16 accuracy, computed by sampling 16 independent responses per problem and averaging their binary correctness scores across the full evaluation set. Implementation Details. We implement all methods with verl (Sheng et al., 2025) and train on 8 NVIDIA H20 GPUs, with 4 GPUs hosting the student and 4 hosting the teacher. All methods are trained for 100 steps, with 256 trajectories per step, one rollout per question unless otherwise specified, and a learning rate of 1 × 10−6 .

8

Calibrated On-Policy Distillation

Preprint

Table 4: Effect of contextual interventions used for TSD estimation on Cal-OPD performance for Qwen3-4B-Thinking-2507→Qwen3-1.7B. Intervention Set Cinst Ceval Cans Csol

AMC23

AIME24

AIME25

AIME26

HMMT26

MATH500

Avg.

83.9 84.5 84.2 81.3

40.8 45.0 40.8 40.6

34.8 37.1 34.8 30.0

34.2 36.3 37.3 30.8

25.2 24.8 23.9 22.0

90.3 90.8 90.9 89.2

51.5 53.1 52.0 49.0

During training, we set response length to 16,384, temperature to 1.0, and top-𝑝 to 1.0; during evaluation, we use neg pos response length of 20,480, temperature 0.6, and top-𝑝 to 0.95. For Cal-OPD, we use 𝑐 eval and 𝑐 eval as contrasting interventions and set the relaxation factor to 𝜆 = 5. Method-specific hyperparameters are provided in Appendix A.6.

4.2 Main Results Table 3 shows that Cal-OPD achieves the highest average performance in both teacher–student configurations: 53.1 for Qwen3-4B-Thinking-2507→Qwen3-1.7B and 69.0 for Qwen3-30B-A3B-Thinking-2507→Qwen3-4B. This corresponds to gains of +3.9 and +2.4 over the student baselines, and +2.3 and +3.1 over standard OPD. Cal-OPD also achieves the best result on 7 of 12 benchmark–configuration pairs, demonstrating consistent gains across model scales and reasoning benchmarks. In contrast, standard OPD improves the 4B→1.7B student only modestly, from 49.2 to 50.8, and even degrades the 30B→4B student from 66.6 to 65.9, despite both teachers being substantially stronger than their students. This shows that raw teacher–student discrepancy is not uniformly beneficial supervision: optimizing all discrepancies also learns teacher-side deviations, whereas Cal-OPD removes the TSD-explained component and retains a more effective signal. Privileged-OPD exhibits the strongest degradation, obtaining the lowest average performance among distillation methods in both configurations, at 49.3 and 64.0, respectively. In the 4B→1.7B setting, it nearly eliminates the gain from standard OPD, while in the 30B→4B setting it falls 2.6 points below the student. Together with our earlier finding that privileged context substantially amplifies TSD, these results suggest that directly distilling the privileged teacher can transfer teacher-side deviation alongside task-relevant information, offsetting the benefit of stronger supervision. Cal-OPD mitigates this effect by filtering out the TSD-explained component before distillation. 4.3 Analysis and Ablations Effect of Contextual Interventions on Cal-OPD Performance. Table 4 compares different intervention sets with 𝜆 = 5 throughout, while Figure 4 shows their calibration dynamics. Ceval achieves the best performance, reaching an average of 53.1. Although evaluative feedback introduces external judgment, it provides no task-specific solution knowledge and yields stronger calibration than instruction interventions, which retain nearly 70% of the teacher–student discrepancy. In contrast, Csol causes the largest degradation, with the average falling to 49.0 while retaining only about 20% of the discrepancy. This suggests that solution-induced TSD contains a larger task-relevant component, such that using it for calibration over-filters useful teacher supervision. Overall, Ceval provides a stronger probe of TSD without the excessive filtering induced by solution-level privilege. Further training dynamics, including entropy and other measures, are provided in Appendix A.7. Effects of Calibration on OPD Training Dynamics. Figure 4 shows the training dynamics of the main Ceval configuration. The retained-discrepancy ratio decreases from 65% to 52%, while the zero-advantage token ratio increases from 27% to 34%, indicating stronger filtering of the OPD signal. Calibration also alters response-length dynamics. Standard OPD expands the average response length from about 9.8K to 11.8K tokens, whereas Cal-OPD ends at only 9.3K after an initial decrease to 8.5K. Despite the additional teacher computation for estimating TSD, the shorter trajectories make Cal-OPD approximately 1.26× faster to train. Together with the main results, these

9

Calibrated On-Policy Distillation

Preprint

60 12K

0.6 0.5

Cal-OPD (inst ) Cal-OPD (eval )

0.4

Cal-OPD (ans ) Cal-OPD (sol )

0.3

Average Response Length

Zero-Advantage Tokens (%)

Retained OPD Discrepancy

0.7

50

Cal-OPD (inst ) Cal-OPD (eval )

40

Cal-OPD (ans ) Cal-OPD (sol )

10K

8K

OPD Cal-OPD (inst ) Cal-OPD (eval )

6K

30

Cal-OPD (ans ) Cal-OPD (sol )

0.2 0

20

40

60

80

100

0

20

40

Training Steps

60

80

100

0

20

40

Training Steps

(a) Retained OPD Discrepancy

60

80

100

Training Steps

(b) Zero-Advantage Token Ratio

(c) Average Response Length

Figure 4: Training dynamics of Cal-OPD under different TSD-estimation interventions for Qwen3-4B-Thinking-2507→Qwen3-1.7B. Avg AMC23 AIME24 HMMT26 MATH500

+1.4

Δ Avg@16

1 +0.2

+0.0

0 -0.5

−1

-1.4

−2

-1.7

Retained OPD Discrepancy

1.0 2

λ=1 λ=5 λ=10

0.8

λ=20 λ=40 λ=80

0.6

0.4

0.2 −3 1

5

10

20

40

80

0

λ

20

40

60

80

100

Training Steps

(a) Impact of 𝝀 on Performance

(b) Impact of 𝝀 on Retained OPD Discrepancy

Figure 5: Effect of the relaxation factor 𝜆 on Cal-OPD performance and retained discrepancy ratio.

dynamics show that filtering the TSD-explained component yields a more effective and efficient supervision signal. A complete efficiency analysis is provided in Appendix A.8. Effect of the Relaxation Factor 𝝀. Figure 5 shows that Cal-OPD achieves optimal performance at 𝜆 = 5, retaining approximately 52% of the teacher–student discrepancy. Increasing 𝜆 beyond 5 sharply degrades downstream performance, eventually dropping 1.7 points below the 𝜆 = 1 baseline at 𝜆 = 80. As 𝜆 increases, the estimated TSD region expands too broadly, reducing the retained discrepancy to roughly 20% and over-filtering task-relevant supervision. We further compare against a TSD-threshold filtering baseline matched in retained teacher–student discrepancy in Appendix A.9. Cal-OPD remains stronger, indicating that its gains are not due to signal attenuation alone.

5 Conclusion We show that the teacher–student discrepancy used by standard OPD contains substantial teacher self-deviation (TSD), which emerges even without task-specific knowledge, remains largely insensitive to intervention semantics and correctness, and concentrates on surface-form tokens. Based on these findings, we introduce Calibrated On-Policy Distillation (Cal-OPD), which estimates the teacher’s self-deviation region and removes the TSD-explained portion of the OPD signal before optimization. Across different teacher–student scales and mathematical reasoning benchmarks, Cal-OPD consistently outperforms standard OPD and its variants while retaining only 52–65% of the original discrepancy and mitigating response-length expansion during training. These results suggest that effective on-policy distillation depends not on indiscriminately learning more teacher–student discrepancy, but on identifying which 10

Calibrated On-Policy Distillation

Preprint

discrepancy is truly worth learning.

References Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In International Conference on Learning Representations, volume 2024, pages 21246–21263, 2024. Hongzhan Chen, Siyue Wu, Xiaojun Quan, Rui Wang, Ming Yan, and Ji Zhang. Mcc-kd: Multi-cot consistent knowledge distillation. In Findings of the association for computational linguistics: EMNLP 2023, pages 6805–6820, 2023. Chengwei Dai, Kun Li, Wei Zhou, and Songlin Hu. Improve student’s reasoning generalizability through cascading decomposed cots distillation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 15623–15643, 2024. Yi Ding and Ruqi Zhang. Does on-policy distillation really distill? from noisy teacher to self-improvement. arXiv preprint arXiv:2608.31046, 2026. Tao Feng, Yicheng Li, Li Chenglin, Hao Chen, Fei Yu, and Yin Zhang. Teaching small language models reasoning through counterfactual distillation. In Proceedings of the 2024 conference on empirical methods in natural language processing, pages 5831–5842, 2024. Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. Minillm: Knowledge distillation of large language models. In International Conference on Learning Representations, volume 2024, pages 32694–32717, 2024. Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. Qiangqiang He, Zhongheng Wu, and ZiJian Wang. Not all tokens deserve equal credit: Counterfactual sensitivity credit reallocation for long-cot reasoning. arXiv preprint arXiv:2607.27888, 2026. Byeongho Heo, Jaehui Hwang, Sangdoo Yun, and Dongyoon Han. On-policy delta distillation for multilingual math reasoning. arXiv preprint arXiv:2608.05802, 2026. Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. Namgyu Ho, Laura Schmid, and Se-Young Yun. Large language models are reasoning teachers. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), pages 14852–14882, 2023. Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, et al. Uni-opd: Unifying on-policy distillation with a dual-perspective recipe. arXiv preprint arXiv:2605.03677, 2026a. ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, et al. Dash: Divergence-adaptive supervision horizons for on-policy self-distillation of reasoning models. arXiv preprint arXiv:2608.06243, 2026b.

11

Calibrated On-Policy Distillation

Preprint

Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. In Findings of the association for computational linguistics: ACL 2023, pages 8003–8017, 2023. Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, et al. Reinforcement learning via self-distillation. arXiv preprint arXiv:2601.20802, 2026. Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, and Junpei Komiyama. Privileged solutions or context-induced teacher behavior? dissecting on-policy self-distillation. arXiv preprint arXiv:2608.09228, 2026. Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. In Findings of the association for computational linguistics: EMNLP 2020, pages 4163–4174, 2020. Woogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei, Yi Zhou, Swanand Ravindra Kadhe, Nathalie Baracaldo, and Kimin Lee. Entropy-aware on-policy distillation of language models. arXiv preprint arXiv:2603.07079, 2026. Simran Kaur, Narutatsu Ri, Yinghui He, Liam Fowl, and Sanjeev Arora. Rethinking on-policy self-distillation for thinking models. arXiv preprint arXiv:2607.05184, 2026. Junlong Ke, Zichen Wen, Weijia Li, Conghui He, and Linfeng Zhang. Respecting self-uncertainty in on-policy self-distillation for efficient llm reasoning. arXiv preprint arXiv:2605.13255, 2026. Yoon Kim and Alexander M Rush. Sequence-level knowledge distillation. In Proceedings of the 2016 conference on empirical methods in natural language processing, pages 1317–1327, 2016. Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se-Young Yun. Distillm: Towards streamlined distillation for large language models. arXiv preprint arXiv:2402.03898, 2024. Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. Symbolic chain-of-thought distillation: Small models can also “think” step-by-step. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2665–2679, 2023. Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, and Nuno Vasconcelos. On-policy self-distillation without any supervision. arXiv preprint arXiv:2608.06296, 2026a. Yuhan Li, Mingxu Zhang, Dazhong Shen, and Ying Sun. Phf: Privileged hidden flow for on-policy self-distillation. arXiv preprint arXiv:2606.29340, 2026b. Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In International Conference on Learning Representations, volume 2024, pages 39578–39601, 2024. Xiaogeng Liu, Xinyan Wang, Yingzi Ma, Yechao Zhang, and Chaowei Xiao. When are teacher tokens reliable? position-weighted on-policy self-distillation for reasoning. arXiv preprint arXiv:2605.21606, 2026. Arindam Mitra, Luciano Del Corro, Shweti Mahajan, Andres Codas, Clarisse Simoes, Sahaj Agarwal, Xuxi Chen, Anastasia Razdaibiedina, Erik Jones, Kriti Aggarwal, et al. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2311.11045, 2023.

12

Calibrated On-Policy Distillation

Preprint

Leyi Pan, Shuchang Tao, Yunpeng Zhai, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, and Lijie Wen. Rlcsd: Reinforcement learning with contrastive on-policy self-distillation. arXiv preprint arXiv:2606.11709, 2026. Emiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste, Laurent Charlin, and Massimo Caccia. Privileged information distillation for language models. arXiv preprint arXiv:2602.04942, 2026. Hejian Sang, Yuanda Xu, Zhengze Zhou, Ran He, Zhipeng Wang, and Jiachen Sun. On-policy self-distillation for reasoning compression. arXiv e-prints, pages arXiv–2603, 2026. Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019. Guobin Shen, Xiang Cheng, Chenxiao Zhao, Lei Huang, Jindong Li, Dongcheng Zhao, and Xing Yu. Anti-selfdistillation for reasoning rl via pointwise mutual information. arXiv preprint arXiv:2605.11609, 2026a. Zhanming Shen, Jintao Tong, Shaotian Yan, Chen Shen, Hao Chen, Wentao Ye, Xiaomeng Hu, Rui Miao, Haobo Wang, Junbo Zhao, et al. Purified opsd: On-policy self-distillation without losing how to think. arXiv preprint arXiv:2607.02234, 2026b. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf framework. In Proceedings of the Twentieth European Conference on Computer Systems, pages 1279–1297, 2025. Samyak Shrestha and Alexander Tessier. Rethinking privileged information in on-policy self-distillation. arXiv preprint arXiv:2608.18271, 2026. Alex Stein, Furong Huang, and Tom Goldstein. Gates: Self-distillation under privileged context with consensus gating. arXiv preprint arXiv:2602.20574, 2026. Zhiquan Tan and Yinrong Hong. Self-supervised on-policy distillation for reasoning language models. arXiv preprint arXiv:2605.17497, 2026. Kanghui Tian, Siyuan Liu, Ziang Yan, Sheng Xia, Shuai Dong, and Yi Wang. Vicur: Visual cues as recoverable privilege for multimodal on-policy distillation. arXiv preprint arXiv:2606.05718, 2026. Xumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye, Zhirong Wu, Yang Wang, Zhijian Xu, Xiao Liang, Junjie Li, Ziming Miao, et al. Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base llms. In International Conference on Learning Representations, volume 2026, pages 49450–49483, 2026. Jianyu Wu, Yizhou Wang, Encheng Su, Chen Tang, and Shixiang Tang. Dapd: Dual-anchored policy distillation. arXiv preprint arXiv:2608.01735, 2026. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, and Yankai Lin. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation. arXiv preprint arXiv:2602.12125, 2026a. Yuxiao Yang, Xiaoyun Wang, and Weitong Zhang. Ogls-sd: On-policy self-distillation with outcome-guided logit steering for llm reasoning. arXiv preprint arXiv:2605.12400, 2026b.

13

Calibrated On-Policy Distillation

Preprint

Tianzhu Ye, Li Dong, Xun Wu, Shaohan Huang, and Furu Wei. On-policy context distillation for language models. arXiv preprint arXiv:2602.12275, 2026. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. Advances in Neural Information Processing Systems, 38:113222–113244, 2026a. Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo Feng, et al. Dopd: Dual on-policy distillation. arXiv preprint arXiv:2606.30626, 2026b. Ke Zhang, Yunjie Tian, Dongdi Zhao, Yijiang Li, Yuanye Liu, Vishal M Patel, and Di Fu. On-policy distillation with best-of-n teacher rollout selection. arXiv preprint arXiv:2605.09725, 2026. Songming Zhang, Xue Zhang, Zengkui Sun, Yufeng Chen, and Jinan Xu. Dual-space knowledge distillation for large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 18164–18181, 2024. Wenhao Zhang. Beyond absolute imitation: Anchored residual guidance for privileged on-policy distillation. arXiv preprint arXiv:2606.10385, 2026. Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, and Aditya Grover. Self-distilled reasoner: On-policy self-distillation for large language models. arXiv preprint arXiv:2601.18734, 2026a. Xuyang Zhao, Liting Zhang, Zichen Xu, Zhihu Wang, Xu Caiyue, Shiwan Zhao, and Qicheng Li. Is more privileged information better? from solution traces to problem-solving structure in self-distilled reasoning. arXiv preprint arXiv:2608.01589, 2026b.

14

Calibrated On-Policy Distillation

Preprint

A Appendix A.1 Related Work Knowledge Distillation and On-Policy Distillation. Knowledge distillation (KD) transfers capabilities from a stronger teacher to a weaker student and has become a standard framework for model compression and capability transfer(Hinton et al., 2015). Sequence-level and pretrained language-model distillation extended this paradigm to generation and compression(Kim and Rush, 2016; Sanh et al., 2019; Jiao et al., 2020), while LLM-specific methods introduced reverse or student-aware divergence objectives and output-space alignment(Gu et al., 2024; Ko et al., 2024; Zhang et al., 2024). For reasoning, teacher-generated rationales and reasoning strategies have been transferred through Fine-tune-CoT(Ho et al., 2023), Distilling Step-by-Step(Hsieh et al., 2023), SCoTD(Li et al., 2023), MCC-KD(Chen et al., 2023), Orca 2(Mitra et al., 2023), counterfactual distillation(Feng et al., 2024), and decomposed CoT distillation(Dai et al., 2024); DeepSeek-R1 further demonstrates that CoT reasoning behaviors can be distilled into smaller models(Guo et al., 2025). Many sequence- and rationale-level distillation methods remain off-policy, training on teacher-generated or otherwise fixed trajectories. On-policy distillation instead provides teacher supervision on trajectories sampled from the student itself, reducing the training–inference distribution mismatch(Agarwal et al., 2024), and has been adopted in large-scale reasoning post-training such as Qwen3(Yang et al., 2025). Recent extensions modify the supervision signal through reward extrapolation(Yang et al., 2026a), entropy- or uncertainty-aware objectives(Jin et al., 2026; Ke et al., 2026), outcome-guided calibration(Hou et al., 2026a), divergence-adaptive supervision horizons(Hou et al., 2026b), and teacher-versus-base delta signals(Heo et al., 2026). On-policy self-distillation further removes the need for a separate stronger teacher through privileged or behavior-conditioned self-teachers(Zhao et al., 2026a; Sang et al., 2026), self-generated supervision(Tan and Hong, 2026), and internal consistency(Li et al., 2026a), while on-policy context distillation internalizes teacher-only context such as prior experience or optimized system prompts(Ye et al., 2026). Privileged and Context-Augmented Distillation. Recent LLM distillation strengthens the teacher with information unavailable to the student at inference time. On-Policy Self-Distillation (OPSD) conditions a self-teacher on verified reference solutions(Zhao et al., 2026a), SDPO conditions the self-teacher on environment feedback(Hübotter et al., 2026), and privileged-information distillation studies forms of training-only information(Penaloza et al., 2026). Related methods distill teacher-only experience or system prompts(Ye et al., 2026), document-grounded evidence(Stein et al., 2026), or behavioral instructions for reasoning compression(Sang et al., 2026). For mathematical reasoning, Anti-Self-Distillation reverses harmful privileged-teacher guidance(Shen et al., 2026a), DOPD dynamically routes token-level supervision to mitigate privilege illusion(Yu et al., 2026b), PHF transfers privileged hidden-state dynamics(Li et al., 2026b), AR-OPD anchors privileged guidance to a locally compatible view(Zhang, 2026), Purified OPSD removes reference-induced non-transferable components(Shen et al., 2026b), and DAPD introduces dual anchoring against information asymmetry(Wu et al., 2026). Recent analyses further show that privileged context can degrade thinking models(Kaur et al., 2026), that references from other problems can retain comparable gains(Ichihara et al., 2026), that correct references do not provide consistent benefits(Shrestha and Tessier, 2026), and that structured privileged information can outperform raw solution traces(Zhao et al., 2026b). Together, these findings suggest that privileged context changes not only the information available to the teacher, but also the teacher behavior induced by that context, complicating the interpretation of privileged teacher likelihoods as transferable knowledge. Reliability of Teacher Supervision. Recent work has also questioned the assumption that teacher supervision is uniformly reliable across tokens and trajectories. Entropy-Aware OPD adapts distillation to teacher uncertainty(Jin et al., 2026), Uni-OPD calibrates teacher guidance using outcome-level order consistency(Hou et al., 2026a), and Position-Weighted OPSD shows that teacher-token reliability is strongly structured by reasoning position(Liu et al., 15

Calibrated On-Policy Distillation

Preprint

2026). OGLS-SD calibrates privileged teacher logits using outcome contrast(Yang et al., 2026b), while BRTS selects teacher rollouts according to correctness and student alignment(Zhang et al., 2026). Related studies further show that reasoning tokens remain relatively stable while stylistic or surface-form tokens are more sensitive to contextual interventions(Pan et al., 2026; He et al., 2026), and that privileged context can induce shortcut behavior(Tian et al., 2026). More recently, Ding and Zhang (2026) reveal substantial, scale-dependent noise in OPD teacher supervision and show that removing noisy supervision can leave student performance largely unchanged. These studies establish that teacher supervision is heterogeneous, but characterize reliability primarily through uncertainty, outcomes, positions, trajectory quality, or specific privileged-reference effects. In contrast, the variation of teacher likelihoods under controlled contextual interventions has not been explicitly modeled as a finite-intervention self-deviation region for calibrating the teacher–student discrepancy in OPD.

A.2 Data Collection and Preparation for TSD Analysis Our TSD analysis is conducted on a collection of student-generated reasoning trajectories derived from DAPO17K (Yu et al., 2026a). Since DAPO-17K provides problems and ground-truth answers but does not include reference solutions, we first construct a solution-augmented subset, then generate on-policy student rollouts, and finally rescore the fixed trajectories under different teacher-side contextual interventions. Table 5 summarizes the configurations used in this process. Reference-Solution Preparation. We use Qwen3-235B-A22B-Instruct-2507 (Yang et al., 2025) to generate one reference solution for each problem in DAPO-17K. The model is used in its no-thinking mode with a maximum response length of 12,288 tokens, temperature 1.0, and top-𝑝 1.0. All generations use the following shared system prompt: System Prompt You are a helpful math assistant. Please solve the math problem. You must enclose your final answer exactly within \boxed{}.

We extract the final answer enclosed in \boxed{} and compare it with the ground-truth answer provided by DAPO-17K. Solutions with incorrect final answers are discarded. After filtering, we retain 15,560 question–solution pairs, covering 90.18% of the 17,255 problems, with 1,695 problems unsolved. The retained reference solutions contain an average of 7,636 characters, a median of 5,679 characters, a 95th percentile of 19,493 characters, and a range of 445–39,695 characters. Student Rollout Collection. Starting from the 15,560 retained questions, we generate one on-policy response per question using Qwen3-1.7B (Yang et al., 2025) in thinking mode. We set the maximum response length to 16,384 tokens, temperature to 1.0, and top-𝑝 to 1.0, while using the same system prompt as above. Rollout generation is stopped once the cumulative number of student-generated response tokens reaches approximately 60 million. This results in 6,528 question–response pairs used for subsequent TSD analysis. The student responses contain an average of 9,270 tokens and a median of 7,758 tokens. Training Data Usage. All distillation methods in our main experiments are trained on the same 15,560 question– solution pairs. Construction of Privileged Contexts. For the answer-level interventions in Table 1, the positive variant uses the neg ground-truth answer associated with the current problem. To construct 𝑐 ans , we randomly generate an incorrect answer with the same number of digits as the corresponding correct answer, thereby controlling for differences in answer length. 16

Calibrated On-Policy Distillation

Preprint

Stage

Model

Solution generation Student rollout Teacher rescoring

Qwen3-235B-A22B-Instruct-2507 Qwen3-1.7B Qwen3-8B

Mode

Max Length

Temp.

Top-𝑝

No-thinking Thinking Thinking

12,288 16,384 Fixed rollout

1.0 1.0 –

1.0 1.0 –

Table 5: Configurations used for preparing the data for TSD analysis. Solution generation and student rollout collection use one generation per question. Teacher rescoring does not involve sampling, since the student trajectory is kept fixed. Positive

Negative

Group Union Sg (τ)

τ = 0.01

Positive

Negative

Group Union Sg (τ)

τ = 0.05

Positive

Negative

Group Union Sg (τ)

τ = 0.10

39.6

29.8

29.1

30 25.8

25.0

34.4

29.4 25.1

25.2

21.9

20.2

20

10

Instruction

Evaluation

Answer

40

29.5

30 26.3

20

17.4

9.8

10

Solution

17.3

16.3 10.9

14.3

Instruction

23.2

13.4

Evaluation

13.9

14.0

Answer

Tokens with Significant TSD (%)

40 37.0

Tokens with Significant TSD (%)

Tokens with Significant TSD (%)

40

30 23.1

20

20.2 10.1

9.2

10 4.9

Solution

5.6

8.3

7.5

Instruction

16.8

10.5

Evaluation

8.2

8.3

Answer

Solution

(a) Qwen3-8B Positive

Negative

Group Union Sg (τ)

τ = 0.01

Positive

Negative

Group Union Sg (τ)

τ = 0.05

Positive

Negative

Group Union Sg (τ)

τ = 0.10

40.3

30

29.4 25.2

37.7

35.5

30.2 25.3

26.7

29.1

29.2

22.0

20

10

Instruction

Evaluation

Answer

40 31.2

30 28.0

24.9

20.9 18.6

18.0

20

14.6

14.3

12.1

17.8

17.8

15.6

10

Solution

Instruction

Evaluation

Answer

Tokens with Significant TSD (%)

40 32.4

Tokens with Significant TSD (%)

Tokens with Significant TSD (%)

40

30 25.6

22.4

20

19.0

14.1 11.2

10

Solution

6.8

11.7 11.7

8.9

Instruction

8.5

11.7

9.6

Evaluation

Answer

Solution

(b) Qwen3-4B-Thinking-2507 Figure 6: Prevalence of significant TSD across thresholds and teacher models. Although the absolute prevalence decreases with stricter thresholds, substantial TSD continues to emerge under task-agnostic interventions, while solution-level privilege consistently broadens the affected set. pos

For the solution-level interventions, 𝑐 sol uses the verified reference solution generated for the current problem. neg To construct 𝑐 sol , we randomly select the reference solution of a different problem whose tokenized length differs from that of the current reference solution by less than 10%. This length-matching constraint reduces the possibility that differences between positive and negative solution-level interventions are driven primarily by contextual length rather than content. Teacher Likelihood Rescoring. We use Qwen3-8B (Yang et al., 2025) as the teacher and evaluate each fixed Qwen3-1.7B rollout under nine teacher contexts: 𝑐0,

pos

neg

pos

neg

pos

neg

pos

neg

𝑐 inst , 𝑐 inst , 𝑐 eval , 𝑐 eval , 𝑐 ans , 𝑐 ans , 𝑐 sol , 𝑐 sol .

(15)

The shared system prompt above is included in every condition and is therefore not itself treated as a contextual intervention. For all nine contexts, the problem 𝑥, student rollout 𝑦, prefix 𝑦 <𝑡 , and evaluated token 𝑦 𝑡 are kept identical; only the additional teacher-side context is changed. The teacher does not generate new trajectories. Instead, we recompute the token-level log-likelihood assigned to every student-generated token under each context, allowing the resulting likelihood differences to be attributed directly to the contextual intervention. 17

Calibrated On-Policy Distillation

Preprint

τ = 0.01

τ = 0.05

Ret i→j (τ)

C inst

100.0%

86.5%

86.0%

98.3%

C eval

88.7%

100.0%

87.6%

98.6%

C ans

87.4%

86.8%

100.0%

98.2%

C sol

74.2%

72.6%

72.9%

100.0%

C inst

100.0%

76.0%

76.5%

96.1%

C eval

81.0%

100.0%

80.3%

96.7%

C ans

77.1%

75.9%

100.0%

95.7%

C sol

56.7%

53.6%

56.1%

100.0%

C inst

100.0%

68.3%

69.5%

94.4%

C eval

74.9%

100.0%

75.8%

95.5%

0.60

C ans

67.1%

66.7%

100.0%

93.4%

C sol

41.2%

38.0%

42.2%

100.0%

C sol

0.80

Source C i

Source C i

Source C i

C ans

1.00

0.80

0.40

C eval

Ret i→j (τ)

1.00

0.80

C inst

τ = 0.10

Ret i→j (τ)

1.00

0.60

0.40

C inst

Target C j

C eval

C ans

C sol

0.60

0.40

C inst

Target C j

C eval

C ans

C sol

Target C j

(a) Qwen3-8B τ = 0.01

τ = 0.05

Ret i→j (τ)

C inst

100.0%

91.0%

92.9%

99.0%

C eval

88.4%

100.0%

93.8%

99.0%

C ans

84.2%

87.5%

100.0%

98.4%

C sol

72.2%

74.3%

79.2%

100.0%

C inst

100.0%

83.4%

86.3%

97.7%

C eval

81.0%

100.0%

88.4%

97.6%

C ans

74.4%

78.5%

100.0%

96.3%

C sol

56.5%

58.2%

64.7%

100.0%

C inst

C eval

C ans

Target C j

C sol

1.00

C inst

100.0%

76.4%

80.6%

96.5%

C eval

73.4%

100.0%

84.5%

96.4%

C ans

64.2%

69.9%

100.0%

94.3%

C sol

42.3%

43.9%

51.9%

100.0%

0.80

Source C i

0.80

Source C i

Source C i

0.40

Ret i→j (τ)

1.00

0.80

0.60

τ = 0.10

Ret i→j (τ)

1.00

0.60

0.40

C inst

C eval

C ans

C sol

Target C j

0.60

0.40

C inst

C eval

C ans

C sol

Target C j

(b) Qwen3-4B-Thinking-2507 Figure 7: Retention of significant TSD across thresholds and teacher models. Less informative interventions retain nearly all affected token positions under solution-level privilege, whereas retention in the reverse direction is substantially lower.

Consequently, the TSD analysis is based on approximately 60 million student-generated tokens, with each token rescored under the baseline context and eight contextual interventions.

A.3 TSD Semantic Consistency Across Thresholds and Model Scales We further examine whether the TSD patterns reported in the main text persist across significance thresholds and teacher scales. Using the same Qwen3-1.7B student trajectories described in Appendix A.2, we repeat the analysis with Qwen3-8B and Qwen3-4B-Thinking-2507 as teachers under 𝜏 ∈ {0.01, 0.05, 0.10}. We consider the prevalence of significant TSD, retention across intervention groups, and semantic consistency across the six teacher–threshold configurations. Prevalence of Significant TSD. Figure 6 shows that increasing 𝜏 reduces the absolute prevalence of significant TSD, while preserving the relative pattern across intervention groups. For Qwen3-8B, the union prevalence under task-agnostic instructions decreases from 29.8% at 𝜏 = 0.01 to 10.1% at 𝜏 = 0.10, yet remains comparable to evaluative feedback at 29.1% and 9.2%, respectively. The same pattern holds for Qwen3-4B-Thinking-2507, where task-agnostic instructions remain close to evaluative feedback across all thresholds. Solution-level privilege consistently produces the highest prevalence, supporting the conclusion that richer privileged context broadens the affected set rather than being necessary for TSD to emerge.

18

Calibrated On-Policy Distillation

Agr g (τ)

Ovl g (τ)

τ = 0.01 100

74.7

71.3

70.7

80.3 79.1 66.1

61.1 60 40 20

Agr g (τ)

SDR g (τ)

100

76.4

73.6

68.0

80

67.5

48.7

40 20

0

Answer

Solution

SDR g (τ)

τ = 0.10

93.6 84.2 76.9

73.7

68.9 59.7

57.1

60

42.6 40 20

0

Evaluation

Agr g (τ)

88.4

82.2

61.1 60

Ovl g (τ)

τ = 0.05

92.9

88.8

80

Percentage (%)

Percentage (%)

80

SDR g (τ)

88.7

83.0

Percentage (%)

Ovl g (τ) 100

Preprint

0

Evaluation

Answer

Solution

Evaluation

Answer

Solution

(a) Qwen3-8B

72.1

SDR g (τ)

Ovl g (τ)

τ = 0.01 100

90.7 77.0

80.0

77.6

81.8

77.7

40

0

Solution

70.2

70.0

Agr g (τ)

94.8

100

66.3 60

τ = 0.10

80.9

79.6

80

66.6

SDR g (τ)

94.5 80.8

40

0

Ovl g (τ)

τ = 0.05

79.3

78.9

60.7

20

Answer

SDR g (τ)

93.7 79.4

60

20

Evaluation

Agr g (τ)

93.0

80

65.4 60

Percentage (%)

Percentage (%)

80

Agr g (τ)

88.4

Percentage (%)

Ovl g (τ) 100

55.5

61.7

67.8

40 20 0

Evaluation

Answer

Solution

Evaluation

Answer

Solution

(b) Qwen3-4B-Thinking-2507 Figure 8: Semantic consistency of TSD across thresholds and teacher models. Contrasting interventions consistently exhibit substantial positional overlap, high directional agreement, and predominantly shared deviation across thresholds and model scales.

Retention Across Intervention Groups. Figure 7 exhibits a consistent asymmetric retention pattern across thresholds and teacher models. Token positions affected under instruction, evaluation, and answer interventions are largely retained under solution-level privilege. Even at 𝜏 = 0.10, their retention into the solution group remains at least 93.4% for Qwen3-8B and 94.3% for Qwen3-4B-Thinking-2507. The reverse retention is substantially lower and decreases as the threshold becomes stricter. This asymmetry indicates that richer privileged context predominantly extends an existing set of TSD-sensitive positions rather than replacing it with a distinct deviation pattern. Semantic Consistency. Figure 8 shows that the semantic-consistency pattern also persists across thresholds and teacher models, and across all three intervention groups. As 𝜏 increases, positional overlap generally decreases because fewer tokens remain significant under both contrasting interventions, while directional agreement remains consistently high. At 𝜏 = 0.10, answer-level agreement reaches 93.6% for Qwen3-8B and 94.5% for Qwen3-4BThinking-2507, with corresponding SDR values of 76.9% and 79.6%. Evaluative interventions show the same trend, whereas solution-level interventions consistently exhibit higher positional overlap but lower SDR. Thus, stricter thresholds change which tokens remain significant without altering the broader finding that contrasting semantics induce strongly aligned and predominantly shared TSD. Across all configurations, the qualitative conclusions of the main analysis remain unchanged. TSD emerges substantially without task-specific knowledge, richer privileged context largely preserves and expands previously affected token positions, and contrasting intervention semantics leave much of the resulting TSD directionally aligned and shared. These results show that the observed TSD patterns are not specific to a particular significance threshold or teacher scale.

A.4 Token-Level TSD Patterns Across Contextual Interventions To examine whether the token-level structure of TSD depends on a particular contextual intervention, we repeat the token-form analysis under all eight contextual interventions using Qwen3-8B at 𝜏 = 0.01. Following the main analysis, we merge tokenizer variants corresponding to the same normalized surface form and retain forms occurring more than 20,000 times. Tables 6–9 report the top-24 and bottom-24 forms ranked by 𝜌 𝜏 (𝑣) for each positive–negative intervention pair, together with their significant-TSD rates and mean deviation magnitudes. Stability Across Intervention Polarity. Positive and negative variants within each intervention group produce substantially overlapping token rankings. Their top-24 sets share 18/24 forms for instruction, 22/24 for evaluation, 24/24 for answer-level privilege, and 23/24 for solution-level privilege. The corresponding bottom-24 overlaps are 19

Calibrated On-Policy Distillation

Preprint neg

pos

Table 6: Top-24 and bottom-24 token forms under 𝑐 inst and 𝑐 inst at 𝜏 = 0.01. The final row reports unweighted means of 𝜌 𝜏 and |Δ| over each 24-token subset. neg

pos

𝑐inst

𝑐inst #

Highest 𝜌 𝜏

Lowest 𝜌 𝜏

Token

𝝆𝝉 (%)

|𝚫|

1 maybe 2 consider 3 try 4 however 5 since 6 if 7 earlier 8 now 9 how 10 think 11 another 12 check 13 because 14 therefore 15 when 16 there 17 let 18 given 19 also 20 use 21 so 22 but 23 we 24 yes

74.4 73.1 66.1 65.2 64.4 63.9 63.2 63.0 62.8 61.8 60.8 60.6 60.2 58.9 58.9 58.7 58.0 57.2 56.8 56.1 55.8 55.7 55.2 55.1

0.066 0.071 0.086 0.086 0.080 0.072 0.066 0.094 0.078 0.070 0.088 0.075 0.080 0.088 0.078 0.073 0.076 0.105 0.083 0.081 0.081 0.077 0.065 0.096

61.1

0.080

Mean–

Highest 𝜌 𝜏

Token 𝝆𝝉 (%)

|𝚫|

_ 9 𝜃 8 frac 6 7 5 { 2 _i )^ + 4 1 3 ’t )/ cdot

1.0 2.1 2.2 2.6 2.9 3.0 3.1 3.2 3.3 3.3 3.4 3.5 3.6 3.6 3.8 3.8 4.0 4.0 4.0 4.2 4.2 4.6 5.1 5.7

0.069 0.071 0.095 0.068 0.084 0.072 0.099 0.076 0.102 0.062 0.101 0.101 0.096 0.067 0.089 0.077 0.078 0.089 0.095 0.088 0.092 0.088 0.077 0.051

–

3.5

0.083

### 0 ◦

√

Lowest 𝜌 𝜏

𝝆𝝉 (%)

|𝚫|

maybe consider however therefore earlier try now how because another if since think let so but this altern. here when there previous it seems

84.4 84.2 82.8 82.1 80.9 80.9 78.5 75.1 74.6 74.4 74.2 74.1 73.9 73.6 72.7 72.0 71.8 71.7 71.2 71.1 69.2 69.2 69.1 68.8

0.110 0.106 0.157 0.170 0.121 0.177 0.165 0.112 0.114 0.127 0.102 0.109 0.108 0.122 0.118 0.113 0.117 0.123 0.115 0.110 0.099 0.086 0.117 0.097

–

75.0

0.121

Token

Token 𝝆𝝉 (%)

|𝚫|

8 frac 6 𝜃 7 5 _i { 2 4 1 )^ 3 + ’t cdot )/

1.5 2.9 3.4 3.7 3.7 3.9 4.0 4.1 4.2 4.3 4.4 4.4 4.5 4.6 4.7 4.8 5.1 5.3 5.4 5.5 5.6 5.7 7.2 7.4

0.080 0.116 0.082 0.083 0.089 0.119 0.120 0.124 0.073 0.125 0.105 0.124 0.119 0.090 0.088 0.110 0.119 0.109 0.091 0.112 0.106 0.101 0.057 0.100

–

4.6

0.102

0 ### ◦

_ 9 √

24/24, 23/24, 24/24, and 24/24. Thus, reversing intervention polarity or correctness generally leaves the token forms most prone to TSD largely unchanged. This is especially true for forms that remain comparatively stable across conditions, while the instruction-level top set shows somewhat more variation than the other pairs. Stability Across Intervention Types. The same structure persists across different types of contextual intervention. Any two of the eight intervention conditions share at least 18/24 top-ranked forms and 22/24 bottom-ranked forms, while 16/24 top forms and 22/24 bottom forms appear in every ranking. Across all conditions, high-𝜌 𝜏 forms are consistently dominated by natural-language surface expressions such as maybe, however, therefore, consider, and think, whereas low-𝜌 𝜏 forms are dominated by digits, mathematical symbols, and notation. This separation is also quantitatively strong: the mean top-24 𝜌 𝜏 ranges from 61.1% to 91.5%, compared with only 3.5% to 8.2% for the bottom-24, corresponding to an approximately 11–18× gap within individual conditions. Context Modulates Strength More Than Token Identity. While the ranking structure remains stable, the magnitude of TSD varies substantially across contextual interventions. Under positive interventions, the mean |Δ| of the top-24 forms increases from 0.080 under instruction and 0.084 under evaluation to 0.182 under answer-level privilege and 0.436 under solution-level privilege. The bottom-24 exhibits the same overall trend, rising from 0.083 and 0.091 to 0.111 and 0.290, respectively. Negative interventions show a similar overall increase: the top-24 means are 0.121, 0.108, 0.244, and 0.265 for instruction, evaluation, answer-level privilege, and solution-level privilege, respectively, while the corresponding bottom-24 means are 0.102, 0.102, 0.108, and 0.199. Thus, richer privileged context can substantially amplify teacher-side likelihood shifts, particularly at the solution level. Crucially, however, this amplification does not substantially alter which token forms occupy the two ends of the ranking: largely the same 20

Calibrated On-Policy Distillation

Preprint neg

pos

Table 7: Top-24 and bottom-24 token forms under 𝑐 eval and 𝑐 eval at 𝜏 = 0.01. The final row reports unweighted means of 𝜌 𝜏 and |Δ| over each 24-token subset. neg

pos

𝑐eval

𝑐eval #

Highest 𝜌 𝜏

Lowest 𝜌 𝜏

Token

𝝆𝝉 (%)

|𝚫|

1 maybe 2 consider 3 earlier 4 however 5 try 6 therefore 7 now 8 since 9 how 10 if 11 think 12 because 13 another 14 previous 15 says 16 when 17 there 18 so 19 but 20 let 21 this 22 seems 23 also 24 altern.

77.5 76.2 72.4 72.2 69.2 67.4 66.5 66.2 66.1 66.1 65.4 65.1 64.5 63.6 62.8 62.7 62.2 62.2 61.9 61.9 61.1 60.7 60.3 60.2

0.079 0.074 0.082 0.102 0.091 0.098 0.102 0.083 0.081 0.077 0.074 0.086 0.091 0.071 0.073 0.084 0.078 0.087 0.082 0.084 0.086 0.073 0.087 0.085

65.6

0.084

Mean–

Highest 𝜌 𝜏

Token 𝝆𝝉 (%)

|𝚫|

_ 9 𝜃 8 6 { 7 5 _i 2 cdot 4 )^ + 1 3 ’t )/

1.0 1.1 2.5 2.8 2.9 3.2 3.3 3.4 3.4 3.6 3.7 3.8 3.9 3.9 4.0 4.2 4.4 4.5 4.5 4.5 4.6 4.8 5.3 5.8

0.072 0.047 0.111 0.058 0.070 0.093 0.076 0.114 0.082 0.119 0.118 0.074 0.114 0.111 0.084 0.103 0.047 0.111 0.087 0.099 0.103 0.104 0.091 0.087

–

3.7

0.091

}{ ### 0 frac ◦

√

Lowest 𝜌 𝜏

𝝆𝝉 (%)

|𝚫|

maybe however consider earlier therefore try says another since think if how altern. previous now because let but here check seems so when there

83.8 83.2 82.6 79.9 77.9 77.1 74.6 74.5 74.3 73.4 73.4 72.9 72.8 72.6 72.5 72.0 71.0 70.8 70.0 69.9 69.9 69.8 69.3 68.9

0.100 0.139 0.091 0.113 0.131 0.134 0.090 0.118 0.103 0.093 0.090 0.104 0.120 0.087 0.149 0.105 0.133 0.100 0.103 0.098 0.090 0.106 0.100 0.095

}{ ### 0 frac _

–

74.0

0.108

Token

Token 𝝆𝝉 (%)

|𝚫|

𝜃 8 6 _i { 7 5 cdot 2 4 )^ 1 3 ’t + }

1.1 1.3 2.9 3.1 3.6 3.7 3.9 3.9 4.1 4.1 4.3 4.4 4.4 4.5 4.5 4.6 4.9 5.0 5.2 5.3 5.4 5.5 5.6 6.7

0.076 0.053 0.124 0.064 0.092 0.068 0.138 0.108 0.099 0.140 0.138 0.087 0.078 0.138 0.130 0.053 0.112 0.127 0.091 0.112 0.130 0.124 0.106 0.069

–

4.3

0.102

◦

9 √

surface-form tokens remain TSD-prone, while numerical and symbolic forms remain comparatively stable. These results indicate that contextual interventions primarily modulate the strength of teacher self-deviation, whereas its token-level concentration remains stable across intervention semantics and privilege types.

A.5 Trajectory-Level Case Study of TSD We complement the aggregate analysis with two trajectory-level examples illustrating how TSD is distributed within complete reasoning traces. We use the same Qwen3-1.7B student and Qwen3-8B teacher configuration and visualize pos TSD under 𝑐 sol . Red and blue indicate decreases and increases in teacher log-likelihood, respectively, with darker shades denoting larger |Δ𝑇𝑡 |; tokens with |Δ𝑇𝑡 | ≤ 0.1 are left uncolored. Since the student trajectory is fixed, all highlighted variation is induced solely by the additional reference solution. Outcome: Correct

Case 1: Correct Student Rollout Problem. Calculate the least integer greater than 5 (−6) (−5) (−4) ··· (2) (3) (4) . Answer. The student rollout outputs 2 , which matches the gold answer. TSD Legend. (0.1, 0.3]

[ −0.3, −0.1) [ −0.7, −0.3) [ −1.5, −0.7) (0.3, 0.7] (0.7, 1.5] (1.5, 3] >3

[ −3, −1.5)

< −3

pos

Student trajectory excerpt under 𝑐sol . Okay, so I need to find the least integer greater than $5^{(-6)(-5)(-4)...(2)(3)(4)}$. Hmm, let me start by understanding the exponent here. The exponent is the product of all integers from -6 up to 4. Wait, let me confirm: the exponent is (-6)(-5)(-4)...(2)(3)(4). So that ’s multiplying all integers from -6 to 4, but I need to check if that ’s correct . Let me count how many numbers are there . Starting from -6 up to 4. Let’s see: from -6 to 4 inclusive. The number of terms would be 4 - (-6) + 1 = 11 terms.

21

Calibrated On-Policy Distillation

Preprint neg

pos

Table 8: Top-24 and bottom-24 token forms under 𝑐 ans and 𝑐 ans at 𝜏 = 0.01. The final row reports unweighted means of 𝜌 𝜏 and |Δ| over each 24-token subset. neg

pos

𝑐ans #

Highest 𝜌 𝜏

Lowest 𝜌 𝜏

Token

𝝆𝝉 (%)

|𝚫|

1 maybe 2 however 3 therefore 4 consider 5 earlier 6 if 7 try 8 seems 9 another 10 because 11 now 12 how 13 says 14 think 15 since 16 but 17 so 18 check 19 let 20 there 21 altern. 22 previous 23 when 24 wait

83.3 82.2 80.6 80.1 79.5 74.7 74.6 74.6 73.9 73.2 73.1 73.1 73.0 72.6 72.1 72.0 71.4 71.1 70.7 70.3 70.2 70.1 69.2 68.1

0.254 0.210 0.210 0.090 0.214 0.135 0.121 0.458 0.211 0.127 0.154 0.158 0.106 0.126 0.125 0.188 0.142 0.122 0.426 0.141 0.181 0.125 0.123 0.221

}{ ### 0 frac _ √

73.9

0.182

Mean–

𝑐ans Highest 𝜌 𝜏

Token 𝝆𝝉 (%)

|𝚫|

6 _i { 5 7 2 )^ 4 cdot 1 3 + ’t }

1.1 1.6 2.8 3.1 3.7 3.7 3.8 3.9 4.2 4.2 4.3 4.4 4.4 4.5 4.5 4.8 5.0 5.1 5.2 5.3 5.4 5.7 5.9 7.0

0.070 0.058 0.143 0.066 0.088 0.120 0.088 0.155 0.165 0.093 0.157 0.082 0.086 0.152 0.161 0.129 0.087 0.141 0.049 0.133 0.137 0.105 0.120 0.090

–

4.3

0.111

𝜃 9 8 ◦

Lowest 𝜌 𝜏

𝝆𝝉 (%)

|𝚫|

maybe however consider therefore earlier if another try says seems how think because since check there now but so altern. let previous when wait

83.7 82.3 81.0 80.6 79.5 75.9 75.6 75.1 74.4 74.0 73.5 73.4 73.4 73.2 72.4 72.2 72.1 71.8 71.4 71.2 71.2 70.0 69.9 68.7

0.362 0.163 0.090 0.358 0.226 0.176 0.431 0.126 0.111 0.487 0.170 0.189 0.136 0.136 0.154 0.200 0.151 0.136 0.183 0.346 1.088 0.113 0.150 0.175

}{ ### 0 frac _ √

–

74.4

0.244

Token

Token 𝝆𝝉 (%)

|𝚫|

6 _i { 5 7 2 4 )^ cdot 1 3 + ’t }

1.1 1.6 2.8 3.1 3.6 3.7 3.9 3.9 4.1 4.2 4.3 4.3 4.5 4.5 4.5 4.8 5.0 5.1 5.3 5.3 5.4 5.8 5.8 7.3

0.071 0.055 0.131 0.062 0.095 0.119 0.146 0.086 0.152 0.087 0.145 0.085 0.085 0.139 0.150 0.120 0.136 0.087 0.049 0.127 0.132 0.104 0.120 0.112

–

4.3

0.108

9 𝜃 8 ◦

But wait, let me check: Numbers : -6, -5, -4, -3, -2, -1, 0, 1, 2, 3, 4. Yes, that’s 11 numbers. But wait, the problem says "(-6)(-5)(-4)...(2)(3)(4)". So the product is from -6 to 4, including 0? Wait, but 0 is included here . However, when you multiply by 0, the entire product becomes 0. But then the exponent would be 0, so 5^0 is 1. Then the least integer greater than 1 is 2. But that seems too straightforward. Maybe I made a mistake here. Wait, but hold on. Let me check again. The problem says " (-6)(-5)(-4)...(2)(3)(4)". Is the product from -6 up to 4, including 0? But if that’s the case, then the product is zero, so the exponent is zero, and 5^0 is 1. Therefore, the least integer greater than 1 is 2. But that seems too easy . Maybe there’s a misunderstanding here. Alternatively, maybe the product is from -6 to 4, excluding 0? But the problem didn’t specify. Wait, the original problem says "(-6)(-5)(-4)...(2)(3)(4)". The ellipsis is between -4 and 2, so maybe the product is from -6 up to 4, but skipping 0? Wait, but how ? If you have -6, -5, -4, ..., 2, 3, 4. So if you start at -6 and go up to 4, but not including 0? But 0 is between -6 and 4. So unless there’s a typo, but the problem is written as (-6)(-5)(-4)...(2)(3)(4). So maybe the product is from -6 to 4, including 0? But then the product is zero. But maybe there’s a mistake here. Wait, but let me check if the product is from -6 to 4, but not including 0. If that’s the case, then the product is (-6)(-5)(-4)... (-1)(1)(2)(3)(4). Let ’s check how many terms that would be. From -6 to -1 is 6 numbers, then 1, 2, 3, 4. So total of 6 + 4 = 10 terms. But then the product would be (-6)(-5)...(-1) * 1 * 2 * 3 * 4. Let me compute that. First, the product from -6 to -1 is (-1)^6 * 6! = 1 * 720 = 720. Then multiply by 1*2*3*4 = 24. So total product is 720 * 24 = 17,280. Then 5^17280 is an extremely large number. Then the least integer greater than that would be 5^17280 + 1, but that seems way too big. But the problem says "the least integer greater than 5^{...}" so maybe the answer is something else. But the original problem is written as (-6)(-5)...(2)(3)(4). So if the product includes 0, then it’s zero, so 5^0 is 1. Then the least integer greater than 1 is 2. But maybe the problem is written with a different intention . Wait, maybe the exponent is from -6 to 4, but not including 0? But how? If you have from -6 to 4, but not

22

Calibrated On-Policy Distillation

Preprint neg

pos

Table 9: Top-24 and bottom-24 token forms under 𝑐 sol and 𝑐 sol at 𝜏 = 0.01. The final row reports unweighted means of 𝜌 𝜏 and |Δ| over each 24-token subset. neg

pos

𝑐sol

𝑐sol #

Highest 𝜌 𝜏

Lowest 𝜌 𝜏

Token

𝝆𝝉 (%)

|𝚫|

1 maybe 2 however 3 therefore 4 consider 5 altern. 6 earlier 7 seems 8 another 9 since 10 try 11 think 12 says 13 wait 14 check 15 let 16 here 17 because 18 now 19 also 20 if 21 similarly 22 but 23 so 24 how

97.4 97.4 95.9 95.4 94.2 94.2 93.6 92.9 91.8 91.8 91.6 90.4 90.0 89.6 89.5 89.4 89.4 89.2 89.2 89.2 89.0 88.7 88.3 88.3

0.594 0.625 0.460 0.380 0.576 0.413 0.408 0.504 0.441 0.404 0.402 0.246 0.415 0.278 0.504 0.364 0.396 0.485 0.470 0.364 0.532 0.439 0.366 0.408

}{ ### 0 9 8 √

91.5

0.436

Mean–

Highest 𝜌 𝜏

Token 𝝆𝝉 (%)

|𝚫|

_ 6 7 5 𝜃 frac { 4 2 3 _i 1 ’t )^ + cdot }

2.8 3.2 5.0 6.8 7.2 7.2 7.3 7.3 7.4 7.7 7.8 7.9 8.0 8.0 8.8 8.8 9.2 9.2 9.4 10.0 10.5 10.7 12.8 12.9

0.251 0.131 0.356 0.384 0.395 0.278 0.170 0.287 0.398 0.389 0.379 0.280 0.225 0.234 0.372 0.323 0.362 0.283 0.330 0.302 0.208 0.276 0.159 0.188

–

8.2

0.290

◦

Lowest 𝜌 𝜏

𝝆𝝉 (%)

|𝚫|

maybe however therefore consider earlier altern. try another seems think says since here if because wait previous how so check also but now let

96.2 95.3 95.0 93.5 92.5 90.3 90.2 90.1 90.0 88.6 88.6 88.4 88.0 86.9 86.7 86.7 86.3 86.2 86.2 86.0 85.8 85.7 85.5 84.8

0.305 0.358 0.351 0.230 0.304 0.309 0.281 0.287 0.229 0.244 0.213 0.247 0.266 0.242 0.246 0.226 0.219 0.264 0.263 0.210 0.271 0.221 0.310 0.267

}{ ### 0 9 frac _ √

–

88.9

0.265

Token

Token 𝝆𝝉 (%)

|𝚫|

6 7 { 5 𝜃 _i 4 2 3 1 ’t )^ cdot + }

2.0 2.9 4.3 5.9 6.0 6.0 6.2 6.2 6.2 6.4 6.8 6.8 6.8 6.9 7.6 7.6 7.7 8.1 8.2 8.8 9.0 9.4 9.5 11.1

0.141 0.102 0.250 0.270 0.110 0.172 0.205 0.271 0.117 0.278 0.266 0.153 0.257 0.215 0.181 0.256 0.219 0.259 0.218 0.238 0.166 0.094 0.203 0.130

–

6.9

0.199

8 ◦

including 0, then the product is (-6)(-5)...(-1)(1)(2)(3)(4). Which is what I thought earlier. But then the exponent is 17280, so 5^17280 is a gigantic number, and the least integer greater than that would be 5^17280 + 1. But that seems like an unreasonable answer. However, maybe there’s a mistake in my calculation. Wait, let me check the product again. If the product is from -6 to 4, excluding 0, then: First, the product from -6 to -1 is (-6)(-5)(-4)(-3)(-2)(-1). Let me compute that: (-6) * (-5) = 30 30 * (-4) = -120 -120 * (-3) = 360 360 * (-2) = -720 -720 * (-1) = 720. So that’s 720. [Remaining trajectory omitted.] Outcome: Incorrect

Case 2: Incorrect Student Rollout

Problem. Alex the Kat has written 61 problems for a math contest, and there are a total of 187 problems submitted. How many more problems does he need to write and submit before he has written half of the total problems? Answer and error. The student outputs 93.5, whereas the gold answer is 65. It incorrectly keeps the total fixed at 187 and treats 187/2 = 93.5 as the required additional count. TSD Legend. (0.1, 0.3]

[ −0.3, −0.1) [ −0.7, −0.3) [ −1.5, −0.7) (0.3, 0.7] (0.7, 1.5] (1.5, 3] >3

[ −3, −1.5)

< −3

pos

Student trajectory excerpt under 𝑐sol . Okay, let ’s try to figure out this problem step by step. So, Alex the Kat has already written 61 problems for a math contest. The total number of problems submitted is 187. The question is asking how many more problems he needs to write and submit before he has written half of the total problems. First, I need to understand what exactly is being asked here. The total number of problems submitted is 187, but I

23

Calibrated On-Policy Distillation

Preprint

think that might not be directly relevant . Wait, actually, the problem says Alex has written 61 problems, and there are a total of 187 problems submitted. Hmm , maybe I need to clarify that. Wait, maybe the total number of problems in the contest is 187? Or is 187 the number of problems submitted by others? The wording is a bit confusing. Let me read it again: "Alex the Kat has written 61 problems for a math contest, and there are a total of 187 problems submitted. How many more problems does he need to write and submit before he has written half of the total problems?" So, the total number of problems submitted is 187. But Alex has written 61. Wait, maybe " submitted" refers to the total number of problems that have been submitted by everyone , including Alex? Or is 187 the total number of problems in the contest? This is a bit ambiguous. But let ’s think . If Alex has written 61 problems, and there are 187 problems submitted in total, maybe the total number of problems in the contest is 187? But then he has written 61, so he needs to write half of 187, which is 93.5? But that doesn’t make sense because you can’t write half a problem. Wait, maybe the total number of problems is different . Alternatively, maybe the total number of problems in the contest is 187, and Alex has written 61. Then he needs to write half of the total, which is 187/2 = 93.5. But since you can’t have half a problem, maybe it’s 94? But the problem says "how many more problems does he need to write and submit," so maybe it’s expecting a fractional answer? But that seems odd. Wait, maybe the total number of problems is not 187. Let me re-examine the problem. "Alex the Kat has written 61 problems for a math contest, and there are a total of 187 problems submitted. How many more problems does he need to write and submit before he has written half of the total problems?" Hmm . Maybe "submitted" refers to the total number of problems that have been submitted by all participants, including Alex? So, Alex has written 61 problems, and there are 187 problems submitted in total. So , the total number of problems in the contest is 187? But then Alex has written 61, so he needs to write half of 187, which is 93.5. But since you can’t write half a problem, maybe the answer is 94? But the problem is probably expecting an exact answer . Alternatively, maybe the total number of problems in the contest is not 187. Maybe 187 is the total number of problems that have been submitted, and Alex has written 61. Then, the total number of problems in the contest is 187? But then he needs to write half of that, which is 93.5. But again, this seems odd. Wait, maybe the problem is that Alex has written 61 problems, and there are 187 problems in total. So , he needs to write half of 187, which is 93.5. But since he can’t write half a problem, maybe we need to round up? But the problem might be expecting an exact answer. Alternatively, maybe the problem is that the total number of problems submitted is 187, and Alex has written 61. So, the total number of problems that have been submitted is 187. But Alex needs to write half of the total problems. Wait, but if the total problems submitted is 187, then half of that is 93.5. So , he needs to write 93.5 more problems? But that’s not possible. Wait, maybe I need to think differently. Maybe the total number of problems in the contest is 187. Alex has written 61, so he needs to write half of 187, which is 93.5. Therefore, he needs to write 93.5 more problems. But since you can’t write half a problem, maybe the answer is 94? But the problem says "how many more problems does he need to write and submit," so maybe it’s expecting a fractional answer? But in reality, problems are whole numbers. [Remaining trajectory omitted.]

TSD Throughout a Correct Trajectory. Case 1 shows a rollout that correctly concludes 2 after recognizing that the exponent contains zero. Despite reaching the correct reasoning early, the student repeatedly revisits its interpretation. TSD is distributed throughout this process rather than being confined to task-critical mathematical steps. Strong deviations frequently occur on discourse and reasoning-management expressions such as Hmm, Wait, But, However, Maybe, and Alternatively, whereas much of the numerical and symbolic content remains comparatively stable. This qualitatively matches the token-level pattern in Table 2. TSD Is Not Necessarily Localized Around Reasoning Errors. Case 2 shows an incorrect rollout in which the student repeatedly treats 187 as a fixed total and reasons from 187/2 = 93.5. The correct relation is 61 + 𝑥 =

187 + 𝑥 , 2

24

(16)

Calibrated On-Policy Distillation

Preprint

where 𝑥 denotes the number of additional problems Alex needs to write, which yields 𝑥 = 65. However, the TSD pattern does not concentrate around this identifiable conceptual error or the repeated occurrence of 93.5. Instead, substantial deviations again appear broadly on expressions such as First, Hmm, Wait, Alternatively, and Therefore. Thus, even when the privileged solution directly resolves the student’s mistake, the induced likelihood variation does not behave as a localized reasoning-error signal. Takeaway. Across both correct and incorrect trajectories, TSD is broadly distributed and particularly pronounced on tokens that organize or redirect reasoning, rather than selectively aligning with mathematical correctness or error locations. These examples provide a trajectory-level illustration of why context-induced teacher variation should not be indiscriminately treated as transferable supervision, further motivating the calibration mechanism in Cal-OPD.

A.6 Training Details and Hyperparameters We implement all methods in the verl framework (Sheng et al., 2025) and train on 8 NVIDIA H20 GPUs, with 4 GPUs hosting the student and 4 hosting the teacher. All methods share a common training recipe unless otherwise specified. We first summarize the shared configuration, then describe method-specific hyperparameters and the algorithmic role each plays. Shared Training Configuration. Table 10 lists the optimization and training-loop settings common to all baselines and Cal-OPD. All methods are trained for 100 steps, with a maximum response length of 16,384 tokens during training, temperature 1.0, and top-𝑝 1.0. We use AdamW with a learning rate of 1 × 10−6 , no learning-rate warmup, weight decay 0.01, a cosine schedule, and gradient clipping at 1.0. Each training step processes 256 on-policy trajectories, with one rollout per question, one PPO epoch, and a mini-batch size of 256. The entropy coefficient and KL regularization are set to zero because the OPD distillation objective replaces the standard RL KL penalty. We use GRPO as the advantage estimator with standard deviation normalization within each group (Yu et al., 2026a). Distillation-Specific Configuration. The distillation module operates on student-generated trajectories and queries the teacher for token-level log-probabilities. We set the student chunk size to 1024 and the teacher chunk size to 128, which controls the number of tokens processed per forward pass and is chosen to balance GPU memory against throughput. The teacher returns log-probabilities over its top-16 tokens by default, providing a dense yet bounded supervision signal. All student-generated tokens participate in the distillation loss without filtering, corresponding to a token selection ratio of 1.0 with random selection. The policy loss is computed in reinforce mode, and the per-token loss is clamped at a maximum value of 10 to prevent destabilizing updates from outlier advantages. Method-Specific Hyperparameters. Table 11 summarizes the hyperparameters that differ across baselines. Below, we describe the algorithmic role of each. ExOPD. ExOPD extrapolates the teacher-derived reward beyond the teacher’s own performance level with a scaling factor 𝜆exo > 1 (Yang et al., 2026a); we set 𝜆exo = 1.25. It maintains both the current policy and the pre-RL base model. To avoid storing duplicate weights, we parameterize the student with LoRA adapters (rank 64, alpha 128, all linear layers) and switch between base and policy by disabling or enabling the adapter. We set loss_max_clamp=null and loss_agg_mode=token-mean. EOPD. EOPD augments the reverse-KL distillation objective with a forward-KL term applied selectively to high-entropy teacher tokens (Jin et al., 2026). The entropy threshold 𝜏ent = 0.8 determines which tokens receive the forward-KL treatment: tokens whose teacher entropy exceeds this threshold are supervised with forward KL to preserve generation diversity, while lower-entropy tokens retain the standard reverse-KL objective. The mixing 25

Calibrated On-Policy Distillation

Preprint

Category

Hyperparameter

Value

Learning rate LR warmup ratio Weight decay LR scheduler Gradient clip

1 × 10−6 0.0 0.01 cosine 1.0

Optimizer

Training loop Training steps Max prompt length Max response length Temperature Top-𝑝 PPO epochs Train batch size PPO mini-batch size PPO micro-batch size Entropy coefficient KL regularization

100 2,048 16,384 1.0 1.0 1 256 trajectories/step 256 1 + dynamic batching 0.0 disabled

Distillation Student chunk size Teacher chunk size Teacher max logprobs Token selection ratio Policy loss mode Advantage and loss Loss max clamp Advantage estimator

1024 128 16 1.0 reinforce 10 GRPO

Table 10: Shared training configuration for all methods.

coefficient 𝛼ent = 1.0 controls the weight of the forward-KL term relative to the reverse-KL term. To compute teacher entropy, EOPD requires one additional log-probability slot beyond the top-16 used by other methods; we therefore set teacher max logprobs to 17 for EOPD only. Uni-OPD. Uni-OPD introduces a dual-perspective recipe that combines student-side data balancing with teacherside outcome-guided margin calibration (Hou et al., 2026a). On the student side, an online correctness-aware filter reshapes each training batch to maintain a target correct-to-incorrect ratio 𝜌corr = 0.5; we implement this by training on 16 questions per step with 𝑛=16 rollouts per question, yielding 256 trajectories per step, and applying sample filtering to achieve the target ratio. The mini-batch size is correspondingly set to 16 questions. On the teacher side, Uni-OPD calibrates token-level teacher margins against trajectory-level outcome rewards. The margin scope is set to group, meaning calibration is performed within each group of rollouts for the same question. The trajectory reduction is mean, the margin direction is spread (which spreads apart the margins of correct and incorrect trajectories), and the target margin is 𝛿margin = 0.4. The loss aggregation mode is token-mean, and the loss is not clamped.

26

Calibrated On-Policy Distillation

Preprint

Method

Method-specific hyperparameters

OPD

None beyond the shared configuration.

ExOPD

exopd_lambda=1.25; loss_max_clamp=null; LoRA: rank=64, alpha=128, target=all-linear.

EOPD

eopd_entropy_threshold=0.8;

Uni-OPD

Batch structure: 16 questions per step, 16 rollouts per question (256 trajectories per step); mini-batch size 16 questions. Loss: loss_agg_mode=token-mean; loss_max_clamp=null. Margin calibration: target_correct_ratio=0.5; margin_scope=group; trajectory_reduce=mean; margin_direction=spread; margin_delta=0.4.

Privileged-OPD

Max prompt length=12,288 to accommodate reference solutions; longer solutions are truncated. Otherwise identical to OPD.

loss_agg_mode=token-mean.

eopd_alpha=1.0;

teacher max logprobs=17.

Table 11: Method-specific hyperparameters. Entries marked “none” use the shared configuration unchanged.

Privileged-OPD. Privileged-OPD conditions the teacher on additional training-time information, specifically the verified reference solution for each problem (Ye et al., 2026; Kaur et al., 2026). Because reference solutions are substantially longer than the problem statements, we increase the maximum prompt length from the shared default to 12,288 tokens to accommodate the concatenation of problem, reference solution, and student rollout prefix. Solutions exceeding this length are truncated. All other settings follow the standard OPD configuration. In our implementation, the privileged context is prepended to the problem and student prefix in the teacher’s input, while the student receives only the problem and its own rollout, preserving the information asymmetry that Privileged-OPD exploits. Cal-OPD. Cal-OPD follows the shared configuration above. It additionally uses the evaluative feedback neg pos interventions 𝑐 eval and 𝑐 eval to probe the teacher’s self-deviation region, with relaxation factor 𝜆 = 5. These interventions are used only to estimate the TSD region and are not directly distilled into the student. The distillation loss then operates on the calibrated advantage 𝐴𝑡Cal in place of the raw 𝐴𝑡OPD , as described in Section 3.

A.7 Additional Training Dynamics OPD vs. Cal-OPD with evaluative calibration. Figure 9 reveals qualitatively different training dynamics between OPD and Cal-OPD with Ceval . OPD progressively increases response length while reducing student entropy, whereas Cal-OPD maintains substantially shorter responses and higher entropy throughout training. This contrast is also reflected in how each method aligns with the teacher’s top-16 distribution. Let 𝐻 denote the student entropy, Top16Tok the top-16 token overlap, and Top16Mass the top-16 probability-mass overlap. The dynamics can be summarized as 𝐻Cal-OPD > 𝐻OPD , Top16TokCal-OPD > Top16TokOPD , (17) Top16MassCal-OPD < Top16MassOPD . In particular, the top-16 token overlap under Cal-OPD steadily increases to approximately 0.693, whereas OPD peaks earlier and then slightly declines. Meanwhile, OPD achieves a higher top-16 probability-mass overlap, indicating stronger matching of the teacher’s probability allocation. Together, these dynamics suggest that OPD increasingly fits the teacher’s dominant probability pattern, whereas Cal-OPD preserves a broader student distribution while covering more teacher-supported token patterns, enabling the student to capture multiple plausible reasoning modes rather than overfitting to a

27

Calibrated On-Policy Distillation

Preprint

0.35

10K

Entropy

Average Response Length

12K

8K

OPD Cal-OPD (inst ) Cal-OPD (eval )

6K

0

20

40

60

Cal-OPD (ans ) Cal-OPD (sol )

0.30

0.25

0.20 80

100

0

20

Training Steps

Top-16 Probability-Mass Overlap

0.90

Top-16 Token Overlap

0.695 0.690 0.685 0.680

0.670 0

20

Cal-OPD (ans ) Cal-OPD (sol ) 40

60

60

80

100

(b) Entropy

0.700

OPD Cal-OPD (inst ) Cal-OPD (eval )

40

Training Steps

(a) Average Response Length

0.675

Cal-OPD (ans ) Cal-OPD (sol )

OPD Cal-OPD (inst ) Cal-OPD (eval )

80

OPD Cal-OPD (inst ) Cal-OPD (eval )

0.89

Cal-OPD (ans ) Cal-OPD (sol )

0.88

0.87

0.86

100

0

Training Steps

20

40

60

80

100

Training Steps

(c) Top-16 Token Overlap

(d) Top-16 Probability-Mass Overlap

Figure 9: Training dynamics of OPD and Cal-OPD on Qwen3-1.7B.

particular one. Cal-OPD with instruction-level calibration. In contrast to evaluative calibration, instruction-level calibration with Cinst yields dynamics that closely resemble standard OPD, indicating that its calibration signal is too weak to meaningfully separate TSD from transferable knowledge. As shown in Figure 9, the student entropy under Cinst follows the same declining trend as OPD and stabilizes around 0.20, whereas the other Cal-OPD variants maintain substantially higher entropy. Its top-16 probability-mass overlap is also nearly identical to that of OPD, stabilizing around 0.886, which shows that it continues to match the teacher’s dominant probability allocation as strongly as the uncalibrated objective. Similarly, its top-16 token overlap remains low at approximately 0.684, again close to OPD and well below the levels reached by Ceval , Cans , and Csol . These trends indicate that, without task-specific information or strong semantic contrast, the estimated TSD region is too narrow to filter the teacher–student discrepancy effectively, and Cal-OPD largely degenerates to standard OPD behavior. Cal-OPD with solution-level calibration. Solution-level calibration with Csol lies at the opposite extreme: it over-filters the teacher–student discrepancy and causes the supervision signal to collapse. As shown in Figure 9, the student entropy under Csol is the highest among all methods and continues to rise throughout training, stabilizing near 0.36, while its average response length is the lowest and steadily decreases to roughly 5.5K tokens. More strikingly, its top-16 probability-mass overlap is not only the lowest among all variants but also falls below its initial 28

Preprint

Entropy

0.35

0.30

Privileged-OPD Cal-OPD (sol )

0.25

0.700

Privileged-OPD Cal-OPD (sol )

0.88

0.695

Top-16 Token Overlap

Top-16 Probability-Mass Overlap

Calibrated On-Policy Distillation

0.87

0.86

20

40

60

80

0.685

0.680 0.85

0

0.690

100

Privileged-OPD Cal-OPD (sol )

0.675 0

Training Steps

20

40

60

80

100

0

20

Training Steps

(a) Entropy

(b) Top-16 Probability-Mass Overlap

40

60

80

100

Training Steps

(c) Top-16 Token Overlap

Figure 10: Training dynamics comparison between privileged OPD and Cal-OPD with solution-level calibration. Method

Transfer

Avg. time per step (min)

Total time for 100 steps (h)

Cal-OPD OPD

30B-A3B-2507 → 8B 30B-A3B-2507 → 8B

50.05 35.19

∼83.4 ∼58.7

Cal-OPD OPD

30B-A3B-2507 → 4B 30B-A3B-2507 → 4B

48.76 31.26

∼81.3 ∼52.1

Cal-OPD OPD

4B-2507 → 1.7B 4B-2507 → 1.7B

14.90 18.80

∼24.8 ∼31.3

Table 12: Efficiency comparison between Cal-OPD and OPD across different teacher–student configurations.

value, stabilizing around 0.862. Its top-16 token overlap remains moderate at approximately 0.688, below Ceval and Cans , but above OPD. These trends indicate that the TSD region estimated from solution-level privilege is excessively broad, so that a large fraction of the teacher–student discrepancy is classified as TSD and zeroed out. As a result, the student receives too few effective supervision signals to follow the teacher distribution, and instead drifts toward a high-entropy, short-response regime. Over-calibration therefore degrades the signal more severely than no calibration at all, which is consistent with its lowest downstream performance among all variants. Failure Modes of Privileged-OPD and Cal-OPD with Solution-Level Calibration. Both methods shorten responses and degrade performance, but for different reasons. All top-16 metrics are computed against the unprivileged teacher. Privileged-OPD keeps student entropy low and stable (around 0.25) while its top-16 probability-mass overlap rises to about 0.880, indicating that the student becomes overconfident and concentrates on a narrow token set. In contrast, Cal-OPD with Csol drives entropy up to about 0.36, and its top-16 probability-mass overlap even drops to roughly 0.859, showing that it fails to match the unprivileged teacher. We attribute this to two distinct failure modes: Privileged-OPD suffers from overconfident shortcut collapse, where the student imitates the privileged teacher’s sharp distribution but loses diversity; Cal-OPD with Csol suffers from signal collapse, where an excessively broad TSD region zeros out useful supervision, causing the student to drift into a high-entropy, short-response regime.

A.8 Efficiency Analysis Efficiency of Cal-OPD. Cal-OPD requires two additional teacher forward passes per step to probe the TSD region, which introduces extra computation. Table 12 compares the wall-clock time per step and the total time for 100 steps between Cal-OPD and OPD across three teacher–student configurations. For the Qwen3-4B-Thinking2507→Qwen3-1.7B setting, Cal-OPD is actually faster than OPD (14.90 vs. 18.80 minutes per step), because

29

Calibrated On-Policy Distillation

Preprint

OPD

Cal-OPD

OPD

Pre-clipping Gradient Norm

Pre-clipping Gradient Norm

120

90

60

30

0

20

40

60

80

100

60

40

20

0

Training Steps

Cal-OPD

20

40

60

80

100

Training Steps

(a) Qwen3-4B-Thinking-2507 → Qwen3-1.7B

(b) Qwen3-30B-A3B-Thinking-2507 → Qwen3-4B

Figure 11: Gradient norm dynamics of OPD training across different teacher–student pairs.

Cal-OPD avoids the response-length expansion that OPD exhibits, and the teacher is not large enough for the extra forward passes to dominate. In the other two configurations, where the teacher (30B-A3B) is substantially larger than the student (8B or 4B), the two additional teacher forward passes become the main computational overhead, making Cal-OPD slower per step (50.05 vs. 35.19 and 48.76 vs. 31.26 minutes). Overall, Cal-OPD remains practically efficient: its extra cost is bounded by two teacher forward passes, and it can even reduce total training time when the teacher–student size gap is moderate.

A.9 Retention-Matched Control on Teacher–Student Discrepancy Figure 11 reports the pre-clipping gradient norm during OPD training. In both teacher–student configurations, the norm is initially large, around 120 and 65, and gradually decreases to approximately 35 and 15. All methods apply gradient clipping with threshold 𝐶 = 1, so the update is computed as   𝐶 . (18) g̃ = g · min 1, ∥g∥ Since ∥g∥ is far above 1 throughout training, the raw gradient is scaled by roughly 𝐶/∥g∥, corresponding to an attenuation of one to two orders of magnitude. Thus, substantial signal attenuation is already inherent to standard OPD, and the gains of Cal-OPD cannot be attributed to a smaller overall gradient magnitude. Cal-OPD instead changes the relative contribution of individual token advantages by removing the TSD-explained component, while leaving the global clipping behavior unchanged. To rule out the possibility that Cal-OPD gains merely from attenuating the optimization signal, we construct three retention-matched baselines that mimic its signal reduction without TSD calibration. • Advantage-Sync (Advantage-Level Synchronization). At each training step 𝑘, we read Cal-OPD’s advantage retention ratio 𝑟 𝑘 , defined as the ratio between the sum of absolute calibrated advantages and the sum of absolute original teacher–student discrepancies. We then scale standard OPD’s token-level advantages by the same factor, 𝐴˜ 𝑡OPD = 𝑟 𝑘 · 𝐴𝑡OPD , so that the global advantage magnitude matches Cal-OPD step by step. • Token-Sync (Token-Level Synchronization). At each training step 𝑘, we read Cal-OPD’s zero-advantage token ratio 𝑧 𝑘 , i.e., the fraction of tokens whose calibrated advantage is zero. We then randomly select the same fraction 𝑧 𝑘 of tokens in standard OPD and set their advantages to zero, keeping the remaining advantages unchanged. This matches Cal-OPD’s sparsity pattern in terms of the number of discarded tokens, but not which tokens are discarded. • TSD-Filter (TSD-Threshold Filtering). For each token, we compute the maximum absolute deviation induced 30

Calibrated On-Policy Distillation

Preprint

Table 13: Retention-matched control on teacher–student discrepancy for Qwen3-4B-Thinking-2507→Qwen3-1.7B. Method OPD Advantage-Sync Token-Sync TSD-Filter (𝜏TSD = 0.1) TSD-Filter (𝜏TSD = 0.05) TSD-Filter (𝜏TSD = 0.01) Cal-OPD

AMC23

AIME24

AIME25

AIME26

82.0 81.7 81.3 83.6 81.7 80.3 84.5

39.6 39.8 38.3 44.0 41.7 38.8 45.0

35.0 35.4 34.6 37.1 35.4 33.5 37.1

34.6 35.0 34.0 36.0 34.8 32.9 36.3

HMMT26 MATH500 22.9 22.2 22.3 24.4 22.9 21.6 24.8

90.6 90.5 89.8 90.9 90.1 89.3 90.8

Avg. 50.8 50.8 50.1 52.7 51.1 49.4 53.1

by the positive and negative privileged interventions, and discard the token entirely if this deviation exceeds a threshold 𝜏TSD . We evaluate three thresholds, 𝜏TSD ∈ {0.1, 0.05, 0.01}, corresponding to increasingly aggressive filtering. Unlike Cal-OPD, which retains the residual discrepancy beyond the estimated TSD region, this baseline removes the token from the loss altogether. These controls match Cal-OPD in global advantage scale, token sparsity, and TSD-based token selection, respectively, but omit the calibrated residual. Table 13 compares Cal-OPD against three retention-matched controls. Advantage-Sync applies the same global scaling to OPD’s advantages and achieves 50.8, identical to OPD, showing that reducing the overall advantage magnitude alone brings no benefit. Token-Sync randomly masks the same fraction of tokens and even drops to 50.1, indicating that matching the sparsity level without selecting the right tokens is insufficient. TSD-Filter with 𝜏TSD = 0.1 reaches 52.7, the strongest control, but still falls short of Cal-OPD; more aggressive thresholds of 0.05 and 0.01 degrade to 51.1 and 49.4, respectively. These results show that neither global attenuation, nor random token removal, nor hard TSD-based filtering can reproduce the gains of Cal-OPD. The advantage of Cal-OPD therefore stems from its soft calibration mechanism, which retains the discrepancy beyond the estimated TSD region, rather than from signal attenuation or token sparsity alone.

31

Record · ID 1006914 · SHA-256 1c12506e920f702f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.