Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Youngrok Park1,* Sangmin Bae1,*,† Hojung Jung1 Jongwoo Ko2 Yunseon Choi3 Young Jin Kim2 Pashmina Cameron2 Aaron Courville4,5,6 Se-Young Yun1 1
KAIST AI, 2 Microsoft, 3 University of Toronto, 4 Mila, 5 Université de Montréal, 6 CIFAR AI Chair
arXiv:2609.08798v1 [cs.LG] 8 Sep 2026
*Equal Contribution, † Project Lead
Abstract: Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher’s policy shift relative to its reference policy on student rollouts and amplifies the component of the student’s verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student’s own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.
1. Introduction
Reasoning (Pass@1)
Math (Mean@16) 50
60
40
40
20
GRPO OPD OPRD
0
60
120
0
0
60
120
Algorithms (Pass@1)
40
40
30
0
150
300
20
Mix-RL MOPD OPRD
20
20 0
Games (Pass@1) 60
40
60
20
30
Logic (Pass@1) 80
0
0
150
300
0
150
300
Figure 1: On-policy reverse distillation (OPRD) enables faster and stronger weak-to-strong generalization across two key settings. (Left) For successive model transfer, a checkpoint from a post-trained 4B-scale model serves as the teacher for an 8B-scale student. We average evaluations conducted every 30 training steps: Mean@16 over AIME'24, AIME'25, HMMT'25, and OlympiadBench for math, and Pass@1 over Knights & Knaves, Quantum Lock, String Manipulation, and Countdown for reasoning tasks. (Right) In multi-domain consolidation, four domain-specialized 4B-scale teachers are distilled into a single 8B-scale student. Training examples are randomly mixed within each batch, with the corresponding domain teacher activated for each example. We report performance every 60 steps for Logic (averaged over Knights & Knaves and Quantum Lock), Algorithms (String Manipulation), and Games (Countdown). The gray dashed lines denote the performance of the corresponding weak teachers.
Correspondence to: {yr-park, bsmn0223, yunseyoung}@kaist.ac.kr. Code is at https://github.com/raymin0223/on_policy_reverse_distillation.
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Knowledge distillation (KD; Hinton et al., 2015) transfers knowledge from a teacher model to a student. For autoregressive language models, conventional distillation on fixed or teacher-generated sequences can create a mismatch between the prefixes seen during training and those visited by the student at inference time. On-policy distillation (OPD) (Gu et al., 2024; Agarwal et al., 2024; Ko et al., 2024) addresses this mismatch by training on student-generated responses and querying the teacher at the prefixes the student visits. Recent work has applied OPD to efficient reasoning post-training (Qwen Team, 2025b; Xu et al., 2025; GLM-5 Team, 2026) and to consolidating capabilities from multiple domain-specific teachers into a single student (Xiaomi Team, 2026; Yang et al., 2026b; DeepSeek-AI, 2026). Standard OPD optimizes the student toward the teacher policy, making it well suited when matching that policy is the goal. However, useful supervision need not come from a model that the student should ultimately match. Weakto-strong generalization has shown that stronger pretrained models can learn from weaker supervisors and even outperform them across language understanding, reward modeling, and reasoning tasks (Burns et al., 2024; Yang et al., 2024; Lang et al., 2024; Zhou et al., 2025). This regime is especially promising in two key settings in modern foundation model development. (i) Successive model transfer (Figure 1, top-left): A post-trained model from one generation can supervise a larger-scale successor, enabling it to inherit prior post-training gains and improve beyond its supervisor. (ii) Multi-domain consolidation (Figure 1, top-right): Domain-specialized policies can be developed independently at smaller scale, enabling efficient iteration on reward functions, environments, and training recipes. Multi-teacher on-policy distillation (MOPD) (Kimi Team, 2026; Ma et al., 2026; Xiaomi Team, 2026) can then consolidate their capabilities into a unified foundation model. Both settings therefore call for reverse distillation that transfers post-training gains from weaker models without limiting the eventual performance of higher-capacity students. Simply applying OPD in the weak-to-strong direction does not resolve this problem. A weak teacher’s final policy combines changes learned during post-training, preferences inherited from its reference policy, and behavior shaped by its limited capacity. Standard OPD matches this entire distribution, transferring all three and retaining the weak policy as the target at each student-visited prefix. Teacher matching can provide useful guidance when the student underperforms the teacher, but can also suppress surprising student behavior when the teacher favors a different solution (Akhondzadeh et al., 2026; Ziheng et al., 2026). Adding reinforcement learning does not remove this tension if teacher matching remains a separate objective, since the matching loss can compete with reward maximization (Xu et al., 2025; Zhang et al., 2026a). Likewise, isolating the teacher’s post-training policy change is insufficient if the student is still trained to match it. This change captures only the improvements realized by the weak teacher, not the full range available to the stronger student, so direct matching can impose the same capacity limitation. The central question is therefore how to exploit weak-model post-training gains without making either the weak policy or its policy change an independent optimization target. We introduce On-Policy Reverse Distillation (OPRD), which uses the policy change learned during weakmodel post-training to accelerate a stronger student’s own optimization. On the student’s on-policy rollouts, OPRD computes the verifier-driven policy gradient and extracts the weak teacher’s policy shift relative to its reference policy. It projects the student gradient onto the direction of this shift and amplifies the projected component, leaving the orthogonal component unchanged. Because this transformation positively rescales only a component already present in the student gradient, it preserves the stationary points of policy optimization in logit space while adding a nonnegative first-order alignment gain. When the teacher shift and student gradient align, OPRD reinforces their shared direction and accelerates convergence; when they oppose, it strengthens surprising student behavior supported by the verifier, allowing the student to improve beyond the teacher. We evaluate OPRD across mathematical reasoning (MAA, 2024–2025; Dekoninck et al., 2026; He et al., 2024) and logical reasoning tasks (Stojanovski et al., 2026) in two main weak-to-strong scenarios. In successive model transfer, OPRD reaches weak-teacher performance with 33–67% fewer student updates than GRPO (Shao et al., 2024) and achieves up to 22.7 percentage points higher performance at early checkpoints. Unlike OPD, it then moves beyond the teacher rather than saturating after the initial transfer (Figure 1, bottom-left). In the multiteacher setting, OPRD distills four specialized smaller-scale teachers into a single stronger student, reaching teacher-level performance with 55% fewer updates than Mix-RL; the resulting student ultimately outperforms all four specialists (Figure 1, bottom-right). With the same number of rollouts per update, these gains reflect improved sample efficiency during student training. The benefit extends to conventional strong-to-weak distillation, where OPRD moves beyond OPD’s plateau through verifier-driven optimization. Together, these results show that weak teachers can accelerate the post-training of stronger models without limiting students to their teachers’ capabilities, opening a practical path to reusing post-training gains across model generations and domains at scale. 2
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Contributions.
In summary, our key contributions in this paper are as follows.
• Weak-to-Strong Generalization. We study how post-training gains from weaker models can be transferred to stronger students in two practical scenarios: successive model transfer and multi-domain consolidation. We identify the central challenge as exploiting these gains without making either the weak policy or its policy shift a separate optimization target. • On-Policy Reverse Distillation. We introduce OPRD, which evaluates a weak teacher’s policy shift relative to its reference policy on student rollouts and amplifies the component of the student’s verifier-driven policy gradient along that direction. Because OPRD only rescales verifier-supported updates, it accelerates the student’s own optimization while preserving its stationary points, allowing the student to move beyond the teacher. • Empirical Evaluation and Analysis. Across successive-model and multi-teacher settings, OPRD reaches the final performance of competing methods substantially earlier and ultimately outperforms both RL and distillation baselines (§3.2, §3.3). We further confirm that these gains extend to conventional strong-to-weak distillation (§3.4). We also compare against recent weak-to-strong methods (§4.1), analyze the design and dynamics of teacher guidance (§4.2), examine practical challenges and mitigations (§4.3), and study student reasoning and response style under teacher guidance (§4.4).
2. Method 2.1. Preliminary Reinforcement Learning with Verifiable Rewards (RLVR). RLVR optimizes a language-model policy using rewards computed by programmatic verifiers, such as exact-answer checks or code execution, and has become central to reasoning post-training (Shao et al., 2024; Guo et al., 2025). For 𝑥 ∼ 𝒟, the student samples 𝑦 ∼ 𝜋𝜃 (· | 𝑥) and visits prefixes 𝑠𝑡 = (𝑥, 𝑦<𝑡 ). Let 𝐴𝑡 denote the advantage assigned to token 𝑡 and z𝑡 the corresponding next-token logits at prefix 𝑠𝑡 . The token-level policy gradient is g𝑡 := 𝐴𝑡 ∇z𝑡 log 𝜋𝜃 (𝑦𝑡 | 𝑠𝑡 ).
(2.1)
OPRD later rescales this gradient while preserving the RLVR objective, so the student’s attainable performance is determined by the verifier objective and its own policy class rather than being bounded by the teacher’s capacity. On-Policy Distillation (OPD). OPD reduces the training–inference distribution mismatch by sampling responses from the student and querying the teacher at each visited prefix, thereby providing dense token-level supervision over the student’s inference-time state distribution (Gu et al., 2024; Agarwal et al., 2024; Ko et al., 2024). A common reverse-KL formulation is [︃ ]︃ ∑︁ ℒOPD (𝜃) := 𝐷KL (𝜋𝜃 (· | 𝑠𝑡 ) ‖ 𝜋𝑇 (· | 𝑠𝑡 )) . (2.2) E 𝑥∼𝒟 𝑦∼𝜋𝜃 (·|𝑥)
𝑡
Equivalently, OPD can be implemented as token-level policy optimization on student-sampled tokens using the teacher-to-student log-probability ratio as the advantage, with negligible empirical differences from direct reverse-KL optimization. OPD is increasingly used in frontier-model post-training for reasoning and capability integration across domains (Ma et al., 2026; Xiaomi Team, 2026; GLM-5 Team, 2026; Yang et al., 2026b). Recent methods combine teacher matching with reinforcement learning to pair dense teacher supervision with outcome-based optimization (Xu et al., 2025; Ramos et al., 2026). Even in these hybrid methods, however, teacher matching remains a separate objective, leaving the teacher policy as a direct optimization target. Weak-to-Strong Generalization. Weak-to-strong generalization studies whether a more capable model can learn from weaker supervisors, such as smaller models or imperfect human feedback, and ultimately outperform them (Burns et al., 2024). Prior work has used weak labels, preferences, and fixed reasoning trajectories to supervise stronger students. Refinement methods help the student exploit its own representations and greater capacity, but often recover only part of the gap to strong supervision (Yang et al., 2024; Somerstep et al., 2025; Dong et al., 2025; Medvedev et al., 2025). OPD instead provides the full next-token distribution 𝜋 ¯𝑇 (· | 𝑠𝑡 ) at each student-visited prefix, where 𝜋 ¯𝑇 denotes either the weak teacher or a target policy derived from it. Under realizability, the resulting KL objective has the pointwise minimizer arg min 𝐷KL (𝜋(· | 𝑠𝑡 ) ‖ 𝜋 ¯𝑇 (· | 𝑠𝑡 )) = 𝜋 ¯𝑇 (· | 𝑠𝑡 ).
(2.3)
𝜋(·|𝑠𝑡 )
3
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Student Policy Gradient
∇zt log πθ(yt | st) Student Rollout
x∼D πθ Student x
y1 … yt−1 yt
st = (x, y<t)
× ✔
At
Directional Gradient Scaling
g̃t = gt + λt Projdt(gt)
= (1 + λt) Projdt(gt) + g⊥t
gt = At ∇zt log πθ (yt ∣ st)
gt
Teacher Policy Shift
θt
πT( ⋅ | st) Teacher
πTref( ⋅ | st)
Amplify
dt
Teacher Reference
Δt = C (zT − zTref), dt = Δt / ∥Δt∥2
πθ0
g̃t g⊥t
(1 + λt) Projdt(gt)
πT
g̃t
ut = ⟨dt, gt⟩
−
π∗
gt
Projdt(gt) = ut dt,
Project
✘
Learning Beyond Teacher
dt πTref
OPRD amplifies the component of gt along dt instead of targeting πT, preserving the stationary points (i.e., g̃t = 0 ⟺ gt = 0).
Figure 2: Conceptual overview of On-Policy Reverse Distillation (OPRD). The figure illustrates OPRD’s gradient correction procedure for a single query. Here, z𝑇 denotes the logits of the weak teacher after post-training and zref 𝑇 those of its reference policy, and 𝒞 denotes mean-centering. Their centered difference Δ𝑡 is the teacher’s policy shift at that student-visited prefix, and OPRD keeps only its unit direction d𝑡 . In practice, we use a simple top-10 truncation under the student policy to focus the correction on its high-probability vocabulary region. The rightmost panel provides a conceptual view of the resulting student trajectory in the optimization landscape, where the student follows the verifier-driven policy gradient g𝑡 with its component along d𝑡 amplified by 1 + 𝜆𝑡 at each token and its orthogonal component left unchanged.
Alternative teacher-derived targets only change which policy the student matches, while adding reinforcement learning yields a compromise between teacher matching and reward maximization (Xu et al., 2025; Ramos et al., 2026). In both cases, the student remains directly optimized toward a policy defined by the weak teacher. OPRD instead extracts the policy change learned during weak-model post-training and uses it only to rescale the stronger student’s own policy gradient. 2.2. On-Policy Reverse Distillation Overview. OPRD transfers the policy change learned during teacher post-training rather than matching the teacher’s final policy. At each student-visited prefix, it extracts the local direction of this change relative to the teacher’s reference policy and uses its alignment with the verifier-driven student gradient to rescale only the gradient component along that direction. Because the teacher signal only rescales the student’s own gradient, it can accelerate verifier-supported optimization without defining an independent optimization target. Positive-alignment scaling is active from the outset to amplify updates supported by both the verifier and the teacher, whereas negative-alignment scaling is gradually increased to reinforce verifier-supported departures beyond the weak teacher. Teacher Policy Shift. The teacher’s final policy reflects the change acquired during RL post-training, preferences inherited from its reference policy, and behavior constrained by the weak model’s limited capacity. Directly matching it would therefore make all of these part of the student’s distillation target. Let 𝜋𝑇 denote the frozen RL-trained teacher and 𝜋𝑇ref its frozen pre-RL reference policy, and let z𝑇 (𝑠𝑡 ) and zref 𝑇 (𝑠𝑡 ) denote their next-token logit vectors at a student-visited prefix 𝑠𝑡 . To extract the RL-induced policy delta, we mean-center the difference between the teacher and reference logits, removing a common offset that does not affect relative 1 token preferences. With 𝒞(v) := v − |𝒱| (1⊤ v)1, we define (︀ )︀ (︀ )︀ ref Δ𝑡 := 𝒞 z𝑇 (𝑠𝑡 ) − zref 𝑇 (𝑠𝑡 ) = 𝒞 log 𝜋𝑇 (· | 𝑠𝑡 ) − log 𝜋𝑇 (· | 𝑠𝑡 ) .
(2.4)
Intuitively, Δ𝑡 captures the change in the teacher’s relative next-token preferences induced by post-training, and the corresponding uncentered log-policy ratio admits an implicit-reward interpretation under KL-regularized policy optimization. However, because this shift is learned within the weak teacher’s policy class, it need not improve verifier reward for the stronger student. OPRD therefore uses only its unit direction d𝑡 := Δ𝑡 /‖Δ𝑡 ‖2 for gradient scaling rather than optimizing toward the shift itself. The student gradient determines whether the resulting correction follows or opposes this direction, independently of its raw magnitude. 4
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Gradient Scaling along the Teacher Direction. At token 𝑡, OPRD decomposes the student’s policy gradient g𝑡 (Eq. 2.1) relative to the teacher direction d𝑡 . Let 𝑢𝑡 := d⊤ 𝑡 g𝑡 denote their alignment coefficient, and define the projected and orthogonal components as Projd𝑡 (g𝑡 ) := 𝑢𝑡 d𝑡 and g𝑡⊥ := g𝑡 − Projd𝑡 (g𝑡 ), respectively. At optimization step 𝑘, OPRD uses a nonnegative scale 𝜆𝑡 : ̃︀𝑡 : = g𝑡 + 𝜆𝑡 Projd𝑡 (g𝑡 ) g = (1 + 𝜆𝑡 ) Projd𝑡 (g𝑡 ) +
g𝑡⊥
amplified
unchanged
.
(2.5)
Essentially, OPRD decomposes the student’s policy gradient into its projection onto the teacher informed direction and an orthogonal component, amplifying only the projected component by 1 + 𝜆𝑡 , while leaving ̃︀𝑡 in place of g𝑡 and use the resulting parameter the orthogonal component unchanged. We backpropagate g gradients to update the student. Learning Beyond the Weak Teacher. Direct teacher matching makes the weak teacher’s policy a target of student optimization, even when moving beyond the teacher would yield higher verifier reward. OPRD instead uses the weak teacher only to rescale the student’s own policy gradient. At token 𝑡, this scaling can be written ̃︀𝑡 = (I + 𝜆𝑡 d𝑡 d⊤ as the linear map g 𝑡 )g𝑡 . For 𝜆𝑡 ≥ 0, the map is invertible and satisfies: (Stationarity) (Alignment Gain)
g𝑡 = 0.
(2.6)
̃︀𝑡 ⟩ = ‖g𝑡 ‖22 + 𝜆𝑡 𝑢2𝑡 ≥ ‖g𝑡 ‖22 . ⟨g𝑡 , g
(2.7)
̃︀𝑡 = 0 g
if and only if
Since Eq. 2.6 holds at every token, the scaling preserves the stationary points of the verifier objective for a fixed rẽ︀𝑡 retains the first-order progress of g𝑡 and adds the nonnegsponse. Eq. 2.7 shows that the transformed gradient g ative alignment gain 𝜆𝑡 𝑢2𝑡 , so greater alignment magnitude |𝑢𝑡 | yields greater first-order progress. For a fixed response, these token-level gains sum into a nonnegative term in the guaranteed one-step ascent (see Appendix A). OPRD can therefore accelerate the student’s optimization without introducing a teacher-defined target. Asymmetric Alignment Scaling. The sign of 𝑢𝑡 indicates whether g𝑡 aligns with or opposes d𝑡 , so scaling reinforces teacher-following updates when 𝑢𝑡 ≥ 0 and verifier-supported departures when 𝑢𝑡 < 0. However, both signals may include reward-irrelevant bias (e.g., g𝑡 = g𝑡⋆ + 𝜖𝑡 , with g𝑡⋆ denoting the reward-improving signal and 𝜖𝑡 aggregating structured bias components). Scaling only one sign can then systematically magnify this bias term, causing it to accumulate over training (see Appendix B.1 and B.2). Because negative-alignment updates are less reliable on initially weak student rollouts, we activate the positive branch immediately and gradually ramp up the negative branch. At optimization step 𝑘, we set the token-wise scaling coefficient as ⎧ ⎨𝜆, 𝑢𝑡 ≥ 0, {︁ }︁ 𝜆𝑡 := (2.8) 𝑘 ⎩𝜆 min 𝐾warm , 1 , 𝑢𝑡 < 0, where 𝜆 is the scaling strength and 𝐾warm is the warm-up horizon. Since 𝜆𝑡 ≥ 0, Eqs. 2.6 and 2.7 continue to hold under this schedule. The negative-branch warm-up gradually mitigates the initial one-sided amplification caused by positive-only scaling. This preserves immediate teacher-aligned transfer while progressively strengthening verifier-supported departures from the weak teacher. We further analyze this design alongside alternative branch-scaling strategies in Appendix B.3.
3. Experiments We evaluate OPRD on mathematical and logical reasoning tasks in two primary settings: weak-to-strong transfer across successive model transfer and multi-domain consolidation with multiple specialized teachers. We additionally evaluate conventional strong-to-weak distillation to verify that OPRD does not depend on a particular teacher–student size ordering.
5
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Table 1: Main experimental results for successive model transfer on mathematics and reasoning tasks. We use intermediate GRPO checkpoints of 4B-scale Qwen3 models as teachers and report Mean@16 for mathematics and Pass@1 for Reasoning Gym. Teacher and initial-student rows report single-checkpoint results, while trained-policy rows average evaluations at steps 30, 60, 90, 120, and 150 to summarize performance over training. The best result in each column is shown in bold. Detailed task-level curves and results with standard deviations are provided in Appendix D.1 and Appendix D.3, respectively. Math Reasoning Policy
AIME'24
AIME'25 HMMT'25 Olympiad
Reasoning Gym Avg.
Qwen3-4B (Teacher) → Qwen3-8B (Student)
Knights
Quantum
String
Count
Avg.
Qwen3-4B-Base (Teacher) → Qwen3-8B-Base (Student)
Teacher
42.50
38.75
21.25
52.15
38.66
57.50
46.58
32.00
42.50
44.65
Student + GRPO + OPD + KDRL1 + OPRD
25.63 46.63 46.92 53.46 66.92
19.58 36.42 38.21 43.13 56.04
12.50 21.96 21.42 26.08 31.42
46.22 52.52 51.24 53.28 53.26
25.98 39.38 39.44 43.99 51.91
11.00 54.20 55.00 53.30 73.30
3.14 36.24 39.62 36.63 51.11
3.00 35.50 35.10 38.10 42.20
3.00 41.30 41.60 49.50 54.10
5.04 41.81 42.83 44.38 55.18
3.1. Experimental Setup Tasks and Models. For mathematics, we train on DAPO-Math-17K (Yu et al., 2025) and evaluate on AIME'24, AIME'25 (MAA, 2024–2025), HMMT'25 (Dekoninck et al., 2026), and OlympiadBench (He et al., 2024). For diverse reasoning tasks, we train and evaluate on four Reasoning Gym benchmarks (Stojanovski et al., 2026): Knights & Knaves (K&K), Quantum Lock, String Manipulation, and Countdown. For each Reasoning Gym task, we construct a fixed pool of 20,000 examples, using 19,800 for training and holding out 200 for evaluation. All teacher and student models are drawn from the Qwen3 family (Qwen Team, 2025b), and the details are given in the corresponding setting descriptions. Each teacher is post-trained with GRPO (Shao et al., 2024) on the corresponding training data and held fixed during student training. Baselines. For the single-teacher experiments, we compare OPRD with GRPO (Shao et al., 2024), OPD (Agarwal et al., 2024), and KDRL (Xu et al., 2025). GRPO performs verifier-only policy optimization, OPD matches the frozen teacher on student-generated prefixes, and KDRL serves as a representative hybrid baseline that combines verifier-based policy optimization with on-policy distillation. For the multi-teacher setting, we similarly compare OPRD with Mix-RL, MOPD (Ma et al., 2026), and KDRL, which serve as the corresponding verifier-only, distillation-only, and hybrid baselines, respectively. Training and Evaluation. For each experiment, OPRD and all baselines start from the same student checkpoint and use the same task-specific training prompts, batch size, rollout budget, and number of policy updates. Teacher-based methods also use the same frozen teacher checkpoint for each task. We evaluate mathematics with Mean@16 and Reasoning Gym with Pass@1. While teacher and initial-student entries report fixed-checkpoint performance, trained-policy entries in most tables are averaged over five checkpoints to capture both learning speed and performance throughout training: at 30-update intervals for single-teacher settings and at 60-update intervals for multi-teacher distillation. Full training and evaluation details are provided in Appendix C. 3.2. Weak-to-Strong Distillation for a Successor Model Settings. To study successive model transfer in a controlled setting, we perform weak-to-strong distillation across scales within the same model family. Under the default configurations in Section 3.1, we pair a post-trained Qwen3-4B teacher with a Qwen3-8B student for math, and Qwen3-4B-Base teachers with separately trained Qwen3-8B-Base students for the four Reasoning Gym tasks. Additional Qwen3-Base results for mathematics and three other reasoning tasks are provided in Appendix D.2. OPRD Accelerates Learning While Continuing Beyond the Weak Teacher. Figure 1 (bottom-left) shows that OPD improves rapidly but plateaus near the teacher average, whereas GRPO progresses more gradually. OPRD
1We adopt KDRL (Xu et al., 2025) as the representative baseline combining distillation with RLVR (see also GKD (Agarwal et al., 2024) and dGRPO (Ramos et al., 2026)). For a fair comparison, the coefficient on the OPD objective is annealed to 0.0.
6
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Table 2: (Left) Experimental results for multi-teacher distillation on Reasoning Gym. We consolidate four task-specific Qwen3-4B-Base teachers into a single Qwen3-8B-Base student. The teacher and initial-student rows report fixed-checkpoint performance, while each trained-policy row averages Pass@1 over checkpoints at steps 60, 120, 180, 240, and 300. Detailed learning curves for each task are provided in Appendix E. (Right) Experimental results for strong-to-weak distillation. We evaluate Qwen3-8B → Qwen3-1.7B on AIME'24 and Qwen3-8B-Base → Qwen3-0.6B on Knights & Knaves. Each trained-policy row averages Mean@16 and Pass@1, respectively, over five checkpoints. Detailed learning curves are provided in Appendix F. The best result in each column is shown in bold. Policy
Knights Quantum
String
Count
Avg.
Policy
4 Teachers (Qwen3-4B-Base) → 1 Student (Qwen3-8B-Base)
AIME'24
Knights Knaves
8B → 1.7B
8B-Base → 0.6B
Avg.
Teachers
57.50
46.58
32.00
42.50
44.65
Teacher
52.29
64.50
58.40
Student + Mix-RL + MOPD + KDRL + OPRD
11.00 64.90 57.10 65.90 80.70
3.14 43.91 37.70 44.09 59.68
3.00 35.90 35.50 35.00 39.70
3.00 46.00 42.10 44.90 55.00
5.04 47.68 43.10 47.47 58.77
Student + GRPO + OPD + KDRL + OPRD
10.00 18.92 29.79 25.33 33.58
5.00 20.60 20.20 33.70 49.40
7.50 19.76 25.00 29.52 41.49
matches OPD’s initial acceleration, quickly surpasses the weak teacher, and reaches GRPO’s end-of-training performance substantially earlier. Averaged over five evenly spaced checkpoints to summarize the learning curve, Table 1 shows gains of 7.92 points on mathematics and 10.80 points on Reasoning Gym over the strongest baseline. The initially similar trajectories of OPD and OPRD indicate that teacher guidance is useful while the student still trails it, but their later divergence suggests that direct policy matching becomes restrictive once the student discovers reward-supported improvements beyond the teacher. 3.3. Multi-Teacher Weak-to-Strong Distillation Settings. Multi-teacher distillation asks whether capabilities acquired by separately post-trained task specialists can be consolidated into a single policy. We use the same Reasoning Gym configuration and method-specific settings as in Section 3.1, but each training batch now mixes the four tasks equally. Teacher-based methods pair each example with its corresponding Qwen3-4B-Base specialist. Because this reduces exposure to each task by roughly a factor of four, we train for 300 policy updates. Despite the longer run, we retain the single-teacher coefficient schedules rather than extending them to 300 updates. OPRD Consolidates Heterogeneous Specialists without Cross-Task Tradeoffs. Figure 1 (bottom-right) shows that MOPD rapidly approaches the specialist average but then plateaus, whereas Mix-RL improves more gradually. OPRD combines this early transfer with continued improvement throughout training. Table 2 reports an average of 58.77, exceeding Mix-RL by 11.09 points and the specialist average by 14.12 points. OPRD also surpasses the corresponding specialist on all four tasks despite their distinct structures and objectives, indicating joint improvement rather than a cross-task tradeoff. OPRD may reduce cross-task interference by amplifying only the component of each task’s verifier-driven student gradient along its teacher-shift direction, rather than matching the full specialist policy. This projection resembles PCGrad (Yu et al., 2020), but is applied between each task’s student gradient and teacher direction rather than between conflicting task gradients. 3.4. Strong-to-Weak Distillation Settings. We evaluate conventional strong-to-weak distillation under the default configurations in Section 3.1. For mathematics, we pair a Qwen3-8B teacher with a Qwen3-1.7B student, and we pair a Qwen3-8B-Base teacher with a Qwen3-0.6B student for Knights & Knaves. Both teachers are taken from step 105 of task-specific GRPO training. Both settings use prompt batch and mini-batch sizes of 64, and we schedule each method’s distillation coefficient over 45 updates. OPRD Does Not Depend on Teacher–Student Capacity Ordering. Table 2 shows that OPRD remains effective in conventional strong-to-weak distillation, outperforming OPD by 3.79 points on AIME'24 and 29.20 points on Knights & Knaves. Recent studies show that standard OPD can fail when capacity or distributional gaps make 7
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Table 3: (Left) Comparison with various weak-to-strong baselines. We evaluate 4B-to-8B transfer using instruction-tuned models for math and base models for reasoning tasks. Teacher and initial-student rows show fixed-checkpoint results, while trained-policy rows show checkpoint averages. The best result in each column is shown in bold. See Appendix G.1 and Appendix G.2 for implementation details and learning curves, respectively. (Right) Ablation Study on Teacher-Checkpoint Quality. For each task, the bars report the performance of Qwen3-4B-Base teacher checkpoints obtained at GRPO training steps 15, 60, 105, and 150, while the lines show the learning curves of Qwen3-8B-Base students trained with OPRD using the corresponding checkpoints. All configurations other than the teacher checkpoint follow the defaults in Section 3.1. AIME'24
Knights
String
Avg.
4B (-Base) → 8B (-Base) Teacher
42.50
57.50
32.00
44.00
Student + GRPO + OPD
25.63 46.63 46.92
11.00 54.20 55.00
3.00 35.50 35.10
13.21 45.44 45.67
+ W2SR-P + S2L-PO2 + OPSD3
45.42 60.54 15.88
62.50 63.40 24.60
38.00 38.70 24.60
48.64 54.21 21.69
+ Direct-OPD + W2S-OPD
35.04 52.96
49.20 59.50
29.20 35.20
37.81 49.22
+ OPRD
66.92
73.30
42.20
60.81
GRPO
OPRD (15)
OPRD (60)
OPRD (105)
OPRD (150)
Qwen3-4B-Base (Teacher) Qwen3-8B-Base (Student) Quantum Pass@1 Knights Pass@1
Policy
80 60 40 20 0 80 60 40 20 0
15
60
105 150
0
30 60 90 120 150
Training Steps
teacher supervision difficult to exploit, so a stronger teacher need not yield a better student (Li et al., 2026; Fu et al., 2026). Consistent with these findings, OPD improves initially in both settings but quickly saturates well below OPRD. KDRL’s lower score further suggests that supplementing policy matching with verifier feedback does not fully resolve this issue. OPRD instead amplifies only the component of the student’s verifier-driven gradient aligned with the teacher’s policy delta. This allows the student to benefit from teacher guidance along reward-supported directions it can realize, without having to reproduce the stronger policy in full.
4. Analysis 4.1. Broader Comparison with Weak-to-Strong Methods OPRD Outperforms Methods Using Off-Policy Generations from Weak Teacher. Table 3 (left) compares OPRD with three baselines that use weak-teacher generations differently. W2SR-P (Yuan et al., 2026) performs SFT on verified-correct teacher trajectories; S2L-PO (Ren et al., 2026) mixes off-policy rollouts from a weak explorer with student rollouts in shared GRPO groups before transitioning to fully on-policy RLVR; and our OPSD variant (Zhao et al., 2026) uses a verified weak-teacher draft as privileged context for self-distillation. W2SR-P and S2L-PO improve over the initial student, indicating that weak-teacher trajectories can provide a useful bootstrap within the same Qwen3 family. However, reliance on off-policy teacher trajectories can create train–inference mismatch (Agarwal et al., 2024) and need not transfer underlying capabilities across model gaps (Gudibande et al., 2023). OPRD instead remains fully on-policy and uses the teacher shift only to rescale the aligned component of the verifier gradient. Empirically, OPRD reaches 60.81, exceeding the strongest alternative, S2L-PO, by 6.60 points and achieving the best score on all three tasks. Rescaling the Verifier Gradient Outperforms Direct Optimization of the Weak Policy Delta. The lower rows of Table 3 (left) compare OPRD with two closely related concurrent works, Direct-OPD (Feng et al., 2026) and W2S-OPD (Yu et al., 2026), both of which derive the student’s objective directly from the weak policy shift. Direct-OPD uses the corresponding log-ratio as a dense reward, whereas W2S-OPD reanchors the shift
2 S2L-PO (Ren et al., 2026) originally uses a smaller base model as the frozen explorer, reflecting the method’s central motivation to exploit policy-level diversity. Here, we instead use the post-RL weak-teacher checkpoint as the explorer. 3 OPSD (Zhao et al., 2026) and SDPO (Hübotter et al., 2026) use correct self-generated rollouts as privileged information. Here, we instead use correct trajectories generated by the weak teacher, while the KL divergence remains computed against the self-teacher.
8
60 GRPO OPRD (OPSD) OPRD (Delta) OPRD (OPD)
40 20 0
0
30
60
90
120
150
110
80
100
60 λ = 0.0 λ = 0.05 λ = 0.1 λ = 0.5 λ = 1.0
40 20 0
0
30
Training Steps
(a) Construction of d𝑡
Angle θ(dt, gt) (°)
80
Knights Pass@1 (%)
Knights Pass@1 (%)
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
60
90
120
Training Steps
(b) Amplification strength 𝜆
150
90 80 70
Aligned (ut ≥ 0) Opposed (ut < 0)
0
30
60
90
120
150
Training Steps
(c) Gradient alignment
Figure 3: (a) Ablations of scaling-direction construction. All OPRD variants use the step-60 GRPO checkpoint as the weak teacher. For OPSD, a verified draft generated by this teacher is provided as privileged context. (b) Ablation of directional amplification strength. We vary 𝜆, the coefficient applied to OPRD’s directional correction term. 𝜆 = 0 corresponds to GRPO. The gray dashed line marks the performance of the teacher checkpoint. All other settings follow the default configurations in Section 3.1. (c) Alignment dynamics between d𝑡 and g𝑡 . On Knights & Knaves, we track 𝜃𝑡 between d𝑡 and g𝑡 during OPRD with a Qwen3-8B-Base student and Qwen3-4B-Base weak teacher. Excluding rollout groups with g𝑡 = 0 (identical rewards within the group), we report token-averaged angles for aligned (𝑢𝑡 ≥ 0) and opposed (𝑢𝑡 < 0) tokens.
to the student’s base policy and distills the resulting proxy teacher. Given the sensitivity of both methods to the relative strength of the transferred shift, we follow the hyperparameter settings reported in the original papers. However, both methods rely solely on the information encoded in the shift. OPRD instead retains verifier supervision through the orthogonal component g𝑡⊥ , allowing the student to pursue reward-supported directions not captured by the weak policy delta. Indeed, OPRD reaches 60.81, outperforming W2S-OPD by 11.59 points and Direct-OPD by 23.00 points on average. 4.2. Design and Dynamics of Teacher Guidance Better-Trained Weak Teachers Provide More Effective Guidance. Table 3 (right) shows that later, betterperforming GRPO checkpoints of the 4B teacher generally lead to faster learning under OPRD for the 8B student on both reasoning tasks. Because Δ𝑡 is normalized before scaling, this benefit cannot be attributed to shift magnitude alone; instead, later checkpoints appear to encode a more reward-informative direction, yielding a larger verifier-gradient component for OPRD to amplify. Notably, the step-60 teacher achieves only 29.0% Pass@1 on Knights & Knaves, yet the corresponding OPRD student rapidly reaches approximately 88%, far surpassing both the teacher and GRPO. Thus, while teacher quality affects the strength of OPRD’s acceleration, the teacher’s absolute performance need not impose a ceiling on the student. Weak Policy Delta Provides the Most Effective Scaling Direction. Figure 3a compares three choices for the guidance direction d𝑡 : the normalized weak policy delta Δ𝑡 , the OPD teacher-matching gradient, and the OPSD self-distillation gradient. The weak policy delta yields the fastest and most sustained gains. Comparing the post-trained teacher with its reference isolates the reward-relevant update, and their log-policy ratio admits an implicit-reward interpretation. OPRD projects the verifier gradient onto this direction and amplifies its aligned component, exploiting the teacher’s reward information without inheriting its capacity ceiling. In contrast, OPD captures the full teacher–student mismatch and offers limited acceleration when the teacher is too weak, while OPSD’s off-policy supervision can restrict exploration of alternative reasoning paths (Kim et al., 2026; Kaur et al., 2026). Although OPD becomes more effective with a better-trained teacher, the weak policy delta is still the fastest and most reliable guidance signal (see Appendix H). Sufficient Directional Amplification Enables Early Acceleration. Figure 3b examines 𝜆, which scales the directional correction and thus controls the strength of teacher guidance. Every 𝜆 > 0 improves final Pass@1 over 𝜆 = 0 (GRPO). Larger values of 𝜆 up to 0.5 also yield faster gains early in training. This systematic relationship between guidance strength and learning speed confirms that OPRD’s directional correction indeed drives the observed acceleration. Performance changes little beyond 𝜆 = 0.5, so precise tuning is unnecessary once amplification is sufficiently strong. We therefore use 𝜆 = 0.5 as the default.
9
GRPO OPD KDRL OPRD
40 20 0
0
30
60
90
120
150
Training Steps
(a) Limited gradient signal
GRPO OPD KDRL OPRD (Ref. 0) OPRD (Ref. 30)
80 60
Response Length
60
Color Cube Pass@1 (%)
Knights Pass@1 (%)
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
40 20
0
30
60
90
120
150
Teacher GRPO Student GRPO OPRD (Ref. 0 OPRD (Ref. 30
1500
) )
1000 500 0
0
30
60
90
120
Training Steps
Training Steps
(b) Task performance
(c) Response Length
150
Figure 4: (a) Results under limited policy-gradient signal. On Knights & Knaves, we transfer a step-105 Qwen3-8B-Base teacher to a Qwen3-1.7B-Base student, with OPRD’s negative-branch scale 𝜆𝑡 warmed up over the first 75 steps. (b, c) Effect of reference-policy selection on length bias. On Color Cube, we transfer a step-105 Qwen3-4B-Base teacher 𝜋𝑇 to a Qwen38B-Base student, using either the step-0 or step-30 checkpoint from the same GRPO run as 𝜋𝑇ref . OPRD’s negative-branch scale 𝜆𝑡 is warmed up over the first 75 steps. The gray dashed line marks teacher performance, while the colored stars denote the mean response lengths of the two choices of 𝜋𝑇r𝑒𝑓 , and the white star marks that of 𝜋𝑇 . All other settings follow Section 3.1.
Teacher Guidance Bootstraps Early Learning but Becomes Less Influential over Time. Figure 3c tracks the mean angle 𝜃𝑡 between the guidance direction d𝑡 and policy gradient g𝑡 . Early in training, the two directions exhibit substantial alignment for 𝑢𝑡 > 0 and opposition for 𝑢𝑡 < 0. This strong directional coupling allows the weak policy delta to bootstrap student learning. As training proceeds, the mean angle for 𝑢𝑡 > 0 increases toward 90∘ , while that for 𝑢𝑡 < 0 decreases toward 90∘ . Since ‖Projd𝑡 (g𝑡 )‖2 /‖g𝑡 ‖2 = | cos 𝜃𝑡 |, this convergence toward orthogonality means that the component of g𝑡 along d𝑡 becomes smaller relative to the full policy gradient. This indicates that the evolving student gradient increasingly follows verifier-supported directions not captured by the teacher shift, so teacher guidance becomes less influential over time. 4.3. Discussion of Key Challenges Vanishing Policy Gradients Limit OPRD’s Teacher-Guided Correction. OPRD requires a nonzero verifierdriven policy gradient. In an additional strong-to-weak experiment pairing a Qwen3-8B-Base teacher with a Qwen3-1.7B-Base student, most Knights & Knaves responses are invalid, so most rollout groups receive identical rewards (i.e., the resulting group-relative advantages and their contributions to g𝑡 therefore vanish). For these groups, the projection onto d𝑡 also vanishes, leaving no component for OPRD to amplify and hence no teacherguided correction. As shown in Figure 4a, OPRD still accelerates learning relative to GRPO and KDRL, although all three remain below 20% Pass@1, whereas OPD reaches 40.5% using dense policy-matching targets that do not depend on verifier rewards. This challenge arises from the student’s initial rollout distribution rather than the absence of a useful teacher signal. Such an extreme regime is less likely in our primary weak-to-strong setting, where the student has greater capacity than the teacher, but may still arise on sufficiently difficult tasks. A short task-specific SFT or distillation warm-up could bootstrap valid on-policy behavior before switching to OPRD. Reference Policy Selection Can Prevent Length Bias from Distorting Teacher Guidance. As discussed in Section 2.2 and Appendix B, both g𝑡 and d𝑡 can contain reward-irrelevant components such as 𝜖𝑡 , which the projection-and-amplification step can magnify. Response length is one example: when it correlates with verifier reward, both signals can encode a preference for longer or shorter responses, even if changing length does not itself improve reasoning quality. As shown in Figures 4b and 4c, the step-0 reference 𝜋𝑇ref produces substantially longer responses than the step-105 teacher 𝜋𝑇 on Color Cube. The resulting shift Δ𝑡 therefore contains a strong shortening component. With this reference, OPRD rapidly shortens its responses and achieves strong early gains. It nevertheless plateaus at 52.5% Pass@1, below GRPO and KDRL, suggesting that the teacher-guided correction overemphasizes shortening at the expense of task-relevant reasoning. A simple mitigation is to move the reference to step 30, after the teacher’s initial length collapse. This excludes some of the teacher’s early gains from Δ𝑡 but substantially narrows the reference–teacher length gap and weakens the associated bias. OPRD then avoids the plateau and jumps to 89.5%, discovering a more effective reasoning strategy. Appendix I shows the same pattern on Binary Matrix, where this reference policy adjustment is likewise effective.
10
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Connectives -1.00 -0.69 +0.73 +1.00
+1
Modality -1.00 -0.85 +0.42 +1.00 Grammar -1.00 -0.42 +0.53 +1.00
0
Punctuation -1.00 -0.56 +0.39 +1.00 Structure -1.00 -0.78 +0.87 +1.00 t r he OPD PRD den ac u t O e S T
(a) Token alignment and reasoning paths
−1
(b) Response style similarity
Figure 5: (a) Visualizing token alignment and reasoning continuations. An AIME'25 response from the OPRD student at update 150. Green and red indicate positive and negative cosine similarity between d𝑡 and the student policy gradient g𝑡 with 𝐴𝑡 = 1, respectively (see Appendix J.1). The token outlined in black, 0 , has the lowest cosine similarity among displayed tokens. The student’s top-1 token 0 completes 2016 directly. Forcing 5 , the top-1 token under d𝑡 , leads the same student to this result through an intermediate sum. The plots show the student’s top-10 token probabilities above and their teacher-shift values (d𝑡 ) below. (b) Measuring similarity to teacher and student response styles. On AIME'24, we compare response styles using 101 standardized features across five categories. Normalized distance differences indicate whether each method’s average style is closer to the weak teacher (red) or the GRPO-trained student at update 150 (green).
4.4. Student Behavior under Teacher Guidance OPRD Can Move Beyond the Teacher’s Reasoning Paths. Figure 5a illustrates how OPRD can exploit an informative teacher shift while allowing the stronger student to follow its own, more direct reasoning path rather than the one favored by the teacher. At the selected AIME'25 prefix, the student’s top-1 prediction is 0 , which immediately completes 2016. By contrast, the top-1 token under the weak policy shift d𝑡 is 5 . Forcing 5 and continuing with the same student produces 252 + 504 = 756, followed by 756 + 1260 = 2016. This detour also reaches the correct result, showing that the teacher shift provides a valid direction that may be useful earlier in training. Here, however, the student can already complete the calculation directly. This is reflected in the highlighted 0 , which has the most negative alignment with d𝑡 among the displayed tokens. OPRD therefore raises the logit of 0 and lowers that of 5 (when 𝑢𝑡 < 0, OPRD amplifies the component of the verifier-driven policy gradient that opposes the teacher shift). The negative-alignment branch thus favors the student’s shorter solution over the valid teacher-favored detour, providing a token-level example of how OPRD can move beyond the teacher. OPRD Remains Stylistically Closer to the Stronger Student. Figure 5b examines how teacher guidance affects response style on AIME'24. We summarize each method’s average response style using 101 standardized features grouped into five categories: connectives, modality, grammar, punctuation, and sentence and paragraph structure. For each category, a normalized distance difference indicates whether the average style is closer to the teacher (negative) or the GRPO student at update 150 (positive) (see Appendix J.2 for details). At update 150, OPD is closer to the teacher in all five categories, whereas OPRD is closer to the GRPO student. This pattern suggests that the OPRD student can benefit from what the teacher learned without inheriting its response style, consistent with using the teacher shift to rescale the student’s own policy gradient rather than matching the teacher policy.
5. Related Work Weak-to-Strong Generalization. Weak-to-strong generalization has been observed across language understanding, reward modeling, and reasoning, although weak supervision typically recovers only part of the gap to strong supervision (Burns et al., 2024; Yang et al., 2024). Analyses attribute the gains to correcting weak pseudo-labels, extending coverage beyond the weak teacher, and differences between teacher and student hypothesis classes or representations (Lang et al., 2024; Charikar et al., 2024; Dong et al., 2025; Xue et al., 2025; Medvedev et al., 2025), while naive fine-tuning can instead overfit weak errors (Somerstep et al., 2025; 11
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Yao et al., 2025; Shi et al., 2025). For reasoning, W2SR-P trains stronger students on verified weak-model trajectories, S2L-PO and related methods use weaker policies to broaden the student’s rollouts, and weak critiques can generate and filter improved responses (Yuan et al., 2026; Ren et al., 2026; Wang et al., 2026a; Jin et al., 2026). These methods change the student’s training data, exploration, or feedback, whereas OPRD leaves all three unchanged and only rescales the student’s own policy gradient. On-Policy Distillation. Knowledge distillation for language generation has moved from matching teacher distributions on fixed or teacher-generated sequences (Hinton et al., 2015; Kim and Rush, 2016) to objectives evaluated on student-generated sequences (Gu et al., 2024; Ko et al., 2024). OPD makes this supervision fully on-policy by querying the teacher along the student’s current rollouts, addressing the mismatch between the prefixes seen in training and those the student visits at inference (Agarwal et al., 2024). It is now a common step in reasoning post-training (Qwen Team, 2025b; GLM-5 Team, 2026), and later work uses the same interface to consolidate several specialist teachers into one student (Ma et al., 2026; Kimi Team, 2026; Xiaomi Team, 2026), to exploit privileged information available only during training (Zhao et al., 2026; Ye et al., 2026), or to extrapolate the reward implicit in OPD beyond the teacher (Yang et al., 2026a). In all of these, the student is still trained to match a token distribution that the teacher defines, so in the weak-to-strong setting the optimum of the objective is the weak policy itself or a target derived from it. Distillation with Reinforcement Learning. Methods that combine distillation with verifier-based reinforcement learning differ in how the teacher signal enters optimization. KDRL and later work add a teacher-matching term to the reward objective (Xu et al., 2025; Ramos et al., 2026). Others modify teacher guidance through policy ratios, reward-based selection, group-level calibration, or token-level interventions (Zhang et al., 2026a; Akhondzadeh et al., 2026; Zhang et al., 2026b; Ko et al., 2026; Jia et al., 2026), and another uses a privileged self-teacher to control the magnitude of token-level credit (Wang et al., 2026b). However, a teacher-matching loss introduces a second objective that can compete with reward maximization when the teacher favors a solution the verifier does not reward. OPRD adds no such objective and optimizes reward alone. Transferring Policy Shifts. Several methods transfer the shift between a post-trained policy and its reference rather than the final policy alone. During decoding, this shift can steer a larger frozen model (Liu et al., 2024; Zhou et al., 2024). During training, it has been used as an alignment target for a stronger model (Zhu et al., 2025), as a proxy teacher built on the student’s base policy in W2S-OPD (Yu et al., 2026), and as a dense reward on student rollouts in Direct-OPD (Heo et al., 2026; Feng et al., 2026). In each case the shift itself becomes an optimization target, and it carries only the improvements the weak teacher realized. OPRD instead uses the weak policy delta to rescale the student’s policy gradient, so the direction transfers without the shift becoming a target. Gradient Manipulation. Multi-task optimization combines objectives at the level of gradients rather than losses. Gradient surgery projects one task gradient onto the normal plane of another when the two conflict (Yu et al., 2020), a moving average of past gradients makes this projection more stable (Hsieh et al., 2024), and auxiliary gradients can be gated by their cosine similarity with the main gradient (Du et al., 2019; Zhou et al., 2022). In these methods, every direction is the gradient of a loss the model itself optimizes, and conflicting components are removed or down-weighted. OPRD is closest to this family in form, but it removes nothing and only amplifies the component of the student’s policy gradient that already points along the teacher direction, so the stationary points of the student objective in logit space do not move.
6. Conclusion We introduce On-Policy Reverse Distillation (OPRD), which transfers a weak teacher’s post-training policy shift by amplifying the aligned component of a stronger student’s policy gradient without making the teacher policy an optimization target. By rescaling rather than replacing the student gradient, OPRD accelerates learning while preserving the policy objective’s stationary points in logit space. Across successive model transfer and multi-domain consolidation, OPRD reaches teacher-level performance in substantially fewer updates than policy optimization alone and continues improving after on-policy distillation plateaus near the teacher. Its gains extend to strong-to-weak distillation, showing effectiveness under both capacity orderings. Qualitatively, the OPRD student’s response style remains closer to the reward-only baseline than to the teacher, consistent with the shift being expressed through the student’s own policy rather than imitation. OPRD thus enables efficient transfer from smaller specialists without defining the student’s optimization target or limiting its performance. 12
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
6.1. Future Works Broader Tasks and Settings. Mathematical and logical reasoning offer controlled settings in which verifier feedback and policy improvement can be measured directly. Broader evaluations should test whether OPRD continues to transfer useful policy shifts under different forms of feedback and interaction. Code generation and agentic environments are particularly informative because feedback arises from program execution or environmental responses, and early actions influence subsequent observations and rewards. These settings would clarify how broadly policy changes learned through post-training can be transferred between models. Scaling to Larger Models. Our results cover Qwen3 models from 0.6B to 8B parameters and both weakto-strong and strong-to-weak capacity orderings. At larger scales, OPRD may be especially useful because learning a policy shift with a smaller model could be substantially cheaper than optimizing the larger model directly from verifier feedback. Larger-scale experiments would test how students with greater capacity use the same teacher shift, how far they can improve beyond the teacher, and whether the gains in update efficiency persist as post-training costs increase. Systems Considerations at Scale. OPRD adds no-gradient forward passes through the frozen teacher and reference policies and a correction of the student’s logit gradient. In weak-to-strong setup, both frozen policies are smaller than the student and require neither generation nor backward propagation, while the correction retains one additional dense logit-gradient tensor. As shown in Appendix K, these additions only increase wall-clock time by 11.9% and peak GPU memory by 10.2% relative to GRPO. But at frontier-model scale, keeping this overhead modest will require efficient placement, sharding, and scheduling of the frozen policies within hybrid parallelism, together with communication-efficient correction across model and vocabulary shards. Toward Recursive Self-Improvement. An important direction for future work is to connect weak-to-strong distillation with recursive self-improvement, where each model generation contributes to the development of more capable successors through training supervision, evaluation, and algorithmic improvements. These successors, in turn, use their greater capabilities to improve subsequent model development. For example, earlier models helped supervise GPT-6 Astra’s training (OpenAI, 2026), while Google reports using agentic loops to recursively evaluate and refine Gemini 3.8 Flash (Gemini Team, 2026). A promising extension is to incorporate reverse distillation into these workflows, allowing earlier generations to contribute not only to training supervision and development but also directly to their successors’ policy updates through their posttraining policy shifts. Building such pipelines would allow us to test whether reverse distillation can consistently improve sample efficiency and accelerate training as each successor becomes a teacher for the next generation.
Acknowledgements We thank Kee-Eung Kim for facilitating access to computational resources through the National AI Research Hub project. We thank Rishabh Agarwal for helpful discussions on related work and algorithm design. We also thank Reza Bayat for feedback on the manuscript.
13
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
References Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In International Conference on Learning Representations, volume 2024, pages 21246–21263, 2024. Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi, Chanakya Ekbote, Sujay Sanghavi, and Aleksandar Bojchevski. Reward-gated on-policy distillation. arXiv preprint arXiv:2607.04037, 2026. Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=ghNRg2mEgN. Moses Charikar, Chirag Pabbaraju, and Kirankumar Shiragur. Quantifying the gain in weak-to-strong generalization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=MyVyH5Jo1l. DeepSeek-AI. Deepseek-v4: Towards highly efficient million-token context intelligence. arXiv preprint arXiv:2606.19348, 2026. Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger, Kári Rögnvaldsson, Ivo Petrov, Chenhao Sun, and Martin Vechev. Beyond benchmarks: Matharena as an evaluation platform for mathematics with llms. arXiv preprint arXiv:2605.00674, 2026. Yijun Dong, Yicheng Li, Yunai Li, Jason D. Lee, and Qi Lei. Discrepancies are virtue: Weak-to-strong generalization through lens of intrinsic dimension. In Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu, editors, Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 14079–14113. PMLR, 13–19 Jul 2025. URL https://proceedings.mlr.press/v267/dong25g.html. Yunshu Du, Wojciech M. Czarnecki, Siddhant M. Jayakumar, Razvan Pascanu, and Balaji Lakshminarayanan. Adapting auxiliary losses using gradient similarity, 2019. URL https://openreview.net/forum?id= r1gl7hC5Km. Shiyuan Feng, Huan-ang Gao, Haohan Chi, Hanlin Wu, Zhilong Zhang, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, and Hao Zhou. Weak-to-strong generalization via direct on-policy distillation. arXiv preprint arXiv:2607.05394, 2026. Yuqian Fu, Haohuan Huang, Kaiwen Jiang, Jiacai Liu, Zhuo Jiang, Yuanheng Zhu, and Dongbin Zhao. Revisiting on-policy distillation: Empirical failure modes and simple fixes. arXiv preprint arXiv:2603.25562, 2026. Gemini Team. Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, September 2026. URL https://blog.google/ innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/. GLM-5 Team. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026. Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. MiniLLM: Knowledge distillation of large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/ forum?id=5h0qf7IBZZ. Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717, 2023. Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, et al. Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3828–3850, 2024. 14
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Byeongho Heo, Jaehui Hwang, Sangdoo Yun, and Dongyoon Han. On-policy delta distillation. arXiv preprint arXiv:2607.15161, 2026. Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. Yu-Guan Hsieh, James Thornton, Eugene Ndiaye, Michal Klein, Marco Cuturi, and Pierre Ablin. Careful with that scalpel: Improving gradient surgery with an EMA. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 19085–19100. PMLR, 21–27 Jul 2024. URL https://proceedings.mlr.press/v235/hsieh24a.html. Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, et al. Reinforcement learning via self-distillation. arXiv preprint arXiv:2601.20802, 2026. Nan Jia, Haojin Yang, Xing Ma, Jiesong Lian, Shuailiang Zhang, Weipeng Zhang, Ke Zeng, Xunliang Cai, and Zequn Sun. Asymmetric on-policy distillation: Bridging exploitation and imitation at the token level. arXiv preprint arXiv:2605.06387, 2026. Can Jin, Tristan J. Li, Rui Wu, Eddy Z. Zhang, and Dimitris N. Metaxas. Weak critics make strong learners: On-policy critique distillation for scalable oversight. In 3rd AI for Math Workshop: Toward Self-Evolving Scientific Agents, 2026. URL https://openreview.net/forum?id=oEfedgUChS. Simran Kaur, Narutatsu Ri, Yinghui He, Liam Fowl, and Sanjeev Arora. Rethinking on-policy self-distillation for thinking models. arXiv preprint arXiv:2607.05184, 2026. Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, and Yuqing Yang. Why does self-distillation (sometimes) degrade the reasoning capability of llms? arXiv preprint arXiv:2603.24472, 2026. Yoon Kim and Alexander M. Rush. Sequence-level knowledge distillation. In Jian Su, Kevin Duh, and Xavier Carreras, editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1317–1327, Austin, Texas, November 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1139. URL https://aclanthology.org/D16-1139/. Kimi Team. Kimi k3: Open frontier intelligence. arXiv preprint arXiv:2607.24653, 2026. Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se-Young Yun. DistiLLM: Towards streamlined distillation for large language models. In Forty-first International Conference on Machine Learning, 2024. URL https: //openreview.net/forum?id=lsHZNNoC7r. Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen, and Pashmina Cameron. Scaling reasoning efficiently via relaxed on-policy distillation. arXiv preprint arXiv:2603.11137, 2026. Hunter Lang, David Sontag, and Aravindan Vijayaraghavan. Theoretical analysis of weak-to-strong generalization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=HOSh0SKklE. Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, et al. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe. arXiv preprint arXiv:2604.13016, 2026. Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith. Tuning language models by proxy. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum? id=dribhnhm1i. Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. Understanding r1-zero-like training: A critical perspective. In Second Conference on Language Modeling, 2025. URL https://openreview.net/forum?id=5PAF7PAY2Y.
15
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, et al. Mopd: Multi-teacher on-policy distillation for capability integration in llm post-training. arXiv preprint arXiv:2606.30406, 2026. MAA. American invitational mathematics examination (AIME), 2024–2025, 2024–2025. URL https://maa. org/maa-invitational-competitions/. Marko Medvedev, Kaifeng Lyu, Dingli Yu, Sanjeev Arora, Zhiyuan Li, and Nathan Srebro. Weak-to-strong generalization even in random feature networks, provably. In Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu, editors, Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 43519–43556. PMLR, 13–19 Jul 2025. URL https://proceedings.mlr.press/v267/medvedev25a. html. OpenAI. GPT-6 Astra: A new generation of intelligence, 2026. URL https://openai.com/index/gpt-6-astra/. Qwen Team. Qwen2.5 technical report, 2025a. URL https://arxiv.org/abs/2412.15115. Qwen Team. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025b. Miguel Moura Ramos, Duarte M Alves, and André FT Martins. A recipe for long-context reasoning in large language models via on-policy optimization and distillation. arXiv preprint arXiv:2605.12227, 2026. Yiming Ren, Yiran Xu, Zicheng Lin, Chufan Shi, Yukang Chen, Dingdong WANG, Tianhe Wu, Junjie Wang, Yujiu Yang, Yu Qiao, and Ruihang Chu. Smaller models are natural explorers for policy-level diversity in GRPO. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/ forum?id=PI2xku6EDA. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng, Qipeng Guo, and Xipeng Qiu. How to mitigate overfitting in weak-to-strong generalization? In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16100–16118, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.784. URL https: //aclanthology.org/2025.acl-long.784/. Seamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov, Mikhail Yurochkin, and Yuekai Sun. A transfer learning framework for weak to strong generalization. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=PeLLMw3wLX. Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Richard Jones, Abdulhakeem Adefioye, Jean Kaddour, and Andreas Köpf. Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2026. URL https://openreview.net/forum?id=GqYSunGmp7. Dayu Wang, Jiaye Yang, Weikang Li, Jiahui Liang, Liwei Qian, Xin Pei, and Jizhou Huang. It takes 8 tokens: Weak-to-strong off-policy rl via auxiliary branches. arXiv preprint arXiv:2607.16205, 2026a. Zechuan Wang, Siyuan Lu, Hongxuan Zhang, Linjian Mo, Chenyi Zhuang, and Leilei Gan. Teach the magnitude, not the direction: Verifier-bounded credit assignment for multi-turn multi-step llm agents. arXiv preprint arXiv:2608.13179, 2026b. Xiaomi Team. Mimo-v2-flash technical report. arXiv preprint arXiv:2601.02780, 2026. Hongling Xu, Qi Zhu, Heyuan Deng, Jinpeng Li, Lu Hou, Yasheng Wang, Lifeng Shang, Ruifeng Xu, and Fei Mi. Kdrl: Post-training reasoning llms via unified knowledge distillation and reinforcement learning. arXiv preprint arXiv:2506.02208, 2025.
16
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Yihao Xue, Jiping Li, and Baharan Mirzasoleiman. Representations shape weak-to-strong generalization: Theoretical insights and empirical predictions. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=ypEW077kle. Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, and Yankai Lin. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation. arXiv preprint arXiv:2602.12125, 2026a. Yuqing Yang, Yan Ma, and Pengfei Liu. Weak-to-strong reasoning. In Findings of the association for computational linguistics: EMNLP 2024, pages 8350–8367, 2024. Zhuolin Yang, Zihan Liu, Yang Chen, Wenliang Dai, Boxin Wang, Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He, et al. Nemotron-cascade 2: Post-training llms with cascade rl and multi-domain on-policy distillation. arXiv preprint arXiv:2603.19220, 2026b. Wei Yao, Wenkai Yang, Ziqiao Wang, Yankai Lin, and Yong Liu. Revisiting weak-to-strong generalization in theory and practice: Reverse KL vs. forward KL. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 2860–2888, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/2025.findings-acl.148. URL https://aclanthology.org/2025.findings-acl.148/. Tianzhu Ye, Li Dong, Xun Wu, Shaohan Huang, and Furu Wei. On-policy context distillation for language models. arXiv preprint arXiv:2602.12275, 2026. Fangxu Yu, Zinan Lin, Xiaodong Liu, Weijia Xu, Michael Xu, Tianyi Zhou, and Jianfeng Gao. Weak-to-strong on-policy distillation. arXiv preprint arXiv:2607.26246, 2026. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, YuYue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, LingJun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Ru Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Yuxuan Song, Xiangpeng Wei, Hao Zhou, Jingjing Liu, Wei-Ying Ma, Ya-Qin Zhang, Lin Yan, Yonghui Wu, and Mingxuan Wang. DAPO: An open-source LLM reinforcement learning system at scale. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=2a36EMSSTp. Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. Advances in neural information processing systems, 33:5824–5836, 2020. Yige Yuan, Teng Xiao, Shuchang Tao, Xue Wang, Jinyang Gao, Bolin Ding, and Bingbing Xu. Incentivizing strong reasoning from weak supervision. In Vera Demberg, Kentaro Inui, and Lluís Marquez, editors, Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7138–7156, Rabat, Morocco, March 2026. Association for Computational Linguistics. ISBN 979-8-89176-380-7. doi: 10.18653/v1/2026.eacl-long.336. URL https://aclanthology.org/2026. eacl-long.336/. Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo Yang, Zhuowen Tu, Wei Xia, and Stefano Soatto. Reinforcement-aware knowledge distillation for llm reasoning. arXiv preprint arXiv:2602.22495, 2026a. Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, and Yuezhi Zhou. Beyond teacher likelihood: Group-calibrated on-policy distillation for long-context reasoning. arXiv preprint arXiv:2608.19181, 2026b. Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, and Aditya Grover. Self-distilled reasoner: On-policy self-distillation for large language models. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id=Jpxfof0EaS. Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al. Pytorch fsdp: experiences on scaling fully sharded data parallel. arXiv preprint arXiv:2304.11277, 2023.
17
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Shiji Zhou, Wenpeng Zhang, Jiyan Jiang, Wenliang Zhong, Jinjie GU, and Wenwu Zhu. On the convergence of stochastic multi-objective gradient manipulation and beyond. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=ScwfQ7hdwyP. Yucheng Zhou, Jianbing Shen, and Yu Cheng. Weak to strong generalization for large language models with multi-capabilities. In International Conference on Learning Representations, volume 2025, pages 11583–11612, 2025. Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao. Weak-to-strong search: Align large language models via searching over small language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=dOJ6CqWDf1. Wenhong Zhu, Zhiwei He, Xiaofeng Wang, Pengfei Liu, and Rui Wang. Weak-to-strong preference optimization: Stealing reward from weak aligned model. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=f7KxfUrRSb. Zhou Ziheng, Jiaqi Li, Huacong Tang, Ying Nian Wu, and Demetri Terzopoulos. Less is more: Early stopping rollout for on-policy distillation. arXiv preprint arXiv:2605.27028, 2026.
18
Contents 1 Introduction
1
2 Method
3
2.1 Preliminary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3
2.2 On-Policy Reverse Distillation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
3 Experiments
5
3.1 Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
6
3.2 Weak-to-Strong Distillation for a Successor Model . . . . . . . . . . . . . . . . . . . . . . . . .
6
3.3 Multi-Teacher Weak-to-Strong Distillation . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
3.4 Strong-to-Weak Distillation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
4 Analysis
8
4.1 Broader Comparison with Weak-to-Strong Methods . . . . . . . . . . . . . . . . . . . . . . . .
8
4.2 Design and Dynamics of Teacher Guidance . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
9
4.3 Discussion of Key Challenges . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
10
4.4 Student Behavior under Teacher Guidance . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
11
5 Related Work
11
6 Conclusion
12
6.1 Future Works . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
13
A Optimization Properties of Teacher-Direction Scaling
21
B Analysis of Asymmetric Alignment Scaling
22
B.1 One-Sided Amplification under Positive-Only Scaling . . . . . . . . . . . . . . . . . . . . . . .
22
B.2 Isolating the Positive and Negative Alignment Branches . . . . . . . . . . . . . . . . . . . . . .
23
B.3 Mitigating One-Sided Amplification through Branch Scheduling . . . . . . . . . . . . . . . . .
24
C Training and Evaluation Details
25
D Detailed Results for Weak-to-Strong Model Transfer
27
D.1 Detailed Learning Curves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
27
D.2 Additional Results Across Model Variants and Tasks . . . . . . . . . . . . . . . . . . . . . . . .
28
D.3 Evaluation Results with Standard Deviations . . . . . . . . . . . . . . . . . . . . . . . . . . . .
29
E Detailed Results for Multi-Teacher Weak-to-Strong Distillation
30
F Detailed Results for Strong-to-Weak Distillation
30
G Detailed Results for Weak-to-Strong Method Comparisons
31
G.1 Baseline Implementation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
19
31
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
G.2 Detailed Learning Curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
32
H Additional Results on Guidance-Direction Construction
33
I
Additional Results on Length Bias in Teacher Policy Shift
34
J
Detailed Analysis of Student Behavior
35
J.1
Token Alignment Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
35
J.2
Response Style Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
35
K Computational Cost and Memory Usage
37
20
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
A. Optimization Properties of Teacher-Direction Scaling OPRD multiplies the token-level policy gradient by I + 𝜆𝑡 d𝑡 d⊤ 𝑡 , which amplifies the component along d𝑡 by 1 + 𝜆𝑡 and leaves the orthogonal component unchanged. Over a full response, the resulting update vanishes exactly where the unscaled GRPO update does, and its one-step ascent bound gains a nonnegative term. Fix a response 𝑦 with 𝑇 valid tokens and let z = (z1 , . . . , z𝑇 ) collect its next-token logits, with 𝜋(· | z𝑡 ) = softmax(z𝑡 ). The advantages 𝐴𝑡 and the directions d𝑡 do not depend on z, and the token gradient in Eq. 2.1 is the gradient of the objective for this response, 𝐽(z) =
𝑇 ∑︁
𝐴𝑡 log 𝜋(𝑦𝑡 | z𝑡 ),
g𝑡 = ∇z𝑡 𝐽(z).
(A.1)
𝑡=1
Each block of the Hessian of 𝐽 is 𝐴𝑡 times that of log 𝜋(𝑦𝑡 | z𝑡 ), whose eigenvalues lie in [− 12 , 0], so 𝐽 is 𝐿-smooth with 𝐿 = max𝑡 |𝐴𝑡 |/2. Proposition A.1 (Stationarity and One-Step Ascent). Let z𝑘 be the current logits and write g𝑡 = ∇z𝑡 𝐽(z𝑘 ) and 𝑢𝑡 = d⊤ 𝑡 g𝑡 for the alignment coefficient, with ‖d𝑡 ‖2 = 1 and 𝜆𝑡 ≥ 0 fixed for this step. With step size 𝜂 > 0, set ̃︀𝑡 = (I + 𝜆𝑡 d𝑡 d⊤ g 𝑡 )g𝑡 ,
̃︀𝑡 . z𝑘+1,𝑡 = z𝑘,𝑡 + 𝜂 g
(A.2)
¯ ≤ 1 with 𝜆 ¯ = max𝑡 𝜆𝑡 , ̃︀𝑡 = 0 for every 𝑡 if and only if ∇z 𝐽(z𝑘 ) = 0. If in addition 𝜂𝐿(1 + 𝜆) Then g 𝐽(z𝑘+1 ) − 𝐽(z𝑘 ) ≥
𝜂 ∑︁ 𝜂 ‖∇z 𝐽(z𝑘 )‖22 + 𝜆𝑡 𝑢2𝑡 . 2 2 𝑡
(A.3)
Proof. Let g = ∇z 𝐽(z𝑘 ) and let P be the block-diagonal matrix with blocks I+𝜆𝑡 d𝑡 d⊤ 𝑡 , so that z𝑘+1 = z𝑘 +𝜂Pg. Each block has eigenvalue 1 + 𝜆𝑡 along d𝑡 and 1 on the orthogonal complement. Hence P is positive definite and therefore invertible, which gives the first claim, and ∑︁ ¯ g⊤ Pg = ‖g‖22 + 𝜆𝑡 𝑢2𝑡 , P2 ⪯ (1 + 𝜆)P. 𝑡
By 𝐿-smoothness,
𝐿𝜂 2 ⊤ 2 𝐽(z𝑘+1 ) ≥ 𝐽(z𝑘 ) + 𝜂 g⊤ Pg − g P g 2 (︁ ¯ )︁ 𝐿𝜂(1 + 𝜆) ≥ 𝐽(z𝑘 ) + 𝜂 1 − g⊤ Pg 2 𝜂 ≥ 𝐽(z𝑘 ) + g⊤ Pg 2 𝜂 𝜂 ∑︁ = 𝐽(z𝑘 ) + ‖g‖22 + 𝜆𝑡 𝑢2𝑡 . 2 2 𝑡
Setting 𝜆𝑡 = 0 in Eq. A.3 recovers the bound 𝜂2 ‖∇z 𝐽(z𝑘 )‖22 of an unscaled step, so the second term is what scaling adds. It grows with the component of the verifier-driven policy gradient along the teacher direction and disappears when the two are orthogonal at every token. Scaling therefore adds to the progress guaranteed ¯ ≤1 at each step without changing where the update vanishes, and the price is the tighter condition 𝜂𝐿(1 + 𝜆) 2 on the step size, since the scaled update is longer. The gain depends on 𝑢𝑡 , so alignments of equal magnitude contribute equally whether the student follows or opposes the teacher. Appendix B.1 analyzes what changes when the two branches use different scales.
21
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
B. Analysis of Asymmetric Alignment Scaling B.1. One-Sided Amplification under Positive-Only Scaling Positive-only amplification is locally well motivated. When 𝑢𝑡 > 0, the component of the sampled student gradient g𝑡 along the teacher-derived direction d𝑡 follows the teacher’s post-training shift. Amplifying this component therefore reinforces an update supported by both the teacher shift and the verifier-driven student gradient. A related positive-gating rule is used by Du et al. (2019), who weight auxiliary updates by the positive part of their gradient cosine similarity. However, applying different scales to the two alignment signs introduces a one-sided effect. Let 𝜆+ and 𝜆− denote the scales applied when 𝑢𝑡 ≥ 0 and 𝑢𝑡 < 0, respectively. The coefficient multiplying d𝑡 in the added correction is 𝑐𝑡 := 𝜆+ 𝑢𝑡 1{𝑢𝑡 ≥ 0} + 𝜆− 𝑢𝑡 1{𝑢𝑡 < 0} (B.1) − − = 𝜆+ +𝜆 𝑢𝑡 + 𝜆+ −𝜆 |𝑢𝑡 |. 2 2 This decomposition separates sign-symmetric scaling from the asymmetry introduced by using different scales for the two branches. The first term symmetrically scales 𝑢𝑡 by the average of the two branch scales. Because it preserves the sign of 𝑢𝑡 , positive and negative contributions can cancel across tokens. The second term depends on |𝑢𝑡 | and appears only when the branch scales differ. In particular, when 𝜆+ > 𝜆− , this term remains nonnegative for either sign of 𝑢𝑡 . It therefore cannot be canceled by changes in the alignment sign, leaving a one-sided coefficient on the local teacher direction. Under positive-only scaling, 𝜆− = 0. If positive and negative alignments nearly balance across sampled tokens and rollouts, such that E[𝑢𝑡 ] ≈ 0, then E[𝑐𝑡 ] = 𝜆2+ (E[𝑢𝑡 ] + E[|𝑢𝑡 |]) ≈ 𝜆2+ E[|𝑢𝑡 |] > 0.
(B.2)
Thus, even when the signed alignments cancel on average, the scalar coefficient 𝑐𝑡 remains positive on average under positive-only scaling. This residual coefficient is governed by the mean alignment magnitude E[|𝑢𝑡 |], rather than the small signed mean E[𝑢𝑡 ]. Importantly, this residual amplification need not reflect only reward-relevant teacher progress. The sign of 𝑢𝑡 reveals whether d𝑡 and g𝑡 agree, but not why they agree. At an individual sampled token, we write g𝑡 = g𝑡⋆ + 𝜖𝑡 , where g𝑡⋆ denotes the underlying reward-improving signal and 𝜖𝑡 aggregates incidental or misattributed components arising from coarse response-level credit assignment, rollout and mini-batch sampling, and estimator-specific effects that need not correspond to actions responsible for higher reward. Similarly, d𝑡 captures all changes induced by teacher post-training, including both reward-relevant progress and incidental behavioral changes. Positive alignment may therefore arise from either useful teacher-acquired progress or an incidental tendency shared by the two vectors. Positive-only scaling cannot distinguish between these cases and amplifies the aligned component regardless of its source. Response length provides one concrete example of such a shared tendency. In reasoning tasks, higher verifier rewards are often associated with longer reasoning traces, so the student gradient g𝑡 may favor tokenlevel changes that prolong generation. The teacher direction d𝑡 may encode a similar tendency acquired during teacher post-training. This tendency may represent useful additional reasoning, but it may also reflect length-dependent effects in the policy-gradient estimate (Liu et al., 2025). When it is shared by both signals, positive-only scaling amplifies it whenever it produces positive alignment.
22
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
B.2. Isolating the Positive and Negative Alignment Branches To examine the branch-specific effects, we isolate the two alignment branches by activating gradient scaling only when 𝑢𝑡 ≥ 0 (positive-only) or only when 𝑢𝑡 < 0 (negative-only), while holding all other training settings fixed within each task. Figure 6 reports evaluation performance and response length during training on mathematics and Knights & Knaves tasks. Across both tasks, positive-only scaling produces rapid early gains accompanied by a sharp increase in response length. Performance then begins to decline as responses grow toward the generation limit. Negative-only scaling exhibits the opposite pattern: responses become shorter, while performance quickly falls to near zero. OPRD-Aligned (ut ≥ 0)
40 20 0
0
50
100
12k
150
Training Steps
(a) Math
8k 4k 0
0
50
100
150
Training Steps
OPRD-Opposed (ut < 0)
80
Response Length
60
Knights Pass@1 (%)
Response Length
AIME24 Mean@16 (%)
GRPO
60 40 20 0
0
50
100
150
6k 4k 2k 0
0
Training Steps
50
100
150
Training Steps
(b) Knights & Knaves
Figure 6: Evaluation performance and response length over training for OPRD variants with only the positive- or negative-alignment branch active. For Math and Knights & Knaves, we use Qwen3-4B and Qwen3-4B-Base teacher checkpoints obtained after 75 and 105 RL training steps, respectively. The two single-branch OPRD variants are trained for 45 steps with 𝜆 = 0.5, while GRPO is shown through 150 steps for reference. The maximum generation lengths are 20K and 8K tokens for the two settings, respectively. The gray dashed lines denote the performance or response length of the corresponding weak teachers. All other training settings follow the dataset-specific default configurations described in Appendix C.
The rapid gains under positive-only scaling suggest that teacher-aligned components provide effective early transfer of the progress acquired during teacher post-training. By contrast, the collapse under negative-only scaling suggests that teacher-opposing components are less reliable early in training, when the student’s rollouts remain weak. Because the verifier provides only response-level feedback, even a rewarded trajectory may contain locally unhelpful token choices whose gradients are negatively aligned with the teacher shift. Applying negative-branch scaling at full strength from the outset can therefore reinforce unreliable token-level updates. The response-length dynamics are also consistent with the shared tendency discussed in Appendix B.1. In both tasks, performance improvements under GRPO are accompanied by longer reasoning traces, suggesting that the student gradient g𝑡 favors token-level changes that prolong generation. The teacher develops a similar tendency during RL post-training, which may be encoded in d𝑡 . Positive-only scaling reinforces this shared tendency and rapidly drives responses toward the generation limit. Negative-only scaling instead amplifies student-gradient components that oppose d𝑡 , counteracting the length-increasing tendency and producing shorter responses.
23
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
B.3. Mitigating One-Sided Amplification through Branch Scheduling The isolated-branch results suggest that the positive and negative branches play complementary roles over training. The positive branch amplifies components supported by both the teacher shift and the verifier-driven student gradient, thereby providing rapid early transfer. The negative branch instead amplifies verifier-supported departures from the teacher direction, which may help the stronger student move beyond the weak teacher. However, these departures are less reliable early in training, when the student’s on-policy rollouts remain weak. This difference motivates controlling the relative strengths of the two branches over training. We compare three strategies for avoiding persistent one-sided amplification. Under the default OPRD schedule, 𝜆+ remains fixed at 𝜆, while 𝜆− gradually increases from 0 to 𝜆. Early in training, the larger positivebranch scale prioritizes teacher-aligned components and provides an effective bootstrap. As 𝜆− increases, verifier-supported gradient components whose projections oppose the teacher direction receive progressively greater amplification. Once 𝜆− = 𝜆+ = 𝜆, the asymmetric term proportional to |𝑢𝑡 | vanishes and the correction coefficient reduces to 𝑐𝑡 = 𝜆𝑢𝑡 . The schedule thus preserves rapid teacher-aligned transfer early in training while gradually introducing stronger departures from the weak teacher. Alternatively, we keep 𝜆− = 0 and gradually decrease 𝜆+ from 𝜆 to 0. This schedule likewise uses positivebranch scaling as an early bootstrap but progressively removes the added teacher-direction correction. Once 𝜆+ = 0, both branch scales are zero, so the transformed gradient reduces to the original verifier-driven policy gradient and training returns to GRPO. As a schedule-free alternative, we also consider fixed symmetric scaling, which sets 𝜆+ = 𝜆− = 𝜆 throughout training. This removes one-sided amplification from the outset but activates the distinct effects of both branches simultaneously. OPRD (λ± = λ)
OPRD (λ + Annealing)
Knights Pass@1 (%)
AIME24 Mean@16 (%)
GRPO
60
40
20
0
30
60
90
Training Steps
(a) Math
120
150
OPRD (λ − Ramp-Up)
80 60 40 20 0
0
30
60
90
120
150
Training Steps
(b) Knights & Knaves
Figure 7: Comparison of three branch-scheduling strategies. We compare the default 𝜆− ramp-up (𝜆+ = 0.5, 𝜆− : 0 → 0.5), 𝜆+ annealing (𝜆+ : 0.5 → 0, 𝜆− = 0), and fixed symmetric scaling (𝜆+ = 𝜆− = 1.0) against GRPO. The two scheduled variants use horizons of 30 updates for Math and 75 updates for Knights & Knaves. We use Qwen3-4B and Qwen3-4B-Base teacher checkpoints obtained after 75 and 105 RL training steps, respectively. Gray dashed lines denote teacher performance. All other training configurations follow Appendix C.
As shown in Figure 7, the two scheduled variants begin with positive-only amplification and achieve rapid early gains, whereas fixed symmetric scaling improves much more slowly despite using 𝜆 = 1.0: it only gradually breaks through on Math and yields limited early gains on Knights & Knaves. Because response-level feedback can reward trajectories containing locally incorrect or incidental steps, the resulting 𝑢𝑡 < 0 components are less reliable on weak early rollouts and can dampen the positive-branch bootstrap when amplified from the outset. Activating only 𝜆+ is therefore the more reliable default for early acceleration. The later acceleration of fixed symmetric scaling on Knights & Knaves suggests that 𝜆− becomes useful once the student reaches a stronger regime and produces more informative on-policy rollouts. At this stage, it can amplify meaningful verifier-supported departures discovered through the student’s own rollouts, helping it move beyond the weak teacher. Although only 𝜆+ annealing shows that returning to verifier-only optimization after the initial bootstrap is also viable, it forgoes explicit amplification of these student-discovered departures. We therefore adopt 𝜆− ramp-up as the default: it preserves the early acceleration from 𝜆+ while introducing 𝜆− later to remove persistent one-sided amplification and support progress beyond the weak teacher.
24
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
C. Training and Evaluation Details Table 4 and Table 5 summarize the default training settings and method-specific configurations for GRPO, OPD, KDRL, and OPRD. Scenario-specific settings are provided in their respective Appendix sections. We train all models on four NVIDIA B200 GPUs using Fully Sharded Data Parallel (FSDP) (Zhao et al., 2023). Table 4: Default training settings for Math and Reasoning Gym. These configurations are shared across all methods. Method-specific settings are provided in Table 5. Settings
Math
Reasoning Gym
Data and Models Training data
DAPO-Math-17K
Knights & Knaves, Quantum Lock, String Manipulation, and Countdown (19,800 examples per task)
Prompt format
Chat template with a system prompt
Chat template without a system prompt
Student policy
Qwen3-8B (non-thinking)
Qwen3-8B-Base (non-thinking)
Teacher policy
Qwen3-4B (step 75, non-thinking)
Qwen3-4B-Base (step 75 for String task, step 105 for the other tasks, non-thinking)
Training horizon
150 policy updates
150 policy updates
Prompt batch / mini-batch
64 / 64
64 / 32
Rollouts per prompt
8
8
Optimizer
AdamW, 𝛽 = (0.9, 0.999), weight decay 0.01, AdamW, 𝛽 = (0.9, 0.999), weight decay 0.01, gradient clipping 1.0 gradient clipping 1.0
Learning rate
1 × 10−6 (constant schedule with 10 warm-up updates)
1 × 10−6 (constant schedule with 10 warm-up updates)
Policy optimization
PPO clipping range [0.20, 0.28], no standard-deviation normalization, no KL or entropy regularization
PPO clipping range [0.20, 0.28], no standard-deviation normalization, no KL or entropy regularization
Training-time decoding
Temperature 1.0, top-𝑝 1.0, no top-𝑘
Temperature 1.0, top-𝑝 1.0, no top-𝑘
Maximum prompt length
2,048 tokens
2,048 tokens
Maximum response length
20,480 tokens
8,192 tokens
Length-based reward
No penalty up to 16,384 tokens, then linear penalty reaching −1 at 20,480 tokens
–
Optimization
Generation
Table 5: Method-specific training settings. All distillation-based methods use the same task-specific frozen teacher checkpoint specified in Table 4. Settings
GRPO
OPD
KDRL
OPRD
Optimization Frozen teacher
–
Task-specific
Task-specific
Task-specific
Teacher reference
–
–
–
Raw Qwen3-4B family
Teacher signal
–
Teacher–student log-probability ratio
K2 signal
Teacher-shift direction d𝑡
Teacher temperature
–
1.0
1.0
1.0
Token support
–
Sampled tokens
Sampled tokens
Sampled ∪ student top-10
Coefficient schedule
–
Fixed at 1.0
𝛽𝑘 : 0.005 → 0 over Math: 30 updates Reasoning Gym: 75 updates
𝜆+ = 0.5, 𝜆− : 0 → 0.5 over Math: 30 updates Reasoning Gym: 75 updates
25
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Table 6 summarizes the default evaluation settings for Math and Reasoning Gym. Across methods, all trained policies are evaluated using the same task-specific settings. Table 6: Default evaluation settings for Math and Reasoning Gym. Teacher and initial-student checkpoints use the same decoding and scoring protocols as trained student checkpoints. Settings
Math
Reasoning Gym
Benchmarks and Metrics Reported benchmarks
AIME'24, AIME'25, HMMT'25, OlympiadBench
Knights & Knaves, Quantum Lock, String Manipulation, Countdown (200 examples per task)
Evaluation metric
Mean@16
Pass@1
Rollouts per problem
16
1
Decoding parameters
Temperature 0.7, top-𝑝 0.8, top-𝑘 20
Temperature 0.6, top-𝑝 0.95, top-𝑘 20
Maximum prompt length
2,048 tokens
2,048 tokens
Maximum response length 38,912 tokens
8,192 tokens
Decoding
Scoring and Reporting Scoring
Exact match after answer normalization Nonempty boxed answer required, K&K: exact match after normalization, Quantum Lock: 1.0 for a reference-length valid path, 0.5 for any other valid path, 0 otherwise, String Manipulation: case-sensitive exact match, Countdown: valid expression using each given number exactly once and reaching the target
Checkpoint averaging
5-checkpoint mean (30-update intervals) 5-checkpoint mean (30-update intervals)
26
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
D. Detailed Results for Weak-to-Strong Model Transfer D.1. Detailed Learning Curves We study successive model transfer within the Qwen3 family (Qwen Team, 2025b), using GRPO-trained 4B-scale models as teachers to accelerate the post-training of larger 8B-scale students. We select intermediate teacher checkpoints whose the performance exceeds that of the initial student but remains below the student’s end-of-training GRPO performance. We also considered cross-generation transfer from Qwen2.5 (Qwen Team, 2025a) to Qwen3. In preliminary experiments, however, the Qwen2.5-3B and 7B checkpoints remain below this target range (around 14% on AIME'24), while obtaining suitable post-trained checkpoints and the corresponding teacher-shift signals would require substantially more compute. We therefore focus on controlled within-family transfer. We compare OPRD with GRPO, OPD, and KDRL on four mathematics benchmarks and four Reasoning Gym tasks, reporting Mean@16 and Pass@1, respectively. Figure 1 aggregates performance across benchmarks, whereas Table 1 averages each trained policy over five checkpoints. Figure 8 and Figure 9 show the corresponding benchmark-level learning curves. Across these benchmarks, OPRD generally retains OPD’s rapid early improvement. Unlike OPD, which plateaus near the weak teacher, OPRD continues to improve beyond it and reaches GRPO’s end-of-training performance substantially earlier. GRPO
OPD
Mean@16 (%)
60
40
20
75 150 Training Steps
54
50
20
40
0
OPRD
30
60
20
KDRL
0
(a) AIME'24
75 150 Training Steps
10
46 0
(b) AIME'25
75 150 Training Steps
0
75 150 Training Steps
(c) HMMT'25
(d) OlympiadBench
Figure 8: Learning curves for successive model transfer on individual math benchmarks. We use Qwen3-4B as the teacher and Qwen3-8B as the student. The gray dashed line denotes the performance of the weak teacher. All training settings follow the default Math configuration described in Appendix C.
GRPO
Pass@1 (%)
KDRL
OPRD
80
80
40
40
20
20
0
75 150 Training Steps
(a) Knights & Knaves
0
60
40
60
60
0
OPD
40 20
0
75 150 Training Steps
(b) Quantum Lock
0
20
0
75 150 Training Steps
(c) String Manipulation
0
0
75 150 Training Steps
(d) Countdown
Figure 9: Learning curves for successive model transfer on individual reasoning benchmarks. We use Qwen3-4B-Base as the teacher and Qwen3-8B-Base as the student. The gray dashed line denotes the performance of the weak teacher. All training settings follow the default Reasoning Gym configuration described in Appendix C.
27
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
D.2. Additional Results Across Model Variants and Tasks The instruction-tuned Qwen3 results reported in Section D.1 exhibit a potential response-length confound. Although thinking mode is disabled, longer responses may implicitly elicit some of the reasoning behavior associated with that mode, leading to abrupt, transient score gains. In Figure 8, for example, OPD briefly surpasses the teacher on both AIME'24 and AIME'25 at step 30 before returning toward a teacher-level plateau. Such behavior can confound comparisons of early learning speed. We therefore evaluate Qwen3-Base models, for which this effect is less pronounced, in Figure 10. OPRD again substantially accelerates weak-to-strong generalization, achieving high performance much earlier than the baselines on all four benchmarks. GRPO
Mean@16 (%)
30
OPD
25
KDRL
12
40
8
35
4
30
20
20
OPRD
15
10
0
75 150 Training Steps
0
(a) AIME'24
75 150 Training Steps
0
(b) AIME'25
75 150 Training Steps
0
75 150 Training Steps
(c) HMMT'25
(d) OlympiadBench
Figure 10: Learning curves on individual math benchmarks using Qwen3-Base models. We use Qwen3-4B-Base as the teacher and Qwen3-8B-Base as the student. The gray dashed line denotes the performance of the weak teacher. All other settings follow the default Math configuration in Appendix C, but we omit the system prompt and reduce the mini-batch size to 32, yielding two optimizer steps per training batch.
Reward gains often conincide with longer responses. For instruction-tuned Qwen3, this makes a potential confound: distillation gains may simply reflect longer responses eliciting latent thinking behavior. To test whether OPRD depends on this effect, we evaluate three more reasoning tasks in Figure 11, where, as in String Manipulation, post-training shortens responses by a factor of three to four relative to the raw checkpoints. OPRD still improves substantially faster than the baselines, quickly reaching GRPO’s eventual plateau while reducing response length. This opposite trend shows that its gains are not tied to response-length growth. OPRD’s gradient scaling can nevertheless magnify length bias in the teacher-shift signal, as discussed in Section 4.3, Appendix B, and Appendix I. We resolve this by using a later teacher checkpoint, rather than the raw model, as the reference policy.
Pass@1 (%)
GRPO
OPD
KDRL
OPRD
80
40
80 60
60
40
30
20
40
0
75
Training Steps
(a) Zebra Puzzles
150
20
20 0
75
Training Steps
(b) Color Cube Rotation
150
0
0
75
Training Steps
150
(c) Binary Matrix
Figure 11: Learning curves on additional three Reasoning Gym tasks. We use Qwen3-4B-Base as the teacher and Qwen3-8B-Base as the student. The gray dashed line denotes the performance of the weak teacher. For Color Cube Rotation and Binary Matrix, we use the teacher checkpoints from steps 30 and 45, respectively, as the reference policies instead of the raw step-0 models to mitigate length bias (see Appendix I for details).
28
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
To complement the detailed learning curves, Table 7 reports checkpoint-averaged results, providing a numerical summary of how quickly each method reaches high performance. OPRD again achieves the strongest results, confirming that it accelerates weak-to-strong generalization across these additional settings. Table 7: Additional results for successive model transfer with Qwen3-Base models on mathematics and three additional Reasoning Gym tasks. We report Mean@16 for Math and Pass@1 for Reasoning Gym. We use intermediate GRPO checkpoints of Qwen3-4B-Base as teachers for Qwen3-8B-Base students. The teacher and initial-student rows report fixed-checkpoint performance, whereas each trained-policy row averages evaluations at steps 30, 60, 90, 120, and 150. The corresponding learning curves are shown in Figure 10 and Figure 11. The best trained-policy result in each column is shown in bold. Math Reasoning Policy
AIME'24
Reasoning Gym
AIME'25 HMMT'25 Olympiad
Avg.
Qwen3-4B-Base (Teacher) → Qwen3-8B-Base (Student)
Zebra
Color
Binary
Avg.
Qwen3-4B-Base (Teacher) → Qwen3-8B-Base (Student)
Teacher
20.00
18.33
8.13
35.55
20.50
31.00
46.00
55.00
44.00
Student + GRPO + OPD + KDRL + OPRD
12.71 20.00 20.00 21.92 24.79
13.54 16.58 17.21 18.88 20.79
3.75 8.33 9.13 9.29 11.75
30.01 38.59 35.58 38.69 40.01
15.00 20.88 20.48 22.19 24.34
25.50 35.40 29.10 35.60 39.10
27.50 48.30 48.30 49.50 72.40
9.00 56.50 62.10 61.40 66.80
20.67 46.73 46.50 48.83 59.43
D.3. Evaluation Results with Standard Deviations To assess the evaluation-time robustness of the comparisons in Table 1, we report response-resampling variability in Table 8. Because multi-seed post-training is prohibitively expensive, we hold the benchmark problems and trained checkpoints fixed and resample only their responses. Each of 1,000 replicates draws 16 responses with replacement from a pool of 32 per Math problem and one from a pool of eight per Reasoning Gym problem. Trained-policy results are averaged over five checkpoints within each replicate, and we report the resulting mean and sample standard deviation. OPRD still surpasses the strongest baseline by approximately 8.0 points on Math and 9.9 points on Reasoning Gym, margins far exceeding the observed response-resampling variability. Table 8: Evaluation results with standard deviations for successive model transfer on mathematics and reasoning tasks. We use intermediate GRPO checkpoints of 4B-scale Qwen3 models as teachers and report Mean@16 for mathematics and Pass@1 for Reasoning Gym as the bootstrap mean ± sample standard deviation over 1,000 response-resampled evaluations, with benchmark items held fixed. In each bootstrap replicate, we sample 16 responses with replacement from a pool of 32 for each mathematics problem and one response from a pool of eight for each Reasoning Gym problem. Math Reasoning Policy
AIME'24
AIME'25 HMMT'25 Olympiad
Reasoning Gym Avg.
Qwen3-4B (Teacher) → Qwen3-8B (Student)
Teacher Student + GRPO + OPD + KDRL + OPRD
Knights
Quantum
String
Count
Avg.
Qwen3-4B-Base (Teacher) → Qwen3-8B-Base (Student)
41.77
36.77
21.44
52.02
38.00
54.92
24.25
19.84
13.04
46.21
25.83
11.71 ± 1.96
± 1.35
± 1.12
± 1.05
± 0.70
46.88
36.61
22.47
52.48
39.61
53.17
34.70
35.51
41.95
41.33
46.62
37.91
21.40
51.30
39.31
55.78
39.91
34.55
42.11
43.09
54.20
42.80
25.66
53.39
44.01
53.91
37.83
38.10
49.54
44.84
67.40
55.69
31.63
53.28
52.00
72.02
49.70
42.11
55.30
54.78
± 1.48 ± 1.09 ± 0.63
± 0.74
± 0.63
± 0.63
± 1.38 ± 1.09 ± 0.55
± 0.59
± 0.59
± 0.64
± 1.15 ± 0.95 ± 0.51
± 0.47
± 0.55
± 0.58
± 0.23 ± 0.24 ± 0.10
± 0.10
± 0.10
± 0.09
± 0.58 ± 0.45 ± 0.25
± 0.27
± 0.26
± 0.27
± 2.93
± 1.21
± 1.31
± 1.14
± 0.94
41.64 ± 2.73
5.00
± 0.93
± 1.19
± 1.12
± 1.02
33.78 ± 1.17
3.61
± 0.59
± 0.73
± 0.58
± 0.57
42.13 ± 1.57
2.83
± 0.64
± 0.78
± 0.75
± 0.75
43.12 ± 1.11
5.79
± 0.45
± 0.51
± 0.48
± 0.42
29
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
E. Detailed Results for Multi-Teacher Weak-to-Strong Distillation We follow the single-teacher Reasoning Gym setting but jointly train one student on domain-mixed batches, pairing each example with its task-specific teacher. Figure 12 shows the per-domain learning curves. Despite heterogeneous task structures and response-length trends—String Manipulation responses shorten as reward improves, whereas those for the other tasks generally lengthen—OPRD accelerates learning and attains the highest Pass@1 across all four domains. MOPD shows signs of cross-task interference, most notably on Quantum Lock, where it falls below the corresponding specialist, while OPRD rapidly transfers the specialist capabilities and continues improving without comparable degradation. Mix-RL
MOPD
KDRL
OPRD
Pass@1 (%)
80
80 60
20 20
20
20 0
40
40
40
60
40
60
0
150 300 Training Steps
(a) Knights & Knaves
0
0
150 300 Training Steps
(b) Quantum Lock
0
0
150 300 Training Steps
0
(c) String Manipulation
0
150 300 Training Steps
(d) Countdown
Figure 12: Learning curves on individual Reasoning Gym tasks under multi-teacher distillation. We use four task-specific Qwen3-4B-Base models as teachers and jointly train a Qwen3-8B-Base student. The gray dashed line denotes the performance of the corresponding specialist teacher. All other settings follow the configurations described in Section C.
F. Detailed Results for Strong-to-Weak Distillation Figure 13 presents detailed learning curves for strong-to-weak settings. On AIME'24, OPRD raises the Qwen31.7B initial student’s Mean@16 from 10.0 to above 41 within 150 updates, whereas GRPO reaches only about 25 at the same point and 35 even after 240 updates. On Knights & Knaves, OPD improves initially but collapses midway through training and remains below the teacher after recovering. In contrast, OPRD rapidly improves the Qwen3-0.6B student and ultimately surpasses the Qwen3-8B-Base teacher’s Pass@1 of 64.5. GRPO
OPD
KDRL
Pass@1 (%)
Mean@16 (%)
50 40 30 20 10 0
75 Steps (a) AIME'24
150
OPRD
60 40 20 0
0
75 Steps
150
(b) Knights & Knaves
Figure 13: Learning curves for strong-to-weak distillation on two tasks. We evaluate Qwen3-8B → Qwen3-1.7B on AIME'24 and Qwen3-8B-Base → Qwen3-0.6B on Knights & Knaves, using teachers from step 105 of task-specific GRPO. Gray dashed lines mark teacher performance. All other settings follow Appendix C, with method-specific distillation coefficients scheduled over the first 45 updates.
30
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
G. Detailed Results for Weak-to-Strong Method Comparisons G.1. Baseline Implementation Details Under the default training and evaluation configurations in Appendix C, all baselines use the same model pairs, teacher checkpoints, prompt formats, and evaluation protocols as OPRD unless otherwise noted. We describe only their method-specific settings below. • W2SR-P (Yuan et al., 2026). We reproduce the seeded prompt stream used by the 150-update RL runs, yielding 150×64 = 9,600 prompt occurrences. For each occurrence, we sample eight responses from the weak teacher and select one verifier-correct, format-valid, non-truncated response, discarding occurrences with no valid candidate. We then fully fine-tune the initial student checkpoint for three epochs using next-token prediction with a global batch size of 64 and a learning rate of 2 × 10−5 . • S2L-PO (Ren et al., 2026). S2L-PO linearly anneals the fraction of weak-model rollouts over the first half of GRPO training. Although the original method advocates using a smaller base model as the weak explorer to exploit its policy-level diversity, we use the same post-RL weak teacher as the other baselines for a controlled comparison. While the original implementation uses 16 rollouts per prompt, we retain its 16-phase schedule with the default group size of eight. Over 150 updates, the weak/student composition transitions from 8/0 to 0/8 during the first eight phases (updates 1–75) and remains at 0/8 during the remaining eight phases (updates 76–150). For each trajectory, we compute the importance ratio using its generating policy as 𝜋rollout , namely 𝜋𝑇 for weak-teacher rollouts and 𝜋𝜃old for student rollouts. We also retain the original KL regularization toward the initial student with a coefficient of 10−3 . • OPSD (Zhao et al., 2026). OPSD is originally a self-distillation method that uses a correct self-generated rollout as privileged information. To adapt it to our weak-to-strong setting, we instead use a verifier-correct weak-teacher rollout as privileged information, falling back to a correct student rollout when the weak teacher produces none. An EMA copy of the student serves as the self-teacher, conditioning on the privileged rollout to provide distillation targets for the original student trajectories and being updated after each step with a rate of 0.05. Whenever valid privileged information is available, we apply the distillation loss to all student trajectories in the group, regardless of whether they are correct or incorrect. We use generalized JSD with 𝛼 = 0.5 over the top-100 student tokens and an additional tail bucket. • Direct-OPD (Feng et al., 2026). Developed concurrently with OPRD, Direct-OPD optimizes the weak policy shift as a dense reward on student-generated trajectories: [︃ ]︃ ∑︁(︀ )︀ ref 𝒥Direct-OPD = E𝑥, 𝑦∼𝜋𝜃 log 𝜋𝑇 (𝑦𝑡 | 𝑠𝑡 ) − log 𝜋𝑇 (𝑦𝑡 | 𝑠𝑡 ) − 𝛼𝐷KL (𝜋𝜃 ‖ 𝜋𝑆,0 ) . 𝑡
Following the original implementation, we evaluate the dense reward over the top-16 tokens of the old student policy at each visited state and use the reported hyperparameters. The policy-shift scale is fixed at 1, while 𝛼, the coefficient of the KL anchor toward 𝜋𝑆,0 , is initialized at 2.5. Before each actor update, 𝛼 is multiplied by 1.01 or 0.99 depending on whether the batch-mean dense reward is positive or negative, respectively, and clipped to [0.5, 2.5]. This KL anchor is computed separately on the sampled response tokens using the low-variance k3 estimator. • W2S-OPD (Yu et al., 2026). W2S-OPD reanchors the weak policy shift to the initial student by defining the proxy teacher as (︂ )︂𝛾 𝜋𝑇 (𝑣 | 𝑠𝑡 ) 𝜋proxy (𝑣 | 𝑠𝑡 ) ∝ 𝜋𝑆,0 (𝑣 | 𝑠𝑡 ) . 𝜋𝑇ref (𝑣 | 𝑠𝑡 ) Following the original implementation, we set 𝛾 = 1 and compute the proxy scores over the full vocabulary before selecting the proxy’s top-32 tokens. We normalize both the proxy and current-student distributions over this proxy-selected support and minimize the reverse KL from the current student to the proxy. This restrictedsupport reverse KL serves as the sole actor objective, with no additional KL anchor or adaptive coefficient.
31
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
G.2. Detailed Learning Curve In Figure 14, we compare OPRD with three methods that leverage off-policy generations from the weak teacher. OPRD exhibits the strongest and most consistent gains overall. Consistent with Yuan et al. (2026), W2SR-P shows that SFT on verifier-correct teacher rollouts can move the student slightly beyond weak-teacher performance. S2L-PO (Ren et al., 2026) remains competitive on the two Reasoning Gym tasks, although its AIME'24 performance deteriorates after weak-teacher rollouts are fully annealed out at update 75 and its checkpoint-averaged performance remains below OPRD. Our OPSD variant (Zhao et al., 2026) performs poorly whether the privileged trace is self-generated or supplied by the weak teacher. This behavior is consistent with recent findings that privileged self-distillation can impair thinking models by shortening or suppressing deliberative reasoning (Kim et al., 2026; Kaur et al., 2026). Accordingly, OPSD provides a modest benefit only on String Manipulation, where higher rewards coincide with shorter reasoning traces, and fails to deliver competitive gains on the other tasks.
60
40
20
0
30
60
90
120
W2SR-P
OPSD
80 60 40 20 0
150
S2L-PO
String Pass@1 (%)
OPD Knights Pass@1 (%)
AIME24 Mean@16 (%)
GRPO
0
30
Training Steps
60
90
120
OPRD 60
40
20
0
150
0
30
Training Steps
(a) AIME'24
60
90
120
150
Training Steps
(b) Knights & Knaves
(c) String Manipulation
Figure 14: Training dynamics of weak-to-strong methods using off-policy generations from the weak teacher. Because W2SR-P performs SFT without subsequent RL, its final performance is shown as a horizontal dashed line. The gray dashed lines indicate weak-teacher performance. See Appendix G.1 for baseline implementation details.
In Figure 15, we further compare OPRD with Direct-OPD (Feng et al., 2026) and W2S-OPD (Yu et al., 2026), two concurrent methods that likewise exploit the weak policy delta. Although these methods use the same transferred signal, their objectives are defined directly by the delta and therefore receive no independent verifier-driven update direction. W2S-OPD can surpass the weak teacher, but ultimately plateaus near teacherlevel performance because its optimization target remains restricted to the policy changes encoded by the weak teacher. OPRD instead uses the delta only to identify and rescale the component of the verifier gradient aligned with the weak shift, while preserving the orthogonal component g𝑡⊥ . Consequently, the delta guides rather than replaces verifier-driven optimization, allowing OPRD to improve beyond teacher-level saturation and achieve the strongest final performance across all three tasks.
40
0
30
60
90
Training Steps
(a) AIME'24
120
150
Direct-OPD
W2S-OPD
80 60 40 20 0
OPRD String Pass@1 (%)
60
20
OPD Knights Pass@1 (%)
AIME24 Mean@16 (%)
GRPO
0
30
60
90
120
Training Steps
(b) Knights & Knaves
150
60
40
20
0
0
30
60
90
120
150
Training Steps
(c) String Manipulation
Figure 15: Training dynamics of weak-to-strong methods using the weak policy delta. The horizontal dashed lines indicate weak-teacher performance. See Appendix G.1 for baseline implementation details.
32
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
H. Additional Results on Guidance-Direction Construction OPRD requires a guidance direction that captures the reward-relevant change acquired by the teacher. Section 4.2 compares three constructions: the weak policy delta Δ𝑡 contrasts the post-trained teacher with its reference policy and isolates the change acquired during post-training; OPD contrasts the post-trained teacher with the current student, so its direction conflates the teacher’s post-training update with the broader mismatch ref between the teacher’s reference policy and the current student (i.e., z𝑇 − z𝑆 = (z𝑇 − zref 𝑇 ) + (z𝑇 − z𝑆 )); and OPSD derives its direction from the discrepancy induced by a privileged teacher draft. The comparison uses the step-60 checkpoint from the weak teacher’s GRPO run. Under this setting, the weak policy delta outperforms both alternatives by a wide margin. However, OPD follows the gradient of a teacher-matching objective, its usefulness as a scaling direction should depend on teacher performance. We test this using the stronger teacher checkpoints adopted in our main experiments while keeping all other settings fixed (Figure 16). For mathematics, we use the step-75 teacher, which is already relatively strong. For Knights & Knaves, we use the step-105 teacher, which achieves 57.5% Pass@1 compared with 29.0% at step 60. With these teachers, the OPD direction performs well on both tasks, although it remains slightly behind the weak policy delta overall. OPSD is less consistent: it finishes above GRPO on AIME'24 but barely improves on Knights & Knaves. OPRD (OPSD)
60
40
20
OPRD (Delta)
Knights Pass@1 (%)
AIME24 Mean@16 (%)
GRPO
0
30
60
90
Training Steps
(a) Math
120
150
OPRD (OPD)
80 60 40 20 0
0
30
60
90
120
150
Training Steps
(b) Knights & Knaves
Figure 16: Comparison of guidance-direction constructions. The variants construct d𝑡 from the teacher–reference policy shift (Delta), the teacher–student mismatch (OPD), or the privileged-context discrepancy (OPSD). Using stronger teacher checkpoints than those used in Figure 3a, we pair the step-75 Qwen3-4B teacher with a Qwen3-8B student for Math and the step-105 Qwen3-4B-Base teacher with a Qwen3-8B-Base student for Knights & Knaves. Gray dashed lines mark teacher performance. All other training configurations follow Appendix C.
These results suggest that the OPD gradient can provide a useful guidance direction when the weak teacher is sufficiently capable. In practice, however, the eventual performance gap between the weak teacher and the larger student cannot be known without fully training the student, making OPD difficult to adopt as a reliable default. OPSD is also less reliable because a privileged draft can constrain the student to a prescribed reasoning path (Kim et al., 2026; Kaur et al., 2026). Using its self-distillation gradient as d𝑡 can then amplify verifier-gradient components aligned with this restrictive signal and hinder learning. The weak policy delta avoids both limitations because it compares the post-trained teacher only with its own reference policy. This isolates the change acquired during post-training without relying on either the evolving student or a privileged draft. We therefore retain the weak policy delta as our default guidance direction due to its stronger empirical performance and greater reliability in practice.
33
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
I. Additional Results on Length Bias in Teacher Policy Shift The teacher policy shift Δ𝑡 (and hence the guidance direction d𝑡 ) may contain reward-irrelevant components such as 𝜖𝑡 alongside task-relevant progress. As discussed in Section 4.3 and Appendix B, OPRD amplifies the projection of g𝑡 onto d𝑡 , and this can also magnify reward-irrelevant components encoded in the guidance direction. Response length provides one observable example: when it correlates with verifier reward, both g𝑡 and d𝑡 may favor shorter responses even when shortening itself does not improve reasoning. Because d𝑡 is derived from Δ𝑡 , the choice of 𝜋𝑇ref determines how much of the teacher’s length change enters the guidance. A step-0 reference uses the base model and therefore includes the full post-training shift, whereas a later reference can exclude a sharp early length collapse.
80 60 GRPO OPD KDRL OPRD (Ref. 0) OPRD (Ref. 45)
40 20 0
0
30
60
90
120
150
Response Length
Binary Matrix Pass@1 (%)
Binary Matrix provides another instance of this behavior. The teacher’s mean response length falls sharply between steps 30 and 45 and then stabilizes. We therefore compare step-0 and step-45 choices of 𝜋𝑇ref : the former includes a large shortening component in Δ𝑡 , whereas the latter excludes most of it. Figure 17 compares both OPRD variants with GRPO, OPD, and KDRL. Both initially improve faster than GRPO but diverge after step 90. With the step-0 reference, the student’s responses continue to shorten and Pass@1 plateaus at 81.5%, below GRPO and KDRL, consistent with the correction overemphasizing length reduction. With the step-45 reference, response length does not exhibit the same continued decline and Pass@1 reaches 96.0%. As on Color Cube, placing the reference after the sharp length transition mitigates this bias while retaining the teacher’s later task progress. However, changing the reference modifies Δ𝑡 as a whole rather than isolating its length-related component. Disentangling structured bias from task-relevant guidance therefore remains an open question. Teacher GRPO Student GRPO OPRD (Ref. 0 OPRD (Ref. 45
3000
) )
2000 1000 0
0
30
60
90
120
Training Steps
Training Steps
(a) Task performance
(b) Response Length
150
Figure 17: Effect of reference-policy selection on length bias in Binary Matrix. We transfer a step-105 Qwen3-4B-Base teacher 𝜋𝑇 to a Qwen3-8B-Base student, using either the step-0 or step-45 checkpoint from the same GRPO run as 𝜋𝑇ref . OPRD’s negative-branch scale 𝜆𝑡 is warmed up over the first 75 steps. The gray dashed line marks teacher performance, while the colored stars denote the mean response lengths of the two choices of 𝜋𝑇ref , and the white star marks that of 𝜋𝑇 . All other settings follow Appendix C.
34
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
J. Detailed Analysis of Student Behavior J.1. Token Alignment Analysis We analyze the correct AIME'25 response shown in Figure 5a using the OPRD-trained Qwen3-8B student at update 150 and the Qwen3-4B teacher at GRPO update 75. We compute the policy gradient g𝑡 assuming a single correct rollout with 𝐴𝑡 = 1, since the magnitude of a positive advantage does not affect cosine similarity. Each token is colored by cos(d𝑡 , g𝑡 ), with green indicating positive alignment and red indicating negative alignment. We compare the original continuation with an alternative generated by the same student checkpoint, keeping the selected prefix fixed and forcing the next token to be 5 , the top-ranked token under d𝑡 . J.2. Response Style Analysis
-0.65
+0.22
+0.71
+0.78
+0.78
+0.73
+0.73
+1.00
Modality
-1.00
-0.85
-0.72
-0.61
-0.54
+0.53
+0.58
+0.39
+0.43
+0.46
+0.42
+1.00
Grammar
-1.00
-0.42
-0.35
-0.33
-0.43
+0.12
+0.67
+0.65
+0.64
+0.59
+0.53
+1.00
Punctuation
-1.00
-0.56
+0.04
+0.15
-0.27
+0.39
+0.32
+0.39
+0.40
+0.45
+0.39
+1.00
Structure
-1.00
-0.78
-0.26
-0.04
-0.73
+0.22
+0.93
+0.85
+0.85
+0.88
+0.87
+1.00
(a) Overall style similarity
t ud en
15 0 D
St
12 0 D
PR
90 D
PR O
O
60 D O
PR
30 O
PR
D
30 PR O
60
PD O
90
PD O
PD
ch e
150
O
60 90 120 Training Steps
Te a
30
0
−1
r
Teacher Side (−)
-0.65
PD
-0.3
-0.74
12 0
0.0
-0.69
O
0.3
-1.00
15 0
0.6
+1
Connectives
PD
OPD Student Side (+) OPRD
O
Δ Energy Distance
We evaluate Qwen3-8B students trained with OPD and OPRD on DAPO-Math-17K at updates 30, 60, 90, 120, and 150. For each checkpoint, we generate 16 responses to each of the 30 AIME'24 problems, yielding 480 responses. We use two fixed references: the Qwen3-4B teacher at GRPO update 75 and a separately GRPO-trained Qwen38B student at update 150. Each response is represented by 101 style features across five categories: connectives (15), modality (10), grammar (64), punctuation (8), and sentence and paragraph structure (4). The first three categories measure relative frequencies of function words, including connectives, modal and negation words, and grammatical words such as pronouns and articles. Punctuation features count occurrences per 1,000 words, while structure features capture the mean and standard deviation of words per sentence and sentences per paragraph. We standardize the features at every checkpoint of each method using a shared mean and standard deviation for each feature, computed from the 960 responses of the two reference models.
(b) Style similarity by category
Figure 18: (a) Comparing response style distributions over training. Energy-distance differences are averaged across 30 AIME'24 problems, with positive values indicating greater similarity to the GRPO student and negative values to the teacher. Shading shows 95% confidence intervals from 2,000 bootstrap resamples of the problems. (b) Measuring similarity to teacher and student response styles. Normalized distance differences compare average styles in five categories, with red indicating greater similarity to the teacher and green to the GRPO student. Both panels use fixed references: the Qwen3-4B teacher at GRPO update 75 and the Qwen3-8B GRPO student at update 150.
Figure 18a compares response style distributions using all 101 standardized features. For this comparison, we combine connectives, modality, and grammar into a group √ of 89√function-word √ features and scale the function-word, punctuation, and structure coordinates by 1/ 89, 1/ 8, and 1/ 4, respectively, to balance the three groups’ contributions. For each problem, let 𝑋𝑘 , 𝑋𝑇 , and 𝑋𝑆 denote the sets of 16 response feature vectors from the evaluated checkpoint at update 𝑘, the teacher, and the GRPO student, respectively. We compute the energy-distance difference ∆ED = ED(𝑋𝑘 , 𝑋𝑇 ) − ED(𝑋𝑘 , 𝑋𝑆 ). Energy distance measures differences between distributions by accounting for both between-set distances and within-set variation. We average ∆ED across the 30 problems, with positive values indicating greater similarity to the GRPO student and negative values to the teacher. Shading shows 95% confidence intervals from 2,000
35
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
bootstrap resamples of the problems. OPD shifts toward the teacher over training, with its mean difference decreasing from +0.135 at update 30 to −0.190 at update 150. OPRD remains closer to the GRPO student at every evaluated checkpoint, with its mean difference increasing from +0.316 to +0.447 over the same period. Figure 18b compares the same checkpoints separately across the five style categories. For each category 𝑐, we average the standardized features across all 480 responses without additional feature-group scaling. Let 𝜇𝑘,𝑐 , 𝜇𝑇,𝑐 , and 𝜇𝑆,𝑐 denote these average vectors for the evaluated checkpoint at update 𝑘, the teacher, and the GRPO student, respectively. We compute 𝑠𝑐 (𝑘) =
‖𝜇𝑘,𝑐 − 𝜇𝑇,𝑐 ‖2 − ‖𝜇𝑘,𝑐 − 𝜇𝑆,𝑐 ‖2 . ‖𝜇𝑇,𝑐 − 𝜇𝑆,𝑐 ‖2
The score measures the difference in distances to the two reference averages, normalized by their separation. Scores range from −1 to +1, with negative values indicating greater proximity to the teacher, positive values to the GRPO student, and zero indicating equal distance. OPD is closer to the GRPO student in all five categories at update 30 but closer to the teacher in all five by update 150. OPRD remains closer to the GRPO student in every category at all evaluated checkpoints, consistent with the overall distribution comparison.
36
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
K. Computational Cost and Memory Usage Benchmark Setup. To isolate method-specific training overhead from response-length differences, we force every generated response to contain exactly 16,384 tokens by ignoring EOS. Both benchmarks use the mathematics setting of Section 3.2, with a Qwen3-8B student and the Qwen3-4B teacher checkpoint at update 75. OPRD additionally loads the Qwen3-4B base policy as its reference. Within each benchmark, all methods receive prompts in the same order and use the same random seed. Wall-clock timing uses 64 prompts, whereas memory profiling uses 8 prompts, with 8 rollouts per prompt in both cases. All runs execute on four NVIDIA B200 GPUs with DP4, rollout TP1, BF16, Flash Attention 2, padding removal, and gradient checkpointing. Optimization uses a global minibatch of 64 responses and one PPO epoch per training step. Dynamic token batching caps each GPU at 36,864 tokens, which yields two complete samples per microbatch under the fixed-length setting. The rollout engine uses gpu_memory_utilization=0.60 and max_num_seqs=128. Sampling uses temperature 1.0, top-𝑝 1.0, and no top-𝑘 truncation. To keep reward-side computation identical, we replace task-specific reward evaluation with deterministic alternating binary rewards within each prompt group. Validation, periodic model saving, external logging, and all non-training diagnostics are disabled. Wall-Clock Time. We run each method in a fresh process, discard one complete warm-up step, and report the mean and sample standard deviation over the following four steps. As shown in Table 9, rollout generation is the largest component of each training step, taking roughly 602 seconds and accounting for 61.6% of the total GRPO time, with small differences across methods attributable to run-to-run variation. OPD and KDRL, each of which evaluates one frozen teacher, incur total overheads of 7.2% and 7.9% over GRPO, respectively. OPRD evaluates the teacher and its reference sequentially, increasing frozen-forward time from 60.46 seconds for OPD to 108.80 seconds. Because rollout generation dominates the step, this additional reference evaluation increases total time by only 4.4% over OPD, resulting in an overall overhead of 11.9% relative to GRPO. Student-forward time is effectively unchanged, while the update containing the teacher-direction projection and scaling increases by just 2.79 seconds over GRPO, equivalent to 0.25% of the full OPRD step. Beyond the teacher evaluation already required by OPD and KDRL, nearly all of OPRD’s additional runtime therefore comes from evaluating the reference policy. Table 9: Wall-clock time per training step. All methods process the same 64 prompts with 8 rollouts per prompt, with every response fixed at 16,384 tokens to equalize the number of generated tokens. Total time is reported as the mean ± sample standard deviation, while the component columns report their means. Rollout includes the complete generation call and the actor-to-rollout mode transition. Student forward is a no-gradient pass that recomputes the old log probabilities of the sampled tokens. The frozen-model column reports no-gradient teacher evaluation for OPD and KDRL and sequential teacher and reference evaluations for OPRD, while GRPO requires neither. Update includes a separate gradient-enabled student forward pass, backward propagation, and the optimizer step, excluding the separately timed frozen-model evaluations. Etc. includes reward construction, advantage computation, batch assembly and balancing, orchestration, and residual boundary costs. Total Method
Time (s/step)
Rollout Overhead
Student
Model Forward Student
Teacher + Ref
Optimization Update
Etc.
304.00 305.38 305.92 306.79
1.03 1.04 1.05 1.06
Qwen3-4B (Teacher) → Qwen3-8B (Student) GRPO OPD KDRL OPRD
977.72 ± 16.83 1048.00 ± 17.49 1055.41 ± 15.94 1094.44 ± 16.17
– +7.2% +7.9% +11.9%
602.20 610.74 618.20 607.39
70.49 70.38 70.08 70.40
0.00 60.46 60.17 108.80
37
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Peak GPU Memory. We measure peak GPU memory during rollout generation and the actor update, while separately recording the frozen-policy evaluation performed within the update. We also report the overall maximum observed during the complete training step. As shown in Table 10, OPRD carries a nearly constant additional footprint throughout training: approximately 14 GiB per GPU relative to GRPO and 13 GiB relative to OPD and KDRL. The two largest identifiable memory requirements within OPRD are the 3.75 GiB frozen teacher and reference parameter shards and a transient 9.43 GiB dense corrected-gradient allocation within the correction hook. Of the 3.75 GiB in frozen-model parameters, 1.87 GiB is additional relative to OPD and KDRL, which already retain the teacher. The 9.43 GiB hook allocation is also specific to OPRD. Because the absolute NVML peaks additionally include shared model and optimization state, allocator caches, and CUDA and distributed runtime state, these quantities identify the main OPRD-specific allocations but do not provide an exact additive decomposition of the observed peak difference. The overall maximum occurs during rollout for every method. OPRD reaches 151.27 GiB, exceeding GRPO by 13.98 GiB (10.2%) and OPD and KDRL by 13.21 GiB. Rollout itself increases memory by approximately 51–52 GiB for all four methods. The difference is already present before generation, where OPRD begins the measured step at 99.97 GiB, 14.93 GiB above GRPO and 13.25 GiB above OPD and KDRL. OPRD’s higher rollout peak therefore results from adding essentially the same generation-time allocation to a higher starting footprint, rather than from rollout requiring more memory. Frozen-policy evaluation is performed within the broader actor-update interval, and their maximum values coincide in our measurements. OPRD reaches 115.19 GiB during both frozen-policy evaluation and the full update, exceeding OPD and KDRL by 13.15 GiB. During the update, it also exceeds GRPO by 14.84 GiB. These differences closely match those observed before and during rollout, indicating that neither frozenpolicy evaluation nor the correction introduces a separate phase-specific increase in the device-memory peak. Within the correction hook, PyTorch-allocated memory grows by 9.43 GiB, matching the largest dense BF16 corrected-gradient tensor. By comparison, the sparse support formed by the sampled action and the student’s top-10 tokens occupies at most 1.38 MiB, and direct gather and sparse scatter avoid an additional 9.27 GiB response-by-vocabulary copy. The hook allocation is already contained within the 115.19 GiB update peak, which remains well below the overall maximum during rollout. Table 10: Peak GPU memory per training step. We profile 8 prompts with 8 rollouts per prompt and fix every response at 16,384 tokens. Each method is evaluated in four independent trials, each launched in a fresh process on four NVIDIA B200 GPUs with DP4/TP1 and BF16. Each trial discards one complete warm-up step and measures the following step. Whole-device NVML memory is sampled every 100 ms, and each Peak entry reports the mean across trials of the maximum usage over time and across the four GPUs within the indicated interval. All values are in GiB per GPU. Because each Peak entry represents the worst-GPU peak, multiplying it by four provides only a rough upper bound on aggregate device memory. Overall is the maximum over the complete step. Under Rollout, Peak − Start is the increase from the step-start baseline to the rollout peak. Teacher + Ref. Peak reports the maximum during frozen-model evaluation, covering one teacher evaluation for OPD and KDRL and sequential sequential teacher and reference evaluations for OPRD. This evaluation is a subinterval of Actor Update. Params. reports the calculated lower bound for the total BF16 parameter shards of the resident frozen models. Actor Update Peak includes resident models, gradients, optimizer states, activations, and runtime buffers, while Hook Growth reports the additional PyTorch allocation during the OPRD correction. Overall Method
Peak
Overhead
Rollout Peak
Peak − Start
Teacher + Ref Peak
Params
Actor Update Peak
Hook Growth
100.35 102.04 102.04 115.19
– – – 9.43
Qwen3-4B (Teacher) → Qwen3-8B (Student) GRPO OPD KDRL OPRD
137.29 138.06 138.06 151.27
– +0.6% +0.6% +10.2%
137.29 138.06 138.06 151.27
52.25 51.34 51.34 51.30
– 102.04 102.04 115.19
– 1.87 1.87 3.75
38