Conceptio › Archive › arXiv CS
arXiv CSopen access

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

arXiv:2609.20784v1 [cs.CL] 17 Sep 2026

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Yan Yu1∗, Zhengxi Lu1∗, Yizhou Liu1 , Yichen Pan1 , Aozhe Wang1 , Qipeng Chen1 , Hua Yang2 , Wenqi Zhang1 , Weiming Lu1 , Qianglong Chen2 , Yongliang Shen1† 1 2 Zhejiang University Alibaba Group {yuy2003, zhengxilu, syl}@zju.edu.cn

Abstract Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher’s success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting. Code is available at https://github.com/ZJU-REAL/SDAR. Aligned -> Conflict Skilled Teacher (Potentially restrictive)

t men Align ter Fas s" Student

Co

Environment (Task-grounded) s!

w t slo t bu r rec

😡 GRPO: Sparse Supervision Student

Environment

Response Token

y!

y"

😊 Efficient & Sustained Environment

Binary y#

…

y$

Student

Dense Supervision

Teacher

Env Adv (Unified)

Response Token

…

y!

y"

Skilled Teacher Guidance Environment Guidance Joint GRPO-OPD

y#

…

y$

😡 OPD: Teacher Limitation …

Joint Adv Guidance Type

Grounded Supervision

Student

Response Token Teacher Adv (Customized)

Teacher y!

y"

Later

Dense y#

… …

y$

Retired Adv

…

Figure 1: Left: GRPO and teacher guidance align early but later diverge, making teacher matching restrictive. Middle: GRPO’s supervision is sparse and task-level, while OPD is dense but teacherbound. Right: RetireOPD uses both early and retires the teacher at the conflict point, surpassing it. ∗ Equal contribution. † Corresponding author.

1

Introduction

Reinforcement learning with verifiable rewards (RLVR) is the standard approach for training large language model (LLM) agents on multi-turn tasks (Singh et al., 2025; Xu et al., 2026a; Shao et al., 2024), where a single outcome reward per trajectory leaves the intermediate decisions of a long interaction unsupervised. On-policy distillation (OPD) complements this sparse reward with dense token-level supervision from a teacher on trajectories sampled by the student (Zeng et al., 2026; Team, 2026; Agarwal et al., 2024; Lu & Thinking Machines Lab, 2025). In agentic training, the teacher is commonly the same model conditioned on privileged context that is available only during training, such as retrieved task skills (Zhao et al., 2026; Lu et al., 2026a). The student is trained to reproduce the teacher’s behavior without access to this context, so that the skills are internalized into its parameters and no additional context is required at inference (Lu et al., 2026b; Wang et al., 2026a). This paradigm rests on two assumptions. The teacher is assumed to be reliably better than the student because it sees the privileged context, and matching the teacher is assumed to remain useful for as long as training lasts. They amount to two questions that current methods answer by assumption or by a fixed schedule: which teacher the student should learn from, and for how long. We examine both on ALFWorld and WebShop and find that neither assumption holds (Figure 2). Fast start, higher ceiling Rapid alignment

Low ceiling

Gap rebounds

Privileged teacher is not consistently stronger.

Slow start

Figure 2: Training dynamics of the 3B student on ALFWorld. Left: success rate of the teacher branch and the student branch under GRPO+OPSD. Middle: success rate under GRPO, OPD, and GRPO+OPD. Right: teacher-student discrepancy under OPD and GRPO+OPD. Privileged context alone does not make a teacher reliable. Under joint Group Relative Policy Optimization (GRPO) (Shao et al., 2024) and on-policy self-distillation (OPSD) (Zhao et al., 2026), where teacher and student are two branches of one shared policy (Lu et al., 2026a), the skillconditioned branch does not consistently outperform the student it supervises (Figure 2, left). The shared parameters are optimized mainly through the skill-free student objective, so the policy is never trained to act on skill context. Model size does not help: a 7B model prompted with the same skills reaches only 23.4% success on ALFWorld (Section 4.4). A teacher must therefore be trained to use its privileged context before it can supervise. The benefit of teacher supervision is stage-dependent. Adding OPD to GRPO removes the slow start of pure GRPO and the low ceiling of pure OPD (Figure 2, middle). The teacher-student discrepancy, however, first narrows and then widens (Figure 2, right). Once the student has internalized the behavior that the skills induce, reward optimization favors actions the teacher does not take, and the two gradients begin to conflict. Continuing to match the teacher beyond this point holds the student near the teacher’s performance ceiling (Section 4.3). A first-order analysis (Appendix A) shows that a stagnating discrepancy implies locally opposed gradients. Teacher supervision is therefore temporary scaffolding, and the question is when to remove it. Existing work answers this question with a schedule fixed before training. Annealing methods decay the distillation weight over training (Tan et al., 2026; Ding et al., 2026), and two-stage recipes switch from distillation to RL at a predetermined step (Ye et al., 2026a; Li et al., 2026a). Neither observes whether the teacher is still useful. Online decisions exist only at finer granularity, per prompt or per token (Ding, 2026; Lu et al., 2026a), and none of them removes the teacher. In our experiments, the point at which the discrepancy stops decreasing varies from step 50 to step 90 across models and tasks (Appendix E), so any single schedule withdraws guidance too early in some settings and too late in others. These observations suggest a simple principle: whether the teacher is still useful can be 2

read from the training signal itself. The discrepancy stops decreasing exactly when the two objectives conflict, and the student’s success rate relative to the teacher indicates whether the skills have been internalized. The transition should therefore be determined online rather than fixed in advance. Motivated by these findings, we propose Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning (RetireOPD), which internalizes privileged skills into a skill-free student and retires the teacher once its supervision no longer benefits the student. RetireOPD first decouples the teacher from the student and trains it with environment rewards under the skill context, so that distillation starts from a policy that has already learned to exploit the skills. The student is then trained jointly with GRPO and OPD, while the two signals identified above are monitored at fixed intervals. Once the discrepancy stops decreasing and the student’s success rate reaches a set fraction of the teacher’s, the teacher is retired and training proceeds with GRPO alone. The student is thus no longer constrained by the teacher, and no further teacher forward passes are required. Across Qwen2.5-1.5B, 3B and 7B on ALFWorld (Shridhar et al., 2020) and WebShop (Yao et al., 2022), RetireOPD outperforms RL, distillation, and hybrid baselines, improving ALFWorld success rate over GRPO by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%. It also surpasses its own skill-conditioned teacher in every setting. Our contributions are as follows. • We identify two failure modes of privileged-information distillation for agents: an unoptimized skill-conditioned teacher is unreliable, and teacher matching conflicts with reward optimization once the student has internalized the teacher’s knowledge. • We propose RetireOPD, which trains a same-capacity skill-conditioned teacher with environment rewards and retires it online once the teacher-student discrepancy stops decreasing and the student reaches a set fraction of the teacher’s success rate. • Experiments across three model scales on ALFWorld and WebShop show that RetireOPD consistently outperforms RL, distillation and hybrid baselines while surpassing its own teacher, which validates its effectiveness on agentic tasks.

2

Related Work

2.1

On-Policy Distillation with Privileged Teachers

On-policy distillation (OPD) trains a student on its own samples with token-level feedback from a teacher (Agarwal et al., 2024; Lu & Thinking Machines Lab, 2025) and is now a standard post-training stage (Xiao et al., 2026; Zeng et al., 2026). Without a stronger teacher, on-policy self-distillation (OPSD) conditions the same model on privileged information, such as a reference solution or textual feedback, and distills the conditioned branch into the unconditioned one (Zhao et al., 2026; Hübotter et al., 2026; Ye et al., 2026b). For agents the privileged information is typically a retrieved or hindsight skill to be internalized (Lu et al., 2026a; Wang et al., 2026a; Zhang et al., 2026; Lu et al., 2026b). These methods keep teacher and student as two contexts of one shared policy and control the teacher signal per token by gating (Lu et al., 2026a), synchronization (Wang et al., 2026a), or advantage shaping (Zhang et al., 2026). Teacher reliability itself has received less attention, although guidance quality depends on compatibility with the student and on capability beyond it (Li et al., 2026b), and only some tokens carry useful signal (Xu et al., 2026b). RetireOPD instead trains the teacher separately with environment rewards before distillation. 2.2

Combining On-Policy Distillation with Reinforcement Learning

RLVR is the dominant approach for multi-turn agents (Shao et al., 2024; Feng et al., 2026; Dong et al., 2026; Yang et al., 2026; Lu et al., 2025), but its trajectory-level reward is sparse, so hybrid methods add a distillation term or a teacher-shaped advantage to the RL objective (Lu et al., 2026a; Zhang et al., 2026; Ding et al., 2026). Because the two signals interfere as the student improves, several methods reduce the teacher’s influence over training: annealing the distillation weight (Tan et al., 2026; Ding et al., 2026), expanding the trajectory horizon exposed to the student (Wang et al., 2026b), withdrawing in-context skills on a decaying budget (Lu et al., 2026b), or running distillation and RL in sequence (Li et al., 2026a; Ye et al., 2026a; Kim & Lee, 2026). In all of these the transition is fixed before training. Signal-dependent decisions exist only at finer granularity, such as distilling on prompts where all rollouts fail (Ding, 2026) or gating the teacher per token (Lu et al., 2026a), and 3

Joint GRPO-OPD Skill Internalization

Teacher Construction Task

Skill

OPD Loss Task

Teacher Observation

Skill

Student

𝜋(⋅∣ 𝑥, 𝑦"# )

Teacher

𝜋(⋅∣ 𝑥, 𝑐 ! , 𝑦"# )

Adaptive Retirement Alignment Progress vs.

𝐷$% (𝜋& ||𝜋 ' )

∆ = 𝑙𝑜𝑔 p ! − 𝑙𝑜𝑔 p"

Action

Environment Update

Reward

Rollout

Student

Relative Competence

GRPO Loss

Env

Traj 𝜏!

𝜏"

…

Adv A!

A"

… A#

𝜏#

vs. #$%&'%()*+$ 𝜂 = #$%&'%()*+$! "

Update

Skilled Teacher

𝐿!"#$%&" = 𝐿'()* + 𝜆𝐿*)+

𝜌, ≥𝛥!𝛿 and 𝜂 ≥ 𝛾

Figure 3: Overview of RetireOPD. (1) Teacher Construction: We optimize a skill-conditioned teacher with environmental rewards. (2) Joint Skill Internalization: the student learns from both GRPO and OPD training. (3) Adaptive Retirement: OPD supervision is removed by current policy. none removes the teacher. RetireOPD decides the global transition online from the teacher–student discrepancy and the student’s relative competence.

3

Method

Our method consists of three stages: (1) Teacher construction, where a skill-conditioned teacher is optimized with environment rewards (Section 3.1); (2) Joint GRPO-OPD training, where the skilled teacher supervises a skill-free student on student-generated trajectories (Section 3.2) and (3) Adaptive teacher retirement, where OPD is removed once behavioral transfer stagnates and the student reaches sufficient task competence (Section 3.3). 3.1

Privileged Teacher Construction

Task Definition We consider a multi-turn agent interacting with an environment over a sequence of decision steps. Given a task input x ∼ D, the agent iteratively generates responses based on the interaction history and receives environment feedback, forming a trajectory τ = (x, y), where y = (y1 , . . . , yT ) denotes the agent responses, and R(τ ) is the scalar environment reward. We initialize the teacher policy πϕ and student policy πθ from the same base model with identical architectures. During training, the teacher receives task-relevant skill context c+ as privileged information, and is optimized with GRPO to obtain a skilled teacher. The student has no access to c+ and remains skill-free at inference time. Our goal is to enable πθ to internalize behaviors induced by privileged skills through fine-grained teacher supervision. The teacher objective is defined as: ϕ∗ = arg max Ex∼D, τ ∼πϕ (·|x,c+ ) [R(τ )] . ϕ

(1)

Environment reward optimization encourages the teacher to turn privileged information into effective task behavior. After training, we freeze ϕ∗ and denote the resulting policy as the skilled teacher πT . The frozen teacher is subsequently used only to provide dense supervision for the student. 3.2

Joint GRPO-OPD Optimization

Following SDAR (Lu et al., 2026a), the student is trained without access to c+ and learns jointly from environment rewards and teacher supervision. The optimization objective is formulated as Lstudent (θ) = LGRPO (θ) + λLOPD (θ). 4

(2)

where λ controls the contribution of teacher supervision. GRPO. For each task input x, GRPO samples a group of G trajectories {τ (i) }G i=1 from the policy and their environment rewards {R(τ (i) )}G . The group-relative advantage for trajectory τ (i) is i=1 Â (i)

(i)

(i)

 R(τ (i) ) − Mean {R(τ (i) )}G i=1  . = std {R(τ (i) )}G i=1

(i)

(i)

(3)

(i)

Let rt (θ) = πθ (yt | x, y<t )/πθold (yt | x, y<t ) denote the token-level importance ratio between the current and behavior policies. We use the standard clipped GRPO objective: 

 |y (i) | G     X X 1 1 (i) (i) LGRPO (θ) = −E  min rt (θ)Â(i) , clip rt (θ), 1 − ϵ, 1 + ϵ Â(i)  . (4) G i=1 |y (i) | t=1 OPD. For each student-generated trajectory, we compare the student and teacher distributions conditioned on the same trajectory prefix y<t . The student predicts πθ (· | x, y<t ), while the frozen skilled teacher additionally conditions on the privileged skill context c+ , yielding πT (· | x, c+ , y<t ). The teacher provides supervision directly on states visited by the student and does not generate a separate trajectory. We define the OPD objective using the reverse KL divergence:  LOPD (θ) = Ex∼D, y∼πθ (·|x) 

|y| 1 X

|y| t=1

  DKL πθ (· | x, y<t ) ∥ πT (· | x, c+ , y<t )  .

(5)

Exact KL computation requires summing over the full vocabulary at each position, incurring substantial overhead. We instead use a sampled-token approximation on student-generated trajectories, evaluating each yt under both the student and teacher distributions: ∆t = log πT (yt | x, c+ , y<t ) − log πθ (yt | x, y<t ).

(6)

Since yt is sampled from the student policy, −∆t provides a single-sample Monte Carlo estimate of the reverse KL divergence at the corresponding position. 3.3

Adaptive Teacher Retirement

Teacher supervision helps early but may later constrain reward-driven optimization. Since this transition depends on student learning dynamics, a fixed retiring step may be premature or delayed. We therefore determine the retiring point adaptively from training dynamics. (i)

We divide training into monitoring windows of W steps. At step n, let yn be the i-th sampled trajectory. Using the token-level teacher-student log-probability gap ∆t defined in Section 3.2, we (i) write its t-th token gap as ∆t and compute the step-level and window-level alignment progress as (i)

|yn | G 1 X 1 X (i) Kn = ∆ , G i=1 |yn(i) | t=1 t

1 K̄m = W

mWX +W −1

Kn .

(7)

n=mW

To capture the dynamics of this alignment while avoiding premature retirement when the student is still weak, we define ρm and relative competence ηm as ρ(K) m =

K̄m − K̄m−1 , K̄m

ηm = 5

SRm + SRm−1 . 2SRT

(8)

where SRm and SRT denote the student and teacher success rates respectively. A negative ρm indicates that the teacher-student gap is still decreasing, whereas a non-negative one indicates that the reduction has stalled or reversed. ηm measures the student’s competence relative to the teacher, with student performance averaged over two consecutive windows to reduce short-term fluctuations. Teacher supervision is retired at the first monitoring window satisfying: n o m∗ = inf m ≥ 2 : ρ(K) m ≥ δ ∧ ηm ≥ γ .

(9)

where δ and γ are the discrepancy-change and relative-competence thresholds, respectively. Their combination avoids premature retirement caused by transient fluctuations when the student is still weak. Once Eq. 9 is satisfied, we remove OPD and continue training with GRPO alone.

4

Experiments

4.1

Experiment Setup

Benchmark We evaluate our method on two widely used interactive agent benchmarks: ALFWorld (Shridhar et al., 2020) and WebShop (Yao et al., 2022). ALFWorld is a text-based embodied environment where agents follow natural language instructions to complete multi-step household tasks, testing long-horizon planning, state tracking, and interaction with the environment. WebShop simulates realistic online shopping scenarios, requiring agents to search, navigate, and select products that satisfy user-specified constraints, thereby evaluating goal-directed decision making and multistep information gathering. Together, the two benchmarks cover complementary forms of interactive reasoning across embodied and web-based agentic tasks. Implementation We conduct experiments with Qwen2.5-1.5B / 3B / 7B-Instruct (Yang et al., 2024). For each model scale, the student and teacher share the same architecture and initialization. For adaptive teacher retirement, we evaluate the retiring criterion every H = 5 evaluation steps, with the competence threshold set to γ = 0.9 and the behavior-gap stagnation threshold to δ = 0 by default. Following SDAR (Lu et al., 2026a), we set the distillation coefficient λ0 to 0.01. For both environments, we use the SkillBank from SkillRL (Xia et al., 2026) as privileged information and adopt the default Keyword Matching strategy used in SDAR (Lu et al., 2026a) for skill retrieval. Baselines We compare RetireOPD with four groups of baselines: (1) Training-Free. Vanilla directly evaluates the instruction-tuned model, while Skill-Prompt provides task-relevant skills only at inference time. (2) Reinforcement Learning. GRPO (Shao et al., 2024), GiGPO (Feng et al., 2026), PPO (Schulman et al., 2017) and RLOO (Ahmadian et al., 2024) serve as representative RL baselines. Skill-GRPO (Xia et al., 2026) performs GRPO with privileged skills and serves as the skill-conditioned teacher. (3) Distillation. OPD (Ye et al., 2026b) and OPSD (Zhao et al., 2026) use dense teacher supervision without reward optimization. (4) Hybrid. GRPO+OPD and GRPO+OPSD (Lu et al., 2026a) combine reward optimization with dense teacher supervision, while RetireOPD additionally removes teacher supervision adaptively. 4.2

Main Results

Overall Performance As shown in Table 1, RetireOPD achieves the best performance across three model scales on both ALFWorld and WebShop. On ALFWorld, RetireOPD outperforms GRPO by 18.8 points on Qwen2.5-3B (93.8% vs. 75.0%), with similarly strong gains of +17.0 / +14.1 points on the 1.5B / 7B models. On WebShop, RetireOPD delivers similarly strong improvements, reaching 75.8% accuracy on Qwen2.5-1.5B compared with 56.8% for GRPO (+19.0 points), with gains of +14.0 / +11.8 points on the 3B / 7B models. RetireOPD also consistently outperforms GiGPO across both tasks and all model scales. These results demonstrate that RetireOPD provides consistent improvements across different tasks and model capacities. Reliable Teacher Construction RetireOPD substantially outperforms GRPO+OPSD, which does not explicitly optimize the teacher policy. This shows that privileged information alone is insufficient to provide a reliable supervision target, as access to additional context does not necessarily translate 6

Table 1: Main results on ALFWorld and WebShop. ∗ means validation with skills. The highlighted suffix of a distillation method denotes its teacher. Best and second-best results are highlighted. ALFWorld Method

Type

WebShop

Pick

Look

Clean

Heat

Cool

Pick2

Avg

Score

Acc

Qwen2.5-1.5B-Instruct Vanilla Prompt Skill-Prompt∗ Prompt GRPO RL GiGPO RL PPO RL RLOO RL Skill-GRPO∗ RL OPD- Skill-1.5B Distill OPD- 3B Distill OPSD Distill GRPO+OPD Hybrid GRPO+OPSD Hybrid RetireOPD Hybrid

11.1 3.4 85.3 94.4 64.8 88.3 66.7 85.7 25.8 26.3 96.4 87.1 94.9

0.0 16.7 53.7 67.5 40.5 52.8 63.6 46.2 64.7 16.7 88.9 66.7 100.0

6.2 12.9 84.5 94.8 57.1 71.0 100.0 73.1 7.1 9.1 89.3 85.7 90.0

0.0 5.3 78.2 94.4 60.6 62.8 100.0 55.6 21.9 6.7 90.5 58.3 100.0

0.0 0.0 58.7 79.8 46.4 66.4 75.0 76.7 0.0 9.1 73.7 85.0 84.2

4.2 5.0 53.5 76.4 47.4 56.9 100.0 53.3 9.1 5.3 82.6 46.2 77.3

5.5 6.2 72.8 86.7 54.4 69.7 81.2 71.1 5.6 14.1 87.5 72.7 89.8

17.8 20.8 75.8 83.1 73.8 73.9 85.2 84.7 32.5 22.3 82.5 79.9 86.6

5.5 1.6 56.8 65.0 51.5 52.1 67.9 69.1 3.9 10.2 69.9 66.8 75.8

Qwen2.5-3B-Instruct Vanilla Prompt Skill-Prompt∗ Prompt GRPO RL GiGPO RL PPO RL RLOO RL Skill-GRPO∗ RL OPD- Skill-3B Distill OPD- 7B Distill OPSD Distill GRPO+OPD Hybrid GRPO+OPSD Hybrid RetireOPD Hybrid

44.4 51.7 91.2 95.8 93.5 87.5 88.9 74.3 67.5 48.8 96.6 97.6 97.6

11.1 66.7 62.5 81.8 76.5 55.6 76.9 80.0 33.3 41.7 90.0 58.3 83.3

6.2 48.4 96.2 96.4 90.5 71.4 100.0 76.9 37.5 16.7 83.3 91.7 100.0

15.4 0.0 61.9 100.0 50.0 54.5 77.8 69.2 43.8 0.0 87.5 92.3 85.7

28.6 4.3 65.0 88.9 81.2 59.3 74.3 81.8 19.0 15.8 66.7 63.6 95.5

12.5 10.0 47.4 86.4 60.0 80.0 62.5 58.8 7.7 16.7 78.9 31.2 83.3

21.9 28.9 75.0 92.2 81.2 72.7 79.7 74.2 38.3 28.1 82.8 78.1 93.8

6.7 0.2 79.8 83.1 76.0 81.8 79.7 79.1 10.4 11.3 84.5 77.8 85.8

0.8 0.8 63.3 69.5 59.4 68.0 64.8 66.4 2.3 3.1 74.2 66.4 77.3

Qwen2.5-7B-Instruct Vanilla Prompt Skill-Prompt∗ Prompt GRPO RL GiGPO RL PPO RL RLOO RL Skill-GRPO∗ RL OPD- Skill-7B Distill OPD- 14B Distill OPSD Distill GRPO+OPD Hybrid GRPO+OPSD Hybrid RetireOPD Hybrid

36.1 51.7 91.2 97.7 92.3 87.6 97.1 100.0 84.2 50.0 100.0 94.3 100.0

22.2 50.0 87.5 82.7 64.0 78.2 100.0 77.8 40.0 60.0 76.9 58.8 100.0

3.1 32.3 96.2 98.8 92.5 87.3 92.3 76.0 77.8 22.7 95.2 81.0 100.0

0.0 5.3 81.0 83.7 89.5 81.3 100.0 100.0 42.1 21.4 100.0 94.4 76.9

0.0 4.3 65.0 89.3 80.3 71.9 85.0 72.2 66.7 17.6 82.6 88.2 96.4

0.0 0.0 57.9 79.2 68.8 48.9 72.7 90.5 47.4 9.5 87.5 50.0 90.9

12.5 23.4 81.2 90.8 80.4 75.5 90.6 86.7 64.8 32.8 92.2 79.7 95.3

5.9 1.7 80.9 84.4 81.4 80.3 88.7 88.9 28.8 4.5 90.3 86.8 91.7

1.6 0.8 72.6 72.8 68.7 65.7 78.9 78.1 4.7 2.3 81.2 76.5 84.4

into better task behavior. Explicitly optimizing the skill-conditioned teacher with environment rewards converts this information advantage into a more consistent behavioral advantage, providing a stronger distillation signal for the student. These results suggest that teacher quality depends more on how effectively privileged information is used than on model capacity or additional context alone. Adaptive Teacher Retirement RetireOPD consistently outperforms GRPO+OPD, which retains teacher distillation throughout training, indicating that keeping supervision can restrict subsequent reward-driven optimization. Moreover, RetireOPD surpasses the corresponding skilled teacher across all three model scales. For example, on Qwen2.5-3B-Instruct, RetireOPD achieves 93.8% success 7

Ours

GRPO+OPD (w/o retire)

Teacher-Student Gap

Success Rate

1.0 0.8 0.6 0.4 0.2

0

50 100 Training Step

150

GRPO+OPD -> OPD

0.0 −0.1 −0.2 −0.3

0

50 100 Training Step

150

Figure 4: Training dynamics of the Qwen2.5-3B-Instruct on ALFWorld under RetireOPD, continued GRPO+OPD, and GRPO+OPD→OPD.

100 73.4

40

73.4

80

60

78.1

78.1

75.8

75.8

GRPO

Train w/ exit

Train w/o retire Train 92.2 w/ retire 91.4 90.6 89.8 86.7 82.8 93.8 79.7 90.6 89.8 91.4

86.7

79.7 82.8

60 28.9 40

20 0

28.9

23.4

20 0

3B

23.4

3B

7B7B Skill-3B Skill-3B Teacher TeacherModel Model

Skill-7B Skill-7B

1.0 1.0

Success Rate

Success Rate (%)

Success Rate (%)

80

Train w/o exit

Teacher

Success Rate

Teacher

100

0.8

0.8

0.6 0.4

Ours LA (Skill-3B)

Ours

LA (7B)

LA (Skill-3B) LA (7B)

0.6 0.4

0.2 0.2 0 0

50 50 100 100150 Training Training Step Step

150

Figure 5: Ablation study of teacher construction and teacher Figure 6: Adaptive Retirement vs. retirement on ALFWorld. The dashed line indicates the GRPO Linear Annealing (LA) on ALFbaseline without teacher supervision. World. on ALFWorld and 77.3% Acc on WebShop, compared with the teacher’s 79.7% and 64.8%. These results suggest that adaptive retirement preserves the benefit of early teacher guidance while allowing the student to continue improving through reward-driven optimization beyond the teacher. 4.3

Training Dynamics

As shown in Figure 2, integrating OPD into GRPO accelerates early-stage policy optimization compared with GRPO, while avoiding the premature performance plateau commonly observed with OPD. As training progresses, the teacher-student gap first narrows and then widens. The initial decrease indicates effective knowledge transfer from the teacher, whereas the later rebound suggests that the student begins to diverge from the teacher distribution. Under OPD, the gap continues to decrease toward zero. This pattern suggests that, as the student improves, reward-driven GRPO updates increasingly conflict with the behavioral alignment objective of OPD. To further verify this conflict and determine which objective should be retained, we continue training from the checkpoint at the retirement point under three strategies. Figure 4 shows that continued joint optimization leads to a performance plateau while the teacher-student discrepancy keeps widening. Switching to OPD further reduces the gap, but fails to improve task performance and can even lead to degradation, suggesting that stronger teacher alignment may constrain further student improvement. In contrast, removing OPD and continuing with GRPO steadily improves the success rate from 76.6% at the retiring point to 93.8%. This pattern suggests that reward-driven GRPO updates increasingly conflict with the behavioral alignment objective of OPD as the student improves. 4.4

Ablation Studies

Teacher Construction. Figure 5 examines how explicit teacher optimization affects teacher quality and downstream student learning. Without explicit optimization, the 3B and 7B teachers achieve 8

Table 2: Retiring mechanism ablation. Setting

Table 3: Parameter ablation.

Retire @

SR

Setting

60 50 100 – 0

92.2 89.0 89.0 82.8 75.0

γ = 0.9, δ = 0 (Ours) γ = 0.8, δ = 0 γ = 1.0, δ = 0 γ = 0.9, δ = 0.03 γ = 0.9, δ = 0.05

RetireOPD (Ours) w/o Alignment Stagnation w/o Relative Competence w/o Retirement w/o OPD (GRPO-only)

Retire @

SR

60 55 95 65 95

92.2 90.6 91.4 89.1 91.4

only 28.9% and 23.4%, indicating that simply increasing model capacity does not improve skill utilization. After optimization with environmental rewards, their performance rises to 79.7% and 90.6% respectively, and consistently leads to stronger student performance. When OPD is retained throughout training, Skill-7B leads to substantially better student performance than Skill-3B (89.8% vs. 82.8%), revealing the student’s dependence on teacher quality. With adaptive retirement, this gap largely disappears (91.4% vs. 92.2%), suggesting that a reliable teacher is important for early guidance, while the retirement mechanism reduces long-term dependence on the teacher. Retiring Criteria. Table 2 shows that each component of adaptive retirement contributes to performance. Retaining OPD throughout training achieves only 82.8%, while the complete method reaches 92.2%, showing that timely removal of teacher supervision is crucial for further optimization. Using either alignment progress or relative competence alone reaches 89.0%, whereas combining both signals achieves the best performance, indicating that the two criteria are complementary. Performance Sensitivity to Retiring Thresholds. Table 3 examines the sensitivity to γ and δ. Varying either threshold shifts the selected transition step from 55 to 95, while success rates remain relatively stable, ranging from 89.1% to 92.2%, indicating reasonable robustness over a broad parameter range. Performance does not vary monotonically with the transition step, suggesting that RetireOPD does not rely on finely tuned thresholds or a single precise switching point. Adaptive Retirement vs. Linear Annealing. To examine whether the transition should follow a predefined schedule, we replace the adaptive criterion with the linear annealing strategy of ATOD (Tan et al., 2026), implementation details are provided in Appendix B. As shown in Figure 6, linear annealing improves rapidly early in training but later approaches a plateau, while the adaptive strategy continues to improve. This suggests that teacher utility does not necessarily decay according to a predefined schedule, and using training dynamics to determine the transition at the appropriate stage better preserves early guidance while avoiding later constraints.

𝛾 ∈ 0.80,0.96 Exit within ±5 steps

𝛿 ∈ -0.10,0.04 Exit within ±5 steps

Figure 7: Sensitivity of the teacher retiring step to the competence threshold γ (left) and signed relative stagnation threshold δ (right) in the 3B setting. 4.5

Retiring Timing Robustness

We further investigate how the two triggering thresholds affect the retiring timing. As shown in Figure 7, the identified retiring point remains stable over a broad range of parameter values. Specifically, the retiring step stays within ±5 steps of the default setting for γ ∈ [0.80, 0.96] and δ ∈ [−0.10, 0.04]. 9

Noticeable delays arise only under relatively extreme threshold values. This suggests that the adaptive retiring criterion consistently identifies a similar transition stage during training without requiring precise threshold calibration.

5

Conclusion

We proposed RetireOPD, which internalizes privileged skills into LLM agents without requiring such information at inference time. It first optimizes a skill-conditioned teacher with environment rewards, then trains a skill-free student jointly with GRPO and on-policy distillation, and adaptively retires teacher supervision once behavioral transfer stagnates and the student reaches sufficient task competence. Across three model scales on ALFWorld and WebShop, RetireOPD consistently improves performance and outperforms pure reinforcement learning, continuous joint optimization of GRPO and OPD and the corresponding skill-conditioned teachers.

References Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In International Conference on Learning Representations, 2024. Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker. Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12248–12267, 2024. Ken Ding. Hdpo: Hybrid distillation policy optimization via privileged self-distillation. arXiv preprint arXiv:2603.23871, 2026. Yifan Ding, Xincheng Wei, Yoshua Y Li, Ziheng Li, Yuquan Lu, Siyu Zhang, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, and Yun Chen. Saf-opd: Stable advantage fusion for on-policy distillation. arXiv preprint arXiv:2607.29209, 2026. Guanting Dong, Hangyu Mao, Kai Ma, Licheng Bao, Yifei Chen, Zhongyuan Wang, Zhongxia Chen, Jiazhen Du, Huiyang Wang, Fuzheng Zhang, et al. Agentic reinforced policy optimization. In International Conference on Learning Representations, volume 2026, pp. 16981–17017, 2026. Lang Feng, Zhenghai Xue, Tingcong Liu, and Bo An. Group-in-group policy optimization for llm agent training. Advances in Neural Information Processing Systems, 38:46375–46408, 2026. Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, et al. Reinforcement learning via self-distillation. arXiv preprint arXiv:2601.20802, 2026. Jaehoon Kim and Dongha Lee. Opsd compresses what rlvr teaches: A post-rl compaction stage for reasoning models. arXiv preprint arXiv:2605.06188, 2026. Boyan Li, Bingsen Chen, Chenghao Yang, Ping Nie, Chen Zhao, and Xi Ye. Sequential beats joint: On the interplay between on-policy distillation and rlvr. arXiv preprint arXiv:2609.04108, 2026a. Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huanang Gao, Wenkai Yang, Zhiyuan Liu, et al. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe. arXiv preprint arXiv:2604.13016, 2026b. Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/ on-policy-distillation. Zhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, et al. Ui-s1: Advancing gui automation via semi-online reinforcement learning. arXiv preprint arXiv:2509.11543, 2025. 10

Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang, Jinyang Wu, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, et al. Self-distilled agentic reinforcement learning. arXiv preprint arXiv:2605.15155, 2026a. Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Chengcheng Han, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, and Yongliang Shen. Skill0: In-context agentic reinforcement learning for skill internalization. arXiv preprint arXiv:2604.02268, 2026b. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768, 2020. Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267, 2025. Qitai Tan, Zefang Zong, Yang Li, and Peng Chen. Atod: Annealed turn-aware on-policy distillation for multi-turn autonomous agents. arXiv preprint arXiv:2606.27814, 2026. Qwen Team. Qwen3. 5-omni technical report. arXiv preprint arXiv:2604.15804, 2026. Hao Wang, Guozhi Wang, Han Xiao, Yufeng Zhou, Yue Pan, Jichao Wang, Ke Xu, Yafei Wen, Xiaohu Ruan, Xiaoxin Chen, et al. Skill-sd: Skill-conditioned self-distillation for multi-turn llm agents. arXiv preprint arXiv:2604.10674, 2026a. Jiaqi Wang, Wenhao Zhang, Weijie Shi, Yaliang Li, and James Cheng. Tcod: Exploring temporal curriculum in on-policy distillation for multi-turn autonomous agents. arXiv preprint arXiv:2604.24005, 2026b. Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, et al. Skillrl: Evolving agents via recursive skill-augmented reinforcement learning. arXiv preprint arXiv:2602.08234, 2026. Bangjun Xiao, Bingquan Xia, Bo Yang, Bofei Gao, Bowen Shen, Chen Zhang, Chenhong He, Chiheng Lou, Fuli Luo, Gang Wang, et al. Mimo-v2-flash technical report. arXiv preprint arXiv:2601.02780, 2026. Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, et al. Deepseek-v4: Towards highly efficient milliontoken context intelligence. arXiv preprint arXiv:2606.19348, 2026a. Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, and Alborz Geramifard. Tip: Token importance in on-policy distillation. arXiv preprint arXiv:2604.14084, 2026b. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024. Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang, Wenqi Zhang, Jintao Chen, and Xuhong Zhang. Reconciling process supervision with outcome-based credit in agentic policy optimization. arXiv preprint arXiv:2608.31077, 2026. 11

Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757, 2022. Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng, Yuhang Mu, Wenchao Du, and Yiming Wang. Opdsearch+: On-policy distillation with rl refinement for search-augmented reasoning. arXiv preprint arXiv:2608.24310, 2026a. Tianzhu Ye, Li Dong, Xun Wu, Shaohan Huang, and Furu Wei. On-policy context distillation for language models. arXiv preprint arXiv:2602.12275, 2026b. Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, et al. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026. Yanfei Zhang, Xu Lin, and Chenglin Wu. Stepopsd: Step-aware online preference distillation for agent reinforcement learning. arXiv preprint arXiv:2605.27140, 2026. Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, and Aditya Grover. Self-distilled reasoner: On-policy self-distillation for large language models. arXiv preprint arXiv:2601.18734, 2026.

12

A

Theoretical Analysis

Setup.

We analyze GRPO-OPD interaction through a local teacher-regularized view. Let J(θ) = Ex,τ ∼πθ [R(τ )],  D(θ) = Ex,y∼πθ 

|y| 1 X

|y| t=1

 DKL πθ (· | x, y<t )∥πT (· | x, c+ , y<t )  , 

Fλ (θ) = J(θ) − λD(θ),

g = ∇θ J,

h = ∇θ D.

Ignoring finite-sample noise and assuming inactive GRPO clipping locally, the expected update is θ+ = θ + η(g − λh), where g is the reward-driven direction and −h reduces the teacher–student discrepancy. Interaction between GRPO and OPD.

A first-order expansion gives

+

J(θ ) − J(θ) = ∇θ J(θ), θ+ − θ + O(η 2 ) = η ⟨g, g − λh⟩ + O(η 2 )  = η ∥g∥2 − λ⟨g, h⟩ + O(η 2 ), ⟨g, h⟩ < 0 ⇒ OPD accelerates reward optimization, ⟨g, h⟩ > 0 ⇒ OPD conflicts with reward optimization. Moreover, if λ⟨g, h⟩ > ∥g∥2 , the joint update decreases the expected return to first order. Discrepancy as a Conflict Signal.

Under joint optimization,   d dθ D(θ) = ∇θ D(θ), dt dt = ⟨h, g − λh⟩ = ⟨g, h⟩ − λ∥h∥2 .

Thus, a non-decreasing discrepancy indicates local conflict between the GRPO direction g and the OPD direction −h. If h ≈ 0, OPD already provides little additional first-order guidance. Discrepancy stagnation therefore signals either emerging conflict or a diminishing OPD signal. Benefit of Teacher Exit.

Suppose GRPO and OPD are locally conflicting, i.e., ⟨g, h⟩ > 0. Consider + θjoint = θ + η(g − λh), + θexit = θ + ηg.

Their expected returns satisfy h i + + J(θexit ) − J(θjoint ) = J(θ) + η∥g∥2 h i − J(θ) + η ∥g∥2 − λ⟨g, h⟩ + O(η 2 ) = ηλ⟨g, h⟩ + O(η 2 ). Since ⟨g, h⟩ > 0, teacher exit yields a larger first-order reward improvement than continued joint optimization for sufficiently small η. Implication for Adaptive Exit. The analysis shows that OPD accelerates learning while teacher guidance aligns with reward optimization, whereas stagnation of the teacher-student discrepancy indicates diminishing OPD benefit or emerging conflict with GRPO. We therefore use stagnation of K̄j as the exit signal and additionally require sufficient student competence to avoid premature exit caused by early training noise. 13

B

Implementation Details

We implement all methods under a unified training setup to ensure a consistent comparison. Table 4 summarizes the optimization, data, and method-specific hyperparameters used in our experiments. Methods with identical configurations are grouped into the same column for compact presentation. Table 4: Hyperparameter settings. G denotes the group size, λ is the distillation loss coefficient, and αKL is the KL-divergence penalty coefficient. OP(S)D collectively denotes OPD and OPSD, while GRPO+OP(S)D denotes GRPO+OPD and GRPO+OPSD. GRPO, GiGPO and RLOO are grouped because they use the same hyperparameter configuration. A dash indicates that the corresponding hyperparameter is not applicable. Method Hyperparameter Learning rate Critic learning rate G λ αKL Train batch size Validation data size Competence threshold Window size Training steps

RetireOPD GRPO/GiGPO/RLOO −6

1×10 – 8 0.01 0.01 16 128 0.9 5 150

−6

1×10 – 8 – 0.01 16 128 – – 150

PPO

OP(S)D GRPO+OP(S)D

−6

1×10 1×10−6 1×10−5 – – 8 – 0.01 0.01 0.01 16 16 128 128 – – – – 150 150

1×10−6 – 8 0.01 0.01 16 128 – – 150

ATOD Linear Annealing Except for the loss-weight schedule, all experimental settings of ATOD are identical to those of our main method. The training objective is defined as LATOD (s) = λRL (s)LGRPO + λKL (s)LOPD , where s denotes the current training step. We define the annealing progress as s  p(s) = min ,1 . 80 The two loss weights are scheduled as λKL (s) = 1 − 0.9p(s),

λRL (s) = 0.1 + p(s).

At the beginning of training, the GRPO and OPD weights are 0.1 and 1.0, respectively, such that optimization is primarily guided by the teacher distillation signal. As training proceeds, the GRPO weight gradually increases while the OPD weight decreases. From step 80 onward, the two weights remain fixed at 1.1 and 0.1, respectively. This schedule smoothly shifts the optimization from OPD-dominated learning to reward-driven GRPO while retaining a small amount of distillation regularization.

C

Case Study

To provide a more fine-grained view of how teacher supervision affects policy learning, we examine token-level teacher scores on representative ALFWorld trajectories. We focus on two complementary cases: how privileged skills improve supervision during early exploration, and how the same teacher supervision can become restrictive after the student has acquired sufficient task competence. Skill-informed supervision during early exploration. Figure 8 compares our skilled teacher with the skill-conditioned branch used in OPSD on the same student-generated trajectory. Most routine actions, such as navigation and object manipulation, receive similar scores from the two teachers. The difference becomes pronounced at decision-critical states that require task-specific knowledge. At Step 8, the agent has reached the sink while holding the target lettuce, and the retrieved skills explicitly indicate that the object should be cleaned before placement. Accordingly, our skilled 14

teacher assigns substantially higher likelihood to the key token clean (−0.826) than the OPSD teacher (−16.625). This example illustrates that privileged information is useful only when the teacher can reliably translate it into the corresponding behavior. Training the skilled teacher explicitly helps convert such high-level guidance into a stronger token-level learning signal for the student. Step State summary

Retrieved skill

Skilled teacher score

OPSD teacher score

0 initial search

Clean before placing

go to drawer 1

go to drawer 1

1 drawer search

Search likely containers

open drawer 1

open drawer 1

2 cabinet search after Search likely containers open cabinet 10 empty drawer 3 moving to likely food Search likely food sur- go to diningtable surface faces 1

open cabinet 10

4 lettuce found on din- Pick target object ing table

take

take

from diningtable 1

from diningtable 1

5 checking object state

examine lettuce 2

examine lettuce 2

go to sinkbasin 1

go to sinkbasin 1

examine lettuce 2

examine lettuce 2

Verify object state

6 lettuce held; cleaning Use sink for cleaning required 7 before cleaning Verify object state

lettuce

2

go

to

1

8 at sink; cleaning Use sink for cleaning clean lettuce 2 should happen Clean before placing with sinkbasin 1

clean

9 after cleaning; place Place after cleaning target object

move

move

lettuce

countertop 1

2

to

diningtable

lettuce

2

lettuce

2

with sinkbasin 1 lettuce

2

to

countertop 1

Figure 8: Early-stage teacher comparison on ALFWorld. Both teachers score the same studentgenerated action tokens. The retrieved skill provides privileged guidance available only to the skilled teacher. At the critical cleaning step, the skilled teacher assigns substantially higher likelihood to the key token clean. Teacher conflict after competence improves. Figure 9 shows a successful trajectory in which the student takes a task-effective action that is strongly disfavored by the teacher. In particular, the transition toward the microwave receives a low teacher score despite leading to successful task completion. This illustrates that, once the student becomes sufficiently competent, continued teacher matching can conflict with reward-driven optimization, motivating adaptive teacher exit. Step State summary 1 searching for tomato

Student action colored by teacher score take tomato 2

2 holding tomato; needs go to fridge 1 cooling 3 holding cooled tomato; go to microwave 1 microwave needed 4 near microwave put tomato in microwave

Env. return 1 1 1 1

Figure 9: Token-level teacher conflict on a successful ALFWorld trajectory. Token scores indicate the teacher’s preference for the student’s generated actions. Despite receiving positive environment feedback, the student’s correct transition toward the microwave receives strong negative teacher supervision.

15

D

Algorithm

The complete training procedure of our method, including privileged teacher construction, joint GRPO-OPD training, and adaptive teacher retirement, is summarized in Algorithm 1.

Algorithm 1 R ETIRE OPD Training Pipeline Require: Base policy π0 ; training set D = {(x, c+ )}; validation set Dval ; environment reward R; group size G; OPD weight λ; monitoring window W ; exit thresholds γ, δ ▷ Phase 1: Privileged Teacher Construction 1: Initialize teacher policy πϕ ← π0 2: for each teacher training iteration do 3: Sample (x, c+ ) ∼ D 4: 5:

+ {y (i) }G Ri ← R(x, y (i) ) i=1 ∼ πϕ (· | x, c ); (i) (i) πϕ (yt | x, c+ , y<t ) Ri − µ G T Âi ← , ri,t ← (i) (i) σG πϕ (yt | x, c+ , y<t ) old

T }) 6: Update ϕ by minimizing LTGRPO (ϕ; {Âi , ri,t 7: end for 8: πT ← StopGrad(πϕ ); SRT ← Success(πT , Dval ; c+ )

▷ Phase 2: Joint GRPO–OPD with Adaptive Exit 9: πθ ← π0 ; E XITED ← false; m ← 0 10: for each student training iteration n do 11: Sample (x, c+ ) ∼ D 12: // Step 1: Skill-free on-policy rollout (i)

(i)

Ri ← R(x, yn ) {yn }G i=1 ∼ πθ (· | x); // Step 2: GRPO objective (i) (i) πθ (yn,t | x, yn,<t ) Ri − µ G 15: Âi ← , ri,t ← (i) (i) σG πθold (yn,t | x, yn,<t ) 16: if not E XITED then 17: // Step 3: On-policy distillation (i) (i) (i) (i) (i) 18: ∆n,t ← log πT (yn,t | x, c+ , yn,<t ) − log πθ (yn,t | x, yn,<t ) 1 PG 1 P|yn(i) | (i) 19: Kn ← LOPD ← −Kn t=1 ∆n,t ; i=1 (i) G |yn | 20: Update θ by minimizing LGRPO + λLOPD 21: else 22: Update θ by minimizing LGRPO 23: end if 24: if n mod W = 0 and not E XITED then 25: // Step 4: Adaptive teacher exit 1 Pn Kq 26: m ← m + 1; K̄m ← W q=n−W +1 27: SRm ← Success(πθ , Dval ) 28: if m ≥ 2 then K̄m − K̄m−1 SRm + SRm−1 (K) 29: ρm ← ; ηm ← 2SRT K̄m (K) 30: if ρm ≥ δ and ηm ≥ γ then 31: E XITED ← true 32: end if 33: end if 34: end if 35: end for 36: return skill-free student policy πθ 13: 14:

16

E

Training Dynamics

We present the full training dynamics of RetireOPD across all model scales and environments in Figures 10–13. RetireOPD on ALFWorld

1.5B / ALFWorld

Success Rate

1.0

0.8

0.6

0.6

0.4

0.4

0.2

0.2 0

100

150

3B / ALFWorld

1.0 Success Rate

50

0.0

0.8

0.8

0.6

0.6

0.4

0.4

0.2

0.2

0.0

0

50

100

150

7B / ALFWorld

1.0

0.0

0.8

0.8

0.6

0.6

0.4

0.4

0.2

0.2

0.0

0

50 100 Training Step

0

150

0.0

50

100

150

3B / WebShop n * = 80

0

50

100

150

7B / WebShop

1.0

n * = 60

Post-exit

n * = 90

1.0

n * = 60

Exit

1.5B / WebShop

1.0

n * = 75

0.8

0.0

Success Rate

RetireOPD on WebShop

n * = 85

0

50 100 Training Step

150

Figure 10: Validation Score when training with Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on ALFWorld and WebShop.

17

Relative Discrepancy Change ρm Relative Discrepancy Change ρm Relative Discrepancy Change ρm

RetireOPD on ALFWorld

Plateau threshold ε = 0

RetireOPD on WebShop

1.5B / ALFWorld

Exit step

1.5B / WebShop

n * = 75

n * = 90

0.1

0.1

0.0

0.0

−0.1

−0.1

−0.2

−0.2

−0.3

−0.3 0

50

100

150

0

3B / ALFWorld

50

0.1

0.0

0.0

−0.1

−0.1

−0.2

−0.2

−0.3

−0.3 100

150

0

7B / ALFWorld

100

150

n * = 85

0.1

0.1

0.0

0.0

−0.1

−0.1

−0.2

−0.2

−0.3

−0.3 50 100 Training Step

50

7B / WebShop

n * = 60

0

150

n * = 80

0.1

50

100

3B / WebShop

n * = 60

0

Post-exit replay

150

0

50 100 Training Step

150

Figure 11: Relative Discrepancy Change when training with Qwen2.5-1.5B-Instruct, Qwen2.5-3BInstruct and Qwen2.5-7B-Instruct on ALFWorld and WebShop.

18

RetireOPD on ALFWorld

RetireOPD on WebShop

Retirement step

1.5B / ALFWorld

Teacher retired

1.5B / WebShop

n * = 75

n * = 90

Policy Entropy

1.0 0.8 0.6 0.4 0.2 3B / ALFWorld

3B / WebShop

n * = 60

n * = 80

Policy Entropy

1.0 0.8 0.6 0.4 0.2 7B / ALFWorld

7B / WebShop

n * = 60

n * = 85

Policy Entropy

1.0 0.8 0.6 0.4 0.2 0

50 100 Training Step

150

0

50 100 Training Step

150

Figure 12: Entropy Curve when training with Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on ALFWorld and WebShop.

19

RetireOPD on ALFWorld

RetireOPD on WebShop

Retirement step

1.5B / ALFWorld

Episode Reward

10

Teacher retired

1.5B / WebShop

n * = 75

n * = 90

8 6 4 2 0 3B / ALFWorld

Episode Reward

10

3B / WebShop

n * = 60

n * = 80

8 6 4 2 0 7B / ALFWorld

Episode Reward

10

7B / WebShop

n * = 60

n * = 85

8 6 4 2 0 0

50 100 Training Step

150

0

50 100 Training Step

150

Figure 13: Reward Curve when training with Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on ALFWorld and WebShop.

20

F

Prompt

We use a unified prompt format for agent interaction across environments. The student receives task observations and action history, while the teacher additionally accesses privileged skill information during OPD. Prompt template on ALFWorld You are an expert agent operating in the ALFRED Embodied Environment. Your task is to: {task_description} Prior to this step, you have already taken {step_count} step(s). Below are the most recent {history_length} observations and the corresponding actions you took: {action_history} You are now at step {current_step} and your current observation is: {current_observation} Your admissible actions of the current situation are: [ {admissible_actions} ]. Now it’s your turn to take an action. You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags. Once you’ve finished your reasoning, you should choose an admissible action for the current step and present it within <action> </action> tags. Figure 14: Student prompt template for the ALFWorld task environment.

Prompt template on WebShop You are an expert autonomous agent operating in the WebShop e-commerce environment. Your task is to: {task_description}. Prior to this step, you have already taken {step_count} step(s). Below are the most recent {history_length} observations and the corresponding actions you took: {action_history} You are now at step {current_step} and your current observation is: {current_observation}. Your admissible actions of the current situation are: [ {available_actions} ]. Now it’s your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within <think> </think> tags. Once you’ve finished your reasoning, you should choose an admissible action for the current step and present it within <action> </action> tags. Figure 15: Student prompt template for the WebShop task environment.

Privileged teacher input template [Privileged Skill Information] {skill_text} {student_prompt} {student_response_prefix} Figure 16: Teacher-side input used for OPD. The teacher receives the same student trajectory prefix, with privileged skill information prepended to the prompt.

21

Record · ID 978456 · SHA-256 afce88a7783f3cee
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.